Here Come the 2024 ZiPS Projections!

Once again, it’s time for me to fire up my computer and crank out the yearly team-by-team ZiPS projections. This is where I’d normally do my shtick, but we have a lot to get to, so imagine a quote from a 19th century personality, an allusion to a 13th century battle, and a 1980s pop culture reference, and then cram them all together for your own haute couture Szymborski pablum! We’ve got business to take care of, so no time for shenanigans.
ZiPS is a computer projection system I initially developed in 2002–04. It officially went live for the public in 2005, after it had reached a level of non-craptitude I was content with. The origin of ZiPS is similar to Tom Tango’s Marcel the Monkey, coming from discussions I had in the late 1990s with Chris Dial, one of my best friends (my first interaction with Chris involved me being called an expletive!) and a fellow stat nerd. ZiPS quickly evolved from its original iteration as a reasonably simple projection system, and now does a lot more and uses a lot more data than I ever envisioned it would 20 years ago. At its core, however, it’s still doing two primary tasks: estimating what the baseline expectation for a player is at the moment I hit the button, and then estimating where that player may be going using large cohorts of relatively similar players.
So why is ZiPS named ZiPS? At the time, Voros McCracken’s theories on the interaction of pitching, defense, and balls in play were fairly new, and since I wanted to integrate some of his findings, I wanted my system to rhyme with DIPS (defense-independent pitching statistics), with his blessing. I didn’t like SIPS, so I went with the next letter in my last name, Z. I originally named my work ZiPs as a nod to CHiPs, one of my favorite shows to watch as a kid. I mis-typed ZiPs as ZiPS when I released the projections publicly, and since my now-colleague Jay Jaffe had already reported on ZiPS for his Futility Infielder blog, I decided to just go with it. I never expected that all of this would be useful to anyone but me; if I had, I would have surely named it in less bizarre fashion.
ZiPS uses multi-year statistics, with more recent seasons weighted more heavily; in the beginning, all the statistics received the same yearly weighting, but eventually, this became more varied based on additional research. And research is a big part of ZiPS. Every year, I run hundreds of studies on various aspects of the system to determine their predictive value and better calibrate the player baselines. What started with the data available in 2002 has expanded considerably. Basic hit, velocity, and pitch data began playing a larger role starting in 2013, while data derived from StatCast has been included in recent years as I’ve gotten a handle on its predictive value and the impact of those numbers on existing models. I believe in cautious, conservative design, so data is only included once I have confidence in improved accuracy; there are always builds of ZiPS that are still a couple of years away. Additional internal ZiPS tools like zBABIP, zHR, zBB, and zSO are used to better establish baseline expectations for players. These stats work similarly to the various flavors of “x” stats, with the z standing for something I’d wager you’ve already guessed.
How does ZiPS project future production? First, using both recent playing data with adjustments for zStats, and other factors such as park, league, and quality of competition, ZiPS establishes a baseline estimate for every player being projected. To get an idea of where the player is going, the system compares that baseline to the baselines of all other players in its database, also calculated from whatever the best data available for the player is in the context of their time. The current ZiPS database consists of about 140,000 baselines for pitchers and about 170,000 for hitters. For hitters, outside of knowing the position played, this is offense only; how good a player is defensively doesn’t yield information on how a player will age at the plate.
Using a whole lot of stats, information on shape, and player characteristics, ZiPS then finds a large cohort that is most similar to the player. I use Mahalanobis distance extensively for this. A CompSci/Math student at Texas A&M did a wonderful job showing how I do this, though the variables used aren’t identical.
As an example, here are the top 50 near-age offensive comps for World Series MVP Corey Seager right now. The total cohort is much larger than this, but 50 ought to be enough to give you an idea:
Ideally, ZiPS would prefer players to be the same age and position, but since we have about 170,000 baselines, not 170 billion, ZiPS frequently has to settle for players nearly the same age and nearly the same position. The exact mix here was determined by extensive testing. The large group of similar players is then used to calculate an ensemble model on the fly for a player’s future career prospects, both good and bad.
One of the tenets of projections that I follow is that no matter what the projection says, that’s what the ZiPS projection is. Even if inserting my opinion would improve a specific projection, I’m philosophically opposed to doing so. ZiPS is most useful when people know that it’s purely data-based, not some unknown mix of data and my opinion. Over the years, I like to think I’ve taken a clever approach to turning more things into data — for example, ZiPS’ use of basic injury information — but some things just aren’t in the model. ZiPS doesn’t know if a pitcher wasn’t allowed to throw his slider coming back from injury, or if a left fielder suffered a family tragedy in July. I consider those sorts of things outside a projection system’s purview, even though they can affect on-field performance.
It’s also important to remember that the bottom-line projection is, in layman’s terms, only a midpoint. You don’t expect every player to hit that midpoint; 10% of players are “supposed” to fail to meet their 10th-percentile projection and 10% of players are supposed to pass their 90th-percentile forecast. This point can create a surprising amount of confusion. ZiPS gave .300 batting average projections to three players in 2021: Luis Arraez, DJ LeMahieu (yikes!), and Juan Soto. But that’s not the same thing as ZiPS thinking there would only be three .300 hitters. On average, ZiPS thought there would be 34 hitters with at least 100 plate appearances to eclipse .300, not three. In the end, there were 25; the league BA environment turned out to be five points lower than ZiPS expected, catching the projection system flat-footed.
Another crucial thing to bear in mind is that the basic ZiPS projections are not playing-time predictors, at least with players without firm possession of a full-time job in the majors. By design, ZiPS has no idea who will actually play in the majors in 2024. ZiPS is essentially projecting equivalent production; a batter with a .240 projection may “actually” have a .260 Triple-A projection or a .290 Double-A projection. But telling me how Julio Rodríguez would hit in a full-time role in the majors in 2022 was a far more interesting use of a projection system than it telling me that he would only play a partial season (in the end, quite obviously, he played a full year). For the depth charts that go live in every article, I use the FanGraphs Depth Charts to determine the playing time for individual players. Since we’re talking about team construction, I can’t leave ZiPS to its own devices for an application like this. It’s the same reason I use modified depth charts for team projections in-season. There’s a probabilistic element in the ZiPS depth charts: sometimes Joe Schmo will play a full season, sometimes he’ll miss playing time and Buck Schmuck has to step in. But the basic concept is very straightforward.
What’s new in 2024? Outside of the typical calibration updates, there’ll be an extra table in this year’s projections. Don’t worry, the 80/20 splits are returning, but I’m adding split projections into the team-by-team rundowns as well. Usually I create these for the benefit of companies using my projections for their baseball games and calculate it sometime in February. But this year, I successfully integrated that model into ZiPS and, after repairing all the things I broke doing so, platoon splits are now being spit out with the usual array of numbers.
Have any questions, suggestions, or concerns about ZiPS? I’ll try to reply to as many as I can reasonably address in the comments below. If the projections have been valuable to you now or in the past, I would also urge you to consider becoming a FanGraphs Member, should you have the ability to do so. It’s with your continued and much appreciated support that I have been able to keep so much of this work available to the public for so many years for free. Improving and maintaining ZiPS is a time-intensive endeavor and reader support has enabled me to have the flexibility to put an obscene number of hours into its development. It’s hard to believe that ZiPS is now 20 years old. Hopefully, the projections and the things we’ve learned about baseball have provided you with a return on your investment, or at least a small measure of entertainment, whether you’re delighted or enraged.
Dan Szymborski is a senior writer for FanGraphs and the developer of the ZiPS projection system. He was a writer for ESPN.com from 2010-2018, a regular guest on a number of radio shows and podcasts, and a voting BBWAA member. He also maintains a terrible Twitter account at @DSzymborski.
Have you studied whether more contemporary comps are at all better? There are all kinds of nutrition, training, etc sports science things that are available to modern players that were not available to older players.
I wonder if it some point you would underweight a 1950s comp relative to a 2010s comp.
Proximity matters too, and I have certain adjustment factors in as well (for example, contrary to conventional wisdom, 1950s pitchers had more time lost to injury).
It’s often interesting to note how little the pitching staffs of great teams resemble stereotypical notions. The 1939 Yankees had no starters who threw 30 starts, and just one who threw more than 200 innings. The 1953 Yankees (the winningest of the five consecutive WS winners, by a pip) were very similar, differing in this framing only in that Whitey Ford got to 30 starts on September 27th.
The Cy Young Award was created in 1956. I wonder if that led to establishment of the stereotypical 5-man rotation by (e.g.) increasing the perceived value of pitcher-wins? Just a guess.
Diamond Mind is still planning to release a season disc this year, correct? Hopefully before June?
DMB and FG are my last solid connections to the favourite sport of my youth. I’d really hate DMB to go under!
In the past you’ve said that, for minor leaguers, production trumps all. But the prevailing sentiment among analysts these days is that certain metrics – zone contact%, chase%, 90th percentile EV, etc. – are more predictive than simply assessing a player’s output. Has ZiPS come around on this idea yet, or is it still holding firm to your initial position?
I don’t think I *quite* said it this way, only that the overall production is super important. ZiPS has long used these numbers in translations where possible.
Does ZiPS weight within the season at all? For example, if 2 players had the exact same overall stats and underlying stats for the full year last but one player was great in the 1st half and terrible in the 2nd half and the other was terrible the 1st half and great the 2nd half, would ZiPS view them differently with all else being equal?
And, if so, does age make any difference? Like if a 20-year old rookie struggled greatly in April and May and was great in August and September, would ZiPS be more willing to assume that player actually improved over the season?
This made me wonder if it’s useful to try to account for less linear within-season performance variance than the original comment is discussing? Does it increase predictive ability if you include some measure to differentiate between two hitters who ended up with 120 WRC+’s and similar stats/rates, but one hitter consistently posted monthly WRC+’s of 110-130 while the other bounced up and down from 50-150 over the course of the season, for example.
This investigation is worthy of an entire article.
“Even if inserting my opinion would improve a specific projection, I’m philosophically opposed to doing so.”
—If I’m going to be wrong then, by God, I’m going to be wrong on principle!
You said the team-by-team projections are coming out. What about individual players? Will FG members have access to platoon splits?
Oh, that’s what I mean! Those team-by-team rundowns.
Everyone will have access to platoon splits, member or no.
I remember in a chat you mentioned that the new rules for 2023 were not explicitly included in ZiPS, but how long would you expect for those impacts to really impact a players projection? E.g., would it be reasonable to expect Acuna’s stolen base total to be around 60? Or do you think it will still be tempered closer to pre-2023 numbers?
Me: “Who could ever get mad at numbers? They’re just numbers!”
Also me (When ZiPS doesn’t unequivocally say my favorite team will win the World Series):
“I hate you Dan”
Decision-type things stabilize quickly; since the decision to attempt a steal is a *decision* then recent history most important.
I have questions about what ZiPS knows and doesn’t know.
I’ve always wondered how ZiPS produces story-telling lineup-dependent stats like runs and RBIs. I’m assuming it doesn’t know who else is in the order with the player, or where in the order the players bats? So does that mean they’re just the result of regression based on the players other stats, e.g. a player with X,Y,Z slash stats typically produces A,B R and RBI? Or does it know about what the player did in recent years and adjust accordingly (without knowledge of changing teammates)?
Is it the same sort of thing for pitchers and W-L? Seems like innings/start might be something ZiPS knows, but surely it doesn’t know whether the pitcher plays for a team with a good offense?
Thank you Dan, very cool!
“Usually I create these for the benefit of companies using my projections for their baseball games”
Are these agreements confidential? Can you say what games utilize your projections? Really, I’m just fascinated and would be interested to know more about this whole process.
Diamond Mind and OOTP.
How long before you integrate AI into the model in order to 1) learn from the model’s previous forecasts and 2) potentially spot correlations previously unknown?
Well, that’s the whole idea behind dimensionality reduction!
I tried some dimensionality reduction but it left me feeling a little flat
Thanks Dan.
You say you want the data to speak, and I applaud, but how far do you go? Do you adjust/control for team? If a Ray and a Rockie are young players otherwise identical, do they get the same evaluation?
While the idea of Dan having a “ROCKIES!” button on the zips machine to manually clunk a dude brings me joy, it seems the Colorado FO is more than capable of doing that themselves, and given the time and effort to make it this far, I doubt he’d duplicate their impressive effort
Yup. Excepte for parks of course.
“Szym has spent nearly half his life working on his numbers so that everyone can get mad at them.”
Oh, we’re not mad at the numbers, Dan; we’re mad at YOU! (jk, very exciting to hear that next year’s projections are ready to be unveiled.)
It can be both.
Congrats on 20 years of ZIPs, Dan! I bet when you made your first set of projections you didn’t imagine it would grow into the beast of a system it’s turned into today. Great work!
I didn’t even imagine they’d turn out to be useful.
Dan, as usual I am really looking forward to these projections! Thank you for all the thought and hard work to produce them in a timely fashion.
Jay
The title says “Here Come the Zips” but if I recall, you roll out the Zips projections on a team by team basis and the entire spreadsheet is not accessible on Fangraphs Leaders’ Board until the last team has rolled out, sometime in February. Is that all correct? Thanks.
The team write-ups are great but hoping we can get access to all players sooner for those of us with fantasy dynasty drafts over the winter.
Dan, you’re my favorite baseball person besides Francoeur. Thanks for being awesome!
Great job as always Dan, thanks for the detailed explanations.
Here’s a free article topic I’m wondering. If a player underperforms in the first month relative to their pre-season zips wOBA, how likely are they to eventually reach that pre-season wOBA by the end of the season?
It would really depend on the level of underperformance, wouldn’t it?
yes, and on whether they get the chance to recover (don’t mind me; I’m just looking straight at Miguel Vargas)
I love FG and I am a happy subscriber. I want you to know that ZiPS projections are one of the reasons I’ll keep that membership going for as long as you’ll take my money. Thank you, Dan.
CHiPs is so classic! Who could forget the Roller Disco episodes. The Marina Del Rey singles scene. And the freeze frame credits at the end!
Thank you for doing this work; it’s enormously interesting and insightful. You mentioned that the latest release will include “the typical calibration updates”. Without giving out more information that you’re comfortbale with, can you tell us which data field contributed most significantly to the calibrations? For example, injuries data, minor league equivalencies, Statcast numbers, etc.?
It’s hard to believe that, next year, ZiPS will be old enough to drink.
I have several questions:
1. When looking at “actuals”, it’s it better to use bWAR or fWAR to compare to a projection or is it something like .5fWAR + .5bWAR or is pitching closer to fWAR while hitting is closer to bWAR?
2. Is framing in your ZIPS model and if yes is it as material as fWAR framing ?
3. Let’s say the model that you publish uses player A at 600 plate appearances and 3 WAR. At the end of the year he has 200 plate appearances and 2.2 WAR. In comparing projected to actual is it fair to simply use proportionality and say “he would have projected at 1 WAR in 200 PA so he over performed by 1.2 OR does playing time impact projection in a different way?
4. When a player has a down season in base running or fielding does ZIPS react the same as as it would for changes in offensive statistics. In other words, Tommy Edman had dramatically lower results in both in 2023. How does that impact his baseline projection ?
Has Corey Seager ever been one of Kyle Seager’s top 50 comps?
Dan, will you be exploring AI to Zips
What concerns do you have?