Let’s Make Sure We’re Honest About Projections
Over at Baseball Prospectus, it’s PECOTA day. Now, that means a whole lot of different things, but one of the things it means is that now BP gets to unveil its projected 2018 standings. Some years ago, I used to get extremely excited about seeing the first standings projections. Then FanGraphs decided to start hosting projections year-round. They’re always right here, and for months, that information has been based on the Steamer projection system. Pretty soon, well in advance of opening day, the ZiPS system will get folded in, to say nothing of remaining transactions. The point being, having projected standings is fun. They serve to keep the mind occupied with thoughts of baseball.
Projected standings aren’t just a toy for the public. I mean, the public ones are, but teams also have their own private projections, that might be better, or that might be basically the same. Team projections drive perceptions, and team projections drive transactions. We feel like we have a pretty good idea of which teams are situated to contend. We also feel like we have a pretty good idea of which teams are far away. Right now, in 2018, based in part on the projections and in part on what just happened a year ago, we have the sense we’re in an era of super-teams, which might be keeping the market slow. Other teams might not feel like they’re close enough to invest.
I love having access to team projections. I use them all the time for analysis and articles. But I feel like I should remind you of the limitations. This is something I could probably write every single year. I’m sure on some level you already know what I’m going to say. Projections are pretty good. They can also end up very, very far off.
I’ll get right to the point. As I’ve noted probably dozens of times, I have a spreadsheet of preseason team projections, going back to 2005. The types of projections have changed over time, so in that sense I don’t have perfect consistency of sources, but all projections are fundamentally similar. So, 13 years of data, for 390 team-seasons. Here are actual wins, against preseason projected wins.
The linear relationship is obvious, and it had better be obvious, because, otherwise, boy, would we all have been wasting our time. It’s perfectly clear that projections can spot good teams and bad teams. But our R2 is 0.36. I don’t know what mark would be “good enough,” but there’s clearly plenty of over- and under-performance. And let’s say you’re of the opinion that it’s not fair to group everything together, since projections have gotten more advanced. Our overall R2, again, is 0.36. It’s been higher than that just once over the past four seasons. The projections on my sheet actually peaked in 2007.
Looking at the overall sample, the average error has been about seven wins, with a median of six. Looking at just the last five years, the average error is still about seven wins, with a median of five. I don’t think we’re yet to the point where we can assign individual teams specific and individual error bars, but we know there are error bars, and they’re fairly long. It might be our own fault for not displaying any. That probably serves to suggest we have more confidence in our projected midpoints. But take a look at our Steamer-projected standings. Imagine that every projection could be off by seven wins. There’s only a 12-win difference between, say, the Red Sox and the Rays. Same with the Indians and the Twins. Fifteen wins would presently separate the Dodgers from fourth place. We know that there are some obvious favorites. They’re just all more vulnerable than they appear.
I have a couple tables to get to, but first, consider that, from 2005 – 2017, there were 44 teams projected to win at least 90 games. Of those teams, 28 actually won at least 90 games, while 36 won at least 85. Four teams fell apart and finished under .500: the 2005 Dodgers, the 2008 Tigers, the 2012 Red Sox, and the 2013 Angels. I’ll note that the 2015 Nationals didn’t quite finish under .500, but they did fall a dozen wins shy of their projection, after seeming like a super-team all offseason long. Things happen. That’s the whole point of baseball.
Sticking with the theme of under-performance, I looked at every team that fell short of the preseason projection by at least ten wins. This table breaks it down by year.
| Year | Overall | .500+ |
|---|---|---|
| 2005 | 3 | 2 |
| 2006 | 4 | 2 |
| 2007 | 1 | 1 |
| 2008 | 5 | 3 |
| 2009 | 6 | 3 |
| 2010 | 4 | 2 |
| 2011 | 6 | 4 |
| 2012 | 5 | 4 |
| 2013 | 4 | 3 |
| 2014 | 5 | 3 |
| 2015 | 8 | 5 |
| 2016 | 3 | 0 |
| 2017 | 4 | 4 |
| Total | 58 | 36 |
We’ve got 58 teams that meet the criteria, for an average of about four and a half per season. The worst year was 2015, when eight teams fell at least ten wins short. The best year was 2007. Meanwhile, we’ve got 36 teams that meet the criteria, after being projected to play at least .500 baseball. That’s an average of nearly three teams per season. Only once has there been a season without such an underachiever. Just last year, there were four: the Giants, Tigers, Mets, and Blue Jays. The Giants’ projection wound up being off by a full 25 wins, which is so far the largest gap I’ve observed.
Flipping things around, we can take a similar look at over-performance. I looked at every team that exceeded its preseason projection by at least ten wins.
| Year | Overall | sub-.500 |
|---|---|---|
| 2005 | 3 | 3 |
| 2006 | 3 | 2 |
| 2007 | 2 | 2 |
| 2008 | 6 | 4 |
| 2009 | 4 | 3 |
| 2010 | 5 | 3 |
| 2011 | 4 | 1 |
| 2012 | 6 | 3 |
| 2013 | 4 | 2 |
| 2014 | 2 | 1 |
| 2015 | 6 | 2 |
| 2016 | 2 | 2 |
| 2017 | 5 | 4 |
| Total | 52 | 32 |
We’ve got 52 teams that meet the criteria, for an average of four per season. In the worst years — for the projections — there have been six such over-achievers. There have yet to be fewer than two. Meanwhile, we’ve got 32 teams that meet the criteria, after being projected to play sub-.500 baseball. That’s an average of about two and a half teams per season. There hasn’t yet been a year without at least one such over-achiever. Just last year, there were four: the Diamondbacks, Brewers, Yankees, and Twins. The Rockies just missed being included. Have to draw a line somewhere.
Team projections serve a valuable purpose. I don’t want to stray too far from that message — they’re good, and they’re important, and they help to explain why the moves that happen happen. Projections inform the probabilities, and the teams with the strongest projections are indeed the teams most likely to keep playing in October. But for all we talk about how various teams are projected, we might not talk enough about how wide the error bars can be. How wide the error bars have been proven to be. It can feel, sometimes, like the playoff picture is already decided in March. Like we know exactly whose Septembers will be interesting. Maybe that’s how it’ll work out this time. Based on the history, though — nope. We’re not as good at this as projected-standings pages might convey. Humans happen — humans play — and they tend to do very human things.
Jeff made Lookout Landing a thing, but he does not still write there about the Mariners. He does write here, sometimes about the Mariners, but usually not.

Projections are designed to be the median outcome from a theoretical list of outcomes, with each outcome assigned some probability of occurring. They aren’t claiming to be the exact outcome of what they predict.
The Giants in 2017, for example, likely performed at a level that was assumed to have a 1 or 2% change of happening. Incredibly unlikely, but it happens. Someone wins the lottery each time, despite the incredibly long odds.
There is this one other thing.
Assume that a deviation from the projection such as 2017 Giants has 1% probability.
The problem is that there are 30 teams.
Assuming that the deviation are independent (which is semi-justifiable),
in any given year, there is 26% chance that at least one of the teams would give such an “unlikely” outcome.
There is a massive difference between the probability that a particular team under/overperforming a lot and at least one team under/overperforming.
Stating the obvious here, but the accuracy of the projections is complicated by in-forecastable things like sequencing.
IE, if I gave you the stats for every player on each team, you’d have error bars on runs scored/allowed. And even if you had that, you’d still have error bars around wins and losses.
Sure, but there’s no forecasting which one of the 30 teams is going to achieve the 1% outcome. If you asked that question before the season, the only reasonable answer is “probably none of them.” The fact that one of them does in maybe a quarter of the seasons doesn’t change that.
And the deviations aren’t independent. When one team loses a lot, the corresponding wins go elsewhere. That changes what teams do in the middle of the season. Teams that might have sat pat at the trade deadline go all in. Others that might have sat pat at the trade deadline decide to sell. This is part of the reason the spread at the end of the season is always much greater than the projections could reasonably suggest.
Yes, you are correct. The projections themselves are more than what’s displayed. This is essentially just a reminder of that. It’s far too easy and tempting to misinterpret the midpoints that get posted on standings pages. There is a lot of room for very different standings outcomes to be achieved, even without too many actual surprises or improbabilities.
I sometimes wish projections were presented as the entire probability distribution for the team, instead of just whittling it down to a “mean” that isn’t anywhere near as informative or useful.
“we know there are error bars, and they’re fairly long. It might be our own fault for not displaying any. That probably serves to suggest we have more confidence in our projected midpoints.”
I mean… yeah. The way you choose to display the data communicates a lot about what you think it should be interpreted to mean — in this case, falsely. I think this is worth more than a couple of sentences and a shrug.
I have been interested in those error bars for a while. I’d be very interested in seeing them. My guess is that they are really, really big, especially once we get into the tails.
Using this data as a crude reference… 4 teams beat their projections by 10+ games and 4 teams will under-perform their projections by 10+. 30 People flipping a coin 162 times would have, on average, 2 people getting 91 or more heads, and 2 people getting 71 or less heads. I’m using this as a comparison, because this would be essentially knowing that a team is a true talent 81 win team (sort of… a true talent 81 win team wouldn’t have every game be 50% chance of winning, but it’s close enough).
Binomial distribution (coin flip) will have standard deviation of 6.4 wins in 162 game sample. Solving the normal distribution function, so that on average 4 people will flip at least 91 heads, the standard deviation becomes 8.6.
Here are the curves for 30 trials… blue binomial (p=.5), red is normal distribution stdev = 8.6.

IMO… I think it’s rather impressive. What I believe most people don’t realize is how high the uncertainty is even if you know the true talent.
Thanks, BigChief. I agree that is impressive, for the reasons you mention.
Well in fangraphs defense, they do publish % chance to make the playoffs, which is data that directly communicates their certainty.
I’ve actually studied projections (for other things) and found a common projection bias is over-extremism. That is, lower projections are too low and higher projections are too high. Can you publish the raw data of the scatter plot? It is easy to verify; you just look for a correlation between the projection error (projected minus actual) and the projection itself. A real positive correlation (0.1 range in a big data set like this?) is a red flag that it is over-extremist. A real negative correlation indicates it is under-extremist. Based on the fact that 62 percent of your ’10 wins worse’ data set was projected ‘winning’ teams and 62 percent of your ’10 wins better’ data set was projected ‘losing’ teams, I would guess it is over-extremist.
if anything, baseball projections are usually the opposite. For example, no team was projected to be as good as the Dodgers/Astros/Indians last year – 100 win projections are exceptionally rare as you can see from that graph, but 100 win seasons happen.
The same is true on the low end, with no team this year being projected for 100 losses: there’s a reasonable shot that one of KC, DET or MIA gets to 100 losses.
Baseball projections are rolling discrete results into a mean; there’s a good chance that median projections would make more sense. To look over at PECOTA for example, Justin Turner is projected for only 3 wins, which would be a lower than recent, but that’s the product of a bunch of forecasts where he’s injured and doesn’t play much, and other forecasts where he’s a 4 or 5 win player.
I bet there’s not actually that many “JT hitting .280 and only being a league-average player at 3B” instances in PECOTA’s forecast. It’s a bunch of 4.5s, 5s, 6s and 0.2, 0.5, and 1s. The 2.5 win seasons are ones where he only plays 1/2 a season, much moreso than they are seasons where his overall performance declines but he plays all year.
Just because it’s the average result doesn’t mean its the most likely result.
And for that matter, that’s why its important to look more at the projected WAR than anything else. To use PECOTA again, PECOTA only projects True Average (BP’s equivalent of wOBA) and turns it into WARP. The component stats are then reverse engineered.
PECOTA projects players to accumulate certain win totals; it doesn’t predict that they will do so by hitting for .040 higher batting average or hitting 10 more home runs.
To your first point…. this isn’t evidence of what the OP is suggesting. Say you had 30 weird ass coins that you knew for sure each had probabilities of landing on heads somewhere between 0.4 and 0.6. Your projections would have the highest number of heads after 162 games be 97 wins. Obviously this projection has zero bias so it is absolutely correct, but I’d be willing to bet my 401k that one of the coins will end up with more than 97 wins.
You may already know this, but I see all the time people comparing the highest win totals each year to the highest projected win totals each year as evidence that projections are too conservative.
Also, are all projections actually calculated this way or is this feature unique to PECOTA? I was under the impression that most projections systems simple weight weight to recent years performance and then apply appropriate regression.
I don’t think that is true.
Looking at the plot, there was only one team projected to win more than 98 games
and only one team projected to win less than 64 games.
During the same time frame, dozens of teams managed to do either.
Yes, part of it is due to mid-season trades, but that can’t be the whole story.
Your last sentence is not a proof of over-extremism since
the distribution of performance is not symmetric.
A team whose median projection is 91 wins is much more likely to get 81- wins than 101+ wins.
That in itself does not mean that the projection is guilty of over-extremism.
Re: Mikejunt, real extreme values will always be higher or lower than the projected extremes, that is basic stats. What I am talking about is bias. Like I said it’s easy to check by correlating error to projection.
Re: Tung_twista, see above regarding your first point. Even a good projection system with no biases will have lets say 10 teams projected for 95-98 wins in this sample, and real values will fall with a standard deviation of about 7 around that number, so expected real values will range around mid-80’s to mid 100’s. So yes that’s a bunch of teams over a hundred with a perfect projection system with no bias. Basics. Regarding your second point, that is absolutely true as well and that is part of the problem. What are we trying to project, the median of the team outcomes, or the mean? Skewness will always have a long tail toward the center (toward 81 wins). We should be careful to project the median and not the mean.
As an aside, one way this type of error could happen is by double-counting statistical wins that won’t translate into real wins. Consider for a second WAR to be discrete events, whereby player A has a great game on day X and uses one of his WAR up. If another player ‘uses’ a WAR on the same day, it is gone to waste.
I don’t see why this is a problem.
Projections are supposed to be the median outcome.
Here is a simple example.
Assume a team’s true win distribution is as follows.
76 wins: 20%
86 wins: 30%
96 wins: 40%
106 wins: 10%
The asymmetrical shape is typical as central outcome is more likely.
This team has median projection of 91 wins.
What is its mean? 90 wins.
So you talk about the danger of projecting the mean instead of the median yet you also suggest over-extremism when projecting the mean would actually over-moderate what the (median) projections should be.
Going back to your last sentence,
“Based on 62 percent of your ’10 wins better’ data set was projected ‘losing’ teams, I would guess it is over-extremist.”
I understood this to mean that since 62 is significantly higher than 50, it suggests over-extremist and I have shown that if the projections are perfect, it is supposed to be over 50.
Now maybe you have an estimate of what the actual value should be (say 55) and then think 62 is higher than 55, therefore projections are over-extremist.
If so, you should tell us what your model is and how you reached that conclusion.
Otherwise, you are not providing any proof, circumstantial or not, of over-extremism.
Yes, without having access to the raw data, I am only suspecting. The issue is in how the projections are created without centralizing the adjustments after the fact. For example, a method to project by minimizing the total error or error^2 tends to be over-extremist. I have found many excellent sports projection systems, some that have higher correlations to actual outcomes than even las vegas lines, but they have this major downfall; the correlation is higher but it is not 1:1 projection to reality. It may be that 85% projection and 15% mean value (as adjustment) makes the model SUBSTANTIALLY more accurate as a real life predictor. All you have to do is play around with the value (the 85% value) until you find that the correlation between error and prediction is zero. Bingo, bias eliminated. For an under-extremist projection system, this value would actually be over 100 percent, but these are rare.
Steamer and ZIPs have the opposite skew, moderation.
Yeah, I’d expect there to be the opposite but not for reasons the others suggested…
I believe the preseason projections are the result of a simulated season based on a team’s skill, which is based on a teams current depth chart.
So the projections wouldn’t be taking into account in season roster changes. This wouldn’t matter if the roster changes were random, but I can’t imagine they would be random. Good teams +500 teams, would be more likely to trade mid-season to get better now, while the bad teams, would be more likely to trade current talent for future talent. So there is a systematic reason to think a change in talent would correlate with current talent…
I may be completely off on what is going on in the projections. It is strange that published data tables here hint at over-extremeism but if what I’m assuming is correct, I’d be amazed if this is the bias.
I’m curious what this looks like if you could incorporate injuries. Remove teams that had “significant” injuries, however that might be defined (regardless of whether they under- or over-performed).
That would help to isolate variance down to differences in production, as opposed to playing time.
After reading this article, I decided to look at Astros SP projection from January 2017 on the wayback machine and compare how well FIPs and innings pitched matched the actual numbers. Team numbers jumped out at me. Astros projected rotation FIP and WAR were 3.93 and 15.2. Astros actual rotation FIP and WAR were 3.95 and 15.2 despite injuries, trades, and having to resort to getting 111 innings out of Peacock in the rotation.
Interesting. Almost ties into toss’s point below – its essentially speaks to the depth. The Mets would likely be on the other end of this.
I suspect this kind of analysis is already done, but am wondering if teams that rely on a few elites have higher chances of under performing, whereas those with fewer holes tend to do as projected.
You could argue its the opposite, because a team that has a stars-and-scrubs roster gets a lot of benefit if one of those scrubs breaks out and is a 2.5 or 3 win player by surprise, but a team that has 2 win players everywhere gets a much smaller marginal benefit from the same improvement.
I love this point. I’ve been thinking about his a lot primarly because of the discussion around Hosmer. Like if Hosmer is truly super volatile like his past suggest, would a team really pay more $/war for a guy who’s gonna give him 2.5 wins each of the next 4 seasons, or someone who will give you seasons with wins like 4, 0 5, 1 wins. I’d like to think that the 4 and 5 win season would be so much more valuable than the 2.5 win season the more volatile guy might actually be worth more.
Its not because you don’t know what you’re getting until you got it. The stable production is more valuable because you can expect it to be there. Its only after the fact that you find you got 0.5 win hosmer and you lost your division by a game, instead of getting 4 win hosmer and winning it by 3.
Is Fangraphs considering showing additional columns with the 90% or 95% confidence interval spreads or similar, as some others do, to be more transparent on this uncertainty?
Not sure this would be too useful if we don’t forecast different uncertainties for different players/teams (AFAIK ZiPS does this though I don’t think Steamer does?). Without this we just get the same +- for each team which isn’t super useful or informative (wouldn’t mind seeing it on the page but for each team is a bit too much imo).
This article needs to be emphasized much more. Not only is there a ton of uncertainty in player performance that can’t be accounted for in any projection system – unexpected breakouts, freak injuries, etc. – there are many other factors that contribute to their poor track record.
First, the data these public projection systems use just isn’t good data. I don’t mean to be rude here, but ZIPS and Steamer (to my knowledge) do not incorporate Statcast or pitch f(x) data into their projections whatsoever. Given how much randomness there is in single season hitting and pitching stats – BABIP, HR/FB, and even BB% are all very unreliable year to year – this is going to give you poor individual projections because you’re ignoring the most relevant pieces of data. And when you have poor individual projections, that just gets magnified on a team basis. And that’s why the “advanced” projection systems haven’t been able to outperform Tom Tango’s “monkey” Marcel system.
Additionally, the other problem faced by projection systems (even good ones) is sometimes even when the individual players on a team perform as expected, the team as a whole over or underperforms due to a fluke sequencing or BABIP phenomenon. The Royals and Orioles did this for a couple years with sequencing their hits and winning close games, and even the Cubs did this in ‘16 by running absurd team BABIPs and LOB% rates in their pitching staff. Eventually, they all regressed, but not in the given season.
To conclude, while in-season randomness, fluke injuries, and actual skill changes likely can’t be incorporated into any projection system, that doesn’t mean we’ve reached the peak of player and team projection. We haven’t, and using better data will make for better projections. We should be capable of beating the monkey projection system, even if Nate Silver couldn’t do it with the data he had at his disposal.
But in the meantime, people need to stop taking the projections as gospel. Personally, I look at the projections and see many areas where my own data strongly disagrees. I think the Cardinals are better than the Cubs, could easily see the Red Sox finishing behind the Rays and Blue Jays, and think the Mets are much closer to the Nationals than Steamer does.
We should be evaluating teams in a more comprehensive way than just looking at Steamer and saying “yeah, the Rays can’t catch up.” Take a deeper dive into the data and you might find that Matt Duffy is actually better than Xander Bogaerts, that Justin Smoak’s breakout was legit, that Randal Grichuk and Ian Kinsler are due for huge seasons. A better version of xwOBA is a good start.
Anyway, the lesson here is that we need to break free from the projection systems and rely on more detailed player analysis.
That’s a bit overstated.
Just because ‘error bars exist’ doesn’t mean that there is a ‘ton of uncertainty.’ There are 162 games, standard error is 7.
Similarly, it’s incorrect to say that ZIPs and Steamer “aren’t good” simply because they are still a few months away from incorporating a new data source that still has small sample size and measurement/calibration issues. They will incorporate statcast. Whether projection models are good or not depends on predictive value, not just whether they have incorporated the latest sources.
The failure to outperform Tom Tango’s elegant models is hardly a pronounced failure- Tango’s work is gold standard. But your right reverence of Tango’s work undermines your statcast point- projection systems are about predictive value (sometimes elegant and using few input variables due to proxy effects and principal component factors) not ‘including X variables that people like and are new.’
Much of the rest of your observations are off the mark as well, BABIP and BB rates aren’t “very unreliable year by year” by any relevant industry usage of those terms. Instead R^2 can be around 0.5+ for such metrics.
The idea that ‘randomness eludes the projection systems because they can’t account for all input variables” is wrong and anti-empirical. Wins, losses, and runs capture everything that affects wins, losses, and runs. Sure, random things may happen in any system, but that doesn’t necessarily limit a projection system to maxing out at 92% predictive value rather than 93. Baseball doesn’t have unique randomness unknown otherwise to the world. Your body has a bunch of crazy stuff going on, and yet a projection from collecting the sweat on your arm can prove highly predictive as to whether your lungs have tuberculosis. The mere fact that ‘the arm is far from the lungs, it is different than the lungs, lots of random things happen in the body, and the swipe test doesn’t incorporate the newest technology” doesn’t necessarily put any cap at all on the maxmimum predictive value of the method. The data speaks on predictive value, not narrative or theory about which X variables are interesting.
Marcel is not an “elegant model”. Tango himself has said this. It’s a basic model that by design is so simple that a monkey can calculate it. And it often outperforms Steamer and ZIPS.
That is very problematic for the “advanced” projection systems, because it shows that the more complicated adjustments they are making don’t actually add any predictive value.
With regard to predicting randomness – things such as winning one run games and sequencing hits well – I don’t see how any projection system would be able to predict this accurately whatsoever. On a theoretical sense I agree that you can technically “predict” nearly everything in a deterministic universe, but I fail to see what data could be used in our measly human models to predict these things. How would we know that the Brewers would outperform their collective BaseRuns significantly? What data tells you that?
I think that Dan S. said in a chat a few weeks that ZiPS uses launch velocity / launch angle data – https://www.fangraphs.com/blogs/dan-szymborski-fangraphs-chat-1-18-18/ . His answer is a bit unclear, but that’s how I interpret it.
It would be useful to separate true collapses from injury-driven collapses.
If you lose a 4 win pitcher in April and a 3 win hitter in May, then a 10 win collapse is really maybe a 4 win decline.
“Looking at the overall sample, the average error has been about seven wins”
That’s as good as it gets. That’s how well a perfect projection system does, one that knows every teams true win% exactly, and then the teams flip coins to those percentages. That’s God’s prediction system as long as he doesn’t cheat in-season.
The standard deviation for 0.500 over 162 games is 6.4 wins.
Predicting player talent, injuries, all that is great, but even perfection on that doesn’t get you more than the math allows. Probably we’ve gotten a little lucky already.
What’s weird to me is that we see a lot more outliers than the overall error would suggest. If normal distribution we should see projections be off by 10 games given std of 7 60/390 times, yet instead we see 110.
This would be true if the projections themselves were perfect. I.e, you’d expect a team to off its true talent level 60/390 times, by luck of coin flips. But the projections themselves have some variance.
I’m basing this off of the 7 wins error for projections Jeff gave not perfect perfections.
Thank you Jeff!
This is why I’ve been so weirded out by the whole idea that “super-teams don’t want to improve because they’re already so super.” I guess this might be true but if so it means someone at the top is getting bad information about the certainty of the projection.
The chances the Dodgers wind up in a wild card is way better than the projections imply.
Disagree- an 11 game projected advantage over SF/Zona is a massive advantage.
A 7 win standard error is very small, less than 5% error (4.3%). Projecting 155 out of 162 W/L outcomes correct, for the average team, is relatively near max theoretic limits.
I hear front offices talk about dividing the season into thirds. I wonder if the projection systems would be considered more accurate if they projected the season in thirds as well.
It would be interesting to see a projection system continuously updated throughout a season to account for roster changes. Once a team is out of contention, they sell off parts and rebuild for the next year and conversely if a team plays into a competitive position for the post season they buy additional parts in June and their projected wins would improve.
I think Fangraphs projections do exactly that.
But a broader point is that teams doing very badly may decide to sell off and then do even worse, and vice versa. So a projection system that does not incorporate that will be off for that reason, projecting bad teams to do better than they really do and good teams to do worse.
Great piece Jeff! Among teams that meaningfully overperformed projections is it possible to check how many had an outsized contribution in terms of bullpen wpa?
I’m surprised it hasn’t gotten more attention, but wouldn’t the team projections look a lot better if you account for in-season trades? We should not expect a projection system to be able to account for trades. The impact is doubled since it simultaneously weakens/strengths two teams vs their projection.
“I have a spreadsheet of preseason team projections, going back to 2005.”
I’m curious what the source is. Are they opening day projections from FanGraphs going back several years and then PECOTA before that? Are they early March projections? I assume they’re all single source; i.e. none of your projections are the average of two sources.
One thing that’s tricky with projections is it’s hard to measure how accurate they were. Right now, FG projects the Twins to be .500, both scoring and allowing ~800 runs. (I picked them because the numbers are easy.) If the Twins score 805 runs and allow 801, but due to sequencing, luck, and distribution of runs finish 90-72 or 72-90, was that a good projection or a bad one?
I thought this was going to be about Luis Castillo 🙂
I love this piece, but I can’t help feel it’s either preaching to the choir or going to be ignored. A lot of people will only read into projections what they want to see. “I think the Reds are better than that win total, so this system is flawed. Also, you’re dumb and I hope you die.”
The presentation of the data is definitely difficult. People want quick headlines that take complicated matters and condense it down to something that is a quickly digestible, but by then it doesn’t resemble the original thesis. See: any research paper discussed in the news, or even more relatable here, any FiveThirtyEight projection.
At what point do you go “well, the R2 being 0.36 indicates we probably shouldn’t spend so much time poring over them” and de-emphasise the coverage they get on sites like this? It’s not so much that they’re not perfect, it’s just so many articles go over and over them as if they’re far far more important than they actually are.