Team Win Projections vs. Actual Win Totals, 2007-Present

Full-season team projections cause some heated arguments. If a team finishes the year with fewer wins than expected, fans want to know why their club underperformed projections. If a team overperforms its projections, meanwhile, those same fans will insist that forecasts in subsequent years lack the ability to detect their club’s particular strengths and are thus useless.

Here at FanGraphs, we have only been doing full-season projections for a couple years, but just about every week I see a mention of the 2015 World Champion Kansas City Royals’ projected record of 79-83. If I search Google for “79-83 Royals FanGraphs,” I get over 11,000 article links. Unsurprisingly, it’s a popular topic. Rarely does a club, following a pair of World Series appearances, then proceed to fail to break even. But that’s what the numbers suggest for 2016.

While FanGraphs has produced team win projections for only a couple seasons, Replacement Level Yankee Weblog (RLYW) has been publishing win projections for years. Since 2007, to be precise. Given this larger sample, I thought that it might be worthwhile to compare the projected win values produced by RLYW to the actual final win values produced by teams. So, with the permission of RLYW editor SG, that’s what I’ve done here.

I hate to disappoint anyone, but there are actually aren’t any great findings in the plethora of graphs to follow. I did find a couple interesting artifacts of the data, but no game changers. Instead, I see the following mainly as an additional data point in many past, present, and future discussions.

To start with, here is how projected and actual values have correlated.

projwin_2007-2015_720

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

The projections weren’t completely off, but a wide variance exists. Now, here is a graph illustrating how much each team’s projections have deviated from the projections over the years.

I’m pretty sure each team’s fanbase has complained about projections in the past. This graph allows those fanbases to observe how justified their complaints have actually been. In the case of the Cardinals (whose median win totals have been six wins better than their projected totals), reasonably justified. In the case of the Brewers (who’ve finished within four wins of their projection in almost every season), not as much.

Of all the graphs included here, this one offers the biggest surprise. It compares the density of projected versus actual wins.

distribution_of_wins_720

The results are pretty much as expected: a narrow band of projected results and a wider, flatter band of actual results. For me, though, the notable observation here is the small decline right at 80 and 81 wins. This could just be an artifact of a small sample. Or it may occur because teams late in the season see themselves as contenders and bump up their win total a bit with trades. The sellers could see their win total pushed down a bit. Or maybe something else.

Finally, here are two graphs showing the difference between projected and actual wins two different ways. First, the simple way.

projection_error_vs_projected_win_1024

This graph offers another interesting point. Look at the average slope line. The differences should have equal points above and below the 0 line. Instead, the low-win projected teams seem to outperform more often than underperform. The opposite is true for the high-win teams. I am not sure why this has been true since I don’t know the exact details of RLYW’s projection creation process. If I were to guess, I bet the stats used are not regressed enough.

With the simple graph of expected versus actual wins out of the way, here is the complex version.

Besides just the individual results, the quartile, median, and extreme (whisker) values are available for each win total. Looking again at the Royals’ projected 79 win total for 2016, it can be seen that these teams overperform by two Wins, with the quartile range sitting at plus or minus eight wins. On the extreme ends, we find +16 wins by the 2015 Royals (and below that, +11 by the 2011 Rockies) and -15 Wins by the 2014 Diamondbacks. Well, the graphs are done and hopefully readers can find a way to use them in future projection discussions.

Preseason team projections are far from perfect, with players over- and underperforming, injuries, trades, and about 1000 other factors. The preseason projections do give everyone a beginning expectation level to which they can anchor their hopes. The expectations can change and teams can make the projections look silly at a season’s end. Hopefully, I was able to get a quick snapshot on how those projections have historically and the realistic chances of a team beating those projections.

Big thanks to Sean Dolinar for the graphs.





Jeff, one of the authors of the fantasy baseball guide,The Process, writes for RotoGraphs, The Hardball Times, Rotowire, Baseball America, and BaseballHQ. He has been nominated for two SABR Analytics Research Award for Contemporary Analysis and won it in 2013 in tandem with Bill Petti. He has won four FSWA Awards including on for his Mining the News series. He's won Tout Wars three times, LABR twice, and got his first NFBC Main Event win in 2021. Follow him on Twitter @jeffwzimmerman.

20 Comments
Oldest
Newest Most Voted
Joe Joe
10 years ago

On teams under-performing their projections, I suspect plate appearances for good MLB players is overestimated in Depth Charts. Also, I suspect Depth Charts has a tough time with rookies and basically gives them little playing time which may give a rebuilding team the look of over-performing. Plus, it is really hard to under-perform a sub 70 win projection by more than ten games even though the Astros tried their hardest for three straight years.

theoriolewayMember since 2026
10 years ago

Interesting. I think the way actual vs. projected wins plays out in the density graph is highly suggestive of the fact that while baseball contains a lot of randomness, it is not random. Teams that lose or win more do so for a host of reasons that are effectively exogenous (health, over / underperformance) but also ones that are endogenous (trades, playing time decisions, etc). I would be really curious to see how ‘projections’ that are run after the season which account for actual playing time relate to actual wins. That is, if we knew exactly how many plate appearances or innings each team would give to each player (some of whom are not even on the preseason roster), how would the preseason projections have performed?

Meddler
10 years ago
Reply to  theorioleway

Just generally, any additional information you could give the model would make the projection density distribution more similar to the actual distribution. In general, in any type of model, projections will typically be more peaked and actual results “flatter” in this type of comparison to account for error.

Your suggestion is interesting in that it would show us how much playing time modification weights account for this error, but it would still be hard to know exactly what we can do with that information. Noise in projected playing time is still something we had to adjust for, and unfortunately, it’s hard to reduce.

theoriolewayMember since 2026
10 years ago
Reply to  Meddler

All true–I’m just trying to control for those endogenous things (like trades) so we can see how much the more exogenous things cause deviations.

atpkinesin
10 years ago
Reply to  theorioleway

ive been looking into exactly that actually. so far ive just compared 2015 preseason PA/AVG/OBP/SLG from steamer-depth to the actual outcome. the average absolute error for PAs for each player per team ranges from 59 PAs (tigers) to 188 PAs (Braves). the average of all 30 team averages is 107 PAs with a 27 PA std dev. my guess is that avg is not a good representation since injuries will cause players to miss chucks of PAs, but i havent looked into it yet. also, i havent addressed players that changed teams, here both teams get some credit/some debit. i converted obp/slg to woba and converted to runs with the scale factor, then found the difference in runs using (actual pa and projected woba) – (projected pa * projected woba). the average absolute error per team ranged from 19 runs to 46 runs (again, for tigers/braves). the average over all 30 teams was 30.6 runs.

Ernie CamachoMember since 2016
10 years ago

Similar to theorioleway’s suggestion, I’d love to know how well FG’s Depth Chart system has predicted team wOBA and wOBA-against, or whatever it is that the final step in the process uses to spit out team run differential and team records. Is most of the error coming from player projections or how the aggregated player performance combines to produce and prevent runs?

Meddler
10 years ago

Sorry to frame this as a series of questions instead of statements, but I’m just learning a lot of this statistical theory now (and in the context of psychology, no less, so not exactly an overlapping field). That caveat out of the way:

1. Doesn’t the error vs projected wins chart indicate a pretty substantial bias in the model?

2. Could it also indicate that the relationship between the projection components and win totals are actually nonlinear? This would seem preferable to just regressing the projection components more, since;

3. If the answer to #2 is “no you’re wrong, idiot,” and it’s just a question of regressing the projection components more, doesn’t that make the actual wins density graph even more peaked with tiny little tails and everyone just kind of hovering around the mean? Wouldn’t that simply mean there’s less projection difference between the top and bottom, and that the model is just kind of generally underpowered?

Meddler
10 years ago
Reply to  Meddler

“doesn’t that make the actual wins density graph” in #3 should read “doesn’t that make the projected wins density graph”. My bad

DavidBowser
10 years ago
Reply to  Meddler

1. Based on a previous article on FG (http://www.fangraphs.com/blogs/10-years-of-team-performance-10-years-of-team-projections/) the original data from RLYW actually comes from multiple projection sources and models.

There have been various looks at the historical accuracy of projections and trying to determine the nature of teams that beat or fall short over an extended (8-10 years) period of time. One analysis on FG of the Chicago White Sox showed that they had the lowest total player games lost to the disabled list.

Meddler
10 years ago
Reply to  DavidBowser

If the scatterplot of input variable to error doesn’t produce a a trendline across the x-axis (so just a linear relationship of y=0) that indicates there’s bias in the model though, not in the data, right? Since we’re looking at the relationship between the observed predictors and the error between observed outcome and actual outcome.

Weighting playing time projections to some kind of team health performance factor does seem like an interesting idea for improving playing time projections.

Are you referring to teams that consistently beat or underperform expectations? I don’t anything super compelling indicating that’s a huge problem here (though there likely are ways to adjust for this). We always expect error though. The point of projection isn’t to precisely predict an outcome, it’s to generate a number that represents the most likely outcome with the least amount of error and (hopefully) without bias. As noted, my understanding is that comparing error to input is a way to detect bias in the model itself, which seems to be indicated here. That could mean the inputs aren’t being regressed enough, but I believe that would also mean we have less power than we’re assuming. I’m also wondering if its possible that the relationship is actually non-linear, and if doing some kind of transformations to account for this (which are above my level of sophistication) would be a way to reduce this bias without losing power.

Cory Settoon
10 years ago

The Red Sox sure have been bi-polar haven’t they?

2012: -22
2013: +15
2014: -19

Meddler
10 years ago

Also, very cool article. I love stuff like this; trying to open up the “black box” to some extent (even if post hoc) and honestly critique the models many of us just take for granted.

Damaso
10 years ago

We really need to get some error bars along with the projections. Some projections just don’t have the same strength as others.

it would be nice and convenient to find out that the teams who whiffed on their projections were mainly those who relied on many players with spotty or nonexistent track records.

TKDCMember since 2016🏆 MVP
10 years ago

I think there is too much effort put into addressing people who complain about projections.

Slappytheclown
10 years ago

It’s not surprising you have fat tails as the projections assume a normal distribution. Just a guess, but in regards to the ‘errors’ my estimation is it boils down to 50% health, 25% sequencing of runs, 10% sequencing of competition (i.e. luck on when you play someone) and 15% outperformance, which is related to health.

So, if you stay healthy there is a 65% chance that you will ourperform you expected record. I’m sure someone can take the top 20 players for estimated WAR on each team since 2007 and decompose total games missed and that would have a high correlation to changing records positive or negative. Plus, if your healthy you have a good chance to outperform your WAR estimates which are pretty heavily regressed.

Bad Hermit
10 years ago

I might be completely missing something, but at first glance it looks like the labels for the first graph are flipped.

Adam SMember since 2016
10 years ago

Wow, that’s a lot of data. My head is spinning but one note.

On the dip at 80 wins for actual results, I assume that teams that have a chance to get to or over .500 (81 or 82) wins, go for it over the last week or few days of the season rather than shutting down better players to let the kids play. They actually try to win the last weekend of the season instead of mailing it in.

Paul Clarke
10 years ago

That dip at 80-81 wins is something Colin Wyers came across a couple of years ago (http://www.baseballprospectus.com/article.php?articleid=21060).

I think the result you’re seeing where teams with low win projections tend to over-perform and teams with high projections tend to under-perform is to be expected. It’s similar to the question of whether someone putting up a .350 BABIP is due to talent or luck – the likeliest answer is usually “some of both”. And if you ask whether a team with a 60-win projection is bad or under-estimated by the projection system, the likeliest answer is, again, “some of both”. So underestimated teams will be over-represented among low win totals, and overestimated teams will be over-represented among high win totals.

Pwn Shop
10 years ago

I have an idea! Should be easy enough to test. Look at a team’s winning percentage for their first 147 games, or so. Then look at their winning percentage for the remaining 15 games. Do good teams get better, on average? Do bad teams get worse? Maybe ‘contending’ teams (say, 0.540 winning percentage thru the first part) dramatically boost their remaining win percentage, but others don’t. If so, that confirms the theory of making trades to get into post seasons, and/or of dumping contracts when you aren’t going to make it anyway. Explains the wonky win distribution with twin peaks. Could be worth looking at.

Adam Maahs
10 years ago

The error of projected wins vs actual wins look to be normally distributed about 0 (qualitatively), do you have the standard deviation for (projected – actual) wins?

I’d be interested in knowing the z-score of some set win threshold (eg. Vegas win total) versus the projected win total to evaluate the implied probability of reaching that number of wins.