The Royals Haven’t Been the Projections’ Biggest Miss
No team has more conspicuously made us look silly than the Royals. Not in the last few years, for all the reasons you already know. Not many things more visible than consecutive trips to the World Series, and when you look at what the Royals did against what the Royals were expected to do, statistically, it’s natural to wonder what’s up. It’s normal to find comments like this one, left earlier today:
Dave, if the Royals once again reach the post season, or even the world series, is it time to re-calibrate the predictive model? In other words weight some of the production measures differently? 4 years in a row isn’t luck.
For some, “projection” is a dirty word, and for others there’s just a certain skepticism. The Royals are the “face” of this feeling, if that makes any sense, because after all, they’re the defending champs, and they were projected to not be very good. There’s absolutely no question the Royals have exceeded statistical expectations the last few years. What might surprise you is another team has done that even more.
The Royals have completed three years of this. Clearly, they surpassed expectations in 2015. But it also happened in 2014, and even in 2013, when they fell just short of the playoffs. The three years prior, the Royals were pretty well pegged, but they only more recently came into their own, and the projections have lagged. One year, you can buy as a fluke. When you’ve got three years, people wonder, and they’re not wrong for doing so.
We’ll play around with three years of information. As I’ve written about before, I have a spreadsheet of projection information stretching back to 2005, but for now I’ll keep the focus on 2013 – 2015. We know the Royals have blown projections out of the water, but then I also just saw this post by Travis Sawchik, on the Pirates. That was good timing on his part. Here’s a plot of how all 30 teams have fared, compared to their preseason projections, the last three years combined.
This shouldn’t be complicated. The Rockies have won 17 fewer games than they were projected to win. The Royals have won 30 more games than they were projected to win. The Pirates have won 33 more games than they were projected to win. (My numbers differ very slightly from Sawchik’s.) So while the Royals have been more conspicuous about this, getting to two World Series and winning one of them, the Pirates have them one-upped in the regular season, leaving aside their early playoff eliminations. And only the regular season gets projected, anyway.
It’s not a coincidence that, the last three years, the Pirates rank first in baseball in bullpen WPA, and the Royals rank second. That doesn’t explain everything, but it’s amazing how many games you can win if you don’t let late leads slip away. The Pirates have also been progressive with their defensive alignments, which numbers have trouble with, and then there’s all their time-sharing and the Ray Searage stuff. If you’d like, you can throw in some health, and some midseason additions that forecasts don’t consider. There are plenty of reasons why a team might beat its projection. The Pirates have done it more than anyone else of late.
In fact, it doesn’t even have to be limited to three years. Four years ago, they beat their projection by seven wins. Five years ago, they beat their projection by two wins, if that counts for anything. Six years ago, things were ugly, but they were also unrecognizable. This is at least a four-year trend. It just looks the strongest if you focus on the three most recent seasons.
Going back to 2005, now, I have a sample of 270 team stretches of three consecutive years. The last three years, the Pirates have exceeded their projected wins by 33. That ranks them in second place for greatest positive three-year difference, behind only the 2012 – 2014 Orioles. The Royals, at +30 wins, rank in sixth place. Obviously the effect has been big, but what could this mean as far as 2016 goes? For fun, look at the following table. I grouped the best and worst three-year stretches, and then found the average win/projected-win differences in Year 4.

| Teams | 3-Year Difference | Year 4 Difference |
| Top 10 | 30 | -4 |
| Top 20 | 27 | 1 |
| Top 25 | 25 | 1 |
| Bottom 10 | -31 | 5 |
| Bottom 20 | -29 | 0 |
| Bottom 25 | -27 | 1 |
To explain: the top 25 averaged +25 wins over the three-year window. In the fourth year, they averaged a difference of +1 win. The bottom 25 averaged -27 wins over the three-year window. In the fourth year, they averaged a difference of +1 win. It strongly suggests the three years aren’t actually predictive, in terms of beating or falling short of projections. If you look only at the top and bottom 10, you see an opposite effect, if anything. The 2012 – 2014 Orioles had the greatest positive three-year stretch, and in 2015 they beat the projection by two wins. After the Angels ripped off a great run of overachieving, the 2010 team beat the projection by two wins. The 2011 Twins, though, followed their stretch by missing the projection by 21 wins. The 2011 Marlins fell short by 10 wins. The 2015 A’s fell short by 15 wins. History, viewed in this way, doesn’t indicate the Royals and Pirates should be considered special in the year ahead.
There’s also common sense: it’s unlikely one or two teams have just figured something this critical out, given the competitiveness of the industry. But you can never completely close the door. After all, every team and every situation is different. And after all, the Royals have done this for three years. And the Pirates have done it for four. History didn’t love Dallas Keuchel, until he unlocked what he could do. People, at least, will be watching the Royals closely. They should also be watching the Pirates closely. In every such situation, there could be something real, and the longer it goes on, the harder it is to dismiss.
We’ve had so many conversations about how maybe there’s something about the Royals. It’s good that we’ve had those conversations, and many of them have been enlightening, but there’s been a sort of bias, because of how far the Royals have gotten in the playoffs. Even now, we know the playoffs are pretty random, and with that in mind, we should be having the same conversations about how maybe there’s something about the Pirates. They’ve earned it just as much. I guess you could say even a little more.
Jeff made Lookout Landing a thing, but he does not still write there about the Mariners. He does write here, sometimes about the Mariners, but usually not.

Fire Robin Ventura
Fire Dick Monfort. Fire AJ Preller. Fire Ruben Amaro Jr. Oh, wait, hire him as 1B coach…
You have been spot-on with the Braves recently, so I trust the projection for 2016 will be accurate! I’ll go see what the Braves will look like in 2016!
… oh…
How much does production from young, unproven, unpredictable players comes into play? Both KC and Pit have had a pretty young core of good players. Throw Houston in the mix as well. How do you really project the McCutchens and Cains of the world? How do you know that Correa will be so good but Singleton will not? KC had a young, unproven core with Cain and Moose and Hosmer, etc… Moustakas looked like a bust and then he broke out. That’s hard to project. Harrison and Kang in Pittsburgh were really good. Should projection systems know that? Hard to know what those guys will really be. How do you project a Joc Pederson or Corey Seager in LA? Seems like maybe the teams who can get the production from the younger players might outperform expectations.
Projections are just most likely outcomes based on past results… they don’t predict unlikely occurences. I know Royals fans are getting bent out of shape and perceive it as a lack of respect… but as you said up until last year it looked like they had a bunch of failed prospects.
Yeah! That algorithm gives the Royals no respect!
“History, viewed in this way, doesn’t indicate the Royals and Pirates should be considered special in the year ahead.”
I wonder how much the rosters changed over for the teams that you mentioned.
For example, the 2014 Orioles had Nelson Cruz at DH, while the 2015 Orioles had Jimmy Paredes, etc.
The 2016 Royals are basically the same as the 2015 Royals, trading in Ryan Madson for Joakim Soria and Jeremy Guthrie for Ian Kennedy. There is no glaring downgrade.
The loss of Greg Holland is quite a downgrade.
He was lost during the year after being ineffective. He missed the playoffs.
Why would that matter? It’s a measure of the projection’s accuracy, not a measure of the change in absolute number of wins.
Well the projections did not project Nelson Cruz hitting for the Orioles in 2015. So unless you are saying that Nelson Cruz has the skill to motivate his team to out play their projections, this doesn’t make any sense.
If this were the case, the Mariners would’ve outplayed their projections. And I’m pretty sure that was not the case.
This whole “the projection system is broken” nonsense can be stopped if everyone just understood the idea of error bars. With the information available now, a win projection of 85 games does not mean the team will win 85 games on the nose. It means that the most likely outcome given the information we have available to quantify is 85 wins with the possibility of plus or minus 5 wins if the impact of the information we currently do not have available to quantify is significant (which is almost always is). The graph here illustrates that nicely in my opinion.
Yeah if you look at the graph, you see that more than half the teams are within a 10 win error margin over 3 years. That seems pretty good. And all but 4 teams are within a 17 win error margin, which is 6 wins a season. That seems really good.
The Royals, Orioles, Cards, and Pirates are the only teams outside of that margin. The fact that all four overperform is somewhat interesting. It does lead you to think that there is something they are doing that is breaking the system a bit. I think Jeff is on it when he said bullpens a while back. And maybe defense. Seems like teams to defend and relieve very well can break the system. We don’t quite know how to measure that well enough.
Yes, there are error bars, but this is looking for systematic biases.
In addition to what the author has noted, I believe the Pirates’ team defense in underrated. The projections here rate their infield very poorly, but a look at the raw numbers shows that in 2015 they were 3rd best in MLB at getting at least one out on a ground ball, and 2nd best in outs per ground ball, given their proficiency in turning double plays.
I would guess that a projection system that does not effectively evaluate defensive production, cannot manage defensive shifts and favors a strike out over a pitch to contact approach will find the Pirates a puzzle because the Pirates are strongly committed to those goals. Huntington also finds enough relief pitchers that the Pirates’ pen is typically strong. If these are indeed problems, then a mean-regression model will further diminish the productivity of the Pirates pitchers and defense.
Depends on how the team actually outperformed there projections.
Not being able to adequately project defense would show up in the runs/game allowed.
If a team outperforms there pythag, I don’t think you can prove that it has anything to do with defense.
Right, most understand that it is unlikely any team will hit their projection on the nose. As far as over projecting wins, I would say that may be misleading as teams who are out of the race start becoming less concerned about wins and more concerned with getting playing time for young players. Also, if major misses were random, I.e. one year the pirates and royals were grossly under estimated, the next year it was two other teams, and the following again two different teams, it could be dismissed. But if it is the same teams over and over, it is then fair to ask if thses teams are doing something underappreciated by the system.
@jayhawkjac Exactly. I wish people would either publish their own projection system that is clearly better than Steamer, or else STFU about how the system is allegedly biased for/against a given front office based on a couple or even one year of results. #samplesize
It isn’t about bias against, or for, any franchise. I don’t think anyone is making that claim. The question is about the system potentially over/under valuing certain metrics. No one is claiming they have a better system either, most folks on here wouldn’t know where to begin to build one. But that doesn’t mean they can’t discuss a projection meant for entertainment purposes about an industry meant for entertainment. Does it?
In fact, a commenter did make a claim in an earlier article that Steamer is biased toward Ben Cherington. Many of the other comments were whining about why FG doesn’t own up to the fact that they have a crappy projection system that chronically over-rates the Red Sox..
So yes, plenty of people on FG are making that claim.
Oh look…The “It’s only entertainment, therefore whatever I say, no matter how inane or false, has blanket immunity from criticism” defense. An internet classic!
I wish people would either publish their own projection system that is clearly better than Steamer, or else STFU
This line of thinking is just ridiculous. Its the same garbage hack film makers trot out when they make a shitty movie or a terrible comedian after a show. “Lets see you do better then!”. Just because someone might not do better doesnt make them immune from criticism. Participation is not a prerequisite for criticism.
You are essentially saying that everything should be immune from criticism unless you can first do better. Which is absurd.
Right. Except the error bar, for a 90% confidence interval, is +/- 10 wins.
The problem is that doesn’t look like much of a projection — team X will win between 71 and 91 games. Yippee!!
It’s not about error bars, it’s about R-squared. If you plot a regression of predicted vs. actual wins over many seasons and then fit a line through the cloud of data points (which Jeff has done in the past), you can calculate the R-squared associated with that fitted line, which represents the proportion of the variation in the data that is explained by your prediction model. I can’t recall what Jeff’s R-square values have been but the bottom line is there is plenty of variation that’s left unexplained (= 1-R^2). If the best model has an R-squared of 0.70, just to pick a (pretty high) number, it’s a good model but you can’t go around talking like it’s 100% accurate (which some FG folks, esp. DC, tend to do); there’s still 30% of the variation that you have no clue about.
What is the 3 year roster turnover rate? Do teams with less turnover have more stable projection beating abilities?
You guys have the Cubs projected to win about 163 games in the regular season this year, so expect the projections to take another hit soon
I don’t see them winning any less than 170 games. This team could blast the shit out of Hurricane Ditka.
Two words, Cubs fans: sophomore slump.
Being a successful MLB player means that the league will adjust to your strengths and you have to adjust to their adjustments. And they’ll adjust to your adjustments and you have to adjust again, and so on…
Not every young stud that takes the league by storm is able to make the adjustments and sustain that success. And even those that do may take a while to make them, esp. the initial ones.
I’ll take the under.
How much of the projection vs actual difference is solely relief pitching? The thing that stands out about the Pirates and Royals, and Orioles, is 3 teams that have great back ends of their bullpens, with multiple relievers who can be called upon and have perform very well in high leverage situations. I feel like this is the biggest thing, that WAR isn’t weighting the leverage value of these bullpens accurately. E.g. Wade Davis, being projected or having a season with a WAR of 1-2 each year. The WPA leaderboards for all pitchers ’14-15 is Davis at #5 followed by two Pirates relievers. The Royals have also had some better than expected performances from their position players, Hosmer, Moustakas, Cain, which isn’t really something you could correct for. But, I think reliever value, and also possibly regressing superior defensive performers too heavily, accounts for a big part of these misses. You can tell me that the difference between Davis and a replacement level back end reliever has been about 2 wins, and projected him for 1.5 wins for next year, but I just don’t believe that’s an accurate reflection of reality.
the problem is that around 90% of games that can be saved have been saved in baseball history. So, saving a game is considered easy by projection systems that do not differentiate by leverage and is therefore undervalued.
Oh and their is a correlation that is statistically significant for defense and leverage. Better defensive teams tend to perform better in high leverage situations than poor defensive teams ignoring the the pitcher.
It’s actually pretty fascinating to look at how RP dominate the WPA charts given the huge differential in salaries between elite starters and relievers.
In the period discussed in the article, Mark Melancon trails only three players on the WPA chart who are worth about a combined billion dollars of salary commitment (Kershaw, Greinke and Scherzer) more than he is.
In related news, turnovers are the most influential stat in the NFL, yet defensive turnover creators don’t get paid much of a premium. Ever wonder why that is?
Probably because it is a volitile stat where it is easy enough to punish players who trun the ball over, but hard to reward individuals who create turnovers as it depends on too many other factors….may have been a tongue in cheek comment but i responded anywy. Go Chiefs
Right. A highly impactful statistic is not necessarily one that can be predicted well. Most NFL turnovers are due to chance and/or the QB — not the defense.
So wondering after the fact why guy who strung together 50 incredible IP is not getting paid a ton of money is similar to asking the same question about the DBs with the most INTs this season.
This just seems like something everyone is accepting as true without real evidence.
You do mention a couple of teams for which it is true. But you’re using them to proof-text you belief…not looking at the data without a bias.
What about the Yankees who we know have had the 3rd best bullpen for the 3 years included in the graph above, yet the projections seem to be getting mostly right?
In fact, 4 of the top 10 bullpen teams have underperformed projections.
So what do we do? Pick and choose which teams to apply the “bullpen magic” adjustment to?
I don’t think everyone is accepting that it is true or we would have different ways of measuring and predicting relief pitcher value. I see a lot of articles on here talking about the projections but I don’t see a lot of research into where those projections are falling short. I think reliever value is something that’s necessarily hard to predict, relievers are volatile in performance, leverage is volatile (e.g. if you put Wade Davis, Melancon on a team that has bad starting pitching and can’t score runs, they’re not nearly as valuable because there aren’t as many leads for them to protect). Just a hypothesis, that a big part of the problems is we don’t have good ways to measure reliever value and predict their future value, mostly because leverage has so many external influences, strength of the team in other areas, ability of the manager to use his best relievers in the true highest leverage situations. But just because this is hard to measure doesn’t mean it’s not actually measurable.
I don’t think you can just look at the Yankees and say “the projection got it right” either, because you have multiple variables. We’re looking at 3 years of aggregate team performance. It’s possible the Yankees position players underperformed their projection dramatically but ended up still meeting their win projection, in which case it’s not necessarily true that their bullpen wasn’t undervalued. With the Royals, their young position players also outperformed expectations, if the bullpen performance is undervalued, it’s another factor pushing in the same direction. I’d be interested to see a breakdown of where the projections are missing. E.g. it shouldn’t be hard to look at projected position player WAR and see what actual WAR was, and see how much that contributed to gaps in projected vs. actual performance and then look at the other aspects of team performance.
Right about Wade Davis. My guess is that he is severely undervalued, especially giving the leverage of his appearances. He’s probably worth a lot more than what is attributed to him.
I’m shocked to see the White Sox have underperformed their projections more than any team other than the Rockies. Wouldn’t you think they would beat the projections because the projections don’t account for their apparent ability to prevent injuries?
I appreciate Jeff writing this, in fact it is my quote he uses at the beginning of the article from earlier today. While i had not looked at actual performance vs projection for other teams, it isnt surprising that the Royals are not alone. Anyway the calibration I spoke of before is not meant as a slight, just as a method we use in my line of work to improve our outcome prognosticatoion. The other thing may be including more of a stochastic element (assuming one does not exist) or publishing the confidence interval at the rejection percentage. But I guess most fans want a point estimate.
Can’t speak for the Pirates, but the Royals suggest that defense is poorly quantified, that the various FIPs don’t encompass all of pitching, and that chemistry has yet to get numbers assigned. For what stats can measure they do a fine job of but there is a great deal more to on-the-dirt baseball than can be found in even the most advanced box score.
Hey Jim you and I had an exchange on a different site lol
Show me one bad team with great chemistry, and I’ll jump on the “chemistry is real” train.
I’d even settle for proof that a team has good chemistry BEFORE they start winning, and I’d start to be convinced.
Was honestly expecting the Red Sox, because it always feels like they get projected to be a top 5 team but end up waaaaay below that.
So “always” [over-rated] means not even being in the top 5 most overrated teams over the last three years? Where’s the Rockies conspiracy love?
Not always, in 2013 they were projected to be pretty meh (roughly an 81 win team iirc) and won the World Series after a 97 win season. If my memory is correct that would mean presumably a 32 game addition to how much the Red Sox have been missed by. I suspect that if the amount missed was calculated on a cumulative basis from each season rather than a combined three year amount the Red Sox would top this.
My thoughts on projections is that nothing does or realistically can adjust for the confidence level of the players playing the game. Players that get on a hot streak feel better about everything in the game and can over-perform for a while and end up surpassing their projections beyond their expected performance and talent level. The same runs in the opposite direction. If teams start winning it injects belief and confidence, if teams start losing it runs the opposite way. These things will always tend to regress to the mean but not all the way to it.
Would be interesting to see how many games the projections miss by over the 3 period using an absolute value on the underachieving seasons to see how much they are off over that period.
Finally, quantifiable evidence that people who cry “You’re biased” the loudest are often (pun intended) projecting.
I’d give it another 24 hours before jdbolick rolls in here to explain how this entire article is nothing but Red Sox propaganda.
heres what i would love to know – how the error breaks down between talent estimation and playing time estimation. I broke it down for the Pirates 2015 season. I used the depth charts archived on march 17 (wayback machine) and compared it to the final numbers. The pirates batters were projected for 24.5 WAR and they ended up with 21.8 WAR. The pirates pitchers were projected for 10.9 WAR and ended up with 21.7 WAR. Thats where the big error comes from so ill just focus on the pitchers. The RPs were +3 WAR and the starters were +7.8 WAR. To identify errors in projected talent vs projected playing time, I used IP and FIP (projected vs actual). The pirates were projected for a FIP of 3.66 in 1459 IP. They finished the season with a FIP of 3.36 in 1489 IP. The preseason projected depth charts identified 15 pitchers that actually pitched for the pirates during the season. They were projected for 1241 IP and a weighted avg FIP of 3.61. Those pitchers actually pitched 1342 IP and had a weighted avg FIP of 3.45. The difference in FIP talent estimation (3.66 vs 3.45) over the projected IP was worth 22 FIP-runs ((3.61-3.45) * 1241/9). The difference in IP estimation given the actual FIP was worth 39 FIP-runs. So, playing time estimation had nearly twice the error compared to talent estimation. Furthermore, there were 216 projected IP for pitchers that did not actually pitch for the pirates at a FIP of (95 FIP-runs). The Pirates also had 140 IP from pitchers that were not on the preseason depth chart (call up/deadline acquisitions) at a FIP of 2.51 (39 FIP-runs). Thus, the vast majority of the projection error was not due to talent estimation, but instead due to playing time estimation and information the system did not have (though it was information that the team did have, ie call ups (and keep downs?), acquisitions/trades, health/injuries, and roster depth).
What’s the rationale for allowing errors above and errors below the projection to cancel each other out? (Rather than counting up the total error? Shouldn’t 10 wins above + 10 wins below = a 20 win error, rather than 0?) Is it because you’re only concerned with the question of whether there’s a specifically positive or negative bias? Because it seems to have the curious effect of making the projections seem more accurate than they are.
It seems to me that you could likewise have a discussion about why the Red Sox have been consistently missing their projection, regardless of the direction. It might even lead more obviously to the answer of ‘it’s just that shit happens’.
The goal was to find potential bias, which is when a prediction systematically misses on one side or another. If the two missed predictions cancel out, it’s an effect of variance, not bias.
I expect we’ll see a cycle where there will always be one or two teams out ahead of the projections, particularly small markets teams who have to find an edge some other way than signing big money free agents. An enterprising team will find something that both the rest of the league and the people designing projecting systems are undervaluing (or struggling to account for) — in the Pirates case, it would be shifting and framing — and beat its projections for a few years. Then the projections will find a way to account for the new wrinkle, but by then some other enterprising team will have found some new thing that’s being undervalued.
Health is a giant factor when dealing with projections in fairly short time periods. If your best players stay healthy and put up 700+ PA it’s much more likely you’ll beat your projections than if the injury bug hits and the better players miss significant amounts of time. Especially if they have a lack of depth behind said players.
My gut would say teams with more depth and more balanced teams would be closer to the middle than teams that have larger discrepancies in talent on their roster where they have bigger swings due to health within a small time frame.
Whoa – somebody tell Damaso about the clear Rockies, White Sox, and Padres bias in the projection system. I’m sure he’ll make a thousand comments about how Steamer systematically overrates the Jeff Bridich philosophy of team building.
Red Sox have come pretty close. Many a 1st place or last place prediction has been the opposite end of the spectrum
I won’t take issue if some projection system doesn’t predict the Royals to be top 5. But if it doesn’t even predict the 2 time AL defending champ to win their own division 2016, I can’t give it that system much credibility to consider its predictions.