Let’s Project the Royals’ BaseRuns Gap
This morning, Jeff Sullivan posted the results of his team projection polls, and not surprisingly, you guys don’t buy into the 77-win forecast that our Playoff Odds are currently giving the Royals. The aggregate projection from the readers in Jeff’s poll put the Royals at 83 wins, and 71 percent of the people who voted believed that our forecast was at least four wins too low. Which is perfectly understandable, given that they just won the World Series and all, and it is no easy task trying to justify why a team that has won the AL pennant two years in a row might now be the worst team in the league.
So I want to follow up on Jeff’s poll, because while he collected the expected win total, he didn’t gather any information about how they’re going to get there. And the how is one of the most interesting parts of the Royals. Last year, they won 95 games, but their BaseRuns expected record was only 84-78, which is one of the primary reasons the projections are down on their 2016 chances. Forecasting systems only project context-neutral performance, and assume that the timing of events — which is what drives the difference from BaseRuns expected record — will be equal for all teams.
Since you guys believe the Royals are significantly better than ZIPS and Steamer believe, I’m curious how much of that is due to the belief that the projections are simply incorrectly forecasting individual performance, or whether you believe the Royals roster has inherent traits that will allow it to beat context-neutral expectations. Because looking at the difference between the forecasts and the FANS projections — created by the collective balloting of readers here on FanGraphs — doesn’t necessarily support the idea of the projections badly missing on the individual performances.
On the offensive side of things, there are at least five ballots for 11 KC hitters; the regular starting nine, plus Christian Colon and Paulo Orlando. If we just take a weighted average of the projected FANS wOBA for that group, it comes to .323; our depth charts project a .313 wOBA for that same group with that same distribution of playing time, so there’s a difference, but it’s basically the same story if we run that for every club; the fans projections are always optimistic, running 15-20% on the high side for a vast majority of players.
Eyeballing the numbers — we’ll have actual standings projections based on the FANS forecasts later this month — it doesn’t appear that the gap in expected individual performance is much larger than we find with other teams. There’s a huge difference on Lorenzo Cain (+5.9 WAR for the FANS, +3.6 for our depth charts), but most of the rest of the players are close enough to their depth chart forecast that it’s basically equal once you normalize the FANS projections to take out the consistent optimism bias.
On the pitching side of things, there’s more disagreement, with almost every member of the starting rotation projected for at least +1 WAR more than their depth charts forecast.
| Player | FANS | Depth Charts |
| Yordano Ventura | 4.1 | 2.8 |
| Edinson Volquez | 2.8 | 1.4 |
| Ian Kennedy | 2.9 | 1.4 |
| Danny Duffy | 2.4 | 0.8 |
| Kris Medlen | 2.1 | 1.2 |
| Chris Young | 1.3 | 0.1 |
| Total | 15.6 | 7.7 |
The FANS projections think the Royals rotation is just fine; our depth charts think its the second-worst in baseball. So that’s a real fundamental disagreement, and likely the source of some of the disagreement. But the FANS forecasts are overly optimistic for everyone, and once we adjust for that bias, it doesn’t explain the entirety of the difference.
So I figured we’ll just follow up with a poll. Knowing that the Royals beat their BaseRuns record by 11 games last year — and eight games in 2014, plus six games in 2013, for a three year running total of +25 wins relative to their context-neutral performances — what do you expect them to do relative to their BaseRuns expected record in 2016? How much of the difference in opinion against the forecasts is due to the belief that the Royals have indeed solved sequencing to some degree, and have a sustained ability to distribute their performances in a manner that leads to more wins?
We’ll look at the results of this poll next week, and depending on the results, may have some follow-up questions about you voted the way you did. The Royals certainly are one of the most fascinating teams in baseball, and if they’ve figured out how to sustainably sequence their events in a way that can be repeated by other teams, this has the potential to cause some big shifts in the sport. But maybe you guys don’t actually believe that the Royals have figured out sequencing, and you think that the individual projections are just too low. I’m curious to see whether you guys think it’s simply the forecasting systems missing on a very good group of players, or if there’s a context-specific skill that you believe the Royals have that the forecasts simply aren’t trying to project.
Dave is the Managing Editor of FanGraphs.
Dave,
When is the trade value update coming?
Thanks
why is it not possible that the Royals beating their context-neutral projections for 3 straight years is just statistical noise?
i find it weird that this possibility wasn’t mentioned at all.
when talking about season-long context neutral projections, isn’t 3 years a small sample size?
I like your overall point, but to play devil’s advocate for a minute…if 3 years is still too small a sample, then I am very depressed about the amount of time I waste discussing and debating baseball.
I would like to see other avenues explored in terms of looking for elements that could cause projections to skew greatly. Like, say, if there are certain team-wide characteristics that can improve upon/subtract from positive sequencing outcome chances.
We’re going to need a bigger sample
That’s what she said!
indeed – recency bias is strong. Survey results look like a typical example of that.
I’m not sure how you could figure out sequencing? Unless you had a team full of peak Barry Bonds. Even then I’m not so sure.
I think it’s possible the Royals combo of elite contact skills and defense help them beat it to a degree though.
And the elite bullpen, as well.
Don’t the projections love the Royals defense though… Maybe they are underrating it slightly because they don’t want to project bonkers numbers, but they are projected this year again to be head and shoulders above the next defense.
There’s another possibility, which is that you might think individual WAR doesn’t map onto team BaseRuns one-to-one. In particular, if you think that the Royals defense is preventing hits on balls in play in ways that don’t show up in UZR, then you’d expect their BaseRuns wins to be greater than the total wins you’d expect from the some of the WAR for the team.
(I don’t know if I think this about the Royals, but I do think something like this may be going on for the Pirates.)
I don’t think the Royals have mastered anything about sequencing. I think over the last three years their bullpen’s been really good. I boldly predict that if their bullpen continues to be a shutdown bullpen in 2016, there’s a high likelihood they’ll exceed their BaseRuns projections by somewhere between a little and a lot.
Unless, of course, they get unlucky.
The idea isn’t so much about the Royals mastering sequencing, it’s more like the collective traits of the players starting to skew sequencing results in their favor.
Basic example: the average KC base-runner is faster than typical, and the average KC hitter has a better chance of putting the ball in play than typical. Combine the two and their chances that a runner on base will score might be higher than the normalized run expectancy numbers, even though such dynamics aren’t really factored into BaseRuns (understandably so).
http://www.fangraphs.com/leaders.aspx?pos=np&stats=bat&lg=all&qual=0&type=1&season=2015&month=0&season1=2015&ind=0&team=0,ts&rost=0&age=0&filter=&players=0&sort=13,d
23rd in UBR last year.
UBR has two big blind spots which are perhaps where the Royals can score excess runs. First, UBR is context-neutral while advancing bases is a conscious decision where a player can use context to make his decision. Maybe the Royals are getting dinged on UBR for doing things like being aggressive in low-risk high-reward situations (e.g. tagging the man at 3rd with 1-out and nobody else on). That strategy could be “too aggressive” and hurt UBR linear weights but might create surplus value contextually (i.e. in WPA terms). Maybe it is sequencing. But maybe the Royals are influencing the sequencing.
Second, teams lose UBR value when the batter is a good runner: e.g. UBR will reward a runner for “stretching” first-to-third on a long single but will punish a runner for “only” going first-to-third on a hustle double.
It’s far from a comprehensive explanation, but this is exactly the kind of area where the metrics could conceivably be short-changing the Royals. The Royals could be leveraging contextual skills (picking good spots to run, picking good spots to use pinch runners, picking good spots to use excellent relievers, picking good spots to bring in defensive replacements) while sacrificing context-neutral skills (giving bad hitters like Escobar and Dyson too many PAs but letting them be in lower-leverage situations (like leadoff or when the team is winning); or giving good hitters like Gordon too few PAs but letting them be in high-leverage situations.
Aggressive base-running does much less for WPA than you might think. Try calculating WPA for a few games yourself and you’ll see.
Seems like it would be easy enough to answer whether or not 3leet bullpens = beating BaseRuns?
I’d be curious to know how much rhe Royals (or any team, for that matter) rely on proprietary models for their personnel decisions, in game strategy, organizational philosophy, etc. Tools that we, the readers and writers at fangraphs, don’t have.
As in, how far behind are we? Surely a gap exists based on the millions of dollars in play in MLB.
Are your projections systems looking at the Royals/the entire league in a way that they would have looked at themselves 5 years ago? Two years? Ten tears?
My gut thinks it’s mostly random variation due to large error bars on this whole process. It would not surprise me at all if the Royals end 2016 with a +6 win differential, and I would certainly cheer for that as a fan (and troublemaker), but I wouldn’t bet on it. After all, 7 teams finished last season with 6+ wins. Heck, three teams had a 10+ differential.
My best guess is that’s there’s some fire with the smoke, but I’m far from ready to jump in with both feet. I wondered if there might be some unmeasured defense, base running, or even bunting.
I’d pick somewhere in the +1-3.
Really? There are people who think a team can pull 6+ wins out of no where? And I don’t think their pen is elite anymore without Holland at his best. Davis is probably the best reliever in baseball but Herrera’s ERA has bounced around for a few years and their next best reliever is Madson who should be pretty good. Their offense is mediocre and their rotation is bad, also it relies on a 37 year old to continue to beat his peripherals by a mile. Their defense is truly elite but I don’t know why they’ve earned so many extra wins just because. We’ve been through this with the Os and other teams. They’re a decent team that blatantly overachieved last year and happened to make it through October.
This is what turns some fans off of metrics – a refusal to accept that the metrics are wrong in certain situations… instead claiming small samples and variance.
I understand the argument but the Royals have now done this for two baseball seasons. At some point you just need to accept that it’s not statistical noise even if it can’t be explained or defined mathematically.
To compare them to the o’s or any other team is absurd. When you watch the Royals you don’t think “wow these guys are lucky and playing beyond their ass.” You think, wow these guys are really good and play the game in a different fashion than anyone else.
Metrics evolve slowly, and there will always be periods that are classified as noise but in raality they are a failure to compensate quickly enough.
If two years is noise, what isn’t noise? If they do it two more years will we still classify it as noise or will we accept that were missing something?
The royals played to the sum of their parts last year. They weren’t lucky or overachievers. They’re just very good at baseball and it’s a team full of guys who have been playing together since they were kids. Maybe that can’t be quantified but I’m a firm believer that it exists.
Signed,
Someone who bet a bundle on the Royals win total under last year because I was too a skeptic. I can just admit I was wrong and not hide behind the statistical noise or variance argument.
“I understand the argument but the Royals have now done this for two baseball seasons. At some point you just need to accept that it’s not statistical noise even if it can’t be explained or defined mathematically.”
Two seasons! In a row! Holy crap that model self-evidently is rigged against the Royals! grumble grumble…
Of course no one is claiming any model, or this model specifically, is infallible. Of course the model makers are looking for ways to tweak and improve them, but in a way that makes them better across all 30 teams in the aggregate, and not for one team in particular. But if we’re expecting pre-season predictive models to be super close for every team, then…well…
Bear in mind, too, that a 79-win projection with a 6-win standard deviation will miss a team winning 91+ about 5% of the time. That model will have “big misses” on 2-3 teams by random variation each year (to say nothing of trades, major injuries, etc. that occur after the models are run). If it’s consistently missing on the same team (or similar types of teams) often that certainly that should be examined. Everyone benefits from better predictive models, certainly the model makers and the websites that run them most of all. But to automatically claim that “the model is broken”, “biased against my favorite team”, etc., is an exercise in hysterics.
My post above is a hyperlink that for some reason does not look like a hyperlink.
Here it is again: http://www.beyondtheboxscore.com/2014/10/17/6986409/true-talent-run-differential-best-record-random-chance-simulation
It’s not just the sequencing: I also give the royals credit for defensive synergy between their fantastic OF defense and FB heavy pitching staff. FIP war is going to underrate Chris Young for sure but also a lot of the other KC pitchers to lesser extents. There’s gotta be something to the compounding effects of great defense plus pitchers that have perfect approaches for the environment. Theirs is not a neutral context.
Follow up thought: I don’t expect the Royals to greatly outperform their individual FIP-war projections but I do expect the Royals pitching staff to have higher RA-9 WAR figures than their FIP-war figures and can see a few wins there which the projections cannot/will not capture. If we recognize that certain pitchers are better judged by RA9 WAR why not staffs/defensive units on the whole?
Another follow-up:
If Steamer/Zips were run with the ERA numbers projected instead of the FIP numbers could we make a projected RA9 WAR figure and see how that compares?
My wife is a scientist. If her models don’t work, she fixes them instead of writing long letters to the peer-reviewers about why they suck and her models are right.
Last year the Royals were, pythagorean expectation, a 90-72 team. Your models that predicted their 79-83 finish don’t seem to understand that.
I think it would better if stopped making excuses and arguing fans and show some humility and fix your models.
Holy shit! I couldn’t agree more with your comment. Fangraphs authors need to realize that statistics are not some mystical black box. There are people that know ANOVA, linear regression, when your models suck, go look in the mirror before blaming others or making excuses. Too many articles lately, for example, the ones on parity – “oh the reasons we have been so far off base the last few years is because there is so much parity in MLB, and because of that, we will continue to be less and less accurate.” This article and so many articles about the Royals as a team or on individual players like Alcides Escobar – and why Fangraphs does not understand why he works at the top of the order for their lineup. Too many lame excuses.
My wife is a scientist. If the model is not working, instead of writing long letters to estimate the despair and the head of it was true.
Last year, the Royal family awaits Pythagoras, at 90-72. 79-83 models that predict its end does not seem to understand that.
I think it would be better to stop making excuses and claim to be fans and show a little humility and to improve their models.
_____________________________________________________________
Cow dung! I with my commentary. Fangraphs the perpetrators must understand that these statistics are not a mysterious black box. I don’t know the analysis of variance, linear regression, if smoking patterns, look at yourself in the mirror before blaming others. Too much recent messages, for example, those on the feast-“Oh, for reasons that are now in the last few years is that there is parity in the NFL, and that’s why we continue to exact the best” this article and a lot of articles about the Royal family, team or individual players such as Alcides Escobar-. and why Fangraphs doesn’t understand why the order of their composition. too bad excuses.
Ask your wife about the difference between a bad model and a model with a large epsilon. Or, just stop using appeals to authority to argue against that which you know little about. If it was so easy to “fix your models” and remove all errors, then Dave Cameron would not bother writing articles for fangraphs, and instead would be lounging in the Caribbean right now.
A team beat their projection one year! Stop the presses! Better scrap the model and fix it!
This…this may be the most ignorant comment I’ve seen in some time. If it were that easy to just “tweak” you don’t think it would have happened ages ago? Or someone would have put forth one that’s clearly better? Like there’s any redeeming value in keeping something that’s “broken”?
You do realize that a projection is more like a normal distribution of outcomes with pretty fat tails? That a 90-win season from a 79-win projection (or vice versa) is well within the realm of possibility, and quite likely for one of 30 teams in a given year?
(Then again my wife is admittedly not a scientist…)
My brother’s wife is a surgeon. Do I get extra points for that? What if I tell you that my father and brothers are accountants, and my grandfather, uncles, and cousins are engineers? Do I get extra points for that as well? What if I tell you I also know some actuaries? And my favorite color is clear…
I voted with the optimistic plurality in both Jeff’s and Dave’s polls. I know that’s irrational, but I still believe KC will be better than projections/base runs.
It may be that, liking the kind of team they have, I just want them to do well.
Oddly, I don’t feel that way about my own favorite team, Tampa. I see them as an average team that could be better. I hope they will be, but I didn’t vote that way.
As I understand it, you calculate WAR using FIP instead of actual ERA.
That may be great for discussions about Hof worthiness or what will happen when a pitcher changes teams. However, actual results are NOT fielding independent. If the fielding is better than average, the pitchers should be expected to beat their FIP estimates.
FIP is a cool concept. It has uses. But FIP is not good for estimating W/L of real teams. Probably not good for Fantasy Rankings unless the result is an outlier or the pitcher has changed teams.
My guess is the people who think all the KC pitchers are better than you do know there is a good defense behind them and take it into account.
Dave Cameron, in a chat some months back, said:
“For reasons I don’t totally understand, many in the baseball community seem to have settled on using context-neutral numbers to evaluate players, but context-included numbers to evaluate teams.”
Well, yeah. As they should.
While each individual MLB hitter or pitcher is presumably at all times competing to the best of his abilities, the same cannot be said about *teams* — because managers are continually if not continuously striking a balance within every single game between doing the most they can to win right now, and making whatever small or large tactical compromises are necessary to increase their chances of winning tomorrow, next week, next month, and even next year.
BaseRuns isn’t the underlying reality. It’s just an alternative reality.
The most obvious thing that managers are doing which is context-dependent is using great relief pitchers in high-leverage situations and mediocre relief pitchers in low-leverage situations, and this will naturally affect the relationship between wins and BaseRuns.
I’m really surprised that Dave said that. The whole reason to use BaseRuns is to give a more accurate measure of how runs are assembled – Tango gets into this in good detail on his blog.
But even the BaseRuns projections themselves are based upon context-neutral projected statistics, and therefore don’t capture any kind of systematic team-level impact of tactical decision-making (reliever roles probably being the most pertinent). I can see the difficulties in turning these into components of a projection, but they really should be a part of the wholistic evaluation of team performance.
As I recall, my impression was that Dave was trying to hint that context-included data should also be used to evaluate individual players, but he didn’t want to just come out and say this because (1) many of the readers, who believe in context-neutrality to the point of having no interest in the outcome of actual games, would have gone ballistic, and (2) he didn’t have any specific measures to propose at the time.
Thanks for the responses, john-ftg.
I checked, and it turns out my memory was slightly faulty. I was mistaken about Dave’s quote being from a chat; it was actually from this article of 22 September last year:
http://www.fangraphs.com/blogs/the-role-of-context-in-determining-the-best/#more-198624
Dave’s point was basically that individuals and teams should be evaluated by the same standards. Some select quotes:
“Now here’s a statement that I’m guessing would be a bit more controversial: the Washington Nationals are having the best season in the National League East.”
“Over the last few weeks, both Jeff and I have written multiple posts about the role of sequencing in a team’s results, and how large an impact the order of events can have on the outcomes of games and seasons. And I’ve found it interesting to see how different the reactions can be when you suggest using context-neutral or context-included numbers to evaluate players or teams; for reasons I don’t totally understand, many in the baseball community seem to have settled on using context-neutral numbers to evaluate players but context-included numbers to evaluate teams.”
“I guess what I’m suggesting is that we strive for logical consistency. If we’re going to say that context-neutral metrics prove that Harper has been the best player in baseball, let’s also use context-neutral numbers to evaluate team performance.”
Dave’s article was a good one, but a shortcoming of the piece was that it ignored the influence of managerial decisionmaking. Because of managers, there is no such thing as context-neutral team performance.
But he did also write in the article:
If I had an MVP vote this year, I’d vote for Bryce Harper. But I think it’s reasonable to consider the situational performances of the various candidates, and to not shout down those who might consider voting for a guy like Rizzo because of the fact that he ended up hitting better in more important situations.
This doesn’t suggest a strong commitment to context-neutral stats. Even the statement that “the Washington Nationals are having the best season in the National League East” was worded as a hypothetical statement, not Dave’s opinion. Given the overwhelming bias in FG for contrast-neutral stats, I took Dave’s article as being a call for being more open-minded. I think that the reason he didn’t mention managerial decisions was that the article was more about analyzing hitters than analyzing pitchers, and managerial decisions are more obviously related to this issue with regard to pitchers.
Fair points, john. (In fact, I posted the link to Dave’s article in the hope that any interested party still hanging out on this thread might peruse the piece in its entirety.)
I remain disappointed in general, though, re Dave and other SABR-minds who sometimes seem extravagantly neglectful of the fundamental baseball fact that the average manager, at dozens of times during every season, is not at all interested in winning by the largest possible margin, or losing by the smallest. BaseRuns ideologues ignore this reality.
Here’s one explicit case in point (though it has nothing to do with BaseRuns per se). A regular contributor to a nationally prestigious analytical site (no need to name names) said this last summer:
“Good managers play to win the game in front of them, and consider the long term only when they have that luxury.”
I would say this respected writer/analyst got things exactly backwards. What he’s describing applies only to October baseball — and perhaps the very end of a very tight pennant race.
Under ordinary day-to-day circumstances, the good manager must perpetually be considering the long term effects of his in-game tactical decisions, esp. when it comes to the physical well-being & effectiveness of his pitching staff.
Considering only the game in front of them? Now THAT’S the luxury. And a pretty rare one.
There are two objective reasons to think that the Royals wins record should be slightly better than their BaseRuns record. You guys came out with a few articles last October or thereabouts showing that Clutch is affected by (1) elite relievers and (2) a low number of strikeouts for the team’s batters. There is a very big difference in leverage between different situations in which relief pitchers are used, and a team which can use great relievers in high-leverage situations and mediocre relievers in low-leverage situations will naturally do better in terms of wins than a team with whose relievers total WAR is the same but they are equally good. And a team whose batters have a very low K% can produce decent results in high-leverage situations even if they get out by advancing runners. This shouldn’t be controversial at this stage.
I think you’ve got two of the big ones right there. Clutch, by its actual design, will be higher for teams whose hitters don’t follow the TTO formula, though the impact isn’t huge (only about one win for the Royals last year).
Probably the bigger factor is the Royals’ disparity between the quality of their high-leverage relievers vs. the low-leverage relievers and starters. While this difference exists for most teams, it was very pronounced for KC last year (Wade Davis in particular being huge). This may be worth several wins.