Before You Vote, Some Other Things to Consider
With less than a week remaining in the regular season, a number of end-of-year player awards appear to lack a decisive winner. With so few games left, a decisive winner is unlikely to emerge.
In the American League MVP race, for example, you couldn’t have a greater contrast between the top two candidates. Jose Altuve is the smallest position player in the majors, Aaron Judge the largest. They possess different offensive skills and different defensive homes. Yet these two very different players had produced exactly 7.3 WAR entering play Monday. Lurking behind Altuve and Judge is the game’s best position player, Mike Trout. After losing time to injury, Trout isn’t the favorite. He’s been excellent when healthy, however.
The American League Cy Young race might be even more fascinating. After Chris Sale seemed to have run away with the award by the end of July, Corey Kluber has made it very much a contested race thanks to a remarkable series of performances since he returned from the disabled list in June. How one chooses between the two might depend on which version of pitcher WAR one consults: either the FIP-based version (denoted at the site just as WAR) or the one calculated by runs allowed (RA9-WAR).
Sale holds a sizable lead by the former measure (8.2 to 7.1), Kluber by the latter (8.3 to 7.5). And Baseball Prospectus has Kluber (8.09 WARP) and Sale (7.88. BWARP) pretty close. BP’s metric employs its Deserved Runs Average (DRA) as its WAR baseline for pitchers.
Sale has hit a rare milestone: 300 strikeouts. Kluber has had as good a stretch of play as any pitcher in the 21st century. Both seasons rank within the top-five all-time for strikeout- and walk-rate differential (K-BB%), in the company of Pedro Martinez and Randy Johnson.
The National League MVP field is muddled. Charlie Blackmon, Kris Bryant, Paul Goldschmidt, Anthony Rendon, Giancarlo Stanton, and Joey Votto are all within one win above replacement of each other.
Good luck with those votes, fellow BBWAA scribes.
Part of the problem, particularly with the MVP award, is the voting criteria. The official definition can be read here and hasn’t changed since the first ballots were cast in 1931. The vague language has been an issue for nearly 90 years of voting.
Dear Voter:
There is no clear-cut definition of what Most Valuable means. It is up to the individual voter to decide who was the Most Valuable Player in each league to his team. The MVP need not come from a division winner or other playoff qualifier.
1. Actual value of a player to his team, that is, strength of offense and defense.
2. Number of games played.
3. General character, disposition, loyalty and effort.
4. Former winners are eligible.
5. Members of the committee may vote for more than one member of a team.
You are also urged to give serious consideration to all your selections, from 1 to 10. A 10th-place vote can influence the outcome of an election. You must fill in all 10 places on your ballot. Only regular-season performances are to be taken into consideration.
Keep in mind that all players are eligible for MVP, including pitchers and designated hitters.
There is, of course, quite a bit of room for subjectivity there — and debates have revolved for decades around the implications of the word “value.”
But I think there’s a consideration that has been addressed less often in this age of better data and better measurements of performance: with our votes, we are awarding (and evaluating) history, considering what has happened. We’re not trying to be predictive, not attempting to determine the most talented player at the present moment, or the best player on the best team. The runs scored and runs allowed and context matter because they help determine, ultimately, the real wins and losses in the standings.
For example, while you probably wouldn’t base any sort of predictive analysis on either Win Probability Added or Clutch, both metrics are meaningful representations of what has happened, of who’s produced in crucial situations. Not all raw value (WAR) is created equal. Consider what FanGraphs’ “Clutch” metric is measuring:
“…how much better or worse a player does in high leverage situations than he would have done in a context neutral environment.” It also compares a player against himself, so a player who hits .300 in high leverage situations when he’s an overall .300 hitter is not considered clutch. Clutch does a good job of describing the past, but it does very little towards predicting the future.”
Much of our analysis at FanGraphs and elsewhere is forward-looking; award voting is the opposite, though. It represents an attempt to more accurately understand history and distribute credit.
WAR is more predictive than WPA and Cutch, sure, but WPA tells us about performance in context and Clutch about how the player performs, relative to his skill, in high-leverage situations.
This was the problem with Bryant’s MVP candidacy last year, which Jeff documented last season.
What this gets at is that his hitting this year has been less helpful than the overall numbers would suggest. That’s not spin; I don’t have anything against Kris Bryant. It’s just how things have happened. And just as with Saunders, this is an easy thing to break down. In low-leverage situations, Bryant’s wRC+ ranks in the 99th percentile. In medium-leverage situations, it ranks in the 85th percentile. In high-leverage situations, it ranks in the eighth percentile. Bryant has done the most damage when the results have mattered the least. That’s all this says.
It’s a problem with Bryant’s candidacy again this year, a matter that Jeff recently revisited.
You can be the most valuable player in a league and the least clutch. Clutch and WPA/LI are just offensive measures. But in close races, voters ought to drill deeper.
Notably, Bryant is the third-least clutch player this year. Judge is the least clutch. Now, WAR also accounts for defense and baserunning value — which aren’t included in the win-probability numbers — but the idea here is to show that not all production is distributed equally. Some players have produced in more meaningful spots than others. Judge hit his 50th home run on Monday, he’s a great player who has enjoyed a remarkable season, but the home run also came with a significant lead late in the game. Altuve has produced similar overall value this year, but more of it has come in critical situations.
Consider the following scatter chart:

Considered in this light, one might arrive at a different conclusion about the identities of the MVP favorites.
We can also approach this from a Win Probability Added standpoint, which is dependent on context. Judge ranks 56th in WPA (+1.74) while Altuve ranks eighth (+3.68). Trout, for the record, ranks first (+5.45).
Another way to evaluate this is simply by considering leverage splits for wRC+:
Altuve
Low Leverage: 157
Medium Leverage: 174
High Leverage: 133
Judge
Low Leverage: 190
Medium leverage: 150
High leverage: 95
Trout
Low leverage: 168
Medium leverage: 198
High leverage: 158
In the Cy Young race, Kluber (-0.21) and Sale (-0.35) have very similar ratings in Clutch. As noted before, however, the differences in WAR have their own implications. FIP-based WAR, for example, possesses more predictive value because it neutralizes the influence of balls in play and sequencing. RA9-WAR, however, might be more useful for voting, taking into account the runs that were actually allowed.
Again, we’re concerned with history. What actually happened. That’s this author’s stance, at least.
WAR has become an important tool, perhaps the most important tool for many voters. It marks a step forward from batting average and RBIs. And it’s designed to give a sense — through different recipes at FanGraphs, Baseball Prospectus, and Baseball Reference — about which players produced the most total value overall. It’s a great tool. It’s trying to boil a lot down into one number, and that makes life easier for voters. But WAR doesn’t account for context or impact on individual games. In a close race, we need more context.
Award voting is a matter of history, judging recent history. It’s not a predictive exercise or an argument over who’s the top talent at the present moment or the best player on the best team. And to complete that historical picture, to look back as accurately as possible, voters ought to drill deeper.
A Cleveland native, FanGraphs writer Travis Sawchik is the author of the New York Times bestselling book, Big Data Baseball. He also contributes to The Athletic Cleveland, and has written for the Pittsburgh Tribune-Review, among other outlets. Follow him on Twitter @Travis_Sawchik.
I’m a huge Trout fan, and would like to think that his WPA makes him a serious candidate. But I just can’t. Blame it on you folks at FG, but WAR to me is still the best measure of value. I completely get what you’re saying about predictive value vs. what actually happened, but everyone knows that what actually happened has a lot of random chance in it. So does WAR, but not as much.
Here’s another way of looking at it. A player with the highest WAR gives his team the best chance of getting value when it’s needed. Unless you really believe in clutch–that some players choke when production matters most, while others may find an extra gear in the same situation–the context dependent stats are driven mostly by noise. WAR is the best attempt we have at filtering out some of the noise. Even if it doesn’t account for team wins as well as a context-dependent stat, it does show which player is “helping” his team the most, i.e., the player you most want at the plate when the game is on the line–not looking back over the season, maybe, but right now.
So if you want “right now” instead of a total season contribution, maybe you mean WAR/PA? That leaderboard is:
Trout 0.013
Judge 0.012
Altuve 0.011
Rendon 0.011
Pham 0.011
Sure, except that Trout has had only about 500 “right now” situations over the course of the season, whereas Judge and Altuve are pushing 700.
WARrior – just to play Devil’s Advocate:
If we’re talking about “who provided the most value this season?”, that’s a descriptive stat. WPA, for example, or some other that measures what actually happened. Context included, with all the noise of luck, time lost to injury, and everything else.
Within that framework, Judge WAS more valuable than Trout in 2017
If we’re talking about “who we can expect to provide the most value is the most immediate future (e.g. right now)?” then a predictive or context neutral stat like WAR is appropriate. But it’s also appropriate to normalize the PAs to a degree. We don’t want to overvalue part time players who may excel in very limited opportunities, but we also don’t want to penalize a guy who bats 4th against a #1 hitter who gets 100 more PAs over the course of a season.
Within that framework, Trout IS more valuable at this exact moment in time.
But it goes to the age old question with awards. is the award for the best player, or is it for the player who had the best season. The 2 are not mutually exclusive always by any stretch of the imagination. And yes, the player who had the best season has a lot of luck involved. But that’s part of the game.
Maybe I wasn’t clear enough, but when I said “right now”, I meant that as a description of every single PA throughout the season, not literally only right this minute. My argument is probably confusing because I’m taking what happened over an entire season, as reflected in current WAR, then projecting that back over every PA during the season. In my OP, I said not over the entire season, but then I was referring to what actually happened, not to what we would expect would happen based on his season WAR, or WAR/PA if you like.
From the information we have now, Trout was the best bet at every PA during the season, regardless of what actually happened in that PA. If you could know at the beginning of the season what his WAR/PA at the end would be, you would prefer him over anyone else for any PA. That to me is the mark of the MVP, and the only reason I can’t see him as that is because he missed so many PA.
your post though makes my point. Should the award be for the best player(which I don’t think anyone would dispute is Trout) or the guy who had the best season(which I think a lot of folks would say Judge or Altuve has). I think a lot of the voters vote for who had the best season. I think most of them do quite frankly… So I think it’s going to be a tough go for Trout quite frankly.
Don’t disagree. Trout did not have the best season, my point is just that the reason he didn’t was because of all the games he missed.
yeah that’s very true. But it is what it is. A lot of the voters vote on who had the best season. Not who was the best player. 2 completely different animals.
If WAR is the best measure of value, shouldn’t the MVP go to Kluber?
You misspelled “Sale”
Only if you are looking at what they are likely to do next season (FWAR) instead of what they actually did (BWAR).
No no no no. bWAR does not measure what they actually did any more than fWAR does, it just does a worse job of isolating pitcher performance among team performance.
I said this below, but the idea that the predictive validity of FIP-based WAR is independent of its construct validity is simply wrong. bWAR does a lousy job of predicting future performance specifically because it doesn’t measure individual pitcher performance well.
I think we’ve had this conversation before but I have to disagree. fWAR uses FIP to gauge a pitchers success at run prevention. Basically the “unit” or FIP is theoretical runs allowed per nine innings. It’s estimated run prevention of pitcher for fWAR vs. actual run prevention of pitcher + defense for bWAR (though bWAR does make something of an attempt to adjust for defense.)
While I love (*love*) FIP, I think it can get especially problematic the further from the center of the spectrum you get, and this is important when considering end-of-season awards. For example, it goes without saying that having a lower BABIP allowed will lower your ERA, but what is less obvious is that if you are as awesome as Chris Sale, having a higher BABIP allowed will actually lower your FIP because decreases it the denominator (which is IP before you add the constant). In other words, a higher BABIP against just gives a good pitcher more chances for Ks. Look at Kluber and Sale’s Ks and BBs over TBF instead of IP. Crazy how similar they are.
You have no idea if bWAR does a worse job of isolating pitching from team performance than fWAR. fWAR doesn’t take into account quality of contact, and it doesn’t take into account context (LOB% and clutchness may not be predictive, but I think it’s pretty hard to argue they’re not relevant when you’re talking about a performance award). Kluber leads Sale in WPA by a substantial margin. bWAR does do something to normalize based on team defensive performance. To me the context avoidance of pitching fWAR makes it relatively useless for something like Cy Young voting. On top of that, you have no idea what level of control the pitcher actual had on the quality of contact that fWAR is ignoring.
You’re confusing xFIP with FIP. xFIP projects going forward (TTO). FIP captures what occurred (TTO).
Since it seems like the “clutch” statistic is going to continue to be a “thing” here let us recount the arguments against “clutch” here.
Just as a reminder, bad players repeatedly wind up near the top of the clutch leaderboards–guys like Alcides Escobar–while really good players wind up at the bottom. There are a few possible reasons for this, and probably many others too.
1) The situation-specific performance is relative to an individual’s overall performance. Hence, a player who is doing terribly can look comparatively very good if they excel in higher leverage spots. This isn’t as big a deal when everyone is running the same wRC+ overall, but it is a problem for the statistic overall.
2) “Clutch” isn’t really an indicator of performing better in high-leverage situations, as much as it tells you how vulnerable you are to a shift. With runners on, the defense might not be able to shift as heavily. Thus, players like Pujols, who are very vulnerable to a shift, is not as vulnerable to this weakness. What this means is “you have a massive weakness without runners on that ceases to be a weakness with runners on” rather than “you have ice in your veins.”
Is this a problem for assessing MVP candidates? Is it better that your weaknesses be more evenly distributed across leverage situations? Or is it so small it doesn’t matter?
3) “Clutch” may simply indicate that you get a relief pitcher in to specifically get you out. Let’s consider Kris Bryant as a hypothetical. A manager quite reasonably thinks that having a tired pitcher throw to Bryant with runners in scoring position will end badly. So instead, he turns to the bullpen and brings in a Chris Devenski or Andrew Miller to get the hitter out.
That is to say, when you are really good and you have a chance to pile up RBIs, you suddenly get a fresh pitcher to punish you. And so clutch is quite likely a selection artifact, where the best players are placed in much more difficult situations in clutch situations than the worse ones.
This is where “clutch” stops making any sense at all. If you’re so good that you’re getting the other team to scheme around you, does not make you less of an MVP?
There are probably others too.
Re (3) I suspect there is an IBB effect in there too: again selection bias.
Also, is WPA park-adjusted? If not then it would benefit Rockies’ hitters making them look more clutch than they are.
COOOOOOOOORRRRRRRRS
You are not entirely wrong.
But bare in mind the context of this article.
We are not comparing Judge’s clutch/WPA to Jed Lowrie to make a point against him.
(Judge wrc+ 169 WPA 1.89 WPA/LI 5.86 Clutch -3.87
Lowrie wrc+ 119 WPA 1.91 WPA/LI 1.92 Clutch 0.08)
We are comparing Judge’s Clutch/WPA to that of other elite hitters like Altuve and Trout.
Yes, statistically speaking, better hitters tend to have a lower clutch value. But Judge’s clutch stat is terrible even accounting for that fact as can be seen from the plot above.
(Note: From 1987 to 2016, players with wrc+ in [164,174] have clutch of -0.4 on average more in line with Altuve’s clutch of -0.63)
That is a thing that did happen and is definitely a fair thing to
take into account when considering ‘value’.
Also, I want to strenuously disagree with the idea that RA-9 judges “what happened” for an individual player. Yes, it judges what happened on a pitcher’s watch very well, but not the individual contribution of the pitcher. For that, FIP-based WAR is much better (although SIERA is better than FIP).
Some people want to talk as though FIP-based WAR’s predictive validity is somehow distinct from its construct validity, but it is not.
Agreed. People here tend to refer to FIP-WAR as a predictive stat rather than descriptive, but that is flat-out wrong. FIP absolutely describes what actually did happen, it just does so in a way that happens to be more predictive than pure RA/9 (or ERA, obviously).
FIP has it’s problems though. It ignores any batted ball that’s not a homer. Sorry but that’s a joke. Double into the gap- nope not included. But it includes intentional walks as punishing a pitcher.
Agree. Let’s petition Fangraphs to switch to a SIERA-based WAR measure.
Go easy on my here, everybody…this is flawed.
I incorporated xwoba into FIP to come up with an xwobaFIP of sorts. It’s crude to be sure, and I couldn’t really tell you why I chose to factor xwoba in it the way I did, but it’s held up better at predicting 2017’s ERAs than FIP has for all pitchers who have thrown 100 innings in both 2016 and 2017 (probably flukey). So I’m not claiming it has validity. In fact, I’m fairly sure it’s dumb, dumb, DUMB. But here’s the “equation” – maybe someone smarter than me can work it out to be better…or maybe there already is an xwoba-based FIP.
xwobaFIP = ((1.33*xwoba*(TBF)+(3*(BB+HBP))-(2*K))/IP+FIP constant)
This year:
Sale: 2.27
Kluber: 2.36
Kershaw: 2.60
Scherzer: 2.70
I KNOW there has to be a problem with multiplying the xwoba by the TBF and then by 1.33…but I couldn’t tell you because I couldn’t tell you why 13 was the factor chosen for HR in original FIP either.
In addition to improving the incorporation of xwoba, I’d like to see HBP factor reduced as it seems to be as much of a hitter’s skill as it is a pitcher’s fault.
Again, sorry if this is dumb..
Travis, I’m dying to know who the two dots in the lower right quadrant are…
Melky Cabrera and Albert Pujols.
Thanks!
100-ish RBI and a sub-.290 wOBA for Pujols this year! What a world.
Its interesting to me you didn’t even mention the NL Cy Young. I thought that Scherzer had a clear lead (just like I thought Sale had a clear lead), but according to Tango (link), Kershaw now has a “stranglehold” on the award, since he leads in wins and ERA, traditionally a sufficient condition. That a stats-based evaluator can come to such a strong conclusion opposite what feels obvious to me makes it exactly as clear as the others you mention: which is to say, not at all.
And then contrast these to the ROY votes, which are certain to be unanimous.
Some yahoo is going to vote for Matt Davidson for RoY and the internet is going to explode.
There’s no way Hawk Harrelson has an AL RoY vote this year, is there?
The NL Cy Young race is really interesting. I mean- you have 2 trackers. 1 is Tango and he’s got that for Scherzer. The other is Bill James, and he’s got that for Kershaw. Supposedly the Tango one is more accurate now.
What’s going to be really interesting with the NL race is how much innings matter to voters. I mean, Scherzer has a hefty 26.1 inning lead right now.
I think he said something to the effect that leading in wins and ERA interacts and provides an extra bonus. This year might test that.
it’ll be interesting to see just how close the awards wind up being. I mean, AL Cy Young. Kluber is up by a win and in the ERA he’s up by 0.48. I know a lot of folks here want Sale to win, but those numbers talk to the voters normally- especially the ERA. AL MVP- you have a guy in Judge who has 50 homers and tied for 2nd in RBI. And he’s closing with a huge bang. I really wouldn’t be surprised if Judge wins comfortably. I see the 2 NL races a lot closer than the 2 AL races. Although I think if Stanton gets 60, he’s going to be really tough to beat. The real race is the NL Cy Young I think. It’ll be fun to see how the votes actually come down. I remember 2 years ago people saying it would be real close with Kershaw, Greinke, and Arrieta. And that became a real 2 horse race, with the advanced metrics person(Kershaw) a distant 3rd.
Yeah, the way Kluber has caught up in innings, and considering momentum, he feels like a safe bet now. It’s tough for Sale. He’s like the AL Adam Wainwright, he’s been good enough, but there is always somebody better. Always a bridesmaid…
Yeah. And the big difference this year is his 2nd half dip wasn’t as pronounced as before quite frankly. 1st half ERA 2.75. 2nd half era 2.76. His FIP worse, but overall not anywhere near as bad as before.
Kluber just has been dominant. An historic 4 month stretch that while we’ve seen in NL recently with Kershaw and Arrieta- we haven’t seen in the AL much at all.
The one aspect of the description for MVP I don’t align with is the idea that the player does not have to come from a playoff qualifier. No playoffs, no ballot vote is the way I go. You can be all over the offensive player of the year ballot, I believe that is the Hank Aaron Award, but not MVP.
well any shot Sale had at Cy Young I think is gone tonight. It’s possible if Kluber pitches well last start that he finishes with a higher fWAR than Sale. He’s got a chance with a good start to finish with an ERA .75 better than Sale does.
What I think is so funny now is folks that a few weeks ago were saying how the last 3 starts would determine the race- and now that Sale got roughed up in 2 of those 3 starts, while Kluber hasn’t given up an earned run- and yet those same folks want to act like it’s a close race.
Yeah, that was brutal.
Keep in mind, Sale is *still* ahead in fWAR. He has actually provided more value over the year than Kluber. But the race isn’t decided by fWAR. Now Kluber has more wins and a lower ERA and he’s finishing strong, while Sale isn’t.
I’m off the Sale-for-MVP bandwagon, and the Kluber-Sale Cy Young award is a lot closer for me than it was before. But it’s probably Kluber because of the narrative.
One of these years, they’re just going to start Chris Sale’s season in June so they can get that insane level of dominance for the playoffs.
yeah it’s been really strange this year. Sale you would think would totally benefit from the schedule Kluber had.
But what’s up with Sale’s homers is amazing. He’s given up 9 this month and in his last 11 starts(since Aug 1)- his hr/fb% is a brutal 22.4%. Only 7 guys for the entire season have a worse hr/fb% with at least 50 innings pitched. And the names he’s with there are brutal.
But lets look at Sale and Kluber in this fashion. their 4 great months and their 2 bad months…
bad months- w-l ERA FIP xFIP b/k GS IP
Sale 2 bad months- 4-4 4.09 3.64 2.65 16/97 11 66
Klub 2 bad months- 3-2 5.06 4.44 3.87 13/41 6 37.1
advantage for Sale- but not as much as you would think..
but now, lets look at their great months-
Sale 4 good months- 13-4 2.37 1.92 2.65 27/211 21 148.1
Klub 4 good months- 15-2 1.62 2.06 2.17 23/221 22 161.1
Kluber will get 1 more start to build on this. So as good as Sale was in his 4 good months, Kluber was even better. And Sale’s 2 bad months were bad enough to make Kluber’s 2 bad months not matter as much.
Your bad “2 months” for Kluber is really 1 month, being that he only had 1 abbreviated start in May and went on the DL
very true…. Sale’s “problem” is that in his bad period- he had 28.2 more innings than Kluber did, while Kluber has had 13 more innings in his good period. And that number should get up near 20 after Kluber’s final start.
Oh and just saw this for those pimping Sale’s historic strikeout total of 308 as to why he should get the Cy Young. Kluber’s ERA+ right now sits at 202. That’s good for #36 of all time in MLB. Sale’s 308 k’s- good for #49(tied). So you could argue that Kluber has had a more historic season than Sale has had. Oh, and the cherry on top of this. If Kluber gets 13 k’s in his final start, he will move into the top 100 all time in single season strikeouts with 275.
He’s ahead in fWAR 7.7 to 7.1. If Kluber hadn’t been out a month, he’d very likely lead here too. But by bWAR, Kluber leads 7.8 to 5.9; it’s not close. You cannot say “Because his fWAR is higher Sale pitched better”. He did a significantly worse job than Kluber of preventing runs, how much of that was the defense behind them, and how much of that was Sale allowing harder contact and/or more inopportune contact, you cannot say. If we go back to 2014 when there was the Kluber vs Felix argument, where Kluber had the higher ERA but also had a horrific defense, Kluber lead in bWAR as well that year.
saw something that’s really interesting with Sale…
Sale has thrown this year 12.1 fewer innings than last year. 56 fewer batters. But the problem is that he’s thrown all of 3 fewer pitches this year vs last year in those 12.1 fewer innings.