Did WAR Ruin the MVP Conversation?

On Wednesday, I wrote about one of my favorite topics: The impact of sabermetrics on the practice and analysis of baseball. Specifically, in this case: How MVP voters behave in the post-Fire Joe Morgan era. And for those of you who got to the end of that 2,000-word post and did not feel sated, there’s good news! This was not the question I actually set out to answer when I started kicking the topic around.
Welcome to Part 2.
The very name of the MVP award invites voters to consider the value of a certain player’s contributions. For nearly 100 years, that was a tricky proposition. How do you weigh differences in position, in playing style, park factors, hitting versus pitching versus fielding versus baserunning? It’s enough to boggle the mind.
One of the earliest and most enduring projects of the sabermetrics movement has been the pursuit of a catch-all metric that captures the entire picture. To measure every player’s production, put it in context, and spit out a single number that says he was worth X number of runs or Y number of wins. The implications of such a number are obvious for an enterprise such as identifying the most valuable player in the league in a given season.
Critics from days gone by — or critics from now, who are so behind the times they only sound like they’re from days gone by — would wield WAR as a straw man against more numerate analysts. “If you believe in WAR, why not just vote the WAR leaderboard?”
The answer is that WAR, while published to the tenth of a win, is a model built on assumptions that might not hold to that level of precision. It’s a good starting point, but it’s not the whole conversation, especially in a close race.
And you can tell that WAR is not settled science because even the sabermetrics eggheads can’t agree on what WAR is. FanGraphs, obviously, has its own WAR, and I use that because I’m a good company man. But there are dozens of MVP voters, and many thousands of people who publish analysis of baseball for public consumption. Unfortunately, some of them get their information from other sources. (But we’re working on that, I promise.)
When people talk about WAR, they’re either talking about FanGraphs WAR, Baseball Reference WAR, or WARP, from Baseball Prospectus. (The “p” at the end is for PECOTA, I think.) Each makes its own assumptions about the role of defense in pitcher evaluation, just to name one point of divergence. So while all three mainstream WARs usually arrive at similar conclusions, they don’t arrive at the same conclusion. Which is what you’d need to make “sort by WAR” a coherent award voting strategy.
I’ve compiled a list of every player since 2000 to finish in the top 10 in MVP voting, and their WAR totals that year for fWAR, bWAR, and WARP — 481 seasons in all. On only one occasion — Justin Turner in 2017 — did all three major flavors of WAR agree on a player’s value to within a tenth of a win. So I feel pretty confident that Turner was worth exactly 5.6 WAR in 2017. Everything else is open to interpretation.
There’s another complicating factor: All three flavors of WAR get updated every so often, as wins above replacement is not some inviolable concept that was considered perfect when it was first willed to humanity by the gods. It’s a statistical model, after all, and those ought to get updated when new information emerges. So the WAR totals you see now might not be exactly the same as they were when voters were looking at the various sites’ leaderboards back in the day.
So there are different ways of answering the question of how closely WAR and the MVP vote line up. From a normative standpoint, you could look at it as the voters getting the question right, or as a check on WAR — do the numbers match what we see with our eyes, or is this math left to run amok?
From 2000 to 2023 (since we don’t have the MVP vote tallies from this year yet), there were 134 individual league leaders in WAR. That’s 24 seasons times two leagues times three WAR brands, minus the five years (2000 to 2004) for which BP doesn’t have published WAR data.
Out of those 134 individual player seasons, the WAR leader — for each league and WAR type — finished in the top 10 in MVP voting 120 times. Here are the exceptions.
| Season | League | Player | WAR(P) Type |
|---|---|---|---|
| 2005 | AL | Victor Martinez | Baseball Prospectus |
| 2006 | AL | Grady Sizemore | FanGraphs |
| 2008 | AL | Nick Markakis | Baseball Reference |
| 2009 | AL | Zack Greinke | FanGraphs and B-Ref |
| 2009 | AL | Justin Verlander | Baseball Prospectus |
| 2011 | NL | Cliff Lee | Baseball Reference |
| 2011 | NL | Brian McCann | Baseball Prospectus |
| 2014 | AL | Corey Kluber | Baseball Reference |
| 2014 | AL | José Bautista | Baseball Prospectus |
| 2016 | NL | Buster Posey | Baseball Prospectus |
| 2021 | NL | Corbin Burnes | FanGraphs |
| 2021 | NL | Zack Wheeler | Baseball Reference |
| 2022 | NL | Juan Soto | Baseball Prospectus |
It should surprise no one that starting pitchers are heavily represented here. MVP voters hate pitchers for reasons that are not entirely clear to me. A starting pitcher can lead the league in WAR and not get the time of day from voters. Greinke, apparently, can lead the league in two different WARs and still be ignored.
On Wednesday, I dragged Max Nichols, the poor soul who cast the one dissenting MVP vote against Carl Yastrzemski in 1967. Today, I’d like to praise another singularly iconoclastic voter: Nick Piecoro of the Arizona Republic in 2018.
That year, Christian Yelich ran away with the MVP vote, winning 29 first-place votes for a season in which he missed the Triple Crown by two home runs and one RBI. He hit .326/.402/.598, which was not only a paradigm-shifting breakout for Yelich himself, but also a transformative one for his team. Yelich, who was traded from Miami to Milwaukee the previous winter, led the unfashionable Brewers to the no. 1 seed in the National League. Narratively, it was not unlike Yaz in 1967: an unbelievable individual season at the center of the league’s best team level story.
And yet Jacob deGrom was so much better. This was the 10-9, 1.70 ERA season that spawned a million memes and ended up, somehow, going underrated by voters. The Mets ace led the league in all three WAR categories and had more than a win on Yelich in all three cases, up to 2.6 WAR over the eventual MVP according to B-Ref. And yet he only finished fifth.
Not every MVP debate has a right and wrong answer. This one did, and Piecoro was the only voter who found it.
Most of the rest of the ignored WAR leaders topped the table in WARP, which can be a bit stingy and occasionally spits out an unexpected result. Nevertheless, it’s a reminder that Soto was really good even in a down year in 2022, and McCann was hugely underrated in general.
I did want to highlight Markakis because, at the risk of being unkind, he’s the last person I expected to lead the league in anything, ever. He led the league in bWAR in 2008 by having, in some respects, a typical Markakis season: He hit around .300, with about 20 homers and 40-odd doubles, and played good defense in right. What was unusual is that he also walked a career-high 99 times, which elevated his OBP to .406, and DRS credited him with 22 runs saved above average, which is a ridiculous number. That was enough to boost a very, very good season to 7.4 bWAR (and 6.1 fWAR, so it’s not like B-Ref was an outlier here), and in a weak season that topped the AL.
This was my favorite moment in researching this piece. Finding out Markakis led the league in WAR once was like finding out Cesar Tovar was actually better than Yastrzemski in 1967. Suffice it to say, the voters did not notice. Not only did Markakis miss out on the top 10 in MVP balloting, he didn’t get a single vote from anyone.
Now that we’ve seen the WAR leaders who didn’t get much love from the MVP voters, let’s look at the other side of the equation: MVPs who didn’t lead the league in any version of WAR. Since 2000, there have been 18 such cases. (Incidentally, that list includes the past two Twins MVPs and the past three Phillies MVPs, so in cases where voters dislike WAR they apparently love sarcasm and smoked and/or cured meats.)
That group of 18 also includes both MVPs in 2000 and nine of the 20 MVPs between 2000 and 2009. That inflection point would seem to hold with the historical proliferation of advanced stats. The Fire Joe Morgan era ended in 2008, and the early 2010s were to the anti-intellectual voter what the Hundred Days were to Napoleon. These included the Jack Morris Hall of Fame debate and the Miguel Cabrera vs. Mike Trout MVP campaigns of 2012 and 2013.
But recent history doesn’t line up with the WAR leaderboards either. From 2020 to 2022, four out of six MVPs — José Abreu, Freddie Freeman, Bryce Harper, and Paul Goldschmidt — failed to top any WAR leaderboard. Of course, all of those races were hard cases: 2020 because of the 60-game season, 2021 and 2022 because of a diffuse and pitcher-dominated WAR leaderboard.
Indeed, many of these supposedly WAR-deficient MVPs only missed out on league leadership by fractions of a win. So let’s narrow that group to MVPs who won in years where a different player led the league in at least two different WAR categories by at least a win in each case.
| Year | League | Actual MVP | bWAR | fWAR | WARP | WAR MVP | bWAR | fWAR | WARP |
|---|---|---|---|---|---|---|---|---|---|
| 2002 | AL | Miguel Tejada | 5.7 | 4.5 | n/a | Alex Rodriguez | 8.8 | 10.0 | n/a |
| 2012 | AL | Miguel Cabrera | 7.1 | 7.3 | 6.5 | Mike Trout | 10.5 | 10.1 | 6.5 |
| 2013 | AL | Miguel Cabrera | 7.5 | 8.6 | 7.0 | Mike Trout | 8.9 | 10.1 | 7.5 |
| 2015 | AL | Josh Donaldson | 7.1 | 8.7 | 7.0 | Mike Trout | 9.6 | 9.3 | 8.0 |
| 2017 | AL | Jose Altuve | 7.7 | 7.7 | 5.4 | Aaron Judge | 8.0 | 8.7 | 8.3 |
| 2018 | NL | Christian Yelich | 7.3 | 7.7 | 5.9 | Jacob deGrom | 9.9 | 9.0 | 7.0 |
This list doesn’t include cases like the 2000 AL race or both races in 2006, where it’s clear that a purely WAR-based voter would not have given the MVP to the player who won in real life, but the top of the leaderboard is close enough that it was less clear which player should’ve won. (I had forgotten, for instance, how good Jason Giambi was in 2001 and Carlos Beltrán was in 2006.)
Every one of these cases has a strong narrative basis. Cabrera won the Triple Crown in 2012. Donaldson, Altuve, and Yelich turned up-and-coming teams into juggernauts. Nobody respects pitchers or A-Rod, and for some reason Trout became a proto-culture war volleyball in the waning days of the sport’s battle over empirics.
From 2000 to 2004, FanGraphs and Baseball Reference WAR had the same leader nine times out of 10 chances. That might have something to do with methodological convergence, but it probably has more to do with Barry Bonds and Alex Rodriguez just dragging everyone else in the league during that period. The voters — most of whom had probably never heard of WAR at the time — gave the consensus WAR leader the MVP five times out of those nine chances, and voted him second on two other occasions.
In the 19 seasons that followed, one player led the league in all three WARs on 15 occasions. The voters agreed 10 times.
| Year | League | Name | bWAR | fWAR | WARP | MVP |
|---|---|---|---|---|---|---|
| 2007 | AL | Alex Rodriguez | 9.4 | 9.6 | 6.3 | Rodriguez |
| 2008 | NL | Albert Pujols | 9.2 | 8.7 | 9.2 | Pujols |
| 2009 | NL | Albert Pujols | 9.7 | 8.4 | 10.4 | Pujols |
| 2012 | NL | Buster Posey | 7.6 | 9.8 | 8.0 | Posey |
| 2013 | AL | Mike Trout | 8.9 | 10.1 | 7.5 | Miguel Cabrera |
| 2015 | NL | Bryce Harper | 9.7 | 9.3 | 8.0 | Harper |
| 2015 | AL | Mike Trout | 9.6 | 9.3 | 8.0 | Josh Donaldson |
| 2016 | AL | Mike Trout | 10.5 | 8.7 | 8.9 | Trout |
| 2017 | AL | Aaron Judge | 8.0 | 8.7 | 8.3 | Jose Altuve |
| 2018 | NL | Jacob deGrom | 9.9 | 9.0 | 7.0 | Christian Yelich |
| 2019 | NL | Cody Bellinger | 8.6 | 7.9 | 6.9 | Bellinger |
| 2020 | AL | Shane Bieber | 3.2 | 3.1 | 2.5 | José Abreu |
| 2021 | AL | Shohei Ohtani | 8.9 | 8.0 | 10.2 | Ohtani |
| 2022 | AL | Aaron Judge | 10.5 | 11.1 | 10.0 | Judge |
| 2023 | AL | Shohei Ohtani | 9.9 | 8.9 | 9.3 | Ohtani |
In three of those cases — Trout in 2013 and 2015 and Judge in 2017 — the consensus WAR leader came second. The other two involved pitchers: deGrom in 2018 and Shane Bieber in 2020, and while neither won the MVP, they are the two most recent full-time pitchers to finish in the top five.
Has WAR turned MVP voting into a leaderboard-reading exercise, as critics predicted? Not really. Thanks to the three publishers’ subtle differences in methodology, I’m not sure that it can. When there is consensus, the voters usually follow suit (again, unless there’s a pitcher involved). But that was basically the case 20 years ago, when on-base percentage was state of the art.
If MVP voting has become predictable in the 2020s, and if it’s the result of pressures and innovations brought on by sabermetrics and online media, I don’t think it’s because everyone is just blindly following WAR. We’re all getting better information now, and people with clubhouse access and award votes are better equipped to use that information. Whether that leads to conformity is of secondary importance to the quality of the analysis.
Even in the 1960s, there was a risk of being brigaded for expressing an unpopular opinion — just ask the ill-fated Max Nichols. And fear of the dogpile has only grown immeasurably since then. It’s an incentive not to disagree with the consensus, sure, but it’s an even bigger incentive to get your facts straight. Surely nobody would place a league-average utilityman over an 11-WAR slugger on an MVP ballot in 2024. (With that said, nothing would make me happier than finding out, in a week’s time, that someone voted for Matt Vierling over Judge for AL MVP. I will throw an actual party if that happens, and you’re all invited.)
So is it that voters no longer have the courage of their convictions, or is it just that they have better convictions these days?
Michael is a writer at FanGraphs. Previously, he was a staff writer at The Ringer and D1Baseball, and his work has appeared at Grantland, Baseball Prospectus, The Atlantic, ESPN.com, and various ill-remembered Phillies blogs. Follow him on Twitter, if you must, @MichaelBaumann.
I will defend the Yelich over deGrom vote. Was deGrom the better player that season? Yes. Was he the more valuable player? I disagree. The Mets had a better winning percentage when deGrom didn’t pitch than when he did. That’s the furthest thing from deGrom’s fault, and it’s not fair. But I don’t think the “value” of the “Most Valuable Player” should be completely hypothetical as it would be with 2018 deGrom.
So tour argument here is “pitcher wins are an important statistic”? That’s…weird, especially in an era where bullpen usage is higher than bygone eras, so starters have less direct control of the 9 definite innings.
I think pitcher wins are the dumbest stat alive, I just think it’s understandable to say a guy whose team lost most of the games he played in wasn’t the most valuable player in the league, even if he was the “best”.
The guy whose team famously lost a lot of games where he gave up 1 run over 7 innings is not an argument against his lack of value, its just an indictment of the rest of his team. And that was the story of 2018 Jacob DeGrom.
In your world, I guess, DeGrom would only have been valuable if he’d matched Orel Hershiser’s consecutive scoreless innings streak, because that’s what he would have had to do for the Mets to win more of the games he pitched – they simply scored 0, 1 or 2 runs in a huge percentage of his games, and won some of them anyway because he was so dominant.
That’s not a fault of Jacob DeGrom.
The games were boring and the bats fell asleep.
I am of the opinion that MVP basically does mean “best.” That’s how I would vote, anyway.
How many of those games would they have lost with someone else pitching?
“Alright cashgod27 thanks for the call. Next up on WFAN is Jimmy from Staten Island, who thinks Juan Soto walks too much”
Say less Poindexter. Go upstairs and make your mom dinner. Have her clean your pocket protector. Talk to a girl. Are.you not old enough to be familiar with the standard response to your ain’t been clever since 1999 insults?
?
Sometimes I wonder how people like you even end up as FG members to begin with. This is one of the most pro-nerd spaces on the internet and yet you pay to be a member here and rail against nerds.
Thanks for funding the site man, please don’t feel the need to comment while you’re here.
I understand why people are downvoting this, but it is really fascinating.
And no, this is not about pitcher wins.
Mets won 44% of deGrom starts in 2018 when he averaged 6.8 innings with 1.99 RA9 and 1.70 ERA.
They also won 48% of non-deGrom starts in 2018 when the starting pitchers averaged 5.4 innings with 4.37 RA9 and 4.12 ERA.
Just looking at the numbers, it would have been surprising if the Mets had “only” 4% higher win probability in deGrom starts, but apparently, 32 games is still a small enough sample size for these kinds of wacky results to happen.
This isnt a surprise, players and teams routinely put up fluky results over whole seasons.
It was probably a fluke, but it’s super weird. In 2016, 2018, and 2019, DeGrom was well below the Mets’ overall run support.
2016
Mets: 4.2, DeGrom 3.5
2017
Mets: 4.6 DeGrom 5.1
2018
Mets: 4.2, DeGrom 3.5
2019
Mets 4.9, DeGrom 4.1
The years before and after, he was average or even above average in terms of run support, but it almost seems like something happened those years that made his team less likely to score runs for him.
he faced the ace of the other team most if the time.
In his first few starts, perhaps. But after a short while most teams will have had a different number of games played due to off days, rain outs, etc. And I don’t think teams will then try to line up their ace against the opposition’s ace in the regular season.
How many games did they win with him starting? Him being the best pitcher on the planet would mean that their record would have been worse had any other pitcher been in the mound on those days he pitched. Any pitcher not in DeGroms level would theoretically mean even fewer wins, more stress on the bullpen for subsequent games, etc. hence, value.
Most people replying to this are not actually considering the argument. The argument is, as I understand it, that the Mets could have replaced Degrom with a completely average starting pitcher and their record would not have changed very much. There is nothing Degrom could do to change that, but it does mean his performance is less meaningful from a W-L perspective.
Put another, weirder, way (couple glasses of wine deep at this point): if I have two bottles of wine, and one is better than the other but my fiancée pours out half of the better bottle, that less good bottle is now more valuable to me.
It’s an argument you can disagree with, but it is way more sound than commenters here are making it out to be. Prof. Cashgod (I assume it’s a teacher’s surname) isn’t saying Degrom “wasn’t valuable”, they are saying his value gets diminished when his team sucks.
Awards demonstrate what the electorate values at the time, and that is their value as historical artifacts. Knowing that Juan Gone won the 98 MVP over A-Rod tells us something about the baseball culture of the time.
While that’s true (and several guys would have been better picks in 1998 than Juan Gonzalez who had a lower WAR than the next 11 guys in the MVP voting), Arod was somewhat hamstrung by the fact that Jeter and Garciaparra also had awesome (and better than Juan Gonzalez) seasons in 1998 on playoff teams:
ARod — Mariners record 76-85
.310 .360 .560 .919 – 213Hits 35Doubles 5Triples 42HR 124RBI 46SB 123Runs 8.5WAR
Garciaparra — Red Sox record 92-70
.323 .362 .584 .946 – 195Hits 37Doubles 8Triples 35HR 122RBI 12SB 111Runs 7.1WAR
Jeter — Yankees record 114-48
.324.384.481.864 – 203Hits 25Doubles 8Triples 19HR 84RBI 30SB 127Runs 7.5WAR
I do think it’s worth noting that WARP is an insane outlier in terms of Position Player WAR, because they regress BABIP on a season-by-season basis without accounting for contact quality or the fact that BABIP stabilizes after 800 – their offensive metric claims that Gleyber Torres has been a comfortably better hitter than Derek Jeter was, on his career.
It’s basically “What if we based WAR on xwOBA, and used a far worse version to do it?”
Agree on this one. I don’t know that anyone considers WARP as a serious alternative to fWAR or bWAR.
My opinion is that if a player has a 1+ WAR lead in both flavors, you’d need an unusually compelling argument for someone else.
To me this isn’t all that different from basing WAR on FIP or any such estimator. Some pitchers get highly rated by FIP like Nolan Ryan, and others by RA9 like Phil Niekro.
Exactly. One version of the stat is meant to measure what has happened and the other is meant to be more predictive. They serve different purposes and that’s great!
I prefer RA9 based WAR for Cy voting, and FIP based for future season planning
I mean…I agree that FIP sucks, but Tango hasn’t published the methods he used for the Strider/Snell article last year across all players yet.
Of course the easiest answer is that “Best” and “Most Valuable” are not synonyms, further that “Value” is a subjective measure. If there has been or ever will be any real convergence in voting behavior it will likely have far more to do with settling those arguments than any particular statistical metric.
Can we also acknowledge that it’s weird that we argue about the meaning of the word “valuable” almost as much as “a well regulated militia?” I doubt the initial creators of the term “Most Valuable Player” intended it to carry so much ambiguity. I also doubt they spent a lot of agonizing over where to call it “most valuable player” or “most outstanding performer” or “player of the year,” or whatever ever other iteration you could think of. At least, I’m sure they didn’t spend a tenth of time on it that we all have over the years.
It’s the best player. Let’s quit hanging up on “valuable” to try to arrive at someone other than the best player
But that’s not how voters have always behaved historically, especially back in the early days of the award. 1941, 1947 say hello. I don’t have the time right now, but I bet with some research I could turn up newspaper articles from the first 20 years of the award with some versions of most valuable vs. best player debates.
I know its not the takeaway from this article but man, Grady Sizemore was really really good. Its such a shame he was ruined by injuries. He was like the pre-Trout, Trout.
(and its a shame my Expos traded him because MLB sucks and hated them)
I would say that WAR has made MVP voting better. The last player to win an MVP with a WAR below 5.0 was Justin Morneau in 2006. And hard to imagine it ever happening again.
It used to be much more common: Juan Gonzalez in 1996 and 1998, Mo Vaughn in 1995, Eckersley in 1992, Dawson in 1987, etc.
(I’m using Baseball Reference WAR here since they have an easy to read chart: https://www.baseball-reference.com/awards/mvp.shtml).
Probably the biggest thing WAR has done is:
– made people realize that high leverage relievers that only throw 60-70 innings are not that valuable.Just not enough batters faced unless they did something really insane, like a ran 50% K rate or something..even then, not sure it’s enough.
If starting pitcher usage decreases any more, I think there could actually be rare cases where a reliever would be good enough on a per inning basis to beat out the best starter in WAR.
Ehhh, that seems pretty unlikely. There is certainly a limit on how many innings a reliever can throw while still being good. The whole reason they are a reliever is because they cant be as effective over a large amount of innings. So at a certain point the more they pitch the less effective they become. The idea that reliever innings are going to keep going up and up with no drop in production is pretty unrealistic IMO. Once they get past the break even point they are just not going to be effective and stop accumulating WAR (or start outright losing WAR)
The ceiling on reliever innings is probably low enough that it would be nearly impossible to rack up 5+ WAR or whatever it would take to be above the top starter.
I also do think if the streams ever did cross, you’d probably need to adjust FIP-based WAR with a common sense/sanity check. Like it simply is more valuable to throw 180 innings while allowing 2.5 runs per 9 than to throw 80 innings while allowing 2.5 runs per 9 even if the latter guy did it with insane strikeout numbers. I guess this gets into a larger issue, which is that I think end-of-season pitcher awards should be primarily based on runs allowed, for the same reason that xWOBA isn’t a major component of MVP voting.
I get the idea behind your thinking in using RA9-WAR for awards purposes, but I question whether that’s a better metric for what you’re trying to measure. It seems like you want to credit results over process with awards voting, which is fine, but the results-based metric here is so dependent on outside factors (fielding, HR/FB rate, etc) that you will inevitably end up overrating worse pitchers.
I could see it, but I think we’d be talking about big changes, like starters workloads decreasing a lot more and then bullpens having more long-ish guys. So maybe some five and dive types become 3-4 inning relievers. Like if the top pitchers maxed at around 150 innings and there were some bullpen guys over 100. But we’re not even close to that yet.
It was partially a tongue-in-cheek comment, but if starter workloads get down to only 140 IP, a reliever *only* needs to be over twice as good in 70+ innings to beat them. There have been a handful of 5+ WAR relief seasons. Combine that will just a crappy year for starters, and it becomes mathematically possible if still very rare. I agree that it’s impossible with current usage unless we get a return to Sutter/Gossage levels of relief ace usage.
In my lifetime (1981) there hasnt been a single RP who accumulated 5 WAR. Only 5 have even broken the 4 WAR mark and none have done it in the past 20 years.
1986 Eichhorn (4.9)
2003 Gagne (4.7)
1990 Dibble (4.3)
1996 Rivera (4.3)
1991 Ward (4.1)
Of those 5 only Gagne did it in less than 100 IP. Eichhorn pitched an insane 157 innings in 1986 without making a single start. If you are curious his 4.9 WAR would have been 8th overall. Mike Scott led pitchers in 1986 with 8.6 WAR. In fact all 5 guys only had around half as many WAR as the top starter, so its not even close.
I think we are more likely to just see innings spread over more relivers than relievers throwing significantly more innings.
I love when we pose an honest question or good topic for additional research and a reader goes and shows their work. It’s one of my favorite things about this site.
The highest reliever WAR since 1994 is Eric Gagne with 4.7 in 2003, when he won the Cy Young. He was 14th in WAR that year, and not particularly close to Mark Prior’s 7.8.
But it’s not super far behind the leader this year, Chris Sale, with 6.4. Last year was Zack Wheeler with 5.9.
If we assume Gagne’s is the best year possible for modern relievers (until proven otherwise), I could imagine a scenario where the league leader among starting pitchers is 5, if they are all throwing 160 or fewer innings.
Eh?
15 years of teams having WAR as a framework have pretty clearly demonstrated that public WAR models aren’t good at valuating high leverage relievers.
Agree.
Regarding pitchers winning MVPs, I think there are two reasons why they don’t get votes. First, and most important in my mind: pitchers have the Cy Young. While there are (and should be) exceptions when the offensive pool is particularly weak or an individual pitcher’s performance is particularly dominant, the MVP is generally thought of as a hitter’s award that parallels the pitcher’s award (since the advent of the Cy Young in 1956, only 12 MVPs have gone to pitchers out of 176 total MVP awards). Generally, people expect to award both the best pitcher and the best position player in each league. 2018 deGrom probably deserved more consideration than he got, but, from the perspective of many voters and fans, he already had his award.
Second, and I suspect less important to most analytically-minded fans and voters (although still relevant to many others): pitchers simply don’t play as much. There’s something that feels counterintuitive about a player who only plays roughly one out of every five games being more valuable than someone who plays nearly every day. Because I am writing a comment on Fangraphs, I like analytics and know that starting pitchers generally influence more plate appearances than any position player (Logan Webb led MLB pitchers in 2024 with 841 batters faced while Jared Durran led in plate appearances with 735) and therefore have as many or more opportunities to create value for their teams. But, if you aren’t an analytics person (or, worse, are skeptical of analytics by default), that counterintuitive feeling is going to make it difficult to see a pitcher as more valuable than a position player without extenuating circumstances (particularly weak offensive pool or particularly dominant pitcher).
For the first reason (and expressly not the second reason), I actually like that pitchers rarely win MVP awards. If they regularly won MVPs, I suspect that the Cy Young award would lose much of its luster since it would no longer be the main prize for being a dominant pitcher.
Same. Re: “MVP voters hate pitchers for reasons that are not entirely clear to me,” the primary reason is the Cy Young, which could easily be renamed Most Valuable Pitcher. Voters tend to restrict MVP to hitters, w/few notable exceptions to celebrate a pitcher’s dominance, like when Kershaw, Verlander, Eck, & Clemens won MVPs.
Separating the hitters and pitchers when assessing value seems to be one thing that possibly transcends eras
The drop in innings pitched certainly plays a part, I think. Even though Webb faced more batters then any hitter had plate appearances, it is nothing compared to the days when pitchers threw 250+ or even 300 innings. Those guys faced 1000-1200 hitters. Easier to make the argument for a pitcher with that amount of volume.
The other thing is the realization that the fielders do a lot of work. Even with 250 innings, of those 750 outs the pitcher gets, how many are attributable to them? How many to the D? Reality is that with analytics that mostly comes down to K’s or K-BB% (which leads to FIP, xFIP, etc) & you’d need a guy with high innings/high K’s to make the case for the pitcher to get it over a hitter.
It is a little strange that the “pitchers should be considered on an equal basis for MVP” has become identified as a stats/sabermetrics/Fangraphs position, when really it has nothing to do with any of that and more to do with people’s preferences for variety in their award winners. Everyone acknowledges and understands that position players and pitchers are doing really different things; the vast majority of articles on Fangraphs (and other analytically-inclined outlets) compare hitters to hitters and pitchers to pitchers for good and obvious reasons. I’m not sure why analytics writers insist on treating this debate as if it’s a misunderstanding about the value pitchers can provide rather than a disagreement about how that should be recognized by an award with a specific history and meaning within baseball culture.
While this perception of once every 5 days versus plays every day is correct in the minds of people, the numbers suggest that a good pitcher faces about 750 batters a season(a couple get into the 800’s every season and guys like maddox and randy johnson back in the day got up to over 1,000 batters faced), which is more than the league leader in PA’s this season(735).
While true, shouldn’t you also give hitters credit with the plays they make in the field?
I know the electorate is different, but I’m still slightly surprised that the Hank Aaron Award hasn’t been viewed as the position player Cy Young equivalent by anybody. I know it’s just an award for hitting without fielding and baserunning, but hitting is the majority of a position player’s value.
Since 2012, 18 of the 24 Aaron Award winners were also the MVP. One of those mismatched instances was because Kershaw won MVP. The Aaron Award and MVP are essentially synonymous at this point if MVP is position-player only. There’s room for the MVP to include all players.
I can see an argument for WPA in MVP discussions or even CPA or whatever it’s called. There’s already a movement to use context in justifying reliever HOF candidacies
I was waiting for something like that to be mentioned in the article.
Like it or not, value COULD mean- Who contributed the most to a winning team? Or who had clutch moments that won a # of games for a winning team?
Fair? Probably not, you’re rewarding someone for having good teammates to a degree..but, the flip side is you could say that 9 WAR on a 70 win team doesn’t bring that much value cuz the team still stunk. I think that did play a part in the 2013 Cabrera-Trout MVP race, for example.
No “correct” answer here because “value” is a subjective measure.
I think the problem people have with that WPA approach is that it doesn’t really correlate with skill. As you correctly point out, its almost as much about who else is on your team. Or just simply how many opportunities you had and/or how lucky you got. In theory someone could have no high leverage PA’s in a season. Which would render him irrelevant to the MVP award no matter what else he did.
Do we want to hand out the MVP to a guy who’s arguably not even the best player on his team and just found himself in important situations alot?
I mean, you could apply the same logic to closers and the Cy Young. They are always going to have higher WPA than a starter because they basically only pitch high leverage innings. Clase had an insane 6.2 WPA in 2024. six times higher than Bibee. All things being equal I doubt anyone is keeping Clase over Bibee though if push comes to shove because we inherently understand the value of SP is more.
RE24 definitely deserves more consideration as a way to determine a player’s offensive value. How players performed in context should enter more into the MVP discussion, even if RISP splits aren’t particularly sticky season-to-season.
WPA is certainly useful as a statistic, especially for relievers, but for hitters it’s very dependent on late-inning performance. Hitting a 3-run HR in a game won by 1 run is just as valuable whether it happens in the 3rd inning or 10th inning, but WPA is going to spit out drastically different numbers.
WPA is just “What if RE24, but useless” for relievers, too.
WPA is just RE24 for gamblers who don’t like baseball.
The P in WARP is for ‘player’ (Wins Above Replacement Player)
Bill Pecota is sad
I suppose the Fangraphs and BBRef versions leave space for replacing someone with some sort of snack food or objet d’art.
I’m not sure if that was the worst attempt at a joke ever written in a baseball article or if Baumann simply didn’t bother looking WARP up either in the BP or Baseball Reference stats references.
Im not sure which would be worse.
I think the Justin Turner line was meant to be glib, but, and I don’t want to sound like a caveman, so please don’t take it this way, I just get the feeling that in 20 years people will be looking at articles like this and laughing about how wrong the consensus was about value (WAR, or whatever it is called then). This isn’t to suggest that we should revert to worse measures like RBI or anything. I think using WAR as a strong starting point is a good idea for voters. We don’t know what we don’t know, and all that. But I just would caution that it is bordering on certainty that we don’t have it all figured out, and honesty it’s possible we aren’t even close. These two articles could age poorly.
I appreciate your overall point, but I actually think the writer did a good job establishing up front that WAR cannot be viewed as gospel or there wouldn’t be three different versions. I never got the impression he was suggesting every one of the non-top-WAR MVPs was the “wrong” choice, per se. Rather, he was trying to determine how swayed by WAR present-day voters have become.
As for analytically-minded fans likely underestimating the error bars on our current WAR formulas? I agree with you there. For every time there’s a disclaimer that WAR should not be viewed as being a “perfectly precise measure of total value,” the proceeding prose/analysis treats it as just that.
Of course, world’s smallest violin playing for that quibble. I’m darn grateful that smart folks took the trouble to develop measures of value that go beyond OPS or ERA+.
Me, too, on that last point.
I also think the first article was maybe a little more strident in conveying “now we’ve figured it all out.”
This isn’t so hard to figure out…the majority of voters see MVP as the position players’ award and the Cy Young as the pitchers’ award. And they’re right to do so, IMO. I realize that formally it isn’t, but historically it is, and that’s fine. It would also be fine with me if the powers that be decided to change the award description to reflect this practice, although I do not want to see an acronym change to MVPP.
Minor side quibble, but sure we can. It’s a scale of measurement, based on the run value of anything a player does, that allows us to unify ways of measuring all the facets of the game. The “flavors of WAR” confusion derives from there being many different ways of trying to measure these different facets, not from the idea of combining them — and this is actually fine, even a good thing. Maybe there should be many more different “flavors” and people should have their own individual ones. Plug in your own catcher defense stats, or have your own idea about how defense-independent pitching stats should work, or whatever it is. WAR is just a way of commensurating all these different arguments, not a claim about which one of them is right.
First of all, I love this analysis. I think it has real value, and also helps to rebut and debunk a lot of the whiners out there bitching about WAR. Clearly vast majority of voters have never just lined up WAR and pushed send.
I have a fairly simple process, just one man’s opinion and thought process. It’s not static, and I’m open to suggestions.
I start by averaging the WARS together. It just helps smooth out the different things being measured, and almost always passes my personal smell test when I’m done doing that. Whether it’s allowing for often vast differences in fielding metrics, or finding balance between actual results and deserved results….i.e. a guy with 2.50 ERA but 4.00 FIP, etc. the averaging just works for me.
Then I create TIERS, typically separated by about 2 WAR. Any kind of deep, objective, analytic dive is almost never going to bridge the gap between say a 6 and 8 WAR player. But it’s possible, even reasonable in some cases, to bridge a gap between of 1 to 1.5 WAR, for example.
So within the tiers I start to look at the context of that performance. It’s the MVP award, not the projection for next year award, or the completely context neutral award.
I look at Win Probability stats, whether it be straight WPA, or RE-24, or baseball references versions of these stats as well as their newer cWPA, etc.
I also look at the context of the player’s team, and how much that player meant to his team’s success, and what his production represented as a portion of his team’s success. (i.e. percentage of a team’s WAR or WPA, etc)
Sometimes that context is enough to bridge gaps, sometimes it’s not. The way I weight those things is not going to allow a 6 or 7 WAR player on a playoff team surpass a 10 WAR player on a non playoff team. But if a guy with 7 WAR has MUCH better win probability numbers than an 8 WAR player, he can pass him in some cases.
Of course it’s subjective to apply this way. But I refuse to ignore context completely when considering this type of award. Narrative, the story of each season matter too. And those stats help tell the story.
At the same time, it’s the BBWAA’s award. As insufficient as it may feel, these are the criteria and guidelines they offer. So I consider them too.
Dear Voter:
There is no clear-cut definition of what Most Valuable means. It is up to the individual voter to decide who was the Most Valuable Player in each league to his team. The MVP need not come from a division winner or other playoff qualifier.
The rules of the voting remain the same as they were written on the first ballot in 1931:
1. Actual value of a player to his team, that is, strength of offense and defense.
2. Number of games played.
3. General character, disposition, loyalty and effort.
4. Former winners are eligible.
5. Members of the committee may vote for more than one member of a team.
You are also urged to give serious consideration to all your selections, from 1 to 10. A 10th-place vote can influence the outcome of an election. You must fill in all 10 places on your ballot. Only regular-season performances are to be taken into consideration.
Keep in mind that all players are eligible for MVP, including pitchers and designated hitters.
WARP actually stands for Wins Above Replacement Player not Pecota. Here is there blurb about it where they say that a team full of replacement level players would be expected to win approximately 50 game on average by their definition of Replacement Player. Proceeds to glare menacingly at Jerry Reinsdorf.
https://legacy.baseballprospectus.com/glossary/index.php?search=WARP
I’m pretty sure it’s “Wins Above Ryan Pepiot” but some sites use “Wins Above Ryan Pressly” or even “Wins Above Rick Porcello”
Hot taek:
My feeling is that Pitchers should be explicitly excluded from MVP voting. They have their own award. And even the CY should be restricted to only qualifying pitchers. Relief Pitchers have their own award. SP affect at the most 35 games. RP maybe 75-80 max. MVP should be reserved for everyday players only.
Hitters have their own award, the Hank Aaron award. Fielders get Gold Gloves.
So, that leaves… pinch runners for the MVP?
This is a common, but misguided, argument. No position player has even remotely the same impact that a starting pitcher has on games. A starter may only play 30ish games but they basically are the direct reason their team wins or loses those 30ish games. Thats true of literally every single starting pitcher. On days they play, they are the main factor in the games outcome. You cannot say the same about literally any hitter. If a batter has a bad day, it often really doesnt mean much. Hell, the very best players fail to do anything around 65% of the time. If a starting pitcher let 65% of batters on base he would be out of a job in a month.
The avg SP faces ~750 batters a year and they are all concentrated into 30-33 games. The avg batter gets less than 700 PA’s and they are spread out over 150+ games. On a per game played basis, SP are insanely more valuable than any hitter. Like, you dont need to look any farther than WAR to see that. Chris Sale contributed more wins in 29 games than any Braves hitter did in 162 games. And it wasnt even close (Sale had 6.2 WAR, Ozuna is the highest position player at 4.7). If you want to go even a step further, Sale and Fried combined for 10 WAR.
So the “they only play 35 games” argument really holds no water when you actually look at it objectively.
A while ago I looked at the top most jobbed MVP votes, in terms of winners’ BRef WARs as a percentage of top BRef WAR. Here were the twenty most egregious offenders:
1979 NL – Willie Stargell – 2.5, Dave Winfield – 8.3
1996 AL – Juan Gonzalez – 3.8, Ken Griffey Jr. – 9.7
1992 AL – Dennis Eckersley – 2.9, Roger Clemens – 8.7
1979 AL – Don Baylor – 3.7, Fred Lynn – 8.9
1934 AL – Mickey Cochrane – 4.5, Lou Gehrig – 10.1
1974 NL – Steve Garvey – 4.4, Mike Schmidt – 9.8
1987 NL – Andre Dawson – 4.0, Tony Gwynn – 8.6
1955 AL – Yogi Berra – 4.5, Mickey Mantle – 9.5
1984 AL – Willie Hernandez – 4.8, Cal Ripken Jr. – 10.0
1974 AL – Jeff Burroughs – 3.6, Rod Carew – 7.5
1947 AL – Joe DiMaggio – 4.7, Ted Williams – 9.5
1935 NL – Gabby Hartnett – 4.9, Arky Vaughan – 9.8
1944 NL – Marty Marion – 4.6, Stan Musial – 8.9
1995 AL – Mo Vaughn – 4.3, John Valentin – 8.3
1950 NL – Jim Konstanty – 4.4, Eddie Stanky – 8.2
1970 AL – Boog Powell – 5.1, Carl Yastrzemski – 9.5
1938 NL – Ernie Lombardi – 4.8, Mel Ott – 8.9
1964 NL – Ken Boyer – 6.1, Willie Mays – 11.0
1958 AL – Jackie Jensen – 4.9, Mickey Mantle – 8.7
1962 NL – Maury Wills – 6.0, Willie Mays – 10.5
The five worst since 2000:
2004 AL – Vladimir Guerrero – 5.6, Ichiro Suzuki – 9.2
2006 AL – Justin Morneau – 4.3, Johan Santana – 7.6
2006 NL – Ryan Howard – 5.2, Carlos Beltran – 8.2
2002 AL – Miguel Tejada – 5.7, Alex Rodríguez – 8.8
2012 AL – Miguel Cabrera – 7.1, Mike Trout – 10.5
Three major patterns emerge:
Also, I forgot how good John Valentin was for a hot minute in the mid 90’s.
> We can’t just keep giving this award to Mays and Mantle, right?
Lol.
There seem to be other “Hey let’s give it to a non-HoF-bound guy who was surprisingly top excellent over a HoF-bound guy who is consistently top excellent (and was better this year)” votes too.
Thanks for this great article-well-written, informative, and interesting. If this were my first time at this site, I would subscribe immediately upon finishing the article.