Putting WAR in Context: A Response to Bill James
Nine years ago next month, we introduced a new stat to the pages of FanGraphs. We called it Win Values, and on the player pages and leaderboards, it went by the acronym WAR. We wouldn’t actually start calling it that, or use the words for which the acronym stood (Wins Above Replacement) for a little while, since we thought Win Values sounded cooler. And as the people who bring you WPA/LI and RE24, we’re clearly the experts on statistical naming coolness.
Over the last nine years, WAR has become something of a flagship metric, not just for us, but for the analytical community at large. Baseball-Reference introduced their own version, while Baseball Prospectus modernized their version of WARP — their version adds the word player to the name, thus the P — to provide something that scaled a bit more like what was presented here and at B-R. Because WAR is a framework for combining a number of different metrics into a single-value stat, there are also quite a few other versions of WAR out there, each with their own calculations.
But while everyone uses different inputs — and therefore arrives at slightly different results — almost all of the regularly updated WAR metrics are built on some version of linear weights, which assigns an average run value to each event in which a player is involved, regardless of what actually happened on the play. If you hit a single, you get credit for hitting a single. It’s worth some fraction of a run, regardless of whether you hit it with two outs and the bases empty in a the first inning of an eventual blowout, or whether it was a walk-off two-run single to give your team the lead. In most versions of WAR, the value of a player’s contribution is calculated independent of the situation in which it occurred.
Bill James is not a fan of that decision.
We come, then, to the present moment, at which some of my friends and colleagues wish to argue that Aaron Judge is basically even with Jose Altuve, and might reasonably have been the Most Valuable Player. It’s nonsense. Aaron Judge was nowhere near as valuable as Jose Altuve. Why? Because he didn’t do nearly as much to win games for his team as Altuve did. It is NOT close. The belief that it is close is fueled by bad statistical analysis—not as bad as the 1974 statistical analysis, I grant, but flawed nonetheless. It is based essentially on a misleading statistic, which is WAR. Baseball-Reference WAR shows the little guy at 8.3, and the big guy at 8.1. But in reality, they are nowhere near that close. I am not saying that WAR is a bad statistic or a useless statistic, but it is not a perfect statistic, and in this particular case it is just dead wrong. It is dead wrong because the creators of that statistic have severed the connection between performance statistics and wins, thus undermining their analysis.
James strongly believes that the metric falls apart by building up from runs, rather than working backwards from wins, since the context-neutral nature of the metric means that what WAR estimates a group of players are worth won’t add up to how many wins their team actually won. In his mind, the decision to make WAR context-neutral isn’t a point on which reasonable people can disagree; it’s just a mistake.
James, of course, is a pioneer in this field, and FanGraphs probably doesn’t exist today if not for the work he did in the 1970s and 1980s, laying the groundwork for most everything that has come since. So when a person with his resume suggests that WAR isn’t just imperfect — which it absolutely is, no matter which version you use — but instead wrongly constructed, I think it’s worth responding to. So let me take a few minutes to talk about the relationship between context and value, and how that informs people’s preferences for building out WAR in this way.
My primary guiding principle on the usefulness of a metric is twofold:
1. What question is it answering, and is the answer to that question interesting? If yes, proceed. If no, ignore.
2. Does the metric answer that question accurately?
There are plenty of statistics that keep an accurate count of things that happened. In many cases, however, these numbers answer only trivia questions. Who got the most hits on Tuesdays in 2017? There’s an accurate and measurable answer to that question, but I have no idea what it is, because it doesn’t matter in any tangible way.
WAR, on the other hand, attempts to address a question that a lot of people seem interested in answering. If the WAR leaderboards were posed as a question, they might be written as something like this:
“What did each player do, as an individual, to help his team try to win games?”
Wins are a team accomplishment, an amalgamation of performances from a large number of different players. The way WAR has generally been constructed means that it attempts to isolate the player’s contributions towards winning games. And when it comes down to assigning value to individuals for the events in which they’re involved, the general consensus in the sabermetric community has been that we want to reward (or penalize) hitters for what they can control. And the context of the situations in which they play is just not something players can create.
To be clear, this decision wasn’t made solely when WAR models started getting calculated online for people to track. Pretty much every single baseball statistic that is used on a daily basis, regardless of how analytically inclined the user is, is designed to be context-neutral.
Batting average, on-base percentage, slugging percentage, home runs, stolen bases, walks, and strikeouts: all of it is counted without regard to the number of runners on base, the inning, or the score. Outside of runs and RBIs, pretty much every mainstream measure of individual hitter evaluation has been designed to ignore the context of the situation in which it occurred. And runs and RBIs only consider baserunner situation, not number of outs, the score, or the inning in which they occurred. On the pitching side of things, ERA includes baserunner performance, but not inning or score.
There are context-dependent metrics, of course, and we’ve tried to do our part to promote Win Probability and Leverage Index as tools to be included in the discussion of value. We have a number of stats that measure different levels of context-specific performance, ranging from RE24 (baserunner/out context included, inning/score excluded) up to WPA, which includes everything. We keep a statistic called Clutch to track the difference in win values between a player’s context-neutral performance and his context-specific performance.
So why don’t we build WAR off of one of these numbers instead of a linear-weights-based method? Would WAR be better if it included the context of the events, rather than just an estimate of the player’s contribution to the result based on historical averages?
I think the answer is that it depends on how you’re using WAR. In the case of MVP voting, I do think there is a case to be made for looking at the circumstances under which a player performed, and I did use context-dependent metrics when I was an MVP voter. WAR is an imperfect tool, and it’s particularly imperfect for things like the MVP award, which is why even those of us who host sites that promote WAR fairly extensively suggest not relying solely on its results when filling out a ballot.
But if one wanted to build a complete and thorough version of WAR that tied back perfectly to a team’s win total, the reality is that it would not be particularly useful for answering many other questions. Because assigning an individual player with the true contextual value of his performance requires far more adjustments than the simple proposed fix in James’s article.
For instance, one of the main reasons we use BaseRuns instead of linear weights as the underpinnings of our expected won-loss record calculations is that run scoring isn’t linear. As you stack more and more good hitters together, the increase in run production will go up faster than a simple addition of run values would suggest. So, if you wanted to do a true valuation of a player’s performance, you’d have to account for his teammates’ performances, as well, and how the combination of the two translated into runs scored.
But if you do that, you’re explicitly giving a player win-value credit for having better teammates. Jose Altuve is great on his own, but what do we do with a metric that says he’s even better because he has Carlos Correa hitting behind him? That metric is no longer measuring each player’s performance but the specific value he created in that specific lineup, based on what other players did in the aggregate.
You can drill down even further if you want to take the context argument to its logical conclusion. If we want to reduce Aaron Judge’s WAR by the amount of wins he cost his team by performing poorly with men on base, do we want to also adjust his WAR (and everyone else’s) by how his teammates did after he reached base? If we’re primarily interested in tying performance back to team wins, so that a bases-loaded hit is worth more than a bases-empty hit, the same logic would suggest that a base hit in front of a home run is worth more than a base hit in front of a double play. In either case, the player wasn’t responsible for what came before or after his individual contribution, but the results were wildly different, and a hit in front of a double play doesn’t contribute to a win any more than a ground out would have.
And that’s just baserunner/out context. Score and inning context is even thornier ground, because no one really wants to conclude that a home run with your team down 10-0 is worthless. If an individual player metric cannot reward an individual player for his own performance because his teammates were so bad that the game was already out of hand, then it is no longer an individual performance metric.
We’ve written a lot of articles over the years about the pros and cons of context-neutral and context-dependent metrics, and offer a range of stats that attempt to measure value in nearly every phase of the context scale.
But WAR became popular as a metric largely because it attempts to isolate just a player’s individual contribution to a team’s wins and losses, and once you start adding in some context, there’s no real reason to stop until you’re at WPA, which tells you that Melky Cabrera‘s 98 wRC+ resulted in a better offensive season than Jose Ramirez’s 148 wRC+. And even then, WPA ignores the after-he-hit events, so even that isn’t really telling you the full story about the value of Melky’s offensive inputs as they relate to wins.
Once you adjust for the full context of a player’s input into wins and losses, you’re left with a version of WAR that is so far removed from his own contributions that I don’t know what question it would answer. And if you just include some context but not others, then you have to justify only going part way? For instance, James’s Pythagorean expected record was used for decades as a shorthand for “team luck,” because it stripped out the sequencing that turned runs into wins. But it didn’t do anything to strip out the sequencing that turned individual events into runs, so it only measured part of a team’s sequencing value. Why? Because any single metric can’t measure all things at all times.
Like every other metric in existence, WAR measures some things well and other things poorly. It is useful to answer some questions, but not all questions.
For the MVP voting, perhaps WAR is less useful than James would like it to be. On that point, I agree, and I used other metrics when filling out my MVP ballot when I was assigned to be a voter. But for many cases, the questions people are attempting to answer by using WAR are better answered by a context-neutral metric. It might not answer those questions perfectly, but it at least aligns with the questions about which fans are curious.
Is there a place for a context-dependent version of WAR? Perhaps. But then again, we’re already accused of undermining the model’s credibility by having multiple popular methods of calculation. And as James himself found when developing Win Shares, tying individual player performance to team wins isn’t quite as easy as one might hope.
I have no problem admitting that WAR as a model contains a number of flaws, or that our specific implementation of the framework is also flawed. There are a lot of areas for improvement. Forcing it to account precisely for the exact number of wins with which each team finished, though, would probably make it less useful overall as a measure of individual performance.
Dave is the Managing Editor of FanGraphs.
Is it impolite to acknowledge that if he hadn’t been a pioneer in the field — which is important; popularizers are important! — nobody would care at all what James has to say these days? On merit, his ideas aren’t really interesting anymore. The world passed him by a while ago.
Actually Tom tango (who definitely isn’t a dinosauer) has sided with bill on this one. it isnt about right or wrong, it is about what you want to measure: skill or result
Bill is not senile, he understands that basic sabermetric stuff very well, his argument is that clutch affects real wins on the field, so why not factor it in.
now there is a problem with that that clutch isn’t really repeatable year to year and could be considered random and it is also affected by team context (albeit you could easily normalize for that).
however we also use other random factors. many still use ERA with pitchers and even FIP includes HR/FB luck.
and many consider wRC+ objective but it is affected big time by luck (babip, hr/fb…). even war is still based on those luck influenced outcomes, chris taylor wouldnt have had 5 war without outperforming his xWOBA by 30 points.
bills argument is very valid: if we use other context dependent stats (facing billy hamilton or mccutchen in center or head or tailwind…) why not go full results based?
I don’t agree with bill here however but I also don’t think that wRC+ or OBP are objective stats, in fact I would prefer to go full xSTats to judge players.
I actually submitted an article on this in the community section last night:).
Whether his argument is right isn’t really my point, though.
It kind of seems like it is though when you say that his ideas no longer have any merit.
When’s the last time he wrote something original that changed how you thought about baseball? Performance vs value certainly isn’t a new thought — we’ve been kicking around these ideas for years.
I dunno, the answer to my original “is it impolite” question is, apparently, still yes. So I’ll be over here in my corner, quietly mystified about why we should really care what he thinks about anything anymore.
James has been working for the Red Sox since 2003, so my guess is most of his original writings are proprietary.
But the answer to your original question is partly in the question, that he is a pioneer in the field. He has probably forgotten more about baseball and statistical analysis than most people will ever know, plus he has a ton of credibility as a pioneer and current employee of a baseball team. Bill James has more than enough credibility that we should care about what he thinks, as opposed to some random cynical internet poster who goes by a mythological Greek creature.
I don’t think you have a point – just a progressive narrative.
To dismiss someone’s argument by saying he doesn’t merit an opinion is the cheapest way of attacking that person’s point without addressing it.
Bill’s argument is correct but not imo because of ‘clutch’ or other batting-context stuff. It is important because position players have two very different functions – batting/baserunning (scoring) and fielding/defense (preventing scoring). As an aside – the former is also mostly individual and the latter is very much ‘team’ as well.
The context of the game dramatically changes the relative importance of one function vs the other. Sometimes offense is important, sometimes defense is important – and context may also change that relative weighting by position during the game too.
The main problem with WAR imo is that it assumes constancy in the relative weighting of defense v offense for each player throughout the game – and for each position relative to other positions throughout the game. That is a simple error – not a ‘disagreement’.
That particular context is a)repeatable and b)skills-based and c)is the best way of resolving the disconnect between runs and wins.
The main problem with WAR is that it gives too much weight to the position listed on the front of the baseball card. If Rfield was measured correctly, then a position adjustment would not be needed.
I feel like most people are only interested in PROJECTED stats. Projected WAR, projected wRC+, projected ERA, etc. Projected stats are, in my opinion, the best measure of talent. Or at least the most useful and interesting measures of talent.
Even retroactively to determine what actually happened (ie: MVP awards) rather than projecting future performance, I think you should look at what you would project them for based on the data from that time period as if the season was played out all over again in the same amount of games played, at the same age. And I think WAR is the best measure of that, by far. Much better than win shares or WPA. Incorporating luck into the equation just is not interesting at all, to me, and it should be removed. Obviously WAR itself includes BABIP-related luck, HR/FB% luck, strand rate luck, and lots and lots of forms of luck. But projected WAR does not. The issue here is that although the most recent season will be most heavily weighted in a WAR projection, it also takes into account prior seasons. So, maybe we make a ZiPS or Steamer or whatever model that uses ONLY one season of data as if the season in question was to be replayed.
He has a point though, which does seem to be missed by many, which is the difference between talent and value. Yes they have a strong positive correlation, especially over a very long sample, but they need not be that closely related over the course of one season.
WAR (especially the fangraphs version) is much closer to a measure of talent than it is of value, but lots of stat-cognizant fans talk about it as if it’s measuring value. You could certainly argue that it’s silly to even have an MVP award, but as long as we do it should be voted on based on value not talent.
My initial go-to if I had an MVP vote would be a version of WAR which replaces the batting and baserunning wins with WPA, while leaving the defense, positional, etc. components the same. It provides a lot of context while not doing something silly like comparing a 1B and SS as apples-to-apples. Using this metric we wouldn’t even be talking about Altuve (6 wins) and Judge (4.6 wins), but Trout and Betts at about 7 wins each.
I think his underlying point is valid, and Dave expressed that he does, too.
People are missing that point because of the harsh, dismissive, and confrontational tone that Bill used. He doesn’t actually say “there is a difference between talent and value” he uses phrases like “dead wrong”, “bad statistical analysis”, and “nonsense”…it’s the tone that makes him sound like a senile old man that has lost touch with reality, whether it’s true or not.
I would love to see a live/recorded debate between Dave and Bill on this topic. Just from reading Mr. James’ comments in the article, he does come across as a bit surly in a get-off-my-lawn kind of way (maybe just my impression). I have seen Bill on MLB TV in discussions with Brian Kenney: maybe he could moderate a debate between Mr. James and Mr. Cameron.
It wouldn’t be a debate – it would be a slaughter. James has lived this life for decades. I doubt Dave knows anything about this that Bill does not. Bill certainly has a much deeper understanding of all the issues involved. I don’t think Dave has a particularly deep understanding of the big picture. No particular offense intended, Dave. Dave is an author about sabermetrics – he is not really a baseball guy from what I gather from his articles. Bill James is both. We should all pay attention when Bill James speaks.
Bill’s pretty much always written that way. Part of his charm, he bruises egos and doesn’t care.
WPA doesn’t tell us anything about value, though, because it argues that a run in the first is worth less than a run in the ninth of a 1-0 game.
Its only valid application is as a fun fact.
That’s actually true in the context of the game, however: a run when you have few remaining chances to score, or your opponent does, is in fact more valuable than a run when all your outs remain.
It is only more valuable due to sequencing, not difficulty (minus details about relievers, platoon advantages, etc), but that is still factually true.
And that’s why it’s such a great fun fact!
I love “Biggest WPA swings” articles…but it doesn’t tell us much of anything about how much a player helped their team win.
I love WPA. It definitely tells you a lot of information in the narrative sense, but that doesn’t necessarily mean it’s a good predictive or valuation metric. It’s important to remember what question you’re trying to answer when you read a stat rather than just trying to find the “best” or “most bottom line” stat.
Sure, but event value is not equal to player value credited for the event if the player isn’t totally in control of all factors associated with the event.
Winning the lottery is more valuable than saving for a 401K, but if I’m evaluating someone’s financial aptitude I’ll credit them more for the 401K.
I don’t know what your definition of value is, but it takes all of a team’s actual wins and losses and metes them out to each of the players according to what they actually did in a game. The fact that it accounts for what happened in the past but not what will happen in the future is a feature to me, not a bug. It aligns with how we experience the game but does so in a methodical fashion.
If we had two players with identical stats but we knew that one of them had hit two walk-off homers during the season we would reasonably argue that the player with the walk-offs had actually provided more value over the course of the season. Now we might not remember that the other player had two leadoff homers in 1-0 games, but it’s not like WAR fixes this problem, because it just assumes that all of the home runs hit by each of the players were done in some average on base environment that never actually exists in a game. Over the course of a season most regulars will have close to the same amount of high-leverage opportunities, so I’m not that concerned about the swings that come from big plays.
There are other flaws with WPA, like the fact that it assigns all of the credit and blame for what happens on defense to the pitchers while it should be split with the fielders (I would expect someone to come up with a version that attempts this soon given the recent advances in ball & player tracking). So yeah it’s not a perfect measure of value, but if it’s not measuring value at all then I don’t know what is.
The argument is that because of that, and Judge being Not-Great in those situations, it wasn’t a contest, or at least that’s what James said; that Altuve was probably twice as valuable as Judge in actual wins and therefore should have easily won the award.
“I don’t know what your definition of value is”
Exactly!
Why do we pretend there is a right answer here?
Player A hits 2 home runs in a loss
Player B goes 1-4 with one home run in a 1-0 win
Player C goes 1-1 with one home run in a pinch hit appearance in a 1-0 win
Player D hits a single with men on second and third and two outs in the bottom of the ninth inning of a game that was 0-1 before the plate appearance started.
Player E pitches a complete game shutout and allows 14 baserunners
Player F pitches a no-hitter in a 9 inning 0-1 loss with 0 Ks and 5 walks.
Player G pitches a no-hitter in a 9 inning 0-1 loss with 27 Ks, his team commits 2 errors.
Player H pitches a no-hitter in a 9 inning win with 27Ks, HE commits 2 errors
etc…
Which player was most “valuable”?
(and these are only some of the variables contained within one game situations, let alone an entire season)
WAR suggests the player with the most value will be the middle infielder who hits like a corner outfielder.
lol, exactly.
And what if the other player hit two homers in one-run games…but hit those homers in the third or seventh?
WPA, like any statistic which includes LI, is a great storytelling tool, but useless as a measure of value added.
Actually, WPA is pretty interesting if you are trying to figure out the likelihood of someone winning the game. That’s what it is good for–as a game-level statistic.
It is NOT useful as an individual player statistic, for all of the reasons mentioned here and everywhere else. And on this, I agree with CLS wholeheartedly.
If your stat for who won games on the field has Altuve significantly behind Trout, you’re doing it wrong. Trout was on the field for 57 wins this year. Altuve was on the field for 94 wins this year.
Base coaches for MVP
I mean, the Angels were 57-58 with Mike Trout on the field. Astros were 94-59 when Altuve played. I can believe that Mike Trout was more valuable than Altuve this season, but not through the method he’s discussing.
I’m not advocating Carlos Beltran for MVP here, Altuve hit .441/.529/.661 in late and close situations. He hit .379/.443/.599 in his 450 PA that occurred when the game was within two runs for a team that won 101 games. If your stats are context dependent instead of context neutral, I’m trying to figure out how Altuve takes a significant step backwards.
That’s fine I guess. Here is your 2017 leaderboard:
#1 Francisco Lindor [#11]
#2T Yasiel Puig [#65]
#2T Edwin Encarnacion [#76]
#4 Carlos Santana [#61]
#5 Alex Bregman [#41]
#6T Corey Seager {#12]
#6T Enrique Hernandez [unranked, not enough PA]
#6T Jose Altuve [#2]
#9T Chris Taylor [#22]
#9T Jose Ramirez [#8]
fWAR rank in brackets — man fWAR sucks, not even close to the truth!
2001 Bret Boone, greatest season 1913 onward.
1998 Chad Curtis with his 90 OPS+ and 2.4 rWAR is tied for 4th.
This is reductive. “If the data analysis comes to a different conclusion than I expected, the analysis must be wrong”.
I also find it more than a bit disingenuous to say he was just being polite all these years and didn’t want to pull up the ladder behind him. Maybe that made sense 10 years ago, but most of the sabermetric world has climbed past him on that ladder a long time ago.
I think James is getting stuck on nomenclature about the “wins” part, which is understandable. WAR is a metric used to estimate a player’s relative contributions to other players with different skill sets across different positions. It is scaled to wins, but is not designed to explain the attribution of a team’s actual wins.
Way to hold yourself in higher esteem than the person that founded this community’s set of values. Very progressive stuff!
I think sabermetrics are useful – thanks Bill James – but misuse of WAR is a problem – thanks Bill James. It is not very progressive to disregard the opinion of a highly informed, experienced opinion… or is it because it is not blind support of a new idea?
I think your answer nailed it. Well done, Dave.
Except for all those lines of not mentioning the fact that WHERE a guy plays on the field means a lot when calculating WAR.
You seem…rather hung up on positional adjustments.
Hitting a hr while up 10-0 also doesn’t change the outcome of the game, but it was probably hit off of a much easier opponent. Does WAR account for this by say normalizing for opponent strength?
As an Astros fan, I’m happy they won the world series, but the whole season I thought their offense was somewhat overrated. There were many games in which the offense would score 8 or so runs before the 5th inning, at which point opposing teams would put in the mop-up reliever or even position players off of whom the Astros would proceed to score another 8 runs. This clearly drives up all of their offensive stats, and I think made them appear to be a much better offensive team than they were.
Obviously the offense was still excellent, but maybe over-valued. On the other hand, I know when watching the Astros-Yankees matchups over the course of the season I never felt stressed when Judge would come to bat with runners on base because I assumed he would strike out, which he did an incredible amount of the time. And when Altuve came to the plate with runners on I felt confident that he would advance them / drive them in on a sac fly, and maybe even just get a hit.
Disregarding how clutch a player is, I think there must be some value in putting the ball in play with runners on 2nd and 3rd and less than 2 outs.
Are there any metrics that try to measure / normalize for these two specific factors (Value added or lost due to contact by situation / Normalize by opponent strength)?
Sure, there’s value in putting the ball in play with men on third and less than two outs, but his a HR in that situation any less valuable? It’s fine to provide contextual isolation, but then you have to explain why one specific skill is more valuable than another.
Also, if you’re going to start introducing contextual weights, how much context do you provide? Perhaps we should examine how many strikes a hitter gets with men in that 2nd/3rd situation? I am guessing Altuve sees a lot more because of Correa lurking, for example, whereas Judge probably sees a lot of sliders out of the zone. Could it be that pitchers would rather have Altuve put the ball in the play than risk giving up a 3-run HR to Judge? Now, at some point Altuve may begin to accord that respect, but if context is vital, how can we ignore that number of hittable pitches each batter sees?
I agree with the whole article above by the way. I guess my point is that we already include some context implicitly: All of Altuve’s stats measured from last season are actually Altuve’s stats when hitting in front of Correa. So if we care about OBP for instance then we would like to measure
p(Altuve_OB),
but we’ve only ever measure
p(Altuve_OB | Correa hitting next).
We would really like to measure the joint distribution,
p(Altuve_OB, any_batter_hitting_next)
and then sum over all possible subsequent
batters to get …
p(Altuve_OB) = \sum_{all_batters} p(Altuve, any_batter)
But we can’t do this because we don’t have the statistics. You could estimate the missing statistics to recover the relevant marginal, but we don’t actually do that as far as I know. Or am I wrong? Shouldn’t we include as much context as we can reliably estimate?
My point is that all the stats we are reporting anyways are all conditional
distributions, presented as marginals and I think that was Bill James’s point as well, though I completely disagree with how he wants to adjust everyone’s stats based on the number of team wins.
I’m fine with pretending that the conditional distributions we estimate are
marginals when we cannot do any better, but we should always try to remove conditional dependencies where possible, which is the whole point behind using ballpark adjusted statistics, expected value of FIP, etc.., and I think that it should also be possible to normalize for the quality of the pitcher faced and normalize hitting stats by the opponent’s strength. For instance, even though we’ve only measured Altuve’s performance when he is hitting in front of Correa, we have measured Altuve’s performance against many different pitchers. Maybe
even enough that we can try to accurately estimate the marginal distribution that we are actually interested in.
P(Altuve_OB) = \sum_{all_pitchers} p(Altuve_OB, pitcher)
Where as right now, I think what we are doing is …
p(Altuve_OB) = \sum_{all_pitchers} p(Altuve_OB | pitcher)
And these are not at all the same thing. So my question was does WAR do what I think it does, which is use the conditionals are marginals, and if so are there reasons we don’t adjust for strength of opponent?
Sorry for the length. I am not trying to argue, I just want to know if WAR includes this or not.
Does WAR account for this by say normalizing for opponent strength?
Baseball Reference’s pitcher WAR has an adjustment for strength of opposition.
Every time I think bWAR is couldn’t be overfitted any more, I learn something like this and shake my head. At this point you might as well just set up a naive data mining algorithm to determine value.
And it will creep closer and closer to base out percentage, or Runs+RBI-HR.
Has someone shown that normalizing by opponent strength doesn’t generalize across years? The extra information can only hurt if we don’t use it well. But I would have thought it would help. But I guess I wouldn’t be too surprised if it lead to some overfitting either.
rWAR for pitchers accounts for opponent strength, and RE24 believes that a groundout with a runner on third is more valuable than a K.
Thanks!
“RE24 believes that a groundout with a runner on third is more valuable than a K.”
Nope, RE24 thinks that advancing the runner is more valuable than not advancing the runner, it is agnostic as to how that occurs.
I was very surprised to read James’ criticism because, well, it didn’t sound like James at all. It almost seemed to be a contrived argument designed to both discredit WAR (perhaps a little bitterness about how it has far exceeded Win Shares in relevance), and demean Judge, or at least his style of player. The rebuttal above does a good job responding to James’ indictment of WAR, but the article also contained other oddities, such as criticizing Judge for batting only .262 with RISP, while ignoring his 1.015 OPS, which was about 200 points higher than Altuve’s. There were other split-based arguments that were similarly dissonant from what you’d expect of “advanced analytics” that I had to stop reading several times and consider whether James was really the author.
Right. The problem with James’ article wasn’t the critique that WAR is a context-neutral stat and that MVP voters should instead consider how the players helped toward actual team wins. Okay, fine. Putting aside that that critique is trite/tired and that many/most voters already do what he suggests by looking at relatively-mainstream saber stats like WPA, the problem was his dependence on silly splits-based arguments (like BA w/ RISP) to make his point.
It’s also impossible to judge Judge (*rimshot*) without also looking at the competition he faced in high-leverage situations. IIRC, Judge faced, on average, the very toughest opposing pitchers in baseball during innings 7 thru 9. If you’re looking for context, you absolutely have to account for that.
“IIRC, Judge faced, on average, the very toughest opposing pitchers in baseball during innings 7 thru 9.”
Hmmm…do you have actual data to back up that statement?
Personally, I’m skeptical that it’s true given that we know some of the “very toughest opposing pitchers in baseball during innings 7 thru 9” were on his own team. And he obviously never had to face them.
You never address a basic point. Are you measuring how good a player is, which should be context neutral since such context (leaving aside “clutch” debates) may be random? Or are you measuring how valuable a player is, which is context dependent, because while it may have a random element it’s obviously true that some hits produce more value than others depending on when they take place.
You seem to acknowledge this basic distinction when you say “I do think there is a case to be made for looking at the circumstances under which a player performed, and I did use context-dependent metrics when I was an MVP voter.” But you then go on to conflate how good a player has been with how valuable that player has been. It’s a simple distinction and is at the heart of saying WAR may be inapt for an MVP voter. In short: if I’m a GM I’m looking at WAR (or its variants) to judge how truly good a player is. But if I’m an MVP voter I’m voting on who produced the most value even if some of that value is due to random selection of when hits/runs were produced.
I think to justify WAR (or the Altuve/Judge equivalency) you need to more directly address the distinction that some value comes out of the random context in which events take place.
To highlight some of the problems that accompany a solely context dependent analysis, it might be helpful to rephrase your questions a bit. Instead of “How GOOD” vs. “How VALUABLE”, how about “How many extra wins DID this player help gain” vs. How many extra wins WOULD I EXPECT a player who played as he did help gain”
The problems that come into play with Context dependent analysis are therefore myriad. For example…
1. What is the “Goal” towards which players are supposed to advance? Are they trying to score runs? win games? make the playoffs? etc.
2. If a player’s actions cause reactions, how do we account for those reactions (if player A hadn’t done what he did, team B would have acted differently, etc)
2b. Does timing count? e.g. is a walk off single more “valuable” than a lead off home run in a 1-0 game?
3. How do we account for the opponents? e.g. is it more “valuable” to get a hit off of Kenley Jansen than to get a hit off Josh Tomlin? If we assume the same launch angle and velocity, Is it more “valuable” to hit a ground ball single in between Jorge Polanco and Miguel Sano than it is to ground out to Francisco Lindor?
If a player has a walk off plate appearance in game 162 for a team that made the playoffs by a half game, should he win the MVP? He is almost guaranteed to have the highest Playoff Chance Added on the season, potentially being worth up to +99% playoff odds, but we have no idea how hard that thing was to do, or how many other players, given the same opportunity, would have done the same.
On the other end of this logical extreme is the case of a team on which the 1-6 hitters always hit a home run, and the 7-9 hitters always make out. Under a context dependent analysis, the #1 hitter will be far more “valuable” than the #6 hitter, solely because of the manager’s lineup decisions.
I think the moral of this story is that MVP voting isn’t as simple as correctly answering any given question like “which player’s presence on his team caused the largest increase in team win total in comparison to how that team would have performed without said player” or “list every player, without whom his team would not have made the playoffs this year, in rank order of how many other teams in baseball each respective player could have caused to have made the playoffs had he played for them this season”… Instead, I think the MVP award asks “who was the best player this season”.
The fact that this question can’t be answered simply by looking at the name on top of any given leaderboard is precisely why we tend to invest so much in the process of deciding.
From this context, James is essentially repurposing the older argument along the lines of “comparable RBI numbers doesn’t indicate equal skill level” and applying it to WAR. Just replace WAR with RBI…
“baseball ref lists the small guy at 86 and the big guy at 81, but in reality they are nowhere near that close. I am not saying that RBI is a bad statistic or a useless statistic, but it is not a perfect statistic, and in this particular case it is just dead wrong.”
What I don’t understand here is what he is arguing… He says Altuve was clearly more deserving of the MVP than Judge… Ok, Altuve nearly unanimously won the MVP, what’s the problem here? This just seems to me like an end run to complain about the fairly obvious point that WAR doesn’t answer every question.
Bill James isn’t just stating that Altuve should be MVP over Judge, but rather that there isn’t even a question of whether Judge should even be MVP. That he definitively isn’t an MVP candidate.
From a win-based metric standpoint this makes all sorts of sense, since the measures that contribute to WAR and evaluation of the Yankees treat them like a ~100 win team instead of the 91 they actually won. If you believe Judge should be awarded for helping the team achieve what should be 100 wins, then he’s either MVP or basically tied for MVP. If you believe he should be awarded for helping his team win 91 games, then there’s really no debate at all and Judge is barely a candidate.
Altuve got 27 1st place votes, Judge got 2, Ramirez got 1.
James proposes no more worthy candidate for second place.
Somebody must receive second place votes.
James is statistically savvy enough to realize that “some of his friends and colleagues” is not necessarily a representative sample of his general audience.
erego, his piece can be translated as…
“garrr, not EVERYBODY agrees with me!, I’m going to publish a simplistic argument against a straw man in order to make sure that Aaron Judge doesn’t get too much credit for being good. Shame on everybody else for thinking that an extremely well constructed stat purposed to show generalized value might be a good starting place for deciding the MVP. Ohbytheway, definitely no bias or saltiness here, its everyone else who is crazy.”
Yup.
My biggest issue with James’ argument isn’t the principle, but the fact that, when you use context to account for how much Altuve and Judge “actually” helped their team score…Altuve nets a grand total of 1.1 runs, still leaving him far behind Judge in FG’s estimation.
Bingo. This is the key. I’m totally on board with RE24 instead of the context neutral stuff for MVP awards, but it just isn’t that big of a difference.
(altough Judge would lose some too, I think)
Absolutely tremendous article.
This is a great article, Dave, and I’m surprised you left out one argument that is particularly important as relates to Fangraphs WAR:
WAR is in part designed to be forward looking. That’s why fWAR uses FIP instead of RA/9, right? It is a better predictor of future performance.
Perhaps with player careers that are long-over, players up for HOF induction, it makes sense to simply look at what happened. Baseball Reference uses RA/9 in their WAR stat, and it happens to be primarily used in backward-looking analysis. It’s good for that.
But with current baseball players, outside of awards season, I’m not just interested in what they’ve done this year: I’m interested in how likely they are to maintain it. WAR helps us do that, especially when we combine it with other individualized ‘luck’ measurements like BABIP, LD% and HR/FB%.
End of season awards voting will always be flawed. There’s a lot of clear bias trends in voting as a whole: not just in terms of looking at bad stats (20 win pitchers), but also awards fatigue for the very best players (how many MVPs should Barry Bonds have actually won? Or Michael Jordan for that matter? More than they did) and the way that we consider those awards so heavily when it comes to HOF cases. Its too bad that we treat these with such import, as we know some fool is eventually going to argue that Joey Votto isn’t a HOF caliber first baseman because he never won an MVP award. Its almost inevitably going to happen in 12 years or so, and it will be dumb as rocks.
It will be a particularly dumb argument since Votto has already won an MVP (2010).
Haha, true! Edit “He didn’t win multiple MVPs, so he wasn’t clearly the best of his era.” .. still will get said, still will be stupid.
Is Votto’s MVP award becoming a Berenstain Bear thing?
Best comment on the thread!
Even though the intention of WAR may be to be more “predictive” by avoiding luck-related stats, is it actually predictive, though? It may be, I just don’t know and that might be a bit hard to prove. That is why I think a descriptive (context-relevant) stat is more useful.
WAR (or more specifically, the components chosen for war, like FIP) are much more predictive of future results than purely descriptive stats are. That’s why we use FIP instead of ERA in the first place; it is a better indicator of true talent level, which is a better indicator of future performance than actual performance would be.
It is also better at isolating individual performance, which is why it is better at predicting future performance. Estimating “true talent level” is more complicated. You’d probably want to triangulate between estimators and scouting reports for that.
War and fip are entirely backwards looking. The reason why fip is used is because it gives credit to fielder for BIP since pitchers have supposedly no control over BIP. It leaves out sequencing to be context neutral but is still backwards looking
I think you misunderstand the metric. FIP is backwards-looking in that it is based on what happened, but it’s used because defense-independent pitching statistics are more predictive than defense-dependent pitching statistics. That’s the cornerstone of DIPs theory in sabermetrics back to Voros McCracken.
FIP matters because it only gives credit to the pitcher for things exclusively under his control, and THEREFORE it is more predictive of future performance than ERA or RA/9, which are at the mercy of more variable factors and therefore are less predictive.
FIP isn’t about telling you how much of the run scoring is the pitcher’s ‘fault’. It’s about giving you a better understanding of the pitcher’s underlying performance so that you can better understand what to predict going forward, because 2017 FIP has a higher correlation with 2018 ERA than 2017 ERA does.
FIP is valuable because it is more predictive. It’s not entirely predictive, but it’s significantly closer than ERA, because it lacks outside influences (or at least, it has less outside influences)
The point of DIPS is to get rid of the defense’s effect and to see what the pitcher’s performance would have resulted in with an average defense. This is looking at the performance and is entirely backwards looking.
In doing so they decide that to isolate defense they ignored any play where the defense had an impact because they believed (somewhat incorrectly) that the pitcher had no control over these and that it was entirely based off of the defense.
FIP is still backwards looking but it strips away any possible impact of the defense.
Yes, and by doing that it is more predictive of future performance than ERA
https://www.baseballprospectus.com/news/article/29898/prospectus-feature-dra-2016-challenging-the-citadel-of-dips/
FIP has a closer correlation with next year’s ERA than this year’s ERA. That’s why DIPS theory works at all.
I don’t disagree that it is a better predictor of future ERA but it isn’t its goal. If one was trying to predict future ERA using K,BB, and HR the coefficients are fairly different.
A stat that avoids context in pitching is not very good at looking backwards. FIP treats a grand slam and a solo home run as if they are the same thing. They are clearly not the same thing in the reality of trying to win baseball games. A pitcher who went 6 innings, gave up 3 singles, 1 walk and 1 run off a solo home run, going against a pitcher who went 6 innings, gave up 3 singles, 1 walk and a grand slam because sequencing, are you really going to say that their value in that game was the same?
The forward-looking part of this story is important as a measure of value too. Predictive validity is an important way to make sure your measurement actually measures what you’re trying to measure. Unless I’m mistaken, it is the only form of content validity that we can use here–this isn’t like a likert scale where you can just ask more questions.
FIP-based WAR for pitchers is predictive because it does a better job of isolating individual-level performance, for the same reason that SIERA-based WAR would be better because it does a better job of isolating individual-level performance.
This may or may not be relevant to RE24, which is likely not repeatable but could more accurately measure performance. (WPA should not be used for individual-level performance, and Clutch is totally busted and shouldn’t be used for anything)
“Predictive validity is an important way to make sure your measurement actually measures what you’re trying to measure.”
Best.
Sentence.
Ever.
I am glad you like it. Should I say it again? They should all be the same thing!
I get where Bill James is coming from when he says we need to dig deeper than just the one catch-all number to determine a player’s true value for the past season, but I wouldn’t be fully on board with the idea of adjusting WAR down, or up, based on a team’s Pythag. After all, Judge and the other Yankees delivered the performance on the field that should have led to 100 wins, were it not for situational factors such as a poor one-run record, and regardless as to whether the wins weren’t actually pocketed, the performance still has predictive value going forward. And when it comes down to it, WAR is an economic statistic that helps teams forecast performance to make business decisions for the future. It is, as noted, far less perfect as a value statistic measuring performance for the purposes of rewarding a player’s impact in the past.
This all said, I wonder whether it would be useful to develop a WAR-type statistic that adjusts for situational (e.g., clutch) stats, where value delivered is heightened in high leverage situations and dampened in low leverage situations. It wouldn’t necessarily have to replace WAR, but could serve as a shorthand companion statistic, to relieve us of having to rigorously examine dozens of split stats to get under the hood of a player’s topline WAR. Call it aWAR, maybe, but give me credit if you do. 😉
The issue with the concept of leverage is that it claims a run scored in the 9th is more valuable than one scored in the 1st, even in a 1-0 game.
And that’s simply not true.
I’d really enjoy a version that used RE24 instead of wRAA, however.
A single run scored in the 9th of 0-0 game would lead to 1-0 game 95% of the time.
A single run scored in the 1st of 0-0 game would lead to a one run game far less than 50% of the time.
A run scored in the 9th at 0-0 is absolutely more valuable than a run scored in the 1st.
That’s only true prospectively. After the fact, we know that both decided the outcome of the game.
What matters is a single run scored in the first at 0-0 is much much less likely to affect the outcome of the game than a single run scored in the ninth at 0-0.
Yes, some of the games in the former end up being one run games, but the point is they are few and far between. WPA, ideally, adjusts for the possible distribution of scores leading from 1-0 at 1st inning.
To give one example at how looking back could lead to weird outcomes, assume a starting pitcher pitches perfect for eight innings and leaves the mound at 0-0. Then his team scores (a) 10 points/(b) 1 point in the ninth.
Are you willing to argue that knowing the results of the game, the starting pitcher’s contribution in (a) is much smaller than (b)?
-1 on all posts just for the fact of referring to runs as points.
It absolutely is true because of the context. A run scored late in a previously tied game when there are few remaining chances to produce additional runs is considerably more valuable than a run scored in a previously tied game when there are many remaining chances to produce additional runs. Your argument that a run is a run is a run would have to mean that a home run in a 10-0 blowout would have equal value to a home run in a tie game, which I doubt that you would agree with. The context of the event affects the value of the achievement.
What CLS is trying to say (I think, and if not, this is what I am trying to say) is that there is no reason to prospectively assign value to a run based on context when you don’t know the whole context. If a game ends in a score of 8-7 (or any other random score), the most important run is the one scored first, second, third…all the way to eighth.
WPA just shouldn’t be used to assess cumulative value because its knowledge of the context is outdated by the time the game is over. The ordering of who scored the run when is ultimately not important, just that you piled up more at the end of the game (when you can properly assess the ultimate effects of a run, which is equal no matter at the first or last inning).
“If we want to reduce Aaron Judge’s WAR by the amount of wins he cost his team by performing poorly with men on base, do we want to also adjust his WAR (and everyone else’s) by how his teammates did after he reached base?”
The answer to this question has to be yes. To suggest Judge on first has the same impact to the hitter at the plate is to suggest Billy Hamilton, Dee Gordon, Ricky Henderson, etc. on first do not impact the pitcher and fielding positions. It’s a non-zero value of impact.
Does the pitcher throw fewer off-speed pitches to give the catcher a better chance to throw Dee out in case Dee is stealing? Do the outfielders shift slightly to account for trying to prevent Hamilton from taking the extra base, this creating a better opening to hit the ball?
These instances exist. Your article suggests they should be ignored for the sake of preserving a metric.
I find Bill’s over(?)-reliance on player performance to team performance a bit startling, to be honest.
For simplification purposes I’ve always judged WAR and WPA as individual performance where one is “forward”-looking and the other “backward”-looking. For MVP awards (especially for things like the World Series MVP) I’ve always viewed WPA as a better indicator, because it explains what actually happened, not what *could* happen.
Bill’s issue seems to be that everybody dropped Win Shares after about 30 seconds, which is ironic because it only came out as equal to actual wins because he forced it to.
Correct. Win Shares is a “top-down” approach instead of “bottom-up” like WAR, although somewhere here on FG there’s been research how accurate the sums of a team’s players’ WAR (plus the replacement constant) is to actual wins.
Personally, I’d instinctively go with the bottom-up approach instead of top-down in valuing player performance.
Context neutral metrics are useful for fantasy, a context neutral game. Context stats are more useful for the game on the dirt played by teams with variables beyond quantifying at the moment.
“Context stats are more useful for the game on the dirt played by teams with variables beyond quantifying at the moment.” This argument makes sense, but I believe the evidence right now is that its incorrect. Context neutral metrics are much better for predicting the future than context stats. Therefore, enough of what context stats measure is random (or “luck”) that it’s not capturing much actual signal rather than noise.
My understanding is that WAR now includes enough of baserunning value that it captures most of the difference between having Judge on first base, as opposed to Billy Hamilton or even an excellent baserunner who isn’t super fast, such as a Cal Ripken or Derek Jeter.
Is “luck” shorthand for “this is a variable we can’t measure yet?” Wasn’t many years ago that an outlier BABIP was luck rather than speed of hitter or placement of hits. WAR is getting better at baserunning but still doesn’t quantify for coaches. Defensive stats can’t factor for pitchers missing their spots. Caught stealing doesn’t seem to weight for a pitcher’s time to the plate.
Stats have their uses but there is still a lot of information that is left out and that is the “context”. That context is what coaches are paid to understand instinctively based on decades in the game.
Good points.
As it relates to MVP voting, I have long thought (and posted questions in chats) that WPA (with a defensive component) would be the best way of actually measuring “Most Valuable”. It may not be sustainable in terms of “Clutch”, but I like how WPA measures the impact on Wins & Losses of actual events that happened.
The problem I have with WPA has a measurement of the player, as opposed to just the event, is it isolates one link from a chain. It also ignores how hitters are treated in high leverage situations. Also not considered is how value can be added beyond the outcome of the current game. For example, a 3-run HR in the top of the seventh to extend the lead from 3-0 to 6-0 might not rank that highly on the leverage scale, but it could allow a team to rest its best relievers, who would then be available to pitch in close game that might occur tomorrow. There is undoubtedly value in that, and WPA ignores it completely.
Sure, and it’ll also give credit where is isn’t “deserved” like a player reaching on a dropped third strike or whatever. But I still like it a lot more than WAR when looking backward to determine which player contributed the most to getting wins.
In that case, Trout absolutely crushed the field in that. 5.58 to 3.9 for Cruz in 2nd. Even adjusting for leverage (WPA/LI), he narrowly defeats Judge. Should this be another year where Trout was “robbed” of an MVP he deserved?
In my opinion, yes. Trout was hugely impactful on actual wins in a way that nobody else in his league could really compare to. And as a good CF, there’s no reason to dock him points for defense. I would have voted for him.
Also note, this is why I think Nolan Arenado got hosed. He was near the top in WPA with Votto, Stanton, & Rizzo but was the only one of those three to play outstanding defense at a difficult defensive position.
I think Bill’s argument holds merit for the specific purpose of MVP voting. WAR as we know it is great, but for a general-purpose stat to be used in so many specific functions is problematic. I’d love to see a win-dependent version on the site to help guide voters to better ballots.
It is not entirely true that players do not have the ability to account for context. Players are very aware of what team they play for and how good their teammates are. The individual strategies that a player should pursue in a given situation do actually depend on the context of the game and the context of their team.
For example, walks are more valuable to Yankees and Astros hitters than they are to Padres hitters. If you play for the Padres and you want your team to win the game you are playing in, it may very well be the case that you should always be trying to hit doubles or home runs with two outs rather than taking a walk since your teammates are rarely good enough to score a runner from first base with two outs.
…of course doing the right thing on your particular team might hurt your WAR because WAR assumes everyone plays on the average team and that the things that win games on your team are the things that win games on every other team and in every situation. These are simplifying assumptions that are not actually true. The extent to which player approach to real game situations changes performance is not well understood.
Isn’t the real argument whether WAR should be determinative, or just another data-point in evaluating the player and the award?
I’m sure at some point a metric could be devised to isolated and place a value on most every single “transaction” that a player enters into during the course of a season. But does it tell the whole story? Does the 4th inning 3 Run HR which makes the 3-0 game 6-0 change the way the manager uses his roster? How about if it’s top of the 9th and you had your closer warming up? James is a little more “get off my yard” than perhaps he used to be…maybe there’s a creeping realization that the revolution he started may have moved away from him.
Well said, Dave.
I found Bill James’ article unhelpful, and the suggestion in it that Hosmer was more valuable than Judge laughable. It is very, very complicated to correlate individual contributions at the plate, with the glove and arm, and on the bases to team wins and losses. No one, to the best of my knowledge, has attempted to measure the clutchiness of good and bad defensive plays, for instance, and incorporate that into a metric.
As Dave says, MVP voters ought to take into account more than WAR totals in filling out a ballot, and certainly clutch offensive performance would be one such thing. There are times when WAR tells a very incomplete picture- with catchers, high leverage relief pitchers and for position players with outlying defensive seasons. The 2017 MVP race wasn’t one where these considerations applied, but rather a case of two fairly evenly matched candidates, one of whom performed well in the clutch (at least offensively) and one who didn’t.
WAR is particularly useful in assessing value over a period of several years (when defensive ratings take on more meaning). And frankly, I have little interest in how a player contributed to wins and losses over a period of years. Red Ruffing’s value when he was traded from the Red Sox to the Yankees was not a function of how his pitching contributed to wins and losses for the Red Sox but how well he pitched. WAR captures that well.
I’m with James on this. WAR, by its name implies that it measures Wins. It is all well and good to wave your hands and say that it is a measure of true talent, or actual performance expressed as Wins, or whatever, but the fact is that what is being measured, by deliberate intent is not how many wins the player produced for the team. If you read 99% of the articles and comments, like stay the one which Cameron wrote about how the writers are doing a better job of picking MVP because their choices align better with WAR, imply that WAR reasonably accurately capture both Wins and Value. If it was called PAR, performance above replacement and measured on a scale of -100 to +100, and Trout was a 92 and Stanton a 76, this exact issue would not come up. Once you try to convert a measurement of something to Wins, then it is perfectly fine to question whether that process is accurate or in fact misleading.
The only statistic that measures wins is Wins. WAR might have units of wins, but it measures value. RAR equivalently measures value but with units of runs.
tl;dr; j/k
Short version
1) WAR is probably the wrong metric to use for MVP discussion. This doesn’t mean there’s a problem with WAR, however, rather that it’s just the wrong tool for the job.
2) Since the notion of what constitutes an MVP is ever changing, based on the whims of the electorate from one year to the next, there is no such thing as a truly right metric anyway.
Now I’m going to spend all day trying to figure who did have the most hits on Tuesday’s in 2017.
Abreu and Bogaerts (36)
Mr.James sports a worlds series champion ring.
One that he richly deserves for helping Boston reach the promised land.
His ideas have real life worth in todays baseball world.
Wins,not runs equal rings.
Count da rings, baby!
You can’t win without scoring (& preventing) runs. So by your own logic runs do in fact equal rings.
Appeal to authority fallacy
You should vote for the MVP based off of what happened. Not what was expected to happen. But if you are a GM (even of your fantasy team) you should base your decisions on what would of happened in a context neutral environment. How are people even debating this?
Because
1. THE EVENTS THAT MAKE UP WAR DID HAPPEN. If a player’s WAR is based on a set of discrete events that happened in the past season, then their WAR is – in a way, albeit a linearly-weighted, context neutral way – “what happened”, not “what is expected to happen”.
2. LOOKING AT WHAT HAPPENED IS NOT A FOOLPROOF WAY OF DESCRIBING REALITY. Some events are more accidental than others. Players really do get lucky/unlucky sometimes. I get that saying this sounds like a cop-out for a player who’s slumping or a hand-waive at a player that is playing well, but it’s a real factor. So we can’t judge our opinion of the past simply based on outcomes; we should also look at the events that precipitated those outcomes, and WAR helps do that.
I get that, but getting lucky is something that happens. Which is exactly my point. It still happened!
As an older baseball fan who experienced the entire analytical revolution growing up, here’s my feelings on not just this MVP issue but also just in general with fangraph/BR/BP direction of analytics.
When I first came into contact with analytics through BP and Rob Neyer, I was hooked. Because the BA/HR/RBI slashline has always bothered me, and in general it felt weird so many stats (notably walks) are rarely talked about and often dismissed. To me it was amazing people were trying to better understand what is a positive for a team and what is a negative beyond BA/HR/RBIs. And with every step I understood more, and all of them made great sense.
However, more recently, I’ve found that stats have become something of a different animal. Instead of saying what is a good thing for the team, we are now saying what SHOULD be good for a team. The most notable representation of this are incorporating BABIP and FIP, which tries to understand what a player SHOULD have done instead of what they actually did. If we’re talking about FA signings or trades, this makes sense. If we’re predicting who will win a playoff series, this also makes sense. If we’re trying to determine what a player HAS ALREADY DONE, this doesn’t make sense.
The reason is simple. I am a baseball fan who watches and follow baseball games. In baseball games, the team that scores more runs win. Not how many baseruns they score, not how many hard hit balls they hit that should’ve been hits, but the actual score. I don’t look back fondly to a series my team lost 3-4 just because my team outscored the opponent 30-20. Because my team lost and I am always bitter when that happens. And no amount of context independent supporting stat changes that.
People touting context independent measures say it’s “fair” but what is it “fair” for? Honestly the only answer is “luck.” It is balanced for luck. But luck is part of the game and has always been, for good or bad, what my and other fans’ viewing experiences are. To take that out just confirms what all the old school people have always attacked analytics on, which is that it’s no longer really about the game itself that we’re discussing.
And fangraphs is definitely helping to perpetuate this by using a stat that equates to winning (WINS above replacement) when it’s not. It’s a rough estimate of how a player should add to a team’s win total all things neutral. It’s like saying if a player does this through simulated 1000 seasons he’ll average out to roughly this many wins added. But people are misinterpreting this, even on this site, as how much this player actually added this particular season.
^^^so much this
Dave, I think this is right on. Well done.
Dave has constructed an answer here that hopefully is considered by mainstream sports talk shows as much as James’ original comments. James isn’t wrong, and Dave willfully admits that, but it’s James’ tone that has me worried that personalities like Chris Russo (or whoever) will use his statements to embolden their own anti-WAR claims without having to consider what James really meant.
Not that it matters, but I think of WAR as a tool to value players for contracts because it takes into account most things that are skill-based and will translate year to year. It’s also a real nice number to look at to get a fairly accurate glimpse of a player’s contributions to his team during the season, but leaves room to look at RE24, WPA, clutch and the like. The disconnect between what WAR is and what people think it’s meant to be is still huge, and it’s a pretty tired argument, but I guess I could probably tell you more about WAR than I could health care issues or real life economic problems, so I shouldn’t get too worked up over people not understanding the purpose of a baseball metric.
I think it’s pretty simple:
If you’re looking for a predictive stat, and you have shown players do not predictively alter their performance based on context, then a context neutral stat like WAR is as good as it currently gets.
If you’re trying to describe what actually happened, context absolutely matters. If 2 players had identical stats, but Player A hit all his HRs with the bases loaded, he’s the MVP. Period.
But if they finish with identical stats, that means the other guy drove in a shit ton of runners too.
It also means the other guy hit the same number of grand slams, ergo, all his HR’s also came with the bases loaded.
Bill James seems to be taking a both/and question and insisting it is an either/or question. That never works.
WAR is an excellent sanity check for MVP. If someone tells me the player with the 10th highest WAR is MVP, I’m skeptical that contextual factors can bridge that gap. But taking a handful of top 3-5 guys? Yeah, context matters. You need to look at both. The trickiest part is trying to figure out how big a WAR gap context can overcome. That’s an interesting question, but the answer doesn’t start with throwing WAR out the window.
I’m on Dave with this one. James’ use of “big guy” and “little guy” got me thinking. If there was an Olympic stepped platform for the top 3 MVP finishers, would Altuve’s head be higher than Judge’s?
This may be a dumb question, but…I get that WAR is context neutral, and something you may use to value that players performance from year to year. I believe I also understand that WPA is relative to the team, and as Dave noted above it’s just odd that Jose Ramirez had a higher WPA than Melky Cabrera…although I guess Maybe Jose just did not get a lot of credit as he was on such a good team? How does/would WPA correct for good/bad teams such as…I don’t know, Ernie Banks playing on dreadful Cub teams would have a lower WPA than XYZ player on good Dodger teams as an example? If you are always coming to bat down 4-0 vs. 0-0, is that your fault? I don’t know, at some point WAR is great, WPA is good, but you have to watch a lot of games and be open minded and thoughtful about awards to make a good choice and it seems like that is being done now. NOW if only we could translate this to the HOF voting…please Grinch, trammel, Whitaker, Tiant, John, Schilling (hate him but), Mussina, etc.
That WPA thing is mostly about clutchness, not about the teams being good.
One feature/bug of WPA is that a handful of performances in high leverage (~50 PAs) matter much more compared to what players do in low leverage (~300 PAs).
Ramirez has hit .196/.204/.326 in high leverage while
Melky hit .372/.426/.767 in high leverage.
It’s not just that high leverage situations will register disproportionately larger positive or negative WPA numbers as compared with low leverage, but also that starting at a win expectancy of less than 50% will yield higher positive WPA totals than negative WPA totals (and vice-versa for starting with a lead). In other words, the statistic has an imbued momentum towards 0 for the entire universe of players involved in the game, and if your teammates are occupying a disproportionate amount of the negative WPA in the universe, then your contribution will be worth more positive WPA as a result.
WPA measures the run environments before and after an event and subtracts them. However, the runs per change in win probability is non-linear. moving your team from -1 runs to 0 runs will result in a higher WPA than moving them from +2 to +3, which will result in higher WPA than -4 to -3, etc.
The result is that players on teams that spend more time losing by a small amount will have more chances to accrue positive WPA than players on teams that spend more time winning by a lot, etc.
As such, we should expect higher WPA totals from, for example, Trout and Betts (lone bright spots on teams that finished close to .500 and 0 run differential) than from Judge and Altuve (players on teams that won a higher percentage of games and ran large positive run differentials), given similar context neutral production.
This isn’t an indictment of WPA, WAR, or any other evaluation method, but the question underscores how important it is to understand what a statistic is actually measuring in order to properly incorporate it into analysis.
Another reason to use re24 over WPA for a context dependent metric
It’s not a “one or the other” thing. RE24 incorporates the base-out situation but not the game situation. They describe different things. If you want to know how rare that 9th inning 4 run comeback was, RE24 doesn’t help. If you want to know how effective a player was at advancing baserunners over a sample of more than 1 PA, WPA doesn’t help you…
Great article describing the differences. Each one is valuable. For MVP, context matters in my opinion. The ultimate MVP stat is cWPA, Championship Win Probability Added which takes each game’s WPA and multiplies it by the team’s change in championship odds from the probabliities in the Fangraphs standings section. The Baseball Guage site does a variant of this.
The higher correlation with future performance is an important strength of WAR and helps answer one of the most important questions – what will the player do next year?
The Baseball Gauge is an excellent site. And if you like Win Shares, it is the very best place to go for information.
Certainly for the purpose of a backward-looking award — who was the most valuable player in 2017? — adding some context is appropriate. For Judge, one piece of context is that he did A LOT of damage when his team was leading or trailing by 5 or more, slashing .382/.500/1.000. By the old runs created formula, that’s 89 total bases multiplied by a .500 on-base percentage plus 2 SB and no CS. That’s 45 runs out of 144, or 31 percent of his output in 17 percent of his PAs.
Dave’s point about ‘which question you’re trying to answer’ can’t be repeated too much or loud enough.
If I’m trying to answer a question about the fundamental underlying quality of a ball player, give me context neutral every day of the week.
When we vote for MVP though, I don’t think we’re looking for fundamental underlying(predictive) quality, we’re looking for a measure of what happened.
I agree, though I’d also add that if I’m voting MVP, I want a measure of what happened minus randomness and good fortune outside of MVP candidate’s control, and I think a lot of the “measure of what happened” stats don’t solve this problem because randomness/fortune is inextricably tied up in “what happened.”
This piece is unnecessarily defensive and makes it seem (inaccurately, IMO) like James was attacking WAR. I don’t think he was — he is really just saying WAR is better at measuring ability level than at measuring how valuable a player was in a given year. He is not saying WAR is a defective statistic. He is saying that its indifference to (certain) context makes it poorly suited to an MVP determination.
Would be interested to know Cameron’s take on the Judge vs Altuve example James addresses. I.e., does he disagree with James that their closeness in WAR obscures a significant difference in how valuable each was to his team *this year*?
I think the whole discussion just boils down tothe name of the stat. The W in War is for Wins, and context neutral is slightly problematic in that regard, looking backwards. I would like the total of a teams individual WARs values to equal the actual number of team Wins (plus the baseline of whatever it is 48 wins?), and it seems like it should. That it doesn’t is a flaw in calling it “Wins” above replacement. I agree with whoever else said it wouldn’t be a problem if it was called performance above replacement or something similar.
In the smallest of samples, 1 game, the team that wins deserved that win by outscoring that team. A 102 win team may have only “deserved” 93 wins on a macro basis, but on a micro basis it earned every single win and loss one way or another through the year. There is nothing wrong with WAR, it just shouldn’t be called WAR.
Dave Cameron is certainly not wrong, but Bill James does have some valid points (and some unnecessarily grouchy language along the way). I don’t think FanGraphs tries to make it look like something it isn’t.
The problem James needs to address is that a team wins approach would have to account for interaction effects, which if I’m hearing Dave correctly is what he’s saying when he notes that Correa hits behind Altuve. In my mind, the model for this fraction of the lineup would be:
team_wins = Altuve + Correa + (Altuve*Correa)
This approach doesn’t allow you to simply award some fraction of the total wins to Altuve and some fraction to Correa because you also have to award some fraction to the interaction between Altuve and Correa. I would argue this is undesirable for your model because it means that the value for Altuve can’t be isolated, and that’s oftentimes the central point of player valuation.
And then, of course, the interactions get really unwieldy when you add in every player on the roster.
Those interactions between the players on the same team are part of Win Shares.
There are two areas of interest(concern) that I have had ever since I became a member of Fangraphs and have been exposed to an education in metrics and statistical analysis. This debate seems focused on whether they are true or relevant. Is there such a thing as a clutch player and does the pitcher have any control over the ball after it is put in play? Bill James is definitely saying that there is a difference in performance in various situations. That is a position that I strongly agree with. Years of watching David Ortiz, and now Mookie Betts, makes you think that way and Derek Jeter had enough special performances in “clutch” situations to add to my perception that it is real to some extent. The pitcher argument that there are only three true outcomes that the pitcher can control holds no water whatsoever. Anyone with two eyes should be able to see that many pitchers are clearly able to induce weak contact, for which I offer Mariano Rivera who wasn’t given a chair made of broken bats for no reason. This is not to say that these differences are huge but it is still significant.
On the topic of stats measuring what ‘should have happened’, I have always had one nagging problem with old school stats. Simple old batting average. How is it that a single, thrown out advancing to second, considered a single? I always suspected that the inventors of this logic considered what ‘should have happened’, that he should have stayed on first base, and gave him credit for a single. But to the team? Same as a ground out (unless runner scored or something). So they have been playing around with this concept for a while, we’re just much better at it now.
This is interesting. Within the modern OBP context, this makes perfect sense, as the player reached base safely, and then made a baserunning out. This might exemplify that what “they” were trying to get at with Batting Average was essentially “how often does the plate appearance end in not an out”. Once the player has reached base, if they still make an out, the hitting isn’t the problem to solve.
I hate WAR. It’s the stat for lazy people that want to discuss baseball without putting in any effort. Nowhere in science do people attempt to blend completely different features of a system into one number.
Are you kidding? This is what every social science tries to do (and baseball is absolutely a social science, not a natural one).
WAR is more or less baseball’s equivalent to GDP. One big summary number that (while inherently lacking nuance) attempts to summarize the overall value of its subject. No, it isn’t perfect. No, it isn’t a single source of knowledge that eliminates any need to look at the underlying statistics. But does it effectively capture the approximate overall value of its subject? Yes.
I’d actually liken WAR to currency…
An Apple, an Orange, and a Peach are each worth a dollar, but there are good reasons I might prefer one over the others.
Tommy Pham, Francisco Lindor, and Max Scherzer were all worth 6 WAR in 2017, but there are good reasons I might prefer one over the others.
I agree that WAR should be context-neutral and we should have a different stat (one with the same goal as Win Shares but better constructed) to apportion credit for wins and losses.
The problem is WAR ISN’T context-neutral. The fielding component depends on opportunities, whose frequency is far from perfectly correlated with playing time.
Fact of the matter is , regardless of what the number is intended to do , so many people try and make it so much more. Be honest, how many times have you read an article where someone used WAR as justification for MVP? In that sense James has done us all a service.
I thought James most interesting point, or shall I call it a confession was here
” The logic for applying the normal and usual relationship is that deviations from the normal and usual relationship should be attributed to luck. There is no such thing as an “ability” to hit better when the game is on the line, goes the argument; it is just luck. It’s not a real ability.
But. . . I have held my peace on this for 20-some years. . .that argument is just dead wrong. There are five reasons why it is wrong.
First, we do not, in fact, “know” that there is no such thing as an ability to hit better or worse in a key situation. We do know that MOST deviations from normal performance in clutch situations are the result of luck, rather than ability, and we cannot prove that those deviations are not 100% due to chance—but we can’t prove that they are 100% due to chance, either. The data would look very much the same as it does whether those deviations were 100% due to chance or whether they were 70% due to chance. We do not, in fact, know which one it is.
I acknowledge that, in the 1970s and 1980s, sabermetrics reached a consensus on this issue, and I acknowledge that I was part of that consensus. But we were wrong. We jumped the gun. We should have remained agnostic on the issue until more convincing analysis is done.”
Basically, hopefully, this puts RBI’s in the discussion for MVP without being mocked by the brainwashed masses. Sure, adjust the numbers for opportunity, park, run environment and recognize luck is a component (lucky hitters are more valuable than unlucky hitters), but remember, MVP is not about best player, its about Most Valuable player within a given year.
Its not rocket science children
Would you guys ever consider having two versions of WAR one based off of WPA and another off wRAA. Another two unrelated ideas for changes in WAR are opponent adjustments and using Voros McCracken’s Base Runs FIP as a much more precise version of FIP.
The one based off of WPA is called WPA.
WPA does not have any adjustment for park, defense, UBR, position or include a replacement value. WPA does not equal a WPA WAR
BRAVO, DAVE!
James’s column made a valid point about Judge-vs.-Altuve, serving as a useful reminder that WAR isn’t everything. But the rest of his WAR critique implies that there COULD be a better one-size-fits-all value metric, if only we’d start from actual wins. That is a fool’s errand, as you amply showed.
Probably all of us who cite WAR have been sometimes guilty of investing it with too much precision. Even so, it’s the best basic tool I know of. And context-neutral works for me far more often than it doesn’t.
If you agree that the shortstop who hits in the top of the first is somehow more valuable that inning than his first baseman teammate who is also batting in the top of the first inning. That’s the flaw behind WAR, it double counts for defense. It allows you to save runs on defense based on how well you play, and then gives a position adjustment based on where you play. That might be useful in an MVP discussion, but that’s about it. Why Dave does not address this remains a mystery.
Player 1: SS with a .250/.330/.410 batting line
Player 2: 1B with a .260/.340/.420 batting line
You have first pick. Which guy do you want?
Depends on what the Mariners are offering me for the second guy.
Is it a SS with a .250/.330/.410 batting line?
Because if so, sign me up for the 1B, making trades is fun!
I think the biggest attraction of the WAR framework is that it can be easily assembled in various forms. It would be cool to have various options at FanGraphs, just like at The Baseball Gauge.
While this was an interesting reply, I found it very curious there was not one word about defensive position adjustment. THAT is what makes any version of WAR useful IMHO, and then only in regards to MVP discussions. And certainly not as a cumulative stat over a long career. The writer went out of his way to suggest that WAR is a lousy metric to use when considering who was the MVP. I love fangraphs and the work the do, but Dave and I really disagree here. MVP consideration is about the ONLY time I would look at WAR, in any version. Career WAR is even worse as rPOS (credit for WHERE you play) adds up. It is mistakenly used for Hall of Fame columns (google em’). So it’s bad for the HOF articles, and now I read that it’s bad for MVP columns… my oh my. We have once again arrived at the question: WAR, what is it good for? (Apologies to Edwin Starr)
I view WAR as a way to put players into buckets based on out of context results.
If I am looking at 2 players and comparing them, WAR is what I would use to tell if they had comparable seasons if they were on a team by themselves, with no teammates. It isn’t something that is realistic, but that is what it does.
I do think to determine who actually had the better season you do have to go back and add context into the season. RBIs are an amazing example of this. Alone, they mean almost nothing. But they can give a little bit about the context of the season being examined. Yes there is a ton of noise in the stat, and if you have amazing teammates, or perfect timing on your hits, you can accumulate a ton of them by accident. But they do also tell if you are doing your job if you look at them in context, since they are a way to measure how many times you got the run in (which is the point of baseball). Yes, a single with a man on third is more valuable than a weak grounder to shortstop, but both of them score the run.
The RBIs are one of the worst examples, but at the end of the day, they do matter a little bit. WAR tells us if players were comparable enough to look at the other numbers against each other, and then go from there.
10 years from now I assume that batting runs will be replaced by a opposition-quality-adjusted xWOBA. Not too dissimilar from what Prospectus is trying to do with DRA. We still have a ways to go and the current iteration of WAR definitely includes plenty of luck and context, but it’s better than WPA, win shares, etc.
Excellent article. I have one gigantic problem, fundamentally, with “top-down” statistics like Win Shares that try to work back from the end result. In economics, there’s an idea called “microfoundations” – that all macroeconomic theories must be underpinned by microeconomic behavior. In other words, macroeconomic theories must be built from the ground up – the theories and models must be based upon the behavior of individuals within the economy.
Now I realize you can probably take this line of thinking either way, but I think the ‘Win Shares’ method fundamentally fails here, as being a ‘top-down’ model rather than a ‘ground-up’ one. To try to build a model for evaluating individual performance by looking at a super-macro level – team wins – and apportioning them out to players on the team, ignores the fact that players are contributing in varying positive and negative ways to both team wins and team losses.
Take the following situation: Player A, a centerfielder, goes 2-for-4 with a home run and two walks, in a game his team happens to lose. Player B, the centerfielder on the opposing (winning) team, goes 0-for-4 with a walk. Which player held more value to his team (assuming defensive quality is equivalent)? Should we discount Player A’s performance because his team lost? Would the fact that, say, a reliever on Player A’s team coughed up a late lead, reflect at all on Player A’s value to his team?
Taking Bill James’ argument to its logical extreme, if the underpinning idea behind the model is to apportion out contributions to team wins, should we give players any positive credit for actions in games in which their team loses? After all, the end-goal was not met, regardless of whether the player went 0-for-4 with a golden sombrero, or hit 3 homers on a day where his teams’ pitchers couldn’t get anyone out. But those two performances are not, and should not be seen as, equal to one another.
Take James’ logic to look at a single game. If Team A wins 6-4, how should ‘win shares’ be apportioned out. Should team A receive 100% of the ‘win shares’ (as James seems to imply), or should they be divided up roughly 60%/40%? Or in some other fashion that looks at players’ performances individually? We can reasonably disagree on that (I would support the third option, myself), but I think giving 100% credit to the winning team completely misses the mark on evaluating player value.
Obviously, that’s not what James is trying to do, but if the season lasted one single game (or a week in which the team went 0-7, or…), Win Shares would essentially say that players on the losing team should inherently have 0 Win Shares, as there are no wins to be apportioned out. And IMO, if the model fundamentally fails at a ‘micro’ level like that, we can’t rely on it at a ‘macro’ level. Not that WAR is at all statistically significant in a one-game sample size (obviously), but it theoretically allows for player value to be accumulated even in the absence of wins, which is a much more fundamentally sound idea.
Ultimately, I think James’ logic fails to appropriately isolate a player’s actions/value from his teammates’, and takes into account team context far too much.
Great article. Both Fangraphs and Bill James have been unmeasurably important in my baseball obsession! Would this be a great example of a situation where metrics and the old fashioned eye test could come together?
James has a point, but it seems his only problem is with the name. Rename WAR to xWAR or something similar. It’s actually not Wins Above Replacement. It’s “total runs/average runs per win across the majors”, which seems to be the crux of James’ argument. Many are using WAR for something it is not.
Fangraphs seems to have a focus on answering the question “how much should a team pay this player to play for it?” And that’s a great question to ask and answer. I think fWAR does a great job at providing a decent answer to that question.
But for MVP Win Shares or another stat generated when the season is done and allocating actually earned wins to various players would be something pretty valuable if you ask me.
Of course if you are evaluating a player for a potential free agent signing, WAR is what you want, but I think context for MVP voting is totally legitimate and worthwhile.
That said I’d still likely use a Win Shares type metric as an MVP tie breaker rather than the starting point. Take the top 10 guys in WAR and dig deeper to pick your MVP and you should come out with a good answer.
There are issues with WAR and I understand where James is coming from but I think he overstates his case a bit. His argument about the Verlander/Porcello WAR issues was much more convincing and argued better.
The bigger issue with WAR is not the statistic itself, but the use of it. It is a rough stat where 1-2 points of WAR are not significant and people tend to use it as an exact science. “Player A had 3.6 WAR and player B had 2.9 WAR so player A is better.” It’s even to the point where people say “He’s a 0.5 WAR player so he’s worth $4.5M a year”. That level of exactitude should not exist for what is still a flawed statistic.
I dunno, I think Dave is applying James’ comments too broadly. It seems from the above quote that James is talking about using WAR (or not) as a metric for determining award winners. I think it is perfectly acceptable to use context (e.g. how did you hit with runners on) when doling out awards. Results matter for awards. It is not only I who think this. I recall reading something like this in either Baseball Prospectus or The Book, one of the two.
As far as James’ argument that WAR is imperfect, he is correct and Dave acknowledges this in the article. I have always thought of WAR as the deserted island of baseball statistics as in: “If you were stranded on a deserted island and could take only one statistic with you, which would it be?” WAR fits this bill.
Isnt James point that in determining who is the most valuable player that you cant remove actual on the field performance as compared to context nuetral performance? Seems quite right to me, you are trying to determine in the past who was the most valuable player not form some sort of predictive tool about what might have happened or what is likely to occur or be repeatable in the future. You are explaining who had the most value based on what actually occured. Im with James on this- I just wont use thus analysis to predict next years performance.
What is wrong with Bill James’ take? He is simply saying that analysis is not an simple as using WAR. He doesn’t dislike it – he doesn’t like the way that it is applied. When the father of the sabermetric revolution is not on board, everyone should probably listen. I am with Bill on this – I rail on use of WAR all the time. If you put a gun to my head and make me chose one statistic to judge a player’s worth, well I will chose WAR. That is a lack of analysis though. The problem is that you get the most simplified form of analysis, WAR, and people treat it like it is an adequate stand-alone measure.
You’re actually describing Dave’s argument, not Bill’s. Dave admits here that WAR is not a good stand-alone measure, and it’s why he factored in many other statistics when he had a vote for the MVP.
Bill is not saying that he doesn’t like the way that it is applied, he is 100% saying that he doesn’t like the statistic itself. He says WAR is “dead wrong” and is based on “bad statistical analysis.” He disagrees with the basic foundation and assumptions that are used to calculate WAR, and that’s why everyone disagrees with his take. (Even you, it seems, since your explanation aligns with Dave’s entire article…)
James has devoted his career to saying whatever necessary to convince people to hate Yankees players. Another pathetic example. Aside from WAR, Judge bested if not crushed Altuve in pretty much every stat out there from G to OPS and beyond to XBH and TB and wRC. All Altuve’s “close and late” “clutch” hitting stats mean is that he was much more human before the 7th inning than after… Pitchers feared and pitched around Judge every single inning, to the tune of 128 BB’s, something that was not the case for Altuve.
I just listened to Dave’s appearance on the “Effectively Wild” podcast and posted this question there, but then realized I don’t think anyone ever reads the comments on Podcast articles, so I’ll put it here (though considering this article is over a day old, and thus already ancient history, and this comment is at the bottom of 150+ others, I can’t expect much of an answer either):
If, as Cameron speculates, the future resolves into one where we have a James-approved context-incorporating retrospective WAR, and a Statcast-derived prospective WAR, where does that leave the current Fangraphs/BBRef WAR, which is neither? Does Fangraphs shift away from its current model, adding both of these “backWARd” and “forWARd” ones which its writers increasingly use? Presumably the Statcast-based one would be used the most, but since the Statcast data is sourced from MLBAM via Baseball Savant, which isn’t entirely open, won’t that mean the “forWARd” calculation will be something of a black box?
This is really interesting. Giving it 2 seconds of thought before I regurgitate a response…
It strikes me that fWAR does stand in a midway point between a pure “xSTATS based WAR thing” and a “win shares” type thing, but that doesn’t make it obsolete.
Win Shares will be better at saying stuff like:
“If not for the contributions of Player X, we would have expected his team to win Y fewer games”
xSTATS will be better at saying stuff like:
“Player X hit tons of hard line drives to the gaps”
“We would have expected a player with this batted ball profile to produce X runs”
“we expect a player with this batted ball profile to produce X runs next year”
but fWAR will still be the best at saying stuff like:
“Player X hit tons of doubles and Home runs”
“we would have expected the results of Player X’s “events”(PAs, Defensive plays, baserunning, etc.) to produce Y runs”
“we would expect a player who produced these results to produce X runs next year”
both xSTATS and fWAR are on the runs level as opposed to the team wins level, but as far as I know, we aren’t sure whether batted ball results are more/less/equally as predictive as actual PA results.
There are 5 different digital inputs in baseball, and we have all kinds of different ways of measuring impact on all 5.
You either make an out (or create an out), or you do not
You either score a run (or allow a run), or you do not score a run
You either win a game, or you do not win a game
You either make the playoffs, or you do not
You either win the championship, or you do not
There is no one “correct” input to measure, and the ability to look at each one from multiple viewpoints is valuable. Even the most egregiously nonsensical stats like RBI, pitcher wins, OPS, and Saves still have at least some value today (as well as good reasons they became popular).
In general though, all of these types of large sample size valuation metrics are more fun tools than useful analysis tools though.
Particularly when you are in a front office, you: 1. need to identify value earlier than other advanced actors, otherwise its pretty useless, and 2. have access to WAY more data, particularly with regard to periods and severity of player injuries. This means that the real goal is to identify “true talent level” and how it fluctuates over relatively short samples (maybe a week or two of games). i.e. Player X has a wrist injury, now his “Power” falls from a 60 to a 35 for 6 games… etc.
Once you know that, you can ask yourself questions like
“can we prevent the wrist injury”
“can we easily improve the true talent level through something we have identified, i.e. increase pitcher velo by tweaking delivery”
“what is his replacement level on our roster if we place him on the DL, and do we have the flexibility to be able to move him back to the roster once we bring up his replacement”
“what types of game situations will he thrive in, and do his strengths compliment other players’ weaknesses”
etc.
When you try to answer any kind of real question, knowing that a player was worth X runs or Y wins or Z playoff odds, even if an extremely accurate measurement, doesn’t tell you much without breaking it down into it’s constituent inputs.
James has spent his life coming up with ways to say “this Yankee sucks.” Pathetic.
I think it’s clear that James is right. What does a baseball team try to do? It tries to maximize wins. Numbers are only relevant to the extent they indicate how a player added to a given team’s win total, not how they “should have” added to a hypothetical, mathematically-modeled team’s total. Context is therefore essential, as it is anywhere in life. It’s funny how people say, “James is right, but only for MVP voting”. Isn’t that what all statistical measures should try to get at-a given player’s value, whether he’s an MVP candidate or not? Also, on repeatable numbers as indications of true skill, and the looking backwards or forwards question: Roger Maris never came close to hitting 61 HR after 1961. Do we then say his power in 1961 wasn’t real? What if I just want to look at 1961, or 2017, in themselves? The Astros 2017 WS win won’t disappear if they win 75 games in 2018. You can’t say Denny McLain “wasn’t really” close to Gibson in 1968 because Denny didn’t keep pitching at that level.
No, because if you were signing a player to a contract for next year, you’d be better off looking at WAR and the elements that make up WAR than any context specific stats. It really depends a lot on the question you want to answer.
“Who had the most valuable season in the past year?” is much different than “Who are the best players for next year?”
Both have their use.
Does a stat yet exist that tries to somehow combine the two somewhat competing goals of talent vs value? Say, c(ontextual)WAR, where each event’s linear weights value is multiplied by some form of leverage? Or is that just what WPA does, in different form? Humanities major here, sorry.
To ask a rude question-Why does skill or talent matter-in itself? The only thing that’s important is valued added towards winning games-the MVP question. Somebody above called James’ point “trite”, which is a quality many obvious truths have. It may be harder to calculate this value than the WS system would have you believe, but James is focused on the right question.
“The only thing that’s important is valued added towards winning games”
I dispute this. A baseball team has 5 digital goals it’s interested in achieving over the course of a season (which is the only time restriction stipulated in the MVP award)
Make Outs/Avoid Outs
Score Runs/Avoid Runs
Win Games
Make the playoffs
Win the World Series
The MVP award excludes the 5th by design (voting occurs before the result is determined). This leaves 4 remaining. Without any identifiable distinction between them, they should all be considered. If you want to argue that only one should be considered in “value”, then I can imagine an argument for playoff odds, i.e. it’s the biggest goal we have. I can also imagine an argument for Runs, i.e. it’s the smallest denomination which is necessary to achieve the biggest goal, making the playoffs. And I can imagine an argument for Outs, or an Outs/Runs blend, i.e. its the only denomination that is entirely within any individual player’s control. But I cannot come up with a valid argument that Win probability should be counted in “value” to the exclusion of all other inputs.
I don’t think the post or the comments deal directly with James’s criticism. WAR is supposed to measure “wins” above replacement. If you add the WAR of Aaron Judge, Gary Sanchez, Didi Gregorius and the rest, you get 102 wins. (Similarly, the Yankees BaseRuns for 2017 was 102 wins.) But they only won 91 games. The sum of the parts, in this case, significantly exceeds the whole. That is problematic. The 2017 Yankees were outliers. No other team had an 11-win gap between BaseRuns and their win total. This suggests to me that the WAR of individual Yankees may be inflated. It may make sense to adjust a player’s WAR to align it with the number of games a team actually won. I know that WinShares is a maligned statistic, but adjusting Judge’s WAR to account for the actual number of games his team won would arguably better reflect his contribution. For most teams in 2017, this adjustment would be small. 26 teams were +/- 5 wins from their BaseRuns total. Linear weights are adjusted each year based on the number of runs scored during the season. Why not adjust WAR to reflect the number of games a team actually won?
Don’t get hung up on the word “Wins”
You don’t change a tool to fit the task you want to accomplish, you just use the right tool.
This argument is akin to “why is this called a tack hammer if I’m using it to pound in framing nails?”
WAR is an effort to approximate an output: wins. Unlike traditional statistics that measure inputs (batting average, RBI, slugging percentage, etc.), WAR seeks to estimate the results of these inputs. To do so, it posits a linear relationship between various traditional statistics and runs. It goes on to posit a linear relationship between runs and wins.
As Dave Cameron acknowledges, those relationships really aren’t linear. There is a good deal of noise in the statistic. Generally speaking, that noise is muted. As I noted above, most teams are within 5 wins of their BaseRuns win total.
But when you confront an outlier, as with the 2017 Yankees, where the disparity between BaseRuns and wins is quite large, red lights should flash.
What you do about that is open to discussion. Adjusting Judge’s WAR to reflect his team’s wins is one way to go. That grates a bit on the people who developed WAR because it was a source of their criticism of WinShares. That’s fine.
If you can’t bring yourself to do that, then you should at least apply additional scrutiny to players on teams who underperformed (or overperformed) relative to their WAR.
More importantly, James’ criticism should inspire an inquiry into whether WAR over (and under) rates certain players, with an eye to refining the statistic or developing a new one.
WAR is an effort to approximate an output, Runs above replacement level. To do so, it posits a linear relationship between various inputs and runs. It then scales this run value estimate to represent some version of expected wins given the number of runs produced. The “Wins” aren’t the point here, the “Runs” are. WAR is a “Run” level value estimator.
What you do about that is not really open to discussion, you can take the snapshot that WAR (fWAR, bWAR, RAWAR, xFIPWAR, etc. there are lots, BECAUSE they measure different things) gives you, and then you can either apply it correctly, or incorrectly. Only by understanding what it is telling you, can you use it correctly, wishing it told you something else will always lead to the confusion you currently have.
It’s being refined every day, but that doesn’t mean you take away the old variations. (See, xSTATS WAR)…
Pitcher wins contain the word “Win”
WPA contains the word “Win”
Pythagorean Win % contains the word “Win”
None of these statistics has perfect correlation with team wins either.
I realize this isn’t the crux of this discussion, but I’m curious what others think:
Which stat would you use to evaluate players from across history? Any stat, not simply WAR or Win Shares.
Dave and others – what do you say about Baseball Prospectus’s WARP? From what I understand, this stat uses a “mixed methodology”, ie combining both context-neutral fixed outcomes and luck/clustering//clutchiness/randomness. Could we compare fWAR, bWAR, and WARP over a given time frame to see which one produced the best predictive value (of future performance)?