I Think Win Probability Added Is a Neat Statistic

We’re in a tiny lull in the baseball season, and honestly, I’m happy about it. July is jam packed with draft and trade talk, September and October are for the stretch run and the postseason, but the middle of August is when everyone catches their breath. There’s no divisional race poised on a razor’s edge, no nightly drama that everyone in baseball tunes in for; it’s just a good few weeks to get your energy back and relax.
For me, that means getting a head start on some things I won’t have time to do in September, and there’s one article in particular that I always want to write but never get around to. I’m not a BBWAA member, and I’ll probably never vote for MVP awards, but I spend a lot of time thinking about them every year nonetheless. When I’m looking at who would get my vote, I take Win Probability Added into account. Every time I mention it, however, there’s an issue to tackle. Plenty of readers and analysts think of WPA as “just a storytelling statistic” and don’t like using it as a measure of player value. So today, I’m going to explain why I think it has merit.
First, a quick refresher: Win Probability Added is a straightforward statistic. After every plate appearance, WPA looks at the change in a team’s chances of winning the game. We use our win expectancy measure, which takes historical data to see how often teams win from a given position, to assign each team a chance of winning after every discrete event. Then the pitcher and hitter involved in that plate appearance get credited (or debited, depending) for the change in their team’s chances of winning the game. Since every game starts with each team 50% likely to win and ends with one team winning, the credit for each win (and blame for each loss) gets apportioned out as the game unfolds. The winning team will always produce an aggregate of 0.5 WPA, and the losing team will always produce -0.5, spread out among all of their players.
That’s just tremendously neat. As the glossary entry for WPA puts it, “We know intuitively that a home run in the third inning of a blowout is less important to that win than a home run in the bottom of the ninth inning of a close game.” We do know that! But the argument against WPA is that thinking about the game that way doesn’t match what actually matters.
Here’s an example. Let’s say that the Giants are up 5-0 in the third inning when Joc Pederson hits a two-run home run. Win Probability Added: minimal. Then, the Giants cough up the lead and trail 8-9 headed into the ninth inning. Joc blasts another two-run bomb, this one decisive. Win Probability Added: massive. But since the final score of the game was 10-9, taking either two-run home run off the board would result in the team losing 9-8. Why are they treated differently?
That’s a compelling argument. If you imagine that the base/out state was the same both times, it gets even better. Both home runs were necessary for the Giants to win the game. One gets treated as nearly worthless by WPA, though, while the other is worth its weight in gold. You can make it feel even more unjust if different players hit the two homers. Joc the Irrelevant, Yaz the Hero? What, by virtue of when they did an equally important thing? Sure seems arbitrary when you put it that way.
Given how many baseball statistics there are in the world, you could account for that if you wanted. There’s WPA/LI, which adjusts every outcome for the leverage going into the plate appearance, so that how you perform relative to what was expected in each situation is what matters, not how important the situation was. RE24 uses run expectancy rather than win expectancy, so everything is on the same scale.
I don’t find those arguments compelling, however, because I think they misunderstand the contingent nature of a baseball game. Runs aren’t created equal. Timing matters. The game unfolds differently based on what has already happened; a team might put in their mop-up guy or go to their closer based on the game state. To reduce the argument to absurdity, consider last weekend’s Mets/Braves tilt. The Mets were down 13-3 heading into the ninth inning and thus sent a position player to the mound. The Braves promptly scored eight runs. Were those runs equally as important as the first eight of the game? I can’t imagine making that argument in good faith.
Here’s another way of looking at it. Imagine, if you will, that the Giants were on the road in our initial example. Further, imagine that they gave up a two-run bomb of their own in the bottom of the ninth to lose 11-10. Were Pederson’s two home runs each worthless to the outcome? Did they go from hugely valuable to of no import because of that subsequent event? That doesn’t feel right either.
The future is always unknowable. In my opinion, that means that evaluating one plate appearance based on how the game unfolded afterwards misses the point. Every time a hitter comes to the plate, all they can do to best help out their team is maximize that plate appearance. WPA handles that quite well, because it explicitly doesn’t care about what happens afterwards.
From a predictive standpoint, none of this matters much. A home run is a home run is a home run; you’re not going to get anywhere by treating different ones differently if you’re trying to figure out how good a player will be in the future. Decades of research have hammered that point home. That’s also true if we’re trying to measure a player’s underlying talent; there’s no evidence that hitters control when they hit their home runs. We all pretty much know that; there’s a reason that the single-season home run record is so famous while no one cares about “number of runners driven in via home run.”
If you want to know who the best player was in a given year, I think WAR answers that question pretty well. Is that what the MVP award is for? I don’t read it that way. Per an FAQ on the BBWAA website, the award considers “(a)ctual value of a player to his team, that is, strength of offense and defense” in addition to other clauses about games played, character, loyalty, and effort. To me, “actual value” carries a connotation that the particular circumstances of each event matter. What’s the “actual value” of a hit or a strikeout? The context in which it occurs surely has to matter at least somewhat.
How does that relate to WPA? I think it’s almost a direct translation. Players can’t control the situations they find themselves in; that’s one of the neat things about baseball. The batters in front of a player determine the base/out situation they face. All a player can do – all that’s in their control each time they step to the plate or face a new batter – is increase their team’s chances of winning that game by as much as possible.
If that sounds a lot like WPA to you, then you’re thinking about this the same way that I am. WPA doesn’t care about how you got there. It doesn’t care about what happens afterwards. It bores in on the individual situation and nothing else. How much actual value did a batter provide? I can’t think of a better way to encapsulate that than by starting with how the game looked before they batted and finishing with how it looked afterwards. Whether the team came back later, whether some future event cheapened or heightened their earlier contribution – that isn’t what we’re talking about here. How much did a player help his team? For me, that’s a close corollary to how much win probability they added.
I don’t mean to say that I won’t consider anything else when looking at who deserves to take home hardware at the end of the year. Sure, WPA sounds a lot like the MVP criteria to me, but it’s not a perfect match. “Actual value” is purposefully nebulous. A ton of home runs is a ton of actual value, even if a team squandered that value by not having baserunners on to capitalize. The same is true for someone who reaches base a bunch; if they end up disproportionately doing it in unfavorable spots because their team doesn’t cooperate, it’s hard to blame the player for it.
Quite frankly, a big disagreement between WAR and WPA doesn’t come up very frequently. This year’s WPA leaders in each league? Shohei Ohtani and Ronald Acuña Jr., the two MVP favorites. Last year’s? That’d be Aaron Judge and Paul Goldschmidt, who both took home the trophy. In 2021, Ohtani and Bryce Harper led their respective leagues, and both won MVP. For the most part, this statistic is telling us what we already know.
I hesitate to mention it, but WPA comes close to making me understand why RBI are still considered a key statistic by a lot of baseball fans. We know, again thanks to decades of research, that RBI aren’t a particularly skill-intensive statistic. They depend a lot on context; who’s on base when someone steps to the plate matters a lot more than the skill of the batter. But if you’re wondering who contributed to a game’s outcome, they’re undeniably important. You can’t win without scoring runs, and RBI inarguably produce runs. That’s why people still love them even though they aren’t predictive of future production or even descriptive of current talent.
In some senses, WPA is just a sharper way of measuring what RBI were attempting to capture. Driving that run in from third base with a sacrifice fly really does have value; it truly isn’t the same as a strikeout, at least in terms of winning the game at hand. But that’s less impressive than a solo homer in a different situation, or even a single to drive home the runner, and WPA can handle that range of outcomes much better than a single binary statistic (did the runner score or didn’t he?). WPA also considers driving a run home from second more valuable than driving one home from third, and doing so in a close game more valuable than doing it in a laugher. It’s what RBI fans think their statistic does. WPA also captures the other side of the coin, setting the table for future hitters, which is an equally important part of winning, and it handles it much better than a raw count of runs scored would.
So if you’re reading this and you’re an MVP voter, here’s my plea: take a look at win probability numbers when you’re compiling your ballot. It probably won’t change your vote, because as I’ve already mentioned, it mostly mirrors how MVP voting already goes. The best players tend to add the most win probability because, well, they’re the best players. But in corner cases and down-ballot tiebreakers, looking at who actually did the best with the opportunities they were given deserves a spot in the conversation. Don’t forget to sprinkle in a little bit of accounting for defense, as WPA only gives credit and blame to the pitcher rather than the fielders behind him, but how much defense matters in MVP voting has always been in the eyes of the beholder anyway.
If you’re reading this and you aren’t an MVP voter, I’d basically ask you the same thing. When you’re thinking about who helped their team out the most this year, spare a thought for WPA. It isn’t always the best at telling you who will be good next year. It isn’t always the best at telling you who was the most talented this year, even. But when you’re wondering who helped their team out the most – who came to the plate down and left ahead, who chiseled into deficits and slammed the door on leads – WPA does a great job of explaining exactly that.
Ben is a writer at FanGraphs. He can be found on Bluesky @benclemens.
Devin Williams +11.42 WPA is the highest of any pitcher, starter or reliever, since 2020.
Jordan Romano (+10.28) and Zack Wheeler (+9.05) are the only other pitchers over plus nine.
That highlights the exact problem with WPA though. Josh Hader has been objectively better than Devin Williams in the innings pitched. It isn’t Hader’s fault that his team hasn’t had leads in close games to bring him in to defend. The Padres are a much better team than the Brewers, but because they are, their average margin of victory is way higher, meaning Hader doesn’t get high leverage opportunities. Williams uniquely benefits from his team being almost exactly league average – they are not good enough to have a bunch of huge leads, but still good enough to have some leads.
But that’s the point of the stat, no? Williams has been in tougher situations so has helped his team win closer games, while Hader is not adding as much value by pitching in lopsided games or getting saves with 3 run leads. Therefore, Williams gets more credit via WPA, even if Hader pitched better via a WAR/context-neural view.
Yes, and that point is why the stat is inherently flawed to the point of being almost meaningless for player-level comparisons. Obviously neither Hader nor Williams are going to win the MVP Award, since I think we have all realized that the best relievers are nowhere near as valuable as even a good but not great position player or starter. But using WPA in any kind of comparison of skill is weird, since the context that fuels WPA has nothing to do with the player.
Josh Hader has been a better reliever than Devin Williams this season. If you have a one run lead with the season on the line, Hader is the guy you want up there, not Williams. The fact that Devin Williams has such a high WPA is because of his team, not because of him. If the Brewers were better OR worse OR just less lucky from a run sequencing perspective, Williams would have a much smaller WPA. He has benefitted from a perfect storm of circumstances outside his control – a team that is ok but not good and has had well above average luck with the sequencing of runs, creating a plethora of save opportunities with fairly small leads.
“But using WPA in any kind of comparison of skill is weird, since the context that fuels WPA has nothing to do with the player.” I’m not arguing that WPA is a measure of skill. It’s a measure of results/performance, which is precisely what people consider for MVPs. WPA also indeed has something to do with the player, since his performance in each event is part of the result. I agree with others in this comment section that the apportionment of the WPA to the right player is likely due for refinement, but there is some truth to the direction it gives.
I’m not disagreeing that Hader is the better pitcher based on pure skill or context-neutral performance. But all WPA is measuring is what the guy did with the situations presented to him. It’s not the end-all-be-all.
WPA is not a measure of individual results/performance though. Josh Hader has performed better than Devin Williams and has posted better results than Devin Williams. Individual performance and results are by definition context neutral – Shohei Ohtani is still the best player ever despite his team losing more than it wins.
WPA is a curiosity, not a serious attempt at measuring player performance.
Josh Hader has pitched better than Devin Williams. What WPA tells us is that, if each player’s appearance had been replaced by an average one in the circumstances the player entered, Devin Williams’s teams would have lost more games than Josh Hader’s would have. It’s not crazy to say that Williams has been more important or valuable to his team than Hader has been to his.
Would Devin Williams’s teams have won even more games had Devin Williams been Josh Hader? Likely yes. And vice versa. Hader’s teams would not have won as many games had he been Williams.
There’s a parallel for championships/playoffs.
The 1927 Yankees won the AL by 19 games. Both Ruth and Gehrig had about 12 WAR a piece. The 1927 Pirates won the NL by 1.5 games. Paul Waner had about 7 WAR.
You replace Ruth with a replacement level player and the Yankees still win handily. You replace Waner with one and the Pirates finish 3rd.
To the question: who put up the better performance? There’s no question: Ruth and Gehrig. Whose absence would have made their team more likely to miss the World Series? There’s also no question: Paul Waner.
As long as the award is the “Most Valuable Player” there’s ambiguity around value to whom, value for what? WPA is helpful in answering one perspective on that question.
(I’m glossing over the issues with WPA wrt fielding and so on–it’s useful, but not perfect.)
Your observation is that Devin Williams has pitched in significantly higher leverage than Josh Hader, which is true, but also not particularly meaningful in the context of whether or not WPA is a “serious attempt at measuring player performance”. Hader has been generally babied by Bob Melvin, and his lack of well-planned usage in high leverage is probably one of the main reasons why the Padres are underachieving so much. Additionally, this phenomenon really only exists for relievers, as managers will tend to bring in better relievers in higher leverage. The deviation in leverage index amongst hitters is absolutely minimal and, while it should be considered, makes little impact and is not something to constantly fixate on whenever someone brings up WPA.
Regardless of the reasons as to why Hader has been put in lower leverage than his reputation calls for, which would be interesting to dive into further, none of that reflects on the validity of WPA as a measurement or a concept. It very much is as advertised: it measures the amount of win probability added in plays that player X was directly involved in. The player has complete control over the result (in the same way he controls his more biased context neutral stats); it’s just a more variant but therefore more reflective statistic.
“The Padres are a much better team than the Brewers”
Since the trade deadline in 2022
MIL: 94-86
SDP: 87-90
Since 2020 (time frame quoted in original comment)
MIL: 275-230
SDP: 263-242
Since 2017 (Hader’s debut)
MIL: 546-446
SDP: 470-521
The Padres are a much better team than the Brewers at spending money and grabbing headlines.
There might not be a greater difference between people’s perception of success and actual success than the San Diego Padres (I love the FG team/thought process/community, but woof, the blind spot for SDP is egregious at this point)
(edit: Manny Machado has a negative WPA this season)
The Padres’ Base Runs record is 9 wins better than the Brewers. Their team wRC+ is 18 points higher. Their FIP is .15 runs lower. The only thing they don’t do better than the Brewers is field, but that is more than made up for by the offensive difference, which is massive.
Well, fielding and closing out games.
2023 Brewers bullpen WPA: +7.16 (1st)
2023 Padres bullpen WPA: -2.80 (27th)
There’s the gap in BaseRuns.
They also don’t win as many games as the Brewers, which is a key data point that needs to be addressed in all those things the Padres are apparently great at doing
I really enjoy this site but posters like ascheff exemplify an attitude that i find maddening. There is nothing – nothing – more important than the results on the field. I guarantee that no player, manager, front office, or owner finds consolation after a real life loss by telling themselves “we were better statistically.”
The 1960 NY Yankees destroyed the Pirates in the World Series that year. They outscored them more than two to one: 55-27. The Yankees OPS was .911! Pittsburgh’s was .656. The Yanks out-homered them 10-4.
Mickey Mantle’s batting line was .400/.545/.800/1.345 OPS with 3 HR. Do you think that helped him feel better when the Yankees lost Game 7? “Who cares that they beat us? We were clearly the better team!”
You know what they say – statistics are for losers.
This argument boils down to what is more predictive vs. what is more descriptive. It is often about using the right tool for the right purpose.
“Consolation” means a lot of things to different people at different times. I imagine (almost) all players and managers are going to feel better about a competitive loss where they put up better stats versus one in which they were blown out and uncompetitive.
That having been said, there are times when you have to simply realize the end results aren’t there and act accordingly rather than continuing to wish-cast about the “back of a guy’s baseball card” (see: 2023 New York Yankees).
I’ve heard lots of derogatory sayings about statistics, but I don’t think I’ve ever heard “stats are for losers”. That one might be even worse than “count the rings”.
I believe that quote comes from Scotty Bowman, the legendary coach of the Montreal Canadians.
“Statistics are for losers. New York and Boston each has six players among the top fifty in the league. We only have three. See what I mean?”
That year the Rangers and Bruins were the favorites in the NHL. Of course, the Canadians took home the Stanley Cup.
Seems like the main thing they don’t do better than the Brewers is win. Damn. Isn’t so inconvenient that they have to actually play the game on the field? That’s where it’s matters and that’s where they have been found wanting.
Amen
“The Padres are a much better team than the Brewers…”
….much? better?
Hmmm
It’s weird to me that reliver WAR has a context component
It’s weird to me too, but I sort of get it.
It sort of shouldn’t.
I cannot believe we are still talking about WPA as though it makes any sense as a measure of “value”. WPA is a game-level statistic that tells us how exciting the play is and not anything else. It tells us how much the game swung in one direction rather than the other. It’s literally a quantifiable measure of excitement or fun. If you’re going to vote for MVP based on fun, okay but let’s not pretend this has anything to do with WPA criteria.
There are two minor problems with WPA and one big one. First, pitching and defense are totally confounded (Ben thankfully doesn’t advocate WPA being used for pitchers here, so that’s good). Second, there’s another problem that often the runners on base are the ones who provided as much value as the ones at the plate (some might say more).
But the biggest problem is that WPA is calculated at the time of the event only. This causes all kinds of problems. Let’s say the Cardinals win a game 5-4. Paul Goldschmidt hits a solo home run in the first inning, and Nolan Arenado hits a solo home run in the ninth inning (the go-ahead run), and everything else is the same. Which one helped their team win more? Obviously it’s the same*. They were responsible for the same number of runs, and the Cardinals wouldn’t have won the game without either one. This is obvious. WPA gets it wrong; Arenado gets way more credit than Goldschmidt.
(*If anything, runs scored earlier in the game are more valuable because they convince the opposing manager to pull their starters and also not to burn their best relievers in a losing effort. What happens in the 9th inning doesn’t influence opposing manager behavior)
You could calculate a WPA-like statistic that would give them the same number; at the end of the game, weight a player’s performance more based on the final margin of the game. This would actually get close to what advocates of WPA say it’s doing. Dave Studebaker has a version where the weight of something in a one-run game is about 1.38, and in a 9 run game it’s about 0.51. Use that. Calculate that. Don’t mess around with WPA which is conceptually about fun.
That said, the most important thing in WPA is the base-state. This is addressed neatly by RE24. It rewards players who performed better with runners on base, but doesn’t get into the nonsense of deciding that a run scored in a tie game in the 9th inning contributed more to victory than a run scored in a tie game in the first inning. There is no evidence that RE24 is repeatable either but that’s not necessary.
If timing doesn’t matter at all outside base-state as I read from your comment, a team should be willing to let another team to rearrange runs scored in all the games they played together in a season. I suspect the another team would be able rearrange the runs in a way that would cause their number of wins and WPA to go up while keeping WAR and RE24 the same.
In other words, yes, it is possible to create scenarios that WPA doesn’t matter because of what happened after the fact or beforehand. It just captures how important an event was at the time (without measuring defense, etc) and giving all the credit to one person (I think giving all the credit to one person is more a problem, than the timing component). That said, I’ll take a 3-run homer in the bottom of the 9th down 2 in one game and a K down 10 in the following game than a K down 2 and a 3-run homer down 10 in the bottom of the 9th anyday.
The original argument was that the timing of runs within an individual game are irrelevant. Meaning that a 5-2 victory with all 5 scored in the 1st inning is no different from all 5 runs scores in the 9th. This is very different from what you make it out to be, in which I could choose for my team to lose 700-0 in game 1 and then go on to win the next 161 straight, with a run differential of 0.
The argument that I would choose to “spend” my runs differently at the time is true but ultimately the same argument as doing it in retrospect, except in the ninth inning instead of the first. It doesn’t make that much of a difference. If I could determine exactly how many runs I needed to score in the 9th inning and score exactly that amount after not spending any in the first eight, yes I would do that. That doesn’t change the fact that mathematically it doesn’t matter when it was scored.
The main problem is that WPA – as – value wants to use context even when it is irrelevant.
Context is not irrelevant: you only can do what you can do at the time you have the opportunity – no one knows the future, and WPA only gives credit relative to what is known at that point in time and what could statistically occur in the rest of the game. Using 20/20 hindsight to rejigger the value of the game’s results is nonsense. Hitting a homer to go up 1-0 in the top of the 1st is less valuable because the other team has 27 outs left to overcome it.
This is nonsense. The runs count the same. The outcome of the game is determined by how many runs you scored. Thus, the contribution to winning is the same. Anything else is irrelevant.
It doesn’t matter that it isn’t possible to know the future, because it is possible to know the past. The MVP is not handed out in the middle of games. Just because something feels more exciting does not mean it contributes more to winning.
You may not like context because it is not repeatable, but jacksonv123’s 161-1 team is in the postseason with a bye and a team that goes 81-81 assuming average context is not in the postseason despite the same amount of WAR and Re24. It may not matter in saying this player is better than that player as I think the problem with WPA is how it is credited to a player and not a team.
Timing is relevant to wins on an intergame perspective even if on an intergame, artificial box view, when the 5 runs scored in a 5-4 game is irrelevant. On a season scale, when the runs scored does matter as 3 runs matter a lot more in the 5-4 game than the 10-0 blow out. 3-run homers in the 9th matter a lot more down 2 than they do up 7.
In other words, the context of WPA may be a lot of luck and not matter regarding when runs score within one game, but across games, when runs score matter regarding wins and losses (not necessarily predictive value) and WPA captures this on the team level even if attributting WPA to an individual batter or pitcher only may not be the best.
Ignore luck. Ignore repeatability. None of that matters. What matters is math. You do not get more runs for scoring later in the game. If you subtract the homer in the first it would hurt you the same amount as the homer in the ninth. Maybe even more.
“Maybe even more”. Interested in how to square that last sentence based on your position above.
Probability is also math. It’s not only arithmetic.
The probability of events that have already happened is 1 or 0. The runs scored is a 1 or 0. You wouldn’t hand out an MVP based on preseason projections, would you? No, you would do it based on what happened in a game.
This completely misses the point – its the probability, at a certain point of the game, of winning said game. It’s predicting the outcome of the game based on a change in the game state.
You wouldn’t win the MVP of a particular game based on a first inning HR unless that HR was integral in the outcome of the game. This is pretty much true of ASG MVPs.
I would agree that you wouldn’t win an MVP in that circumstance either. That’s why you would wait until the end to figure out if it matters.
Ultimately it doesn’t matter what the prediction is, because we know the outcome before we vote.
Respectfully, did you read the article? The 2 HRs in different gamestates argument is addressed explicitly.
That doesn’t mean you have to like it or be persuaded by it, of course. You may still disagree! But this article explicitly addresses and responds to your argument and you should at least consider that response.
I saw that he acknowledged it (well a version that was close enough) but I did not see any arguments against it. The argument against it is “not adjusting is worse.” Except sometimes it is! If contingency matters, it’s in the opposite direction. It’s a bad argument.
But it’s not actually addressed! It’s just mentioned and then waved off as “not feeling right.”
He literally spent 5 paragraphs parsing this point and providing scenario examples before the comment you are referring to. The argument against it is summarized therein, including here:
“I don’t find those arguments compelling, however, because I think they misunderstand the contingent nature of a baseball game. Runs aren’t created equal. Timing matters. The game unfolds differently based on what has already happened; a team might put in their mop-up guy or go to their closer based on the game state. To reduce the argument to absurdity, consider last weekend’s Mets/Braves tilt. The Mets were down 13-3 heading into the ninth inning and thus sent a position player to the mound. The Braves promptly scored eight runs. Were those runs equally as important as the first eight of the game? I can’t imagine making that argument in good faith.”
But that example is bad and makes no sense. It is an example of why scoring runs at the beginning of the game is better—because it induces the other team give up earlier. And aside from that it doesn’t actually address the issue.
Agreed, that’s not the best example to highlight where WPA reflects the changing value/impact of an event, as we better understand, say, walk-offs as having the bigger impact and therefore higher WPA. There are diminishing returns, per WPA, of tacking on more runs in a blowout, is his point.
I don’t think Ben is saying getting a big lead early is worthless – that lead by the Braves thru the 8th gave them a 99.9% win probability, and they picked up 19% in the 4th inning alone. Those early runs generated value and the hitters got credit for it accordingly, considering the base state at the time. In that sense, you are right, it doesn’t matter when you score, it matters how much you contribute to getting a win, and you can contribute more per PA by generating big early leads or big plays late in (typically, close) games.
A lot of folks here are talking about the strategy aspect – a big early lead makes the opponent take out their starter, or a big lead can avoid using your own best relievers. That’s all well and good, but that’s not what WPA is trying to measure, so lets be careful not to ascribe strategy implications here. What WPA typically shows is the plays with the biggest WPA generation are ones where a team makes a play that significantly alters the potential outcome, and often leaves the opponent with little or, in the case of a walk-off, no recourse.
This argument boils down to whether teams use their good relievers or not. Maybe that makes a tiny difference. With no attempt to measure that difference, it’s not a very good argument. A better solution would be to just exclude all PAs against non-pitchers.
All of these arguments boil down this example. A player hits a game tying grand slam that takes his team from 5% to 50% likely to win the game. The teams are in a hurry, so they decide to flip a coin to determine who wins instead of playing it out. Do you need to know the outcome of that coin flip to know how valuable that play was? You say yes, I think you’re wrong. I’m not compelled by your arguments and you clearly aren’t by mine.
Y’all either settle this in the gas station parking lot or (more likely) with a grand and glorious POGS battle.
Loser loses their POGS, loses this argument, and buys Yoohoo for everyone
I don’t accept POGS duels. I do, however, accept duels by Jenga, Croquet, water guns, arm-wrestling, and who can hold their breath the longest.
So best of 5, then.
Only if you calculate WPA along the way and use it to determine the winner rather than how many times a person actually won.
I think this exact 5-event battery was on ESPNOcho last weekend
Undercard to the Elon Musk-Zuckerberg fight is Ben Clemens vs Sadtrombone
.
Make it happen!
So what you’re saying is…it will never happen.
Given your summary (and your article), I’m not sure you actually understand his argument well enough not to be compelled by it. Let’s return to the simple point about a 5-4 game with a low-WPA solo homer in the 1st and another, this time high-WPA, solo homer in the 9th. (If you want an even clearer case, you can stack the deck against WPA a little more and put the 4 runs scored by the losing team in the top of the 1st, so the solo homer in the 1st will have a really low WPA compared to the walk-off homer.)
Do you actually get the problem with WPA as a measure of value (rather than coolness/fun/etc) in this scenario — which is that both homers were equally necessary to the eventual win? All the woolly-headed hand-waving in this article about time being linear, the future being unknowable, the nature of contingency, etc., really just seems like a way of dodging this question.
Love the casual condescension, but I think you too are missing the point. You’re talking about contingent value. Both home runs happened to be necessary to the win in that game, but lots of home runs that are hit in the first inning don’t end up mattering. Meanwhile, solo homers in tie games in the ninth end up changing the outcome a lot.
When you’re talking about the value that a single player provides to a team, the outcome of that coin flip can’t be relevant. That’s the ‘the game ends 5-4’ part of it. We don’t know that at the time. It seems very clear to me that you think a home run to go up 5-4 in the eighth is worthless if you end up losing 15-5. I don’t think that. I don’t think that future events squandering the value a player created change the fact that the value was created. When you tie an individual event’s value to things that happen after and which are uncorrelated, I think you’re making a big mistake.
It’s not hand-waving. Your definition is inherently incapable of measuring individual player contributions independently of the rest of their team. So sure, were the two home runs equally necessary to win the game? Yeah. But did the two change the likelihood of the game’s outcome by the same amount? Very clearly not. I’m capable of understanding that these are not the same thing.
WPA is inherently incapable of measuring individual player contributions independently of the rest of their team because the actions of the rest of the team determine the initial win probability and leverage index of a plate appearance.
When you tie an individual event’s value to things that happened before and which are uncorrelated, I think you’re making a big mistake.
Exactly! The apparent blind spot here — and based on these responses it’s a true Westworld “I don’t see anything there” kind of blind spot — is that WPA isn’t actually measuring “individual player contributions.” Everything that goes into the “context” part of the walk-off homer’s WPA is also an aggregate of past individual players’ contributions, which aren’t being credited to them.
Yep, this here is my issue. I like real world examples so I’m looking at the August 5th, 2001 game when Cleveland came back from a 14-2 deficit to beat Seattle 15-in extra innings.
Cleveland slowly chips away at the lead (which of course WPA says doesn’t matter) but is still down 14-9 in the 9th inning. They cut it to 14-11, have the bases loaded and two outs. Omar Vizquel then hits a triple to tie the game up.
WPA gives Vizquel +47% whereas the other 5 batters who got on base get +1-3%. Sorry I’m not buying that. Here were not talking about distant events during the game. We’re talking about events that happened IN THE SAME INNING. Yet WPA says that only one of them mattered, even though Vizquel doesn’t even get to bat in the 9th if the other 5 don’t get on base before him.
Here’s another real world example. On May 17, 1979 the Phillies beat the Cubs 23-22. Dave Kingman, playing for the Cubs, hit 3 home runs and drove in 6 runs (plus drew a walk). Yet he ended up with negative WPA because he made 3 outs in crucial situations whereas all of his homeruns came with the Cubs down by 6 or more runs.
Now let’s look at the Cub with the highest WPA for the game. It was Barry Foote. Foote certainly had a good game (3-6 with a double and an RBI) though it hardly stands out in the context of a 23-22 game. The reason why Foote has the highest WPA? It was his single in the 8th inning that drove in the run that tied the game at 22 apiece.
But now let’s look just at just positive WPA events and ignore the negative ones. If we do that, we see that Foote’s single that drove in one run has a higher WPA than Kingman’s 3 homeruns that drove in 6 runs COMBINED. And it’s not even close. I’m sorry but that’s absurd. If a measure says that a single/1 RBI is more important/valuable than 3 HRs/6 RBIs than we should definitely question the value of that measure. I would seriously love to see Ben or someone else try to defend that.
This is the crux of the problem with your argument: We don’t need to know how much they changed the likelihood of the game’s outcome. The game is over. We know the outcome. They changed the actual outcome of the game by the same amount. If you want to make it so contribution to winning a specific game matters, I can understand that (although it seems to me to be unnecessary in most cases). But intentionally using a measurement that adjusts by the wrong context? I cannot agree to disagree on this. Someone is wrong on the internet!
To extend off your example – let’s imagine in the next Cardinals game, Goldschmidt again hits a solo home run in the first, and Arenado again hits a solo home run in the ninth. But those are the only two runs the Cardinals score as they lose 13-2. In that case, WPA is going to rate Goldy’s dinger as about the same as the one in his previous game (assuming the score was 0-0 in both, and his spot in the batting order was the same in both, it would actually be identical). Arenado’s home run in the second game would be rated much, much lower.
In some ways, WPA isn’t saying that dingers in the ninth are more valuable than dingers in the first. It’s saying that there is just a much wide range of value to ninth inning homers than first inning ones. Which is pretty undisputedly true. If five times a season, a manager could add one to his teams run total, those uses would basically always happen in the ninth inning or in extras, in tie games or game where you are down 1. You’d never see them used in the first, and you’d never see them used in games where your down 12 or up 12.
Right, in that case, it would be biased the opposite way. Neither one matters that much, and it would assign much more to one than the other. I’m not sure that makes it better. You can make a pretty strong assertion that it all averages out over the course of a season but if that’s true then that kind of screws up the argument that the context matters.
If we’re going to go down this road of “well, that run was worth less because it was at the end of a blowout”, don’t we then also logically have to control for the importance of the game itself? All games need to be further weighted by change in playoff probability, therefore anything that happens in games played by eliminated teams or teams that have clinched is worth zero, etc.
This leads right back to only the teams on the playoff bubble can truly have a “most valuable” player.
It’s turtles all the way down
I agree that World Series Championship Probability Added (WSCPA) is the logically consistent way to go with this kind of metric, but I’m not sure I agree about who it would value the most highly. It’d definitely be interesting if someone actually worked it out, but my knee-jerk guess is that by WSCPA you’d end up with the World Series MVP as the league MVP almost every year. A few game-winning hits or pivotal WS innings pitched would end up swamping the entire season of star players on most non-WS teams, or even playoff teams. You’d end up reverse-engineering and rigorously quantifying the attitude of most talk-radio callers.
Baseball-reference has this stat. It’s called cWPA (championship WPA) and, as you might imagine, stuff like Bill Mazeroki’s HR in WS Game 7 1960 is absurdly high.
EDIT: The single highest cWPA for a single play comes from that game, but it is Hal Smith’s 3-run HR in the bottom of the 8th that took the Pirates from down 7-6 to up 9-7.
Context is not entirely irrelevant. Managers make decisions based on context constantly. It’s pretty much their job. Players apply more or less effort constantly, based on the existing game situation. A run scored in the 1st is not identical to a run scored in the 9th
You just reiterated the same silly argument that Ben addressed in his article. Calculating at the time of event is why it is good. Let me ask you: how do you think the linear weights for wOBA are determined?
WPA for position players only gives value for offensive component and gives all the credit of run prevention to the pitcher. Adding the WPA of a player’s pitching and hitting together that is a DH and a pitcher seems to be giving too much credit to that player. Ohtani is good enough not to need help.
That is my main objection to WPA. A version that could at least somewhat account for defensive contributions — perhaps by comparing the results of similarly hit in-play balls, allocating the change from previous to “expected” to the pitcher, then the delta from “expected” to “actual” accounted to the fielder — would go a long way to making it more “useful”. Even a emprically based “split” on plays would make it better.
I think the most egregious example is a HR robbery. Assuming that we hold the batter value at (next state – previous state), the defense would have to recieve the opposite value. But, based on what we know, without the effort of the fielder (or a mediocre fielder), the pitcher gave up a not-in-play ball, a HR, which I think most of us would assume they should shoulder the entire “blame” for. The fielder would then have to get the difference (post-home-run state – actual-next-state) to balance it out.
WPA “works” as is because it’s simple, relatively easy-to-explain (as long as you can get someone to explain the basics of WE), and can be used to “account” any game that we have full PBP for.
Chris Gilligan had a similar idea a few months ago, I think it would make WPA much better: https://blogs.fangraphs.com/what-is-a-web-gem-worth/
I’m not sure about the sabermetric value of this but it would awesome in entertainment value
If you want to account for on-field results in MVP voting, it really is as simple as swapping RE24 for wRAA in the WAR calculation.
This is a great idea. Now I’m curious how those adjusted WAR standings would look.
I took a look after the Acuna article, and it essentially netted Freeman an extra .5 WAR against the other two!
Because bats after Betts
Yeah I think this gets into the inherent problem — there’s so many different factors that need to be accounted for to determine how value should be split. How Betts should be valued for that is a big one that WPA doesn’t really account for as far as I understand, but there’s obviously going to be many more in a game that allows for as many different situations as baseball.
It’s likely that someone could come up with an incredibly comprehensive solution that effectively solves for a good portion of those issues, but there’s no incentive to do it for a purely descriptive statistic. Front offices wouldn’t care to when it won’t help them win more games, and a website like this just doesn’t have the resources to handle such a specific and time intensive project that may not actually pay any significant dividends in terms of new readers.
Imagine a great fielding play or error. Robbing a home run. A triple play. Billy Buckner in the 1986 World Series.
I still don’t understand how measuring that works in any meaningful sense.
I certainly hope that there’s nobody out there who legitimately thinks you can just measure everything with enough cameras and computers. You need the human eye to judge whether a baserunner was deaked by an outfielder or was just careless. Computers can only do so much to tell you how good a diving play is or how bad a bungled double play is.
However, humans also need to know their limitations. And I simply don’t believe any voters are going through every game and looking at every situation and assigning values to plays that computers cannot. Voters that eschew advanced stats and whatnot do so because they are emotional creatures, and they have a hunger to better understand why something is so. RBI and batting average are easy to understand and they happen right in front of you. No ballpark adjustments needed. They’re more comfortable so on an emotional level it’s easy to see why so many revert to them or cling to them.
Depends on what you mean.
The example of a baserunner being deked, for instance. Neither a full technological setup that had enough cameras and computers to record a full continuous 3D model of every player on the field and their locations (plus the ball, but that’s obvious) nor or a human observer can truly, accurately tell whether someone was “deked”, because that requires full knowledge of the baserunner’s internal mental state. The concept of deking is that the actions of the fielder misleads the baserunner into believing that the action on the field is different from reality.
Now, a relatively knowledgable person with experience watching baseball can tell with a reasonably high degree of accuracy because we develop heuristics that correlate with the baserunner being deked — perhaps the change of a baserunner’s behavior that lines up with a specific set of actions of the fielder. It’s certainly possible for those heuristics to be programmed into a program that analyzes the feeds of the aforementioned cameras and determine whether or not a baserunner was sufficiently “deked”.
The question isn’t of possibility, the question is of feasibility. I don’t think we’ve reached the point where we could cover every game to that level of detail yet, from a logistical and technological standpoint, but I think we’ll be there sooner than most people would expect.
As far as how “good” or “bad” a play is, the question is functionally unanswerable as “good” and “bad” are very nebulous, subjective terms that could mean one (or more) of many aspects: on-field value, aesthetics, emotional response (by fans, teammates, or both), etc. We already have technology that can at least do some level of comparison between plays in terms of the difficulty of execution. Aesthetics and emotional response are things that require knowledge of the viewer’s internal mental state, and thus, under at least what’s currently viewed as feasible, would be unmeasurable. The subjectiveness means there isn’t a right “answer” to the question; it’s functionally a values-based judgement of each individual person.
I think you just basically said the same thing I did, but you just sounded a little smarter saying it. 🤓
Sick article.
On the article’s point about RBI being a simple statistic devoid of nuance, has anyone ever done an analysis measuring something like an RBI-above-average rate (“RBI+”?) that measures how the player accumulated his RBIs, based on his opportunities, versus a league average hitter in the same situations? I always go back to Ryan Howard in his prime on this, where he hit in a great lineup and accrued a ton of RBIs, but once your account for his vast opportunities he looked more pedestrian.
I’ve thought about this a lot. And just keep coming back to BA w RISP.
Or I sometimes do a weighted BA.
Corey Seager has been by far the best hitter with RISP this year. (He’s absolutely crushing in general as well). You want RBIs? He’s got 73 in 78 games batting 2nd.
Yordan Alvarez, Matt Olson also pretty good.
And some guy named Luis Arraez.
The thing with Howard is that if you swap RE24 in place of his batting runs, then his WAR totals skyrocket and much better match how he was regarded and how he “feels”.
None of that is to say whether he was a *good* player or what his true talent was. Just that he did take advantage of the opportunities he had and did advance and score a ton of runners and runs in a very stacked lineup. These are different things!
Howard always brings to mind this classic piece from Dave Studeman: https://tht.fangraphs.com/postseason-probability-added/
It’s called RE24.
But RE24 still overly disproportionally rewards (or punishes) players with more RBI opportunities.
It does not differentiate between a player who was good with many RBI opportunities and a player who was exceptional with few RBI opportunities
WPA is neat.
It is literally my favorite stat.
I like it but prefer WPA/lb. Those little 130-pound David Eckstein types are doing more with less than their more rotund Rod Beck-ian counterparts
Christopher Morel is pretty damn fun!
This is absolutely fantastic and gets at how and why context stats are useful.They don’t have projective value and they’re luck + context dependent and… so what? They tell a different story for a different purpose.
Question: Would it make sense to create a context-inclusive version of WAR for position players? Because I think so!
DRS, UZR, and FRV are simply the delta in run probability for defense, no? So it would logically pair with using RE24 for position players on offense, as that’s simply run probability added. Swapping RE24 for wRAA would give us a contextualized fWAR for position players, akin to (tho not the same!) as RA9-WAR for pitchers. Makes sense, no?
Plus, it would make conversations about awards + honors (MVP, HOF, etc.) even more fun, especially for players with a significant divergence.
Yep – RE24 scales exactly to wRAA, so it really is that simple!
Enjoyed the argument and article. Reminds me of Andy Baggarly putting LaMonte Wade Jr 10th on his MVP ballot in 2021, despite his 1.6 fWAR that year. Baggarly got a little roasted for it, but in his justification article he pointed out that “Late Night” LaMonte was 5th in the NL in WPA that year, because he hit .407 with 2 outs+RISP and 13-23 in the 9th inning! “Clutch” hasn’t proven to be a repeatable skill for Wade, but it sure was valuable that year https://theathletic.com/2960813/2021/11/18/value-in-context-why-giants-shortstop-brandon-crawford-got-my-vote-for-nl-mvp/
Love it!
Winning the world series doesn’t mean you’re any more or less likely to win next year. But it does mean winning the world series.
Really interesting article. How would trout ever win MVP with that pitching staff over the years? They were out of so many games he played in….
I love the RBI, WPA, discussion!
And yet, in the three years he won it, he led the league in WPA. His performance made a huge difference to the games they won and lost.
Interestingly, the NL winner in each of those years did not lead the league. Bellinger was 2nd (Yelich), Kershaw was 2nd among pitchers and 4th overall (Stanton, Cueto, McCutcheon), and Bryant, well, uh, Bryant wasn’t even close.
Man, he was so good those years….thanks!
WPA has always been very interesting to me. For many players, it tends to fluctuate randomly, but match their overall skill over their careers. But then there are cases like Marcus Semien (111 career WRC+, -0.77WPA) and Joey Gallo (110 career WRC+, -1.80 WPA). Is this just luck? I can’t understand what has caused them to create such little WPA relative to their above average offensive numbers.
I am also interested in what WPA means for relievers. Elite closers get tons of WPA because they are almost always used in high leverage. Sometimes, they lead the league in it. Are they being massively undervalued by WAR? Are 50 innings from a closer with a 2.00 ERA more valuable than 50 innings from a starter with a 2.00 ERA?
I’m interested if anyone has any thoughts.
I think these are great questions that show actual value of WPA.
We would all prefer 50 relief innings with a 2ERA. Because nearly all of those innings will be close, pivotal moments with little room for mistakes. The starting innings might be close or could turn into blowouts.
Joey Gallo’s career is a great example of the value of context. He will punish bad pitchers. Control issues? He’ll be patient and walk. Make a mistake, hang a slider? He’ll crush to the upper deck. But good pitchers are likely to strike him out. High leverage situation at the end of a close game? He’ll be facing the team’s best reliever.
Marcus Semien may be the opposite, though. Probably just shows how luck/timing are a factor. -0.77WPA over as many PAs as he has is pretty close to average. He’s a leadoff hitter who played most of his career in Oakland – not as many opportunities for critical RBIs, and then likely to single/walk.
But it may be possible that he’s just “not clutch” or may be that he’s weak against the hardest throwing relievers.
Thanks for this great article.
I like all sorts of advanced stats, and I like WPA too. I’ll never understand why WPA is such anathema to a certain subset of stat-head fans, but I think we’ll always disagree.
There are all sorts of interesting questions that I have when considering WPA along with all the other stats for players. Is there something that the leverage adjustment of WAR misses for truly great relievers like Mariano Rivera and Trevor Hoffman, whose WPA greatly outpaced their leverage adjusted WAA? Is there something about such an extreme 3-true-outcome hitter like Joey Gallo that leads to such a consistent trend of WPA < REW < Batting Wins over his career.
I know that the answers are usually still come down mostly to random luck, which is especially tough to parse for something like WPA that probably doesn’t even stabilize to a “true talent” level over an entire career. I still think WPA is a great quantification of events on the field that can lead to some interesting inquiries though.
This is just a transparent attempt to be asked on to MLB Now, Ben 😂
There are some holes in Ben’s reasoning here:
“All a player can do – all that’s in their control each time they step to the plate or face a new batter – is increase their team’s chances of winning that game by as much as possible.”
“as much as possible” means hitting a home run regardless of context-whether close and late or blowout, a HR makes the biggest possible difference. Every other outcome can be measured against hitting a HR by comparing the RE24 of the actual outcome to the RE24 of a HR in that situation.. WPA does not tell how well the batter maximized the plate appearance (relative to hitting a HR), it tells how much the PA altered the probabilities of potential outcomes to the game.
I agree that WPA should have some weight in MVP voting. It’s the difference in seeing the MVP as the most valued player – how much I would have paid for that production if I had known about it at the beginning of the season- or is the MVP the player who contributed most to winning the games as played. The former is WAR and the latter is WPA.
Thank you, thank you, thank you. The MVP should be about the value actually produced and is decidedly NOT context neutral. If John Smith had mediocre stats overall but by some pathological chance got a hit every time he came up with a runner in scoring position then he produced immense value for his team. Similarly, if his brother Robert Smith hits 75 HRs but every last one of them was hit late in the game during a blowout loss for this team then he didn’t provide much value. I’d be much more likely to sign Robert to a long term deal but his brother John might be much more worthy of MVP votes.
I would take Robert in a heartbeat, but it is REALLY hard to compare those two guys though. John Smith had just 44 plate appearances spread over three seasons for the Baltimore Marylands, Baltimore Canaries, and New Haven Elm Citys from 1873-1875, whereas Robert Smith had four straight 1000-yard rushing seasons for the Vikings before abruptly retiring in 2000. I can’t see a good argument for John having more value regardless of the context.
“Value” is a squishy and poorly defined word.
Is it an intrinsic attribute of a player? His ability? Is the MVP the best player in the game?
Is it an economic attribute? Is a good player on an affordable contract more valuable? Or a player at a more scarce position? Must it be a guy on a contender where his contributions really matter?
Or is it what was provided, when it was provided, that directly helped the team to win? If this, well, that’s WPA more or less.
“Value” is inherently subjective, but can be measured when you have a defined goal. If the goal is to win as many games as possible, WPA gives some useful info. If its best production for the contract, then $/WAR is helpful.
I think it’s pretty easy, at least in my mind. We’re going to replay the 2023 season (and only the 2023 season, without considering contract status). We’ll use players’ actual 2023 stats and have a draft. Which player would get drafted first? That’s the MVP.
“But that’s less impressive than a solo homer in a different situation,” i have to disagree on that one.
As a close follower of the Angels, anecdotal they must have the league lead in solo homers if not an MLB record. Ohtani himself has 22 solo. And yes he’s struck out many times in crunch situations w men on base. And good pitchers limit the damage w solo homers. Meanwhile every Astros or Braves homer seems to have traffic. Is that just luck? I don’t think so.
Every time I see the Halos bust out that home run hat on a solo shot, I feel like saying’ maybe save that for a 3 run homer or something?”
Is someone working on DWPA? Based on the hit probability for the batted ball and the game situation, you could estimate the value of a clutch defensive play. Or you could just give all of the credit to Derek Jeter.
I do like WPA, and it’s good to have a look at that column (and the Runs and RBI column), but you will still sort your Hitters table by WAR.
About WPA: When someone hits a 2- or 3-run homer in the middle innings when his team already leads by (2, 3, or 4), there is a shift in the way the leading team will deploy their pitchers. The “good” middle and short guys will get to rest, and the team’s #11 and 12 pitchers will likely pitch instead at the end. Un-quantifiable effect, of course, but likely positive for the team.
In the ten or so years that I have enjoyed this blog, one of the main thoughts I have come away with is the fact that statisticians, who make up the huge majority of people on Fangraphs, have a very difficult time trying to figure out where context fits into their equations. This article exemplifies the conundrum that people who attempt to find a true solution are faced with. When is of absolutely paramount importance as what even if the ensuing events make it irrelevant. A perfect example occurred in last night’s Red Sox-Nationals game when Pablo Reyes, from out of nowhere, tied the game with a HR in the 8th. Wow, what an AB, but 10 minutes later it was meaningless. Does that lessen the quality of Reyes performance, absolutely not! He was rewarded with a ,275WPA. What that actually means IDK. Compare that with a HR that Austin Riley hit off of Danny Mendick in the 9th inning of that 21-3 laugher against the Mets. Riley received .036WPA for that. That is more 10% of what Reyes received but had absolutely ZERO influence in the outcome of the game. Riley did hit a ?HR? off a position player but, because of the context of the situation, for him to get any credit at all for influencing the outcome of the game is absurd.
I agree with a lot of what you said, but the Riley HR doesn’t have zero impact – it had a negligible affect on the outcome. This is all probabilities, remember. Just because a game went from 20-3 to 21-3, as long a the other team has ABs remaining, there’s always a chance, even if a laughably small one. https://youtu.be/KX5jNnDMfxA
I can’t crank up and write code to simulate the probability of winning a game in the 9th inning when down by 17 runs, but I’m pretty sure that probability is as close to 0 as it can be before the fat lady sings…
The Red Sox scored 17 in one inning against the Tigers in 1953. There’s a chance!
Has any team ever scored more than 15 runs in the bottom of the 9th to win a game? Being the dyed in the wool Red Sox fan that I am, I know the Sox scored 8 to beat the expansion Senators, 13-12, in 1961 I would like to ask the group here to find the most runs ever scored by the home team in their last opportunity and win.
Trivia! Only one player has ever scored 3 runs in an inning. It happened in that game. Who is it?
I know this answer, but only because I already Googled to find this game in the first place.
Followup question: This inning also featured one of two players to have three hits in the same inning in the modern era. Who are they?
Hints: Both are Red Sox, neither are the answer to the prior question.
Let’s use that for a back of the envelope calculation:
Let’s say there have been about 200,000 games played in the modern era (completely spitballed, but I think the order of magnitude is correct. That means ~ 1,600,000 innings.
1/1,600,000 = 0.0000625%.
There’s a chance. But I’d rather get home to my family early in case a meteor strikes the earth, wiping out civilisation.
I think you forgot to divide by an additional two in case they lose after that comeback.
I hope I have enough neurons left to understand and agree with your comment but a more realistic WPA for Riley would be closer to .000000000000000000000000036 which i would have no trouble accepting.
I don’t know where you got that number — in the WPA play index on this website, Riley’s home run in the 9th is credited with a .000 WPA, as is every other play in the 8th and 9th innings of that game (and everything after Olson’s homer in the 6th was +/-.004 at most).
Baseball Reference. I went back there and the entire WPA for the 9th inning had been altered with all getting the appropriate 0.
I have a question about the basic assumption in calculating WPA: at the start of the game, each team has a 50/50 chance of winning. This is clearly not the case; does changing this assumption affect the argument?
This instantly caught my attention also. Perhaps, for mathematical reasons, it starts out that way but to claim that a game between the A’s and the Braves starts off at 50-50 puts the claim in question.
I’d love to have a stat for WPA above replacement. Then you could start with the actual situational odds down to the opposing team/batter/pitcher, and then roll these detailed numbers up to wins above replacement but with every context related component reflected.
Just for fun, of course.
I have thought about this quite a bit, too, but that would basically be something along the lines of a “second order” WPA that adjusted the context for the individuals involved. The (very hard) question is what values do you plug in for an individual play? A team baseline? Does that value move once we know more about the quality of teams/players in that time period?
I pulled up the cWPA list (from bbRef) for individual plays and this really stood out to me. Tony Womack’s 2B in the bottom of the 9th of Game 7, 2001 WS without accounting for the players/teams involved is already the third highest most championship-altering play in baseball history, but Degree of Difficulty was off the charts — you’re going up against the thrice-in-a-row WS champs with Mariano Rivera on the mound.
Ben, I like this way of thinking about what MVP really means. I remain uncomfortable with the concept of W.A.R. It still “feels” like it is based on an educated guess of what a mythological replacement player might do in a given situation vs. what an actual player does do in a real situation. I still like the “eye test,” even though it is also flawed. I think those who watch games know a great player and a great season when we see one. Just like we know a great work of art when we see one.
There are over 2400 individual games (2430 for a 162-game, 30 team league, but it can fluctuate if postponed games don’t get made up) in a single regular season across MLB. If you were to physically watch all of them (using 2 hrs/game as a rough estimate if you were to excise inning flips and pitching changes), that’s 4800 hours (or 300 days if you slept 8hrs/day and did nothing else but watch baseball). That’s why we use heuristics to gauge things like the quality of a player’s season.
WAR is just a specific heuristic that allows us to (more) quantitively assess a player’s performance in the 99% of baseball games we don’t visually watch. Like you, I don’t know that I specifically agree with the replacement level “baseline” (a team that would functionally math out to a .294 winning percentage). It’s also very true that there assumptions and simplifications in the model, but that’s what a heuristic is. There are reasons one could consider using a different model (say, baselining it at league average or zero or some other value).
As this article (and its many different threads of comments) illustrate is that as human observers, we’re going to come into it with different experiences and opinions, and what you might find “great” might not correlate with what I find “great”. It is extremely helpful to have a common language. Someone might say “Derek Jeter was great at shortshop because he did an excellent job of converting the plays he got to into outs.” Someone else might say, “Derek Jeter was terrible at shortshop, because there were so many balls he didn’t get to.”
As with many things in life, we have to make compromises. We can look at a player play and have a gut feeling like they are great, and some people can certainly value how a player’s play aesthetically appeals to them moreso than the specific amount of value they bring to the team’s “bottom line” (see: Elly De La Cruz). That’s cool. As much as it may sound, I’m not trying to convince anyone to dismiss their own eyes in their own value of baseball; rather, I’m promoting the value of having a common language to discuss baseball with. I’m pretty sure most people (especially those on this site) would argue with at least some idea about how to calculate WAR conceptually, but it’s still incredibly “valuable”.
Does this mean we should also consider context of each game within the season? Does that mean a walk off HR hit by Bobby Witt Jr, to increase the Royals’ World Series winning odds from 0.0% to a slightly higher 0.0% is worth less than Bucky Dent’s HR in 1978 that gave him his middle initial? Should WPA really stand for World Series Win Probability Added?
There’s a stat for that already. We’re really just talking about scale here – probability to win a game, make the playoffs, or win the world series: each play impact each of those to some degree.
By the way Fangraphs: Get cranking on some “There’s a stat for that” t-shirts.
I can’t find the leaderboard for it though
The best part of WPA is that it is as unbiased as any individual measure in sports can be. What ultimately matters is winning, and if you were to simulate each player’s career (or peak, or whenever you want to evaluate) an infinite number of times, on an infinite number of teams with an infinite number of player combinations, the only constant being the individual player, then the average win total of those teams would be a perfect measure of how valuable that player is in a given role.
That obviously isn’t possible, and there is way too much noise involved to even be close to that approach providing any insight on a player’s ability. WPA is a way of, over the long run, seeing the error in the current linear calculations we have of wOBA. In terms of “Clutch” rating, as calculated by the difference in expected WPA (idk exactly how this is calculated, but this is basically wOBA), and actual WPA, there are some examples of “overachievers” that are not surprising given their playstyles. Ichiro was more than half a win better than expected over his 10 year prime in Seattle, and this trend also continued in his later years, albeit with a much lower baseline. I’d have to think of the right way to run some t-test on this to see if this is significant, but my guess is that it is.
Whatever that significance may be, I do remember running an analysis on clutch rating and finding that a high strikeout rate is pretty heavily negatively correlated with it, same with power ratings like ISO, meaning that the big boppers consistently tend to perform worse than what the linear weights would expect. The only offensive traits that were positively correlated with Clutch rating were hit-related measures like BABIP, batting average, etc.
Point is, if you really want to dive into any “true” measure of past production, then WPA is the stat to go with. Improvements in tracking data can certainly help finetune it, assign credit to defenders, baserunners, etc., but as it is it just does the exact job that it sets out to do, which is an important one.
If you’re talking about “Clutch” as used on bbRef, it’s a player’s WPA / the average “leverage index” of their plate appearances – the sum of (WPA/LI) for all plays, the latter of which basically neutralizes the overall “game context” of an individual play.
That was my best guess, but it doesn’t seem to be clearly defined anywhere. WPA/LI is basically a more overfit version of wOBA. However it is calculated, the point is that there are more stable variables that can predict Clutch rating and there are players who sustainably have higher Clutch ratings, although they are admittedly a little rare. The current mainstream models for evaluating baseball players are very accurate and compelling but they are still biased and imperfect, I don’t see why there needs to be this vehement opposition to any metric that has an element of intangibles involved. Intangibles aren’t tangible because they don’t exist, they’re intangible because they do exist but they’re hard to directly measure.
“Since every game starts with each team 50% likely to win”
huh? I’m not even asking about changing this based upon the relative skill level of teams, but what about home/away? Without fans, it’s still in a vacuum clearly superior to bat in the bottom of innings yes?
Yes.
I’m all for finding a better way to represent ACTUAL wins in way that the WAR stat (wins above replacement) doesn’t actually do. But this Still doesn’t take into account, actual wins. I think it would be better to look at the stats that players accumulate in games where their teams won.
When a player hits for 2 home runs in a game that their team loses don’t mean anything when it comes to wins.
WPA is a great stat for awarding MVP of the All Star game, when you care who got the hit or retired the hitters in a key moment. It’s likewise pretty good for “player of the game” for any game or MVP of a LCS or World Series.
But over the season, context matters too much in WPA. If player A’s 25 HR were more valuable than Player B’s 30 HR, we’re measuring value wrong.
How is it wrong to measure value in comparison to actual wins for a team. Isn’t it important what players do to lead their teams to victory. In other sports you hear about players getting stats in “garbage” time and they refer to them as padding their stats. Their individual stats go up, but what is their actual value in those situations? MVP is Most Valuable Player, that is the issue with the award, what valuable actual means to people.
I wanted to look at this a little bit, and since I am already spending way too much time on this, I’m just pulling out a simple real world example from this season of two players OPS’:
Player A: in team wins: .983, in team losses: .580
Player B: in team wins: .983, in team losses: .728
They’ve both played more or less every game for their team, and play middle IF. Ignoring a bunch of caveats about defense, etc.; if this is the only information you have, which player is more “valuable”? Are they equally valuable?
Player A’s overall OPS is 14 points higher than Player B’s. Why? Player A plays on a division leading club, Player B plays on a “garbage” team.
On one hand, Player B is clearly in the position of his team losing despite his best efforts. On the other, Player A’s performance is more impressive because it comes over more games than Player B.
Full disclosure:
Player A is Marcus Semien, overall OPS .826
Player B is Bobby Witt, Jr., overall OPS .812
The OPS may be the same in wins, but wins and totals aren’t equal. The positive outcomes of stats and wins are happening much more often for Player A, which is why he is much more valuable. Witt is a really good player, but this year, his stats aren’t as valuable.
If Player A wins “Player of the Game” in 25 of his team’s wins with his 25 HR, while Player B wins Player of the Game in 10 with his 30 HRs, then I’d take Player A over Player B, every time*.
*Acknowledging that the above example doesn’t speak to the nuances of a real life situation, where player A might have been player of the game in more games, but “Goat of the Game” in others, or that player B might have been “2nd/3rd best player of the game” in 20 games, etc. What I’m really suggesting is that if we had a comparable per game measure of what “MVP Shares” are on a seasonal measure, then I would be absolutely comfortable using that to decide season MVP. And as you point out, that’s kind of what WPA does.
Yes, WPA has issues.
(Neglects the role of defense, overvalues relievers, etc.)
And nobody is saying lining up players by WPA is how you should measure the value of players.
But the bottom line is that no other stat comes close to measuring the feeing of joy/disappointment you feel watching the game each moment as well as WPA does. WPA helps explain why fans think a certain star was “clutch” more than simple “good players are good” and it is silly to dismiss it because in a one run game, the run scored in the first inning and the ninth inning is equally valuable when the spectators’ and the players’ reaction to those runs are not equal.
Ben, please join the BBWAA to increase the signal to noise ratio of thoughtful votes! Maybe in 20 years the HOF voting will be sensible again…
That’s funny because I had the exact opposite reaction. When seeing WPA being used, I always think, ‘isn’t this the whole reason we stopped caring about RBIs and GWRBIs and the like?
IN addition to WPA’s other problems, adding context is such a slippery slope. Player X hit a home run in the first inning but Player Y hit one in the 8th with the score tied, so Player Y’s HR is more valuable. But wait, Player X’s team came back in the 9th and won by 1 run. So now his HR contributed to a win and Player Y’s didn’t. Aha, Player X’s HR is more valuable.
And what’s more Player Y’s team won 100 games and made the playoffs. Hooray! That win was even more valuable. But oh no, they won their division by 20 games. So any individual win isn’t really that valuable. But they go on to win the World Series! So it’s… more valuable?
This is one of the reasons that so many advanced stats remove this kind of context.
I know it’s silly comparing any measure or even concept of a measure between the two sports for a lot of reasons, but in the NFL, I would expect that kickers have a huge gap between a WAR equivalent and a WPA equivalent. Very replaceable and could easily take a huge share of WPA for a game winning field goal. Especially volatile for such a small sample size season; any given kicker hitting two game winning field goals and not missing any potential game winners (or the opposite) in < 20 games and landing under (or above) replacement level is super plausible. Explains why fans (especially those into the “storytelling”) will so quickly call for pretty much any kicker’s head.
Seen with the browns last year: Cade York hit a game winner and was a clutch hero greatest kicker of all time arguably MVP of his very first pro game, but he ended the season with a pretty low accuracy overall and more sober, “analytical” opinion on him has soured by this preseason because (if you are to believe n=32 FGA is stable and evaluative) he may be replaceable by any other pro kicker. Compare to any browns quarterback who is given an n=2000 PA benefit of the doubt over seasons
In a 2-1 game in which all the runs were scored on HR and the winning HR was hit in the bottom of the 9th, can we really say that HR contributed more to the team’s win than the 1st HR? Both were the difference. WPA credits for timing, a thing which they have essentially no control. It’s a fun, but it is not weighting player contributions to actual win/losses. It’s weighting those contributions to our knowledge of the final outcome. The fact that we didn’t know at the time that a 1st inning HR was going to be the difference maker doesn’t make it less valuable.
Given Poe’s law, I recommend Fangraphs adopt the ability to post in the sarcasm font. The writer should be able to apply it as needed, be it for a comment or an article.