Build a Better WAR Metric: Timing Buckets
On September 1, 2015, the Nationals and Cardinals played a game where the Nationals took a big lead, only to give most of it back almost immediately. The Nationals kept trying to hold on, until the end, when the Cardinals won the game on a 3-run HR.
Source: FanGraphs
Let’s look at that ninth inning. First up was Jason Heyward. He grounded out. That context-neutral run value of making an out is -0.25 runs (or -.027 wins). Making an out to start the inning with the bases empty is only worth -0.225 runs (or -.024 wins). Therefore, the base-out timing value of the out is +.025 runs (or +.003 wins). It looks like this:
-.027 wins: Heyward’s out
+.003 wins: low impact timing of out with bases empty
But we know more information. It was a 5-5 game to start the bottom of the 9th. This is a higher leverage situation than random. Heyward’s out actually reduced the chance of winning by .050 wins, not .024 wins. That is, the impact is felt twice as much as a random leadoff situation. So, there’s yet another .026 wins to account for. This is what it looks like:
-.027 wins: Heyward’s out
+.003 wins: low impact timing of out with bases empty
-.026 wins: high impact timing of out in 9th inning of tied game
The question to ask yourself (not to me, but to yourself), is how much do you want to credit Heyward for making an out in this situation: do you want to just credit him with a random out, because he was just plucked into this situation, or do you want to credit him with making an out as the leverage was lower impact (bases empty) or even high impact (9th inning of a tied game)? Is an out an out, or does the out depend on the situation?
Let’s continue. Yadier Molina also got an out. Going through the above machinations gives us this:
-.027 wins: Molina’s out
+.010 wins: low impact timing of out with bases empty
-.019 wins: high impact timing of out in 9th inning of tied game
Now the fun begins. Cody Stanley doubled.
+.081 wins: Stanley’s double
-.056 wins: low impact timing of double with two outs
+.043 wins: high impact timing of double in 9th inning of tied game
So, in a random situation, a double with two outs is not that valuable. It’s less valuable than a random walk. That’s why we have a huge -.056 win value to account for its low impact. But at the same time, this puts the winning run on base in the bottom of the 9th. This is enormously high impact. How you approach valuation will decide how you want to credit Stanley and his double.
Tommy Pham walked with first base open and winning runner already on base.
+.032 wins: Pham’s walk
-.020 wins: low impact timing of walk with 1st base open
-.009 wins: low impact timing of walk (run is useless)
Let’s pause here. The double put the winning run on base, and left 1B open. The walk is in fact practically useless. The win value changed by +.003 wins, which is pretty close to zero. The batter and pitcher know this, which is why we see a NEGATIVE impact of the walk in the 9th inning of a tied game, even though we are in a high leverage situation. This is unlike the double which had a huge POSITIVE impact. The entire sequencing of the situation matters. Given that the batter and pitcher are aware of the situations as they develop, the entire timing values noted above make perfect sense.
Finally, the HR by Brandon Moss.
+.150 wins: Moss’s HR
+.137 wins: high impact timing of HR with 2 runners on
+.114 wins: high impact timing of HR to win the game
In the end, the Cardinals went from a 61.4% chance of winning to 100%, adding +0.386 wins. Adding up the above, and we get:
+.209 wins: all the events in a random situation
+.074 wins: high/low impact timing for base-out situations
+.103 wins: high impact timing of inning/score (except walk)
So, how do you, the reader, want to evaluate each of these plays? How much do you want to assign to the batter (and pitcher) and how much do you just want to have some general “timing” buckets, not linked to any particular player?
I think WAR and leverage, at least for hitters should generally be separate. I don’t think someone’s WAR should be reduced further because he got an out in the 9th as opposed to the 6th. If I want to know how someone did in a specific leverage situation, I’ll look for that. When I use WAR I want to know how a player performed, not how he performed in a given scenario.
I like this because it gets to the heart of the issue – WAR can either measure context or not, but can’t do both. And as much as “one number to rule them all” is nice, it’s plenty fine to analtze players with multiple numbers.
I like “WAR” for raw talent and something like “Clutch” for situational context and the more story-telling aspects. Story-telling being “how did the talent (WAR) result in a winning or losing game (Clutch).
Two parts of a grand whole that is baseball.
I added a poll.
In this case, what I’d like to say is count the situation for the walk, but not for the homer or the double. But that wasn’t an option. (I went context independent as the poll answer.)
The walk is something the pitcher has a lot of control over, and had almost no impact on the outcome of the game. Pitcher giving up a “free” walk shouldn’t be context independent. The double and the homer both presumably happened when the pitcher was mostly trying to get an out and the hitter simply trying not to get an out.
If you think batters can trade off batting average for power, then I might slightly dock the batter for the double (he should have been going for the homer) and slightly dock the later homer (his result was no better than a good single, so maybe he should have been swinging for a single). The opposite would apply both ways for the pitcher to the extent that you think he has control over quality of contact or grounder/flyball.
But I believe even in my ideal world both of these adjustments would be very very small. The walk I might almost entirely discount as a probable “intentional unintentional walk”.
I want both! An “intrinsic” WAR which is context-neutral and could be applied looking forward, and an “extrinsic” WAR which is win probability dependent and more backward looking. And isn’t the delta a measure of the elusive “clutch” performance? (Base-out context alone is the least interesting to me).
The more I think about it, the more I feel that a WAR model without context accounted for is incomplete. The primary argument against including context is that the batter has no control over the situation he’s in. But WAR is not a predictive stat, at least it isn’t and in its current form, shouldn’t be used that way. It is designed to describe what happened, and leverage stats are about as real as they get.
As far as I can tell, there is no WAR model that really acknowledges context for hitters. I think that would be more useful and/or interesting than another context-neutral WAR model, of which we have a few.
I like this, just to see how much a difference would context make overall.
I am a little unclear about what we are trying to accomplish. There have always been two approaches to evaluating hitting, those that reward batters very highly for taking advantage of event sequencing, and those that are essentially sequence neutral. RBI and BA back in the dark ages, and the more nuanced and logical WPA and WRC nowadays. WAR has always been an attempt to sum up the sequence-neutral contributions of hitting/pitching, fielding, and baserunning to estimate a player’s value in the ultimate currency of the sport, wins. The “build a better WAR” posts have been interesting and have made us think through things–thank you for that. But most of them have pitted the sequence-aware and sequence-neutral approaches and asked us which we like better in a given situation. Right? Yet it seems to me that both perspectives have to be taken into consideration to evaluate overall player performance. We can’t mix and match them in some conglomerate statistic, I don’t think. “Building a better WAR” implies tweaking our estimates of different contributions. Especially it would take a better understanding of the value of fielding, and how best to measure it. Am I misunderstanding something?
Context neutral, no one batter has control over what the leverage of their at bat is.
“no one batter has control over what the leverage of their at bat is.”
That’s not a very good statement imo. The batter is in the situation whether or not they chose to be there. The questions is what do they do in the situation they end up in? Does it matter?
Given that no batter can choose the situation they’re in, you could have two players of equal skill being perceived differently because one was given more “valuable” plate appearances. I think that WAR should be used in a way that would make it easy to compare players, and adding context to their situations reduces your ability to do so.
you are just asserting that you want your WAR stat to represent context-neutral skill levels only. it is certainly possible that someone else might want their WAR stat to incorporate some level of attribution for ACTUAL value added toward (or from) a win…
What we are describing here is what we have seen in baseball video games dating back as far as I can remember and that is the stat Clutch. Lets be honest there are times when you would love to see a certain batter up at a certain time because for whatever reason you feel they will come through. Realistically there should be a distinction between a un-leveraged WAR stat (uWAR) and a leveraged WAR (lWAR). uWAR tells how a player performed and nothing more. Basically how many singles, doubles, stolen bases, etc. lWar shows how said player had an impact. What did the player do to contribute more to the win of the team rather than the individuals stats.
This distinction would help to appease both camps of opinions.
Performance vs. Impact
It is fun (and necessary) to see how an individual performs but in the end the job of statistics is to see how a TEAM wins. Lets be honest, nobody cares if Bryce Harper goes on to hit .400 BA with 93 Hr’s a year for his career and never makes it out of second place (okay we would care but you see the point). The whole point is to see Wins… hence WINS above replacement.
Just my 2 cents
YES! I wrote a similar reasoning above (I called it Talent vs. Storytelling) but like your phrases.
As cute internet meme would say – Why Not Both? Two stats measuring different things, with clear distinction.
I DON’T think WAR should be context-neutral. Standard metrics say he had a MVP season in 2005. 95R, 51HR, 128RBI. And he had 6.7 WAR that season. Atlanta definitely benefited little bit with his MVP like season. But 2005 Andruw Jones had batting average of 0.207 in RISP. If we want context-neutral WAR, we are not taking account for his struggle in RISP and his WAR would have been even higher than 6.7. He definitely does not deserve WAR higher than 6.7.
@Tangotiger
aren’t there a few questions in these series of post that are more nuanced than simply “context neutral” versus “run preservation?” it seems like all of them come back to your concept of defense and offense canceling one another out, but here are a few:
– all outs are in fact not created equally. (the question is how granular to make the WAR engine to address this)
– does context-neutrality ignore a batter or defense ability to tilt the expected run value their way for unique outcomes in unique base-out situations (e.g., the fly ball for a SF or the bases loaded walk)
I am sure there are others, but these two seemed to stand out…
thanks for doing these — can’t wait to see where we are headed!
The comments section is to discuss the more nuanced aspect. When it comes to a poll, I have to section it off to something more obvious, even if it’s less accurate.
I voted for preservations of wins but I think what’s even more important is the quality of contact given the game situation. Maybe Heyward and Molina scorched the ball but hit it right at someone. Maybe Stanley hit a blooper to the opposite field which landed on the foul line.
I think the biggest issue is that context impacts the run matrix which impacts player goal functions. At the end of the day you’re going to need three stats to describe what happened; one that is completely context-neutral, one that takes into account base state, and one that takes into account game-state and base-state.
The run values listed all describe the situation ‘correctly’, but the problem with context specific metrics is not that they are inaccurate, but rather that players don’t control how many opportunities they get in different leverage situations, so two players of the same skill level could have drastically different scores in the leverage specific stats. This isn’t a problem with the statistics themselves, but simply with sample size.
The more I think about it, the more I lean toward some statistic that ‘normalizes’ the win percentage delta in a given scenario. Simply take the fully-context specific change in win percentage, and divide it by the total range in outcomes. For example, in a high-leverage situation, say that the outcome from a given pitch can change the pitcher’s team’s win percentage by somewhere between +.2 and .-4. Simply take the actual outcome (say -0.2), and divide it by the total range, gives you a -0.33 score. This normalization lets you take into account the change in goal functions of the players, while not amplifying the timing of the event that is out of the player’s control.
You should look at WPA/LI. It deleverages the events so that each PA counts as “1”, while allowing for the context-specific impacts to be felt.
That is, a bases loaded walk and a HR in bottom of 9th of a tie game will have identical values, but without the amplified aspect of the leverage.
That’s exactly what I was thinking, kind of figured it should already exist.
I was thinking the same thing while reading the article, but if WPA/LI is my preference how do I answer the poll?
It’s does preserve wins or runs, but it’s not context neutral either.It’s sort of all of the above, factoring in base/out state and game state while weighting every plate appearance the same.
My only issue with the base-state is the fact that hitters will sometimes fundamentally change approach when say trying to drive the runner in from third.
In that situation, the hitter himself may be TRYING to make an out (well, they’d still want it to fall for a hit) and in that case a context-neutral stat will penalize the hitter for a productive approach he was purposely trying for.
Granted, I think the scenarios where “purposely change approach based on context” are fairly minimal, but it’s still an issue (and a fun one to talk about!)
I definitely prefer the context-independent approach in all or virtually all situations, i.e., a single or double or whatever has the same value, regardless of base-out situation, leverage, etc. We do have WPA and other metrics to deal with context. Moreover, there seems to be precious little evidence for clutch, i.e., repeatable/predictable ability to perform better in high leverage situations (and it never made sense to me, anyway, as it implies that a player is not trying as hard in lower leveraged situations, which are virtually by definition more common).
Unless strong new evidence emerges that clutch or whatever you want to call it really does exist for at least some players, I don’t too much of a point to context-dependent stats. They do tell us how much value the player actually had in real situations, but so what? It seems to be mostly chance that he not only got in those situations (WPA) but performed well in them (WPA/LI).
So here’s what I think is the true problem with applying context and leverage to any theoretical WAR-like metric, and funnily enough it’s in the name itself. Wins Above REPLACEMENT.
WAR isn’t measured against some nebulous arbitrary concept of zero-level performance. It is measured against replacement level, which is an actual thing that represents the highest level performance expected from a player not on a major league roster. And therein lies the issue, because you can’t apply anything other than context-neutral information to replacement level. Replacement level, in theory, describes the raw performance of the replacement player, but it doesn’t and in fact CAN’T describe exactly to what degree the replacement player’s performance helps or hurts the team’s win probability in a given game.
With that in mind, how does it make sense at all to apply context to a WAR calculation when the level being measured against is, by necessity, context-neutral? Some might say “Yeah we know leverage and game situation can’t be controlled by the player, but maybe we should take it into account anyway.” I disagree with that line of thinking emphatically. In fact, applying context to the WAR calculation leads to a contradiction. You’re asserting by implication that context IS controlled by the player even though the opposite is clearly true. To make the whole thing work you’d also have to apply non-assumed context information to the replacement level, which is impossible.
Now, contrary to everything I just said, I think there is value in trying to understand how a player’s performance translates to actual runs and actual wins. I just don’t think the WAR paradigm is the place to do it. WAR doesn’t need to be all-encompassing, and it doesn’t need to be used as a universal catch-all for assessing all types of player performance no matter what Dave Cameron might make you think. It simply needs to do the job it was designed to do, nothing more, nothing less.
very interesting point — I like it.