Build a Better WAR Metric, Part 5
When the home team enters the top of the 9th with a 3-run lead, they will win that game 98% of the time. That happens mostly because they get to pick and choose the reliever they want. If they chose a random reliever, they’d win 97% of the time. If they chose a poor reliever, they’d win 96% of the time. It’s pretty tough to mess up a 3-run lead, especially when the home team gets one more crack at it in the bottom of the 9th.
So, we have a SP that went 8, and he hands off to the reliever this 3-run lead. The ace reliever comes in. Let’s call him Armando Benitez. He walks the first batter, allows a HR to the second, then strikes out the side. The game ends, and his team wins. Armando even gets a “save”, whatever that is supposed to imply.
Since he was given a 3-run rope, and he only used 2-runs, he was able to turn a 96% or 98% chance of winning into 100%, all without the help of his fielders. Incredibly, things could have gotten worse, which does happen 2 to 4 percent of the time. In this case, he pitched just bad enough to win.
I’d choose something between the two options.
In all these polls, it’s understood that most people who want some sort of “in-between”. But if I give that as an option, we don’t move the discussion forward.
I’m more interested to see which way you are leaning. If these are your only two options, you choose the one that is closest to what you want.
Again, I feel like the context here shouldn’t matter. Giving up 2 runs in any inning is bad, regardless of how many runs of cushion your offense has to work with. Granted, some pitchers may affect how they approach different batters based on the score, but giving up 2 runs at any time can’t really be considered anything other than Poor.
I feel like this whole exercise could’ve been replaced with a single poll asking “should WAR be context-neutral”, in which a majority of the fangraphs audience would’ve voted “yes”.
I think the audience portion here is really what’s impacting it.
If I take this to a broader set of baseball folks that don’t read Fangraphs (as ubiquitous as Fangraphs might seem), I bet the responses would be very different, or at least not so broadly uncontested and obvious. As such, I think this exercise is a pretty great didactic tool to explain why the current line of thinking and player valuation in baseball works like it does, even if it doesn’t actually “build a better WAR” by the end because the Fangraphs audience will largely be pseudo-hidebound and just reinforce the concepts we already have enshrined in WAR.
This is an excellent post, thanks.
I wouldn’t want to be so preumptive as to have taken as fact the guess that the Fangraphs readers would respond as they did. Especially some of these that went 60-40 or 65-35.
In any case, I’m having fun. And given that there’s been some 3000 votes placed across the 5 polls, I think others are having fun too.
So, I learn something, and readers are having fun. And it’s out in the corner on Instagraphs so it’s not in anyone’s face for too long.
I’m in the minority vote here, which is somewhat disappointing.
The context does matter. Take into account Tango’s setup: The pitcher chosen matters, and is chosen on context. That 98% includes every pitcher who gives up two before ending the game.
Penalizing this pitcher more for giving up two overcompensates the value of ending the game with the lead. If you want to measure this value, look at variance in WAR by AB, and recognize that higher variance WAR pitchers would be the ones we’d prefer in this situation. It’s the right pitcher at the right time, who did exactly what he was put in to do.
Sure, we value the pitcher with a lower variance more as a shutdown guy, but that’s why he a. Gets chosen in closer games and b. Gets a higher WAR share for them.
But if the argument is that context shouldn’t matter, why use that metric as WAR? It’s not WAR, because wins can’t be produced without the context (including that there’s actually a competitive game actually occurring, versus a bunch of folks taking turns at batting practice). It’s not as though there aren’t other, non-WAR metrics that are stripped of context.
This is an excellent argument.
What the Fangraphs readers are choosing, almost overwhelmingly, is either a base-out driven model, or a runs-driven model (as is the case in this poll).
The “wins” part is simply a “conversion” of runs to wins. By slapping the wins term on it, it makes it SEEM as if we are using wins as some sort of context to the whole stat. But we are not.
In this particular example, the pitcher gave up 2 runs, or 1.5 more runs than expected (we expect half a run per inning). Hence, the abysmal part. Those 1.5 runs would translate to RANDOMLY 0.15 wins. This pitcher’s performance should be translated as -0.15 wins.
Even though before he showed up his team had a 96-98% chance of winning, and after he left it was 100% and none of his fielders were involved.
Even if he had a shutout inning, he’d have given up 0.5 runs fewer than expected a value of +.05 wins in a RANDOM inning. But in this case, we didn’t have .05 wins even available. There was only around .03 wins available.
So, we’d have to assign -0.02 wins due to timing in order to balance it out.
But we can get there the same way by leaving everything in terms of runs, not wins. In order to convert it to wins, if you are going to be complete about it, you need to account for this timing bucket.
Here’s a slightly different version of the same scenario that I think makes the “context matters” point seem a little more reasonable: a two-run save opportunity during the course of which the closer gets three outs but also gives up a solo home run. The reason why there’s more of an argument for “this is not a bad performance” is that it’s a cardinal rule of pitching in that kind of scenario that you should go after the hitters very aggressively, because a walk is basically as bad as a home run. The scenario as given above includes a pitcher who first violated the “thou shalt not walk” commandment and then gave up a home run. In fairness neither of these sins were nearly as bad as they would have been given a two-run lead, but it’s still tough to make the case that this was successful situational pitching that led to giving up unimportant runs sort of on purpose.
Not sure if there’s any very administrrable way to distinguish between the two, though.
You read my mind. I was actually going to use the 2-run lead as my next example. I wanted to see how readers responded to the 3-run lead first.