Build a Better WAR Metric, Part 6
In the previous post about the two-run reliever with the three-run lead, the Fangraphs readers leaned 80-20 toward “a run is a run”, compared to “he did his job”.
Let’s make this one similar, but tougher. We have our ace reliever entering the bottom of the 9th inning. This time, our ace reliever gives up a leadoff HR, followed by striking out the side.
We have two different relievers:
(a) we have Billy Wagner entering the game with a 2-run lead, and so, managed to squeak out of it to end the game. If you are interested, his team had a 90% chance of winning before he showed up. And the Astros won.
(b) we also had Trevor Hoffman face the same scenario with a 1-run lead. He leaves the game tied making the Padres go into extra innings, turning an 80% chance of winning into a 50% chance.
And so, you have pushed the context button located deep down in FG readers. I’m a bit surprised it is here–I would have thought double plays, the one stat FG publishes that is unabashedly and wholeheartedly not context neutral.
I’d say they both did a poor job, but giving up a solo homer (or something similar) with a two run lead may well be a result of a different approach than the same pitcher would have with a one run lead.
I’d have no trouble with applying context to possible sacrifice fly AB on the opposite side. Hitting it long and high with a man on third and less than 2 out is (usually) good baseball, even if the ball is caught.
But I don’t want to apply context except in cases where I think the context impacts performance. With a 2 run lead in the ninth I may well prefer to risk a solo homer rather than give up a walk. With a man on third I may well swing differently because a long outfield fly actually is productive.
This one will be fascinating to watch, as I’ve been hoping to get one that will make the readers’ heads explode.
Wagner enters the game with a 90% chance of winning, is the only one involved and exits with the game won. That’s +.10 wins to go to… somewhere. All to Billy? But he gives up a run, which in a random situation would be worth -.05 wins.
So, that’ -.05 wins for his performance, but an extra +.15 wins for tailoring the performance to the context (hence +.10 wins overall).
And it’s those +.15 wins that are at the heart of the question. The reader might even think to allow Billy to get +.05 or +.07 of those wins, for “tailoring”, but not give him ALL of it. Will the Fangraphs readers go that far, to give him a neutral or even positive for his 9.00 ERA?
Stay tuned!
My head did explode on this one!
I voted for the context-neutral choice here, but only to be consistent with how I’ve voted so far (context within innings but not across). But Doug’s point about a possible difference in approach due to the score and inning is what me hold my nose when I clicked the voting button.
I think this doesn’t become so much of a contradiction if the context-specific value of the inning is calculated simultaneously but kept separate from the context-neutral value. Hoffman and Wagner would then end up with the same context-neutral values, but very different context-specific values.
Option 3…both were poor performances, but Trevor should be penalized slightly more due to allowing his home run within his specific context.
From a game theory standpoint, one could argue that Wagner’s primary job with the bases empty was to not allow a walk, making his home run allowed slightly more forgivable than normal (though still bad obviously).
Similarly, one could argue that Trevor’s primary job with the bases empty but a one run lead was to not allow a home run to the first batter, making his slightly less forgivable than normal (but objectively equally “bad”).
Again, I think these should be relatively small factors in evaluating performance (arbitrarily, 20-30% of the context-neutral performance of high-leverage relievers, which is more than other positions should get), but I think they affect strategy/approach enough to be considered.
Yeah, my limitations as a fan might be getting exposed. In isolation, Wagner comes away better because he entered in a more favorable (lower leverage?) scenario even though he was exactly as effective on a base-out basis as Hoffmann. But of course, he also knew he had the buffer and so would almost certainly have pitched differently, favoring trying to cause poor contact rather than chance a walk. One out of four hitters knocking it out of the park is ultimately meaningless if nobody gets on base with a 2-run lead.
But is that what we and teams (when evaluating talent) care most about? Over the course of a season, both guys would have a multitude of different leverage situations and outcomes. Wouldn’t we care most about how, over many innings and appearances, a pitcher fared in the aggregate in all counting stats, rate stats and, at least for an ace reliever, leverage situations? I would assume the answer is “yes”… and if so, then what is the best method to achieve that? I would lean toward some level of context-neutral accounting because otherwise, the aggregate would seem to come away skewed by prioritizing team outcome over player outcome.
I tend to lean very-heavily on the batter when it comes to a home run. In both instances, I tend to think that Hoffman/Wagner had about 10% of the responsibility for the homer, and the batter like 90%. So in both cases, I tend to just say that, sometimes you give up a homer, it happens.
If one of the two was significantly more homer-prone, just due to his style of pitching, then I would shift those numbers to reflect that, but I almost never find that it is accurate to say “this pitcher allowed a homer.” Any individual homer is always the batter’s “fault”, in my view, and only in very large samples do we see a pitcher’s true level susceptibility to homers.
To speak to context specifically, one might argue that, knowing a solo homer is worse to allow with a one-run lead than with a two-run lead, that Hoffman should have pitched in a way to shift the probabilities to make the homer less likely… but I don’t think that is how pitchers generally operate, or if they do, they don’t do it for solo homers.
I think, in both game situations, each pitcher faces the first batter intending to just maximize their chance of getting the guy out. Would there be a measurable benefit in the long-run if Hoffman, in his situation, always pitched to reduce the likelihood of a solo homer, at the expense of some chance of getting the guy out? So, say, Hoffman pitches down in the zone, reducing the likelihood of a homer, but also reducing the chance of a strikeout, and increasing the chance of a single? Does that really pay off in the long run? I would assume, mathematically, you could show that some minor adjustment of strategy would be beneficial, but my guess is the effect is incredibly minuscule.
I know that someone like Mets announcer Ron Darling talks about this sort of situational pitching stuff all the time. Perhaps it’s no longer emphasized as much, but there’s no reason in principle why it couldn’t be, and therefore no reason why it’s wrong-in-principle to account for it, I’d say.
That’s the fine line that’s driving me nuts on these. There are certain types of situations where you can argue that we should credit the player for adjusting well to that situation. For example, it’s easy to argue that some pitchers lose a lot more effectiveness pitching with runners on than with the bases empty (impact of pitching from the stretch, ability to hold runners etc). But the situation in this poll is a lot murkier for the reasons you mentioned.
Will one of the upcoming questions deal with the quality of the opponent (pitcher/hitters/team) as part of the context?
Feel free to offer questions or scenarios more specifically, and I’ll see if it fits.
How about whether or not we treat these ace relievers the same?
a) Wagner strikes out Tulo, Bautista, and Encarnacion to preserve a 1-run lead in the bottom of the ninth.
b) Hoffman strikes out Nick Ahmed, Tuffy Gosewich, and Josh Collmenter (because all the position players are used up) to preserve a 1-run lead in the bottom of the 21st inning.
How about the value of a “run producer” who hits .300/.300/.450 with RISP vs. a “base clogger” who hits .250/.400/.450 in the same situation (i.e. a higher wRC+ but a lower % RISP driven in)?
Probably just overly signaling my own priors here, but WAR is not WPA. So I just keep voting along with the current WAR-consistent idea because that’s ingrained in my mind with how it should work, rather than opting for the “no, we want a stat that’s basically WPA that measure player value.”
Which, don’t get me wrong, WPA with defense added in would be fantastic, but it wouldn’t be useful except in a descriptive sense. (That is, while WAR’s predictive capabilities of future WAR probably isn’t great, it’s likely much better than WPA’s predictive capabilities of future WPA.)
Side note: I think the minute leverage adjustments in current WAR would make it so giving up the run in a higher leverage situation punishes Hoffman a bit more than Wags, but even so, I don’t think one should much more credit than the other given that they both gave up a homer in the same amount of work.
Of course I’m more relieved as a fan to come away with the win, but in terms of performance evaluation, it’d be best to view each event without the context. It’s supposed to be a talent evaluating stat, not a story stat.
Not fair to penalize Hoffman extra when it was his team only got him a run’s worth of padding. Unless you say a pitcher has a magic homer free pitch to throw in one run situations, seems like this should be weighed the same
He does have a magic homer free pitch. He can throw an intentional walk.
But he’ll NEVER throw that pitch with 2 out and no one on in the bottom of the ninth with a two run lead. Change that to a one run lead and he just may, peak Barry Bonds or Mike Trout comes to the base and there’s no one much behind him.
Because with a 2 run lead he’s willing to give up a solo homer, 3-0 count to a good hitter, still throw a strike. 1 run lead to a good hitter, 3-0 count, BB.
Giving up the homer isn’t really controlled by the pitcher, but it can certainly be influenced.
Maybe this was covered in an earlier post, but why is the only context under consideration game state (inning and score)? Why are you not considering, for example, the quality of the opposition? I think it is more telling if we know that Wagner gave up the home run to Neifi Perez while Hoffman gave it up to Barry Bonds.
These are getting really tricky! But, I have to stick to my figurative guns of context-neutrality. Both should be penalized equally.
This is fun!
(smashes mug on floor)
Another!
This is slightly tangential to the specifics of this individual post, but just on the subject of the broader enterprise: why make just one implementation of WAR? FanGraphs already has two different models for pitchers, and plenty of people have advocated a version for hitters that replaces wRAA with RE24. But you could also have a version, for the hardcore partisans of clutchiness, that used straight-up WPA for that component. That way you have a “no context” WAR, a “some context” WAR, and a “full context” WAR. We’ve already got something similar for pitchers, with ERA being “some context” and whatever your favorite peripherals-based metric is as the “no context” option. Maybe you could add a WPA-based option for pitchers, or I’m thinking in particular for relievers (I doubt many people really think that game context, as opposed to inning context, matters much for starters).
And then potentially you can let people do a little choose-your-own-adventure. Maybe FanGraphs could even make it such that registered users could set their personal WAR preferences and then they’d see different numbers as the default (though all of them would be displayed somewhere on a player’s page).
That’s not to say that this series of polls isn’t a worthwhile endeavor; it’s got me commenting regularly on here for basically the first time ever. And these polls and discussion threads might, among other things, reveal more interesting forms that the different versions could take than just wRAA/RE24/WPA/ERA/FIP/whatever. Just that it might be neat to formally recognize the complexity subjectivity involved by providing multiple different takes.
People already do what you are doing, informally and manually.
What you are suggesting is to make it more mechanized and is something I’ve thought about, that the user would create his own “a la carte WAR”, picking and choosing from a menu what he wants, and then always seeing that.
I think we will get there at some point. I don’t know if we’re there yet.
It is impossible that in both scenarios there was a 90% chance of winning.
Both scenarios imply bases empty, no outs in the bottom of the ninth with a 2-run lead and a 1-run lead respecitvely. Hence there cannot be the same chance of winning (given we are talking about two pitchers in the same run environment/league/era)
Sh!t, thanks for catching that. Padres had an 80% chance of winning.
When i voted, 62% of the respondents had chosen the context-neutral answer. It is becoming increasingly clear that fangraphs readers are generally more interested in fantasy baseball than actual baseball. It would be nice if the intelligent people who write for fangraphs and spent so much time researching baseball would devote their energies to real baseball as opposed to fantasy baseball.