Build a Better WAR Metric, Checkpoint
In trying to summarize the responses to the three questions, so far, what we have in terms of preference is:
– the event, regardless of the context
– the event, within the context of the whole game state (inning, score, base, out)
– the event, within the context of the base-out state
– and far down the list, the event as it ultimately affects the inning
What the responders therefore are gravitating toward is a purely
content-neutral metric. But, to the extent that we do want to measure the context-specific impact, that should be kept separate, and perhaps not even tied to the player at all. Just a general “timing” bucket.
If we take the case of the triple in the previous thread, in either case, Hamilton and Dyson will get +1 run, because that’s the context-neutral value of the triple, according to Linear Weights.
We immediately add a -0.1 runs because a triple with the bases empty and 0 outs is worth +0.9 runs. So, they don’t want to penalize either guy for getting the triple when they did, and so, to make things add up, we need “-0.1” runs for timing.
Then the three outs, they each get -0.25 runs, as is the standard weight.
So far, we have this:
+1.0 Hamilton
-0.1 timing: limited impact triple
-0.25 batter1
-0.25 batter2
-0.25 batter3
That’s a total of +0.15 runs. But since the inning started at +0.5 runs of expectancy, and we get 0 runs scored, the total has to be -0.5 runs. So, we add another item:
-0.65 bad timing: leaving runner on base
As for the other scenario:
+1.0 Dyson
-0.1 timing: limited impact triple
-0.25 batter1
-0.25 batter2
-0.25 batter3
But, since we actually scored a run, that should come in at +0.5 runs. We need another:
+0.35 good timing: scoring the runner
For a minority, a vocal minority, those “timing” impact runs should be given to the players involved. Looking at the Hamilton one, whereas a generic out is worth -0.25 runs, an out with a runner on third is more costly. So, that -0.65 runs has to be distributed to the three out-makers, for those readers part of the vocal minority. For the readers in the majority, those runs are an after-thought. Maybe they should be considered, so the thing adds up. But, it shouldn’t fall on the shoulders of the players involved. Just a general team bucket to capture the various plays affected by timing.
So, that’s how you build your WAR:
For each player, figure his context-neutral impact as one value, and his “timing” as another value.
Then, the reader can choose whether to include the timing value or not.
Now, on to the pitchers and fielders!
But how do you split the timing runs between Dyson and the guy who hit the sac fly?
What if, instead of applying the “timing” factor to individual players, you simply added it up by team? The result would be a sort of coefficient to temper our context-neutral WAR valuations. For example, Trout may be a context-neutral 9-win player, but the Angels (and luck, the park, the weather, etc.) may have dropped the ball so often that they had a substantially below average “timing” score, and we would then be able to understand that the value of Trout’s performance was diminished by some measurable context-driven number.
I would actually take a slightly different approach to evaluating the players who bat after the lead-off triples. One key part of the point is that a sacrifice fly is not the same thing as a strikeout. Nor is this simply a matter of luck: with a runner on third and zero or one outs, there’s a real sense that the batter is supposed to prioritize putting the ball in play to get the run in. So I would give a substantial bonus to the guy who hits the sac fly, perhaps the whole +0.65 or perhaps not quite.
Conversely, I’d like to penalize the guys who fail to get that runner in. An ordinary out really is worse there than, say, with the bases empty; among other things, the threshold for achieving a positive result is a lot lower. A generalization of this approach would be to reward/punish players based on their propensity for getting “productive outs.” This isn’t just a matter of luck, since players do actively try for them and some are presumably better than others. And of course in an ideal world you’d put some of the credit/blame on the runner’s speed, in certain situations, but there’s probably no realistic way to implement that.
I’m less sure how to handle the guys who bat with that runner at third and two outs already. A single there feels a lot more meaningful than an ordinary single, or even a single with the runner on third and no outs. Partly, again, that’s because there’s an intermediate possibility with less than two outs, where you get one benefit of the single (driving in the run) but not the other two (baserunner, lack of out). Partly it’s because you’re on the precipice of losing that run forever, whereas the guy who strikes out right after the triple has only lost the first of three chances at it. Certainly these events feel very clutch, but it’s less clear how they should actually be valued.
Is the out in that situation (runner on 3rd, <2 out) really worse than one with the bases empty? It's just a matter of looking at which run you're trying to score, runner ahead of you, or yourself. Bases empty outs I'd say are worse contextually, because you're doing nothing positive for the lineup spots ahead of you or behind you.
I mean, the point is really that there are two kinds of outs you can make with a runner on third and <3 outs. They're the same as concerns scoring a run yourself (you won't). They're very different as concerns scoring the runner ahead of you. Since this difference is far from arbitrary or a matter of dumb luck, I'd argue it should be accounted for.
Now, where does the out with the bases empty fall in between those two kinds of outs with the runner on third? I dunno, but I'd incline to say "somewhere in the middle," among other things since the category includes all of the events from both results with the runner on third (the deep fly-out is just like a strikeout with the bases empty).
<2 outs, whoops.
WAR is fine leave it alone
The framework of WAR is fine. It is not being touched.
There are many implementations of WAR (Fangraphs, BR.com, BPro, openWAR, et al). Which one is “fine”.
Yes, I was going to ask about the out, but seeing the overwhelming support to NOT looking at hits and walks based on context, it is interesting that that thought won’t necessarily apply to the out!
So, you want the triple to always be worth +1 run, regardless of it was hit with bases empty 0 outs or bases loaded 2 outs, or whether the runner scored or not, etc.
But, the same reasoning doesn’t apply with the out. That there are “not-so-terrible outs” and “terrible outs”. If an out moves a runner, or better scores a run, then that should be handled differently.
And the idea is that the batter responded to the situation… if he made an out. But if he didn’t make an out, the batter did NOT respond to the situation. You only know if the batter responded to the situation after you see the response?
Is this what I’m hearing?
Personally, I adjust WAR (for MVP purposes, etc.) by replacing players’ wRAA with their RE24, and averaging their DRS and UZR.
I’m not sure I agree with this interpretation of the polling results so far. We’ve established that (a) different events which have the same contextual effect should be treated differently, and that (b) events should not be treated differently based on what happens after them.
What we haven’t established is that the same event should be treated the same regardless of what came before it, and thus of the context in which it happened. You haven’t polled, that is, on the two-out RBI single versus the two-out bases empty single.
Let’s say that I ask between a bases loaded walk and a bases empty walk. My contention is that people will respond with “a walk is a walk” based on their past responses.
Your contention is that MAYBE they’ll prefer the bases loaded walk to the bases empty walk. Let’s say that that’s true, that the poll goes that way.
How about if I ask for bases loaded walk v bases empty single. How do you think that one will go? Don’t you think people will prefer the single to the walk?
So, we have a bases empty single that has a run value of .25 runs. We have a bases empty walk that has the identical run value of .25 runs. And we have a bases loaded walk that has a run value of 1.0 run. Based on the responses, they aren’t sticking to that.
A neutral single is .45 runs. A neutral walk is .30 runs. Where could based load walk land? It has to be under .45 runs. And it’s likely going to poll at walk is a walk.
Having said, I will poll in a few days and let’s see what we get.
That’s fair.
To a certain extent, this is turning out to be a referendum on whether the respondents are more interested in actual baseball or fantasy baseball, and depressingly enough it’s looking like they’re more interested in fantasy baseball, or at least the statistics on which fantasy baseball is currently based. This would account for the focus on the event and disinterest in the context, since fantasy baseball is based entirely on events with no context. In principle it would be a good idea to ask respondents to include with each response the relative extent to which their answers reflect an interest in actual baseball and to what extent they reflect an interest in fantasy baseball, but I have the feeling that this wouldn’t really clear things up. At some point in the evolution of statistics it’s going to become apparent that these interests are clearly divergent and statisticians will frame their research questions accordingly, but we definitely aren’t there now.
Hi Mr.Tango,
I am just trying to make sense of this: “We immediately add a -0.4 runs because a triple with the bases empty and 0 outs is worth +0.6 runs.” Is this saying that triple with the bases empty and 0 out in an inning, during the game of history, led to 0.6 runs on average?
Sh!t. I was thinking one out. With 0 outs, it’s nearly 0.9 runs. Too many numbers floating in my head. I’ll update it, thanks for pointing it out.
Mr.Tango,
I want to get a complete grasp of this, so I want you to see if I got this right: Inning events were single, walk, strikeout, strikeout, groundout.
+0.73 single
-0.33 limited impact single
+0.76 walk
-0.1 limited impact walk
-0.64 strikeout
-0.50 strikeout
-0.45 groundout
-.53 runs
Based on: http://www.tangotiger.net/lwtsrobo.html
The run value of a single is roughly 0.47 runs. So, you would have:
+.47 single
-.07 limited impact single
Then we have:
+.32 walk
+.34 high impact walk
-.27 strikeout
-.37 high impact strikeout
-.27 strikeout
-.23 high impact strikeout
-.27 groundout
-.18 high impact groundout
Total: -0.53 runs
Oh nice. We did it different way, but same conclusion.
Where do you get the high impact and low impact values?