Archive for BuildABetterWAR

Build a Better WAR Metric, Checkpoint

In trying to summarize the responses to the three questions, so far, what we have in terms of preference is:
– the event, regardless of the context
– the event, within the context of the whole game state (inning, score, base, out)
– the event, within the context of the base-out state
– and far down the list, the event as it ultimately affects the inning

What the responders therefore are gravitating toward is a purely
content-neutral metric. But, to the extent that we do want to measure the context-specific impact, that should be kept separate, and perhaps not even tied to the player at all. Just a general “timing” bucket.

If we take the case of the triple in the previous thread, in either case, Hamilton and Dyson will get +1 run, because that’s the context-neutral value of the triple, according to Linear Weights.

We immediately add a -0.1 runs because a triple with the bases empty and 0 outs is worth +0.9 runs. So, they don’t want to penalize either guy for getting the triple when they did, and so, to make things add up, we need “-0.1” runs for timing.

Then the three outs, they each get -0.25 runs, as is the standard weight.

So far, we have this:
+1.0 Hamilton
-0.1 timing: limited impact triple
-0.25 batter1
-0.25 batter2
-0.25 batter3

That’s a total of +0.15 runs. But since the inning started at +0.5 runs of expectancy, and we get 0 runs scored, the total has to be -0.5 runs. So, we add another item:
-0.65 bad timing: leaving runner on base

As for the other scenario:
+1.0 Dyson
-0.1 timing: limited impact triple
-0.25 batter1
-0.25 batter2
-0.25 batter3

But, since we actually scored a run, that should come in at +0.5 runs. We need another:
+0.35 good timing: scoring the runner

For a minority, a vocal minority, those “timing” impact runs should be given to the players involved. Looking at the Hamilton one, whereas a generic out is worth -0.25 runs, an out with a runner on third is more costly. So, that -0.65 runs has to be distributed to the three out-makers, for those readers part of the vocal minority. For the readers in the majority, those runs are an after-thought. Maybe they should be considered, so the thing adds up. But, it shouldn’t fall on the shoulders of the players involved. Just a general team bucket to capture the various plays affected by timing.

So, that’s how you build your WAR:

For each player, figure his context-neutral impact as one value, and his “timing” as another value.

Then, the reader can choose whether to include the timing value or not.

Now, on to the pitchers and fielders!


Build a Better WAR Metric, Part 3

Before we talk about baseball, let’s talk about the other three major sports. You’re on your own 20 yard line, you march down field on a series of plays, but ultimately, you punt. Or, you march down field and move far enough for a tough field goal that gets made. All those running and passing plays aren’t considered any differently based on the results. The RB got 25 yards on 4 running plays, and no one matches it up to the end result.

In hockey and basketball, a great pass that doesn’t ultimately lead to a goal or basket goes away like a fart in the wind. No one tracks it, and if they do, it’s not considered anything close to the impact of an assist that led to a score.

Why the difference? I think it’s because of the stop-start nature of football, that the “sequence” ends after each play, and the whole drive is 2-5 football minutes or 5-15 human minutes. In hockey and basketball, turnovers happen often enough and each drive lasts 10-30 seconds in sport or human minutes. I think that’s the reason.

So, let’s talk about the leadoff triple. Billy Hamilton gets on third base, and the next three batters strike out. He’s stranded there, no runs score. A fart in the wind triple? Or something much more tangible? Jarrod Dyson gets on third, the next batter hits a medium fly ball out, far enough to let Dyson score. The next two batters strike out. One run scores. That triple is obviously tangible.

How do you see these two triples?


Build a Better WAR Metric, Part 2

Ok, you guys have spoken, and you don’t want a bases loaded walk to count the same as a solo HR. That even though the base-out state before the event and after the event remain unchanged, and that the number of runs now in the bank are the same, the WAY it happened matters to most of you. Therefore, we are NOT trying to preserve the runs, we are not trying to make sure the runs add up. You have been clear on that.

Now, let’s talk about “preservation of wins”. It’s a 0-0 game, the bottom of the 9th, the bases are loaded with two outs. Historically, at this point in the game, the batting team would end up winning 68% of the time. It’s a high stakes situation, a Leverage Index of 6.4. And the batter walks. The batting team wins, game is over. Ooops, I meant the batter hit a single. No, wait, it was a Grand Slam. No, wait it should have been a Grand Slam, but Robin Ventura decided to abandon the bases after he reached first base. Regardless, the game is over, and the batting team won as soon as the batter touched first base.

Your question:


Building a Better WAR Metric

Wins Above Replacement (WAR) has as its genesis Bill James, even if Bill might not necessarily take the credit (or blame based on some readers) for it. But make no mistake, Bill provided the plumbing for it. For those interested, you can read Brandon Heipp’s account on that backstory.

When you put all the plumbing together, you can create a framework. And that’s what WAR is, a framework to provide an estimate. Wins Above Replacement is an estimate of… something. What that something is is different for every person. While the currency is wins, it’s not clear what those wins represent. There are reasonable choices you can make along the way. And for every fork in the road you take, you may diverge yourself from the next guy. This is why WAR can never be one thing.

As a framework, WAR leaves little room for discussion. Whether it’s what you see at Baseball Reference or at FanGraphs or openWAR or (to some extent) at Baseball Prospectus, they have as their framework the WAR that was championed on my old blog, which culminated with this article. But a framework is not the same as an implementation. 95% of the cars on the road all follow the same core design. That’s the framework. But a Chevy is different from a Lexus. Those are implementations. And there are as many implementations of WAR as there are baseball fans. Whereas Baseball Reference and FanGraphs and the others provide a consistent, systematic implementation, most fans have their own personal mish-mash of arbitrary, biased, and capricious combination of stats, which can change as their mood fits.

This series of articles, of which there may be a dozen(*) is an effort to try to come up with a WAR metric that will satisfy the Straight Arrow readers.

(*) I have no idea. This is the first one I’ve written.

***

I’ll ask you a series of questions, starting now. The openWAR guys talk about “preservation of runs”. That is a good starting point, and a great way to describe the concept. So, the question centers around whether we want to make sure that everything adds up at the play level. If you get a bases-loaded walk, do we want to make sure that exactly one run is accounted for or not?

If you care about “talent”, you just want to account for around +.30 runs for offense (and -.30 runs for defense), because you don’t want to be concerned with the specific base-out state. (We’ll talk about “preservation of wins” in a later question.) Similarly, is a bases-empty walk and bases-empty single the same thing or not? And if you want to preserve runs, are you ready to accept a bases-loaded walk and a solo HR as being the exact same thing?

So, have a discussion, and then answer this poll question:

There are plenty of other discussion points that go into building an implementation of WAR, and we’ll get to those in the future. For this post, I’m interested to hear what you guys think about this issue specifically.