Build a Better WAR Metric, Part 4
First thanks so much for the tremendous responses, both in the comments and just the participation in the polls. There’s been 700 to 1200 votes in each poll. Just overwhelming responses.
For now, let’s start the second inning by leaving aside the hitters and talk about defense. Now, when I refer to defense, I mean pitching+fielding. Remember, defense is the whole team, the pitchers and the fielders. We’ll worry about how to separate fielders from pitchers in a soon-to-be-asked question. Just not now. Cool?
Let’s say that one team defense allows 10 hits with 3 walks. But they are all scattered, and so actually end up with a shutout. Another team defense allows the exact same number of hits and walks. They even allow them in the exact same way. The only difference is the timing. They allowed them bunched up, and so resulted in eight runs allowed. From a team defense perspective, how do you see them in terms of assigning value?
After 63 votes, it’s a 60-40 split. Looks like we may finally get one where we’ve found two camps of equal size.
After nearly 400 votes, it’s exactly a… 60/40 split.
If Team A consistently allowed fewer runs/H + BB, I’d consider timing or sequencing as a factor. But assuming it doesn’t, then H + BB seems the better way to go.
Is this a question of a pitcher’s ability to pitch out of the stretch?
Don’t read into the question anything that not there. It’s simply one game, sh!t happens. Sometimes you get a shutout. And sometimes you allow 8 runs. I’m asking how you ultimately view this game from the perspective of the defense.
Usually I’m in the camp of splitting up the game into its individual events, but I voted for the run prevention choice, because I feel like pitchers who tend to be poor out of the stretch(which is a skill) tend to have those sorts of outings ( 5 runs with 10 base hits), whereas, the pitcher who is very consistent out of both the stretch and the windup will have the same hit numbers, but fewer runs allowed, which is the point of baseball.
This one is difficult because the options are bipolar but there seems to be more room for shades of grey on this.
Yeah, I don’t see this as a huge difference, but “runs allowed count, but not as much as the events” simply wasn’t an option.
I went with the runs allowed answer to the poll, but either reply feels like a lie.
The concentrated runs should have seen a shutdown reliever brought in. Someone screwed up not to do so (or the shutdown reliever was brought in and melted down).
Meanwhile, things like inducing ground ball double plays and not allowing as many homeruns when there’s a man on base probably really are at least slightly a matter of pitcher skill.
I lean toward the sequencing neutral answer, but I feel strongly that this is a more nuanced issue that an A/B response allows. In actual implementation, I think a mix of these would be preferable.
A mix is always preferable. The question is the degree to which you have the mix. By figuring out how many people we have in each extreme, that will tell us how much to the middle we have to move.
If I say “a mix”, 90% of people would choose that, and I’d lose important information as to why way people are leaning.
A couple folks commented on whether this boils down to pitchers who are comfortable from the stretch, but I think this boils down to how much credence you’re giving to the claim that “Another team defense allows the exact same number of hits and walks. They even allow them in the exact same way.” Obviously it’s difficult to imagine a pitcher spreading out those hits and walks and throwing the same number of pitches from the stretch, but this whole thing seems to be an exercise in cultivating petri dishes of unlikely hypotheticals until they grow into something monstrous and worth commenting on. If timing truly is the “only difference” I’m saying these scenarios have the same value.
Right, I’m trying to focus on single things, as best as I can, trying to limit the variables, so we can isolate the issues. That’s why I’ll have more scenarios and more questions.
Just FYI, I voted as if it was about the stretch. Based on the comments, you’ve got a fair amount of contamination here.
THe way i have looked at these questions was determining how people view context as opposed to isolating events themselves. When it comes to this question though the context is important because it is a persistent event compared to the others which can easily be viewed as discrete occurrences. A walk is a walk and a home run is a home run. Allowing eight runs is more than one single event and thus the context of what happened does matter. If the objective is to say all the hits ere the same between the two situations, that is fine, but obviously the situations did have significant differences in how those plays transpired as the outcomes were different. Perhaps a single into RF occurs in both cases, but in the second example there is a guy on second who tries to make it home. Does the RF throw home to try to stop him, or play it safe and keep the hitter on first and limit potential future runs?
All of the other events were easy enough to consider as discrete and individual events. In this case yes there might be reason to think “randomness” occurred in leading to the sequencing of events, but the context as a whole matters when comparing the situations. The situations can only be so similar before the differences have to be considered. In this case I would rather not respond, but I would still have to consider the overall context and determine that the sequencing is unfortunate but necessary to consider.
Assume that there was one or two hits+walks in each inning for one team. That’s how they managed to score 0 runs. You got 13 LOB, or less as you got DP, etc.
The other team bunched up their 13 hits+walks in say tbree innings. 8 runs scored, and 5 LOB, or less as you got DP.
The question is centering around sequencing, and how much it matters in evaluating these two scenarios.
It isn’t just pitching out of the stretch. It’s pitchers reaching back for a few extra MPH in a really important situation (which they would blow themselves out if they did 20-30 times a game), and it’s the ability of a team to bring in particular strong relievers in really important situations (including the ability to induce double plays). It’s also giving more relative weight to the ability of infielders to make plays when the infield is drawn in and to make double plays. It’s also giving more relative weight to outfielders’ arms, since the strength and accuracy of outfielders’ arms are more likely to be important with runners on base, since this involves throwing to 3rd or home, whereas with the bases empty it’s almost always just trying to stop a hitter from stretching a single into a double, which is a significantly shorter throw (other than left-field-to-3rd-base).
I voted for the runs allowed version. I’m really starting to lean towards extending the WAR framework to include context dependent events. From a certain perspective, we know explicitly how much ‘value’ each team has during the season – we know exactly how many runs they scored and allowed, and how many wins they had. The runs and wins calculated in the WAR frameworks are abstract things, and don’t necessarily match the real-world equivalents. So we end up with the following equation for each team:
WAR_existing + = Wins
If the WAR framework was extended to calculate the context-dependent portion of each play as well as the context-neutral portion, then you could improve the above equation:
WAR_existing + WAR_context + = Wins
For me, this provides several benefits. First, I think extending the framework in this way provides clear direction for making WAR (both components) better (ex. work to minimize the unknown parts). Second, it provides a clear method for judging improvements/changes (or at least exposes the tradeoffs involved in a change) to the various current WAR valuations. Third, it ties the framework directly to the reality on the field, which I think could be very helpful in getting more buy in from people outside the sabr community.
tangotiger, are you taking poll suggestions?
Sure, go ahead.
Two lines of thought
1) Does the pitcher/batter matchup matter? For example, is hitting a home run vs Jake Arrieta worth more/same than a home run vs Ian Kennedy? Is giving up a homer to Ben Revere different than Chris Davis? Is this part of the context that WAR metrics should attempt to neutralize?
2) If the matchup matters, then does the specific pitch matter? Good pitchers occasionally throw bad pitches, and vice-versa. Is there a difference between hitting a home run off of a flat, down-the-middle slider from Chris Sale versus this http://www.fangraphs.com/blogs/a-home-run-that-must-be-discussed/? Is this part of the context that WAR metrics should attempt to neutralize?
My vote was that sequencing matters. Decisions are made about how to set up the defenders and what pitches to throw based in large part on the current base/out state. Those decision influence the result of the play. Impossible to suss out? Probably. But in an ideal WAR calculation we’d get it figured out.
This is basically asking about FIP-WAR vs. RA9-WAR. And the correct answer is somewhere in between. Although I would weight the individual events more heavily than the runs allowed, I still chose the second answer in the poll because I think there needs to be some weight for runs allowed due to teams having some control over sequencing even though not a lot.
This is one to me where context-neutral is a problem. Part of a team’s defense (and mostly in this case on the fielders as a team – not the pitcher) is understanding the context of the game and adjusting to it so you prevent runs. Looking at the game-events one-by-one – each of them (eg runner on second, one out, power lefty at the plate, etc) creates a context which a good defense will adjust to and succeed and a bad defense won’t.
Assigning the results to an individual player might be problematic because each game event and each adjustment creates different burdens on different players/positions (before the pitch is thrown) and you can’t end up knowing which player/position really performed its job the best based purely on the outcome of the event.