Build a Better WAR Metric: Neutralizing Players
Larry Walker is a great hitter.
He’s a great hitter at Olympic Stadium. He’s a great hitter at Coors. He’s a great hitter at Busch. He’s a great hitter at any ballpark named after a beer.
Whereas the average hitter might create 120 runs per 162 games at Coors, Walker would create 190. That’s +60%.
Whereas the average hitter might create 85 runs per 162 games in every non-Coors park, Larry Walker would create 110. That’s +30%.
When you evaluate Larry Walker, you have two choices:
1. Neutralize Larry Walker by giving him 268 plate appearances in each of the home ballparks of the 30 MLB teams. His 2501 PA at Coors? Now we only count about 11% of that. His 32 PA in Oakland? We have to figure out how he’d have done if he got 268 PA. And so on.
2. Take a league average hitter, and put him in the same playing conditions as Larry Walker. Walker had 2501 PA at Coors? Great, let’s count it all. But let’s compare him to a league average hitter who also got to bat 2501 times at Coors. He came to bat 32 times in Oakland? Then the league average hitter also came to bat 32 times.
So, what are the strengths of these two options. In Option 1, we don’t allow Larry Walker to take “unfair” advantage of a park he might be ideally suited for. Whereas most hitter would increase their runs created by 40% at Coors relative to a non-Coors park, Larry Walker increased his by 70%. Given that he got 2501 PA at Coors, Walker ends up shining more than he would otherwise. It’s like letting Mariano Rivera come in 1- and 2-run games, while letting Trevor Hoffman and Billy Wagner only enter blowouts. They are all suited to close games, but if only Mo gets to leverage that, he’ll be the one getting all the saves. Is that fair? I dunno.
In Option 2, we deal with what actually happened. We don’t have to play a game of what-if. We simply accept what the player did, and that he was able to leverage (or not leverage) his unique playing conditions. All we do is make the comparison level all those thousands of players who also played in those exact same conditions, at the same frequency as our player. Walker@Coors is compared to average player at Coors, and Mo in high leverage situations is compared to the average reliever in high leverage situations, and so on. Everyone gets to keep what they did.
So, how do you want to see it?
“How do I want to compare players” for what purpose? Option 1 seems obviously better for predictive forward-looking talent evaluation/projection, Option 2 seems obviously better for retrospective comparison. Do we need to pick one option for both? If so, why?
Someone people have only one purpose, or prefer one purpose.
Though I wouldn’t be so quick to say that Option 1 is better for projections. You’re severely underweighting about half of the sample size while expanding the relative weight of the rest. That could be a huge issue for players with just a small number of plate appearances at any given ballpark – if Walker only had a .113 wOBA for his 32 PA in Oakland, that’s a meaningful distortion.
You wouldn’t necessarily just pro-rate.
* regress
The results here run contrary to all the other polls.
For the other polls, people did not want to reward better performance in higher leverage situations, or who tailor their performance to their environment. They just wanted to inject the player in standard conditions.
Here, at a 64-36 split, we get the opposite reaction. We want to reward Walker for taking advantage of Coors more than other hitters have taken advantage of it. We do NOT want to inject the player in a standard park.
***
Overall, it basically comes down to that you want to have two different paths, and the reader will get to choose how to weigh the two based on however it is they see the issue at the time they see the issue.
Maybe 25% always see it one way, and 25% always see it the other way. And 50% want some sort of mix in some sort of way.
I would think it’s because many saber-inclined fans have a strong anti-clutchiness bias. The non-existence of clutch hitting is such a major talking point, I know that, at least when these polls began, it drove the majority of my decisions.
In its case, it’s a question that we don’t often talk about, and most people are less likely to vote based on preconceived biases.
I think I’m being consistent.
The reason I want to take out leverage in most situations is because I don’t think hitting well in high leverage is a measurable skill.
If you put Larry Walker into a neutral context by averaging what’s “expected” in various fields rather than what he actually did, you’re admitting that hitting well at Coors ACTUALLY IS a measurable skill.
Put measurable skill into WAR, try to eliminate luck.
I think this is a great perspective. So, it’s not as if the voter is being inconsistent.
He’s simply asserting that for those things that the batter or pitcher has a skill associated to that context, then he wants that included. If he doesn’t have a skill associated to it, then don’t include it.
With respect to the park, it seems that the readers believe there is more skill to leveraging a park than there is to leverage a crucial situation.
It’s not that I believe leveraging a park is possible, it’s that the very idea of correcting for park conditions ASSUMES it’s possible, and then wants to discount that skill.
Otherwise “neutral conditions” is meaningless since any apparent advantage to Walker in Coors is simply coincidence and his “true skill” would be reflected in average park adjusted numbers treating all at bats equally.
Thank you for posting this followup comment. For a little bit there I thought I was losing my mind.
In EVERY single one of the other polls, I voted with the majority (w/o looking at results first) and then this one I ended up on the other side. I had to reread the whole document to see if I missed something or had it backwards in my head or something.
Anyway, yes interesting that result seems backwards than the previous polls. (Or maybe you are just getting better at writing the situation to get closer and closer to 50/50 splits 🙂
I really think the problem here is imprecision about what “better” means — if you’re fuzzy about “better to what end” you get fuzziness in the answers reflecting different unstated priorities or implicit prejudices about that. Maybe this is a fuzziness built into WAR but it’s worth talking about as a separate question from “context” and “leverage” rather than conflating them all into the same question and then calling your poll respondents (rather than the premises of the questions you’re asking) inconsistent.
Good point, and I just responded to the previous reader to that effect.
Basically, I like wRC+. If a guy like Dustin Pedrioa is specifically suited for fenway so be it.
The basic problem I have, in addition to Lampert’s statement above (to which I would add the possibility of a player not only being suited to an environment, but tailoring his approach to that environment), is that players tend to perform better at home vs on the road, if I’m not mistaken. Whether or not it is tied up in being better suited to one’s home park or due to something else, like a slightly higher rate of missed pitches on the road, I want it counted.