Statcast and the Future of WAR

Over the weekend, I had the fortune of attending the Sloan Sports Analytics Conference, and participating on the baseball panel with Mike Petriello, Harry Pavlidis, Patrick Young, and Brian Kenny, which was a lot of fun. While the baseball panel was my only actual obligation at the conference, Petriello was doing double duty, having just presented — along with Greg Cain, one of the lead engineers at MLBAM — the latest update to Statcast, and introducing two new public metrics for 2017, Catch Probability and Hit Probability. These are the kinds of numbers people have been hoping for, and are one of the first steps in moving from collecting interesting single data points into providing more valuable calculations based on the combination of factors the system is measuring.

To help promote the new metrics, Jeff Passan wrote a piece on Statcast over at Yahoo, focusing mostly on what Statcast could do in the future.

Sometime soon, there is going to be a new version of Wins Above Replacement available, and its goal, aside from encapsulating a player’s value into one tidy number, is simple: Don’t be scary. The plan does not involve dumbing down the metric that serves as the flashpoint between those who yearn for a catch-all and those who lament it. On the contrary, as with almost everything it does, Major League Baseball Advanced Media wants to make it so smart people can’t help but like it.

That’s part of the excitement: Defensive WAR has been more guesswork than exact science. Statcast exists for exactitude. Even better, Statcast takes only 10 to 12 seconds to give a play’s precise details, meaning before the next pitch anyone who cares to will be able to contextualize just how good – or at least rare – a catch really was. BAM’s data warehouse then can be queried to provide context, and highlight clips of similar or better catches can be compared and contrasted on demand.

This is, undoubtedly, an exciting future, and the idea of a Statcast-based WAR system is very intriguing. The current versions of WAR still struggle with the difficulty in separating run prevention credit (and thus value) between the pitcher and the fielder. Statcast’s tools seem likely to bridge that gap, and with hit probability and catch probability — though it should be noted, the latter is outfield only right now, as infield calculations are more complicated — we are now closer than ever to being able to build metrics that directly measure the quality of contact a pitcher allowed, and adjust both the pitcher and the fielder’s contributions to the play made (or not made) based on that important variable.

So, yeah, Statcast is going to improve WAR calculations in a significant way, and should allow us to move past the FIP/ERA divide in the not-too-distant future. But perhaps more interesting is Passan’s mention that the guys at MLBAM are dreaming of their own WAR metric, and what that might look like down the line. The potential of a Statcast-based WAR model brings up a fascinating question; how granular should WAR get?

As Tom Tango has said on a number of occasions, WAR is a framework, giving the basic building blocks of adding hitting, baserunning, pitching, and fielding together in a systematic way. But the actual values that go into those components can be determined in a number of different ways, depending on the type of calculation that is being attempted, and more importantly, the question being asked.

Different numbers answer different questions, and have varying uses for determining what happened in the past or what might happen in the future. Often times, metrics are grouped into “descriptive” and “predictive” buckets, depending on whether they are trying to account for what did happen or what we may expect going forward. WAR, generally, is a descriptive metric; it is trying to measure the value a player produced in a given season, not tell you what his value will be next season.

Which makes the idea of a Statcast-based WAR model pretty interesting, because a lot of the presumed value of Statcast’s data is to allow us to say “okay, that result happened, but based on the more granular data, we’d have expected this other thing to happen.” Hit probability, for instance, is going to let us say that a particular batted ball might have been caught by a diving outfielder’s spectacular effort, but that 85% of the time, that ball lands, and the hitter got robbed by something out of his control. Right now, every publicly available version of WAR simply records the play as a negative-value event for the hitter, despite the fact that he just got screwed by a great defensive effort.

Certainly, knowing that the ball normally lands for a hit is valuable information in evaluating that player’s performance. But is there any value produced to a team in hitting a ball that probably should have, but didn’t, land for a hit? Do we want to credit a hitter for what was just under his control, or do we want to credit him with what happened on the field during his at-bat?

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

Or, let’s think about it from the opposite perspective. During his presentation on Saturday, Petriello showed this Andre Ethier home run from last year’s NLCS.

According to the new Statcast Hit Probability calculation, a ball hit at that exit velocity and launch angle is an out 95% of the time. Because of a favorable wind and the fact that the ball was hit in one of the few stadiums where a 353-foot fly ball to left center would clear the fence, Ethier actually got his first home run off a left-hander in three years.

In terms of value, what do you do with that play? Ethier’s home run put a run on the board for the Dodgers, so — ignoring the fact that we don’t have postseason WAR right now — our calculation would give him the same credit for that play as if he had launched a 500-foot bomb onto Waveland Avenue. And our pitching WAR would correspondingly crush Jon Lester, who gave up the home run, even though Lester did his job and induced contact that is almost always an out.

The easy answer is to say a home run is a home run, and we don’t care about what should have happened, only what did happen, and what did happen is that Ethier rounded the bases. In general, WAR is an attempt to isolate player performance from the influence of his teammates, but it is not designed to strip luck out of the equation. And if that’s what we want WAR to do, then perhaps a Statcast-based version wouldn’t be so dramatically different from what is already out there, beyond separating pitching and defense in a more accurate way, anyway.

But it’s actually more complicated than that easy answer would suggest, because right now, there isn’t a publicly available of WAR that really is calculating “what really happened”. The versions published here, at Baseball Reference, and at Baseball Prospectus all use context-neutral run values at the event level, so while Ethier’s home run really added one run to the Dodgers ledger, he’d get 1.4 runs worth of credit in WAR for hitting that home run, because we don’t think it’s his fault that there weren’t any runners on base when he hit the home run, and in general, the average home run produces about 1.4 runs worth of value for an offense.

If we took the “measure what really happened” argument to its logical conclusion, then there’s a good argument to be made that something like RE24 — which gets run values from the base/out state, not the overall average — should be the foundation for the hitting component of WAR. And once you go to base/out context included, you can continue down that path to including inning and score, and argue for WPA instead, since if we’re giving a hitter credit for the situations he hits in, a walk-off grand slam does more to help a team win than a solo homer down by 10.

The reality is that, with almost every component in WAR, you have to decide how much situational context you want to include, and the more context you include, the more credit you give to a player for something he had nothing to do with. And that brings us back to the Ethier home run. He didn’t really have control over the wind carrying his weak fly ball into the seats. So there is some logical consistency in saying that if we’re not including base/out/inning/score context because those things are out of the player’s control, perhaps we don’t want to measure a player’s contribution based on luck-based outcomes.

So an entirely Statcast-based WAR, that measured value solely on the granular data and probability that we think are within the realm of the player’s control, could be fascinating. I don’t know how popular a model that gave Ethier negative value for hitting a home run in a playoff game would be, but it would be a really interesting departure from every other version of WAR out there. And it would be the only model that stripped luck out of the picture, giving us perhaps the best view of a player’s actual contribution to an outcome.

But it would also be a radical departure from what people have said they generally want WAR to be. In responding to the question about how much value to give to a player who hits a triple and gets stranded versus a triple where the next batter drives him in, 95% of our readers said they wanted those two triples to be given the same value, and they didn’t want the run value to be based on the future sequencing after the event occurs. Given that Tom Tango, who now works at MLBAM, ran that series, I would imagine he’s going to be influenced by those ideas when helping craft MLB’s version of WAR, whenever that is.

This tension between what happened, how it affected the scoreboard, and how much credit to give to players for things they don’t control is a difficult thing to resolve. And now that we’re getting even more granular metrics about a player’s contribution to outcomes, these questions are going to continue to be relevant. Knowing the guys at MLBAM on a personal level, I’m pretty comfortable with the fact that they’ll handle these questions thoughtfully, and if (or when?) they do produce an MLB.com WAR model, it will be with all of these questions answered as well as they feel they can.

But while Statcast holds a lot of promise for improving the pitching and defensive sides of the components, getting ever-more granular hitting data might force us to again ask what we want WAR to be, and what the goal of the model is. There is no obvious right answer here, and that’s one of the reasons there will always be multiple ways of calculating WAR.





Dave is the Managing Editor of FanGraphs.

62 Comments
Oldest
Newest Most Voted
Inline Feedbacks
View all comments
Oblarg
9 years ago

It would be really nice if they’d tell us how they’re actually calculating these statcast-based “probability” stats. Early last year, back when I had more time on my hands, I had a go at this and wrote a blog post that covers most of the concerns with what they’re doing:

http://45percentmental.blogspot.com/2016/05/on-development-of-expected-ball-in-play.html

It’s not clear to me what the choice of kernel or bandwidth is on these stats – that’s kind of important. The descriptions make it seem as if it’s constant-bandwidth, which is…not the best choice, given the uneven distribution of batted balls in trajectory-space.

(I’d love to go calculate my own version of these stats with a full year of data, now, but unfortunately calculating the bandwidths for the kernel regression on my desktop would take something like two weeks – O(n^2) problems are unforgiving!)

Chase Hampton
9 years ago
Reply to  Oblarg

From the article on “Hit Probability”:

“What is the percentage based on?
It’s based on the Major League average for the combination of exit velocity and launch angle over the two seasons of Statcast™, and it includes a smoothing process to include larger samples. For example, for the individual pairing of “100 mph and 30 degrees,” we looked at all balls within 4 mph and 4 degrees, with proportionately greater weight to those balls closer to the 100/30 pairing.”

Oblarg
9 years ago
Reply to  Chase Hampton

Yes, I saw that.

So, this tells us that the bandwidths at that particular point are 8mph/8degrees, but doesn’t tell us if they’re constant over the whole trajectory-space, or how they came to those bandwidths. Worst-case scenario is they just “guessed” at a pair of constant bandwidths, which is a pretty terrible way to do it – bandwidths should always be chosen through some systematic data-driven algorithm, and variable bandwidth is quite clearly the way to go here.

“Proportionally bigger” could be interpreted in two ways – either the kernel is triangular (i.e. there’s a linear increase in weight up to the point itself), or else it falls off like 1/d where d is the distance. The former is sort of an odd choice of kernel, but should work reasonably well. The latter would be *extremely* strange. I’m pretty sure they mean the former.

jw757Member since 2024
9 years ago
Reply to  Oblarg

I believe you’re over-thinking this one. It’s likely as simple as assigning a 0 for hit and a 1 for a catch for every play in the database. Then just setting bins, likely in 0.1 second increments and 1 ft increments for distance traveled, and wah-la divide the sum by the total for each bin. Code it or make 2 pivot tables in excel and determine the probabilities in 10 minutes. Of course we don’t have the distance-traveled (due to OF starting point) or hang time as part of the public data set. Of course, just like you did, we can sub in angle and velocity and do a similar analysis. Personally I think the angle/velocity catch probability is a better measure than this one using hang-time and distance anyways. Since hang-time is determined by angle/velocity (might as well have the insight 2 variable offers instead of their combination in hang-time) and OF distance really just adds in more problems than it fixes (who is responsible for OF starting position and how does that impact the play? I.e. he’s so good at positioning he never had any long runs or he’s only good at positing because it’s a team determined thing and that OF, on his own, would not position himself in that way).

Oblarg
9 years ago
Reply to  jw757

Naively tiling the space with bins is a really, really bad way of doing it for a number of reasons – you get nonsensical discontinuities at the boundaries between the bins, for one, and if you use a uniform bin size you’re almost guaranteed to horrendously overfit in regions where the data are sparse.

carterMember since 2020
9 years ago
Reply to  Oblarg

It is as if you are speaking a different language.

Oblarg
9 years ago
Reply to  carter

That’s a pretty common phenomenon when reading stuff about technical fields one isn’t familiar with.

jw757Member since 2024
9 years ago
Reply to  Oblarg

you asked what they did. I’d wager it’s only slightly more complicated than I outlined for you. You can disagree with that methodology for any number of reasons, you’ve posed some problems you think there are with that approach. But it’s very likely what they did to calculate the probabilities (which “It would be really nice if they’d tell us how they’re actually calculating these statcast-based “probability” stats” is what you asked). As the hit article, and the earlier reply noted, the only thing they did above what I mentioned was to make slightly larger bins and use some weighted-average to smooth them so as to address some of your aforementioned concerns. Maybe I’m wrong, and some very similar approach is not what they’ve done.. but I personally doubt it.

piddy
9 years ago
Reply to  jw757

What you’re describing is a naive approach to kernel regression without implementation details. The concerns that Oblarg is posting about are exactly those implementation details.

jw757Member since 2024
9 years ago
Reply to  piddy

https://baseballsavant.mlb.com/statcast_catch_probability

the math is shown. it’s pretty much exactly as I described. tiles based on increments of .1 seconds and 5 feet. (Balls Landed) / (Batted Ball Events) = Probability

Note this is not a regression equation of any kind. Despite all the mini-nate silvers we’ve got commenting just dying for it to be.

You may not like the approach and have a lot of issues with how you would go about calculating the regression equation, but I’m not describing creating a regression equation which would have the issues you noted. ANd a regression equation is not what the creators have done here.
instead we’re talking about a table. A literal table based 100% on actual data. Not a function of distance and hang-time to generate a probability.

Bottom line, I clearly mis-took Oblarg’s post to be an honest question about how they did the math when it was really just a faux-question aimed at giving him the opportunity to talk about how he has tried and enjoys building a statistical model based on Kernel Regression to output the probability in lieu of just going to the appropriate location on a table of historical rates. That’s a fine and interesting discussion, but my mistake for not recognizing the comments true intent earlier.

Justinw303
9 years ago
Reply to  Oblarg

Can I get a ELI5 on this?

orval
9 years ago
Reply to  Oblarg

The right way to do these calculations was presented at hardballtimes a while back and uses exit speed plus launch and spray angle in 3-D to get wOBA with cross-validation used to get the optimal bandwidths. The result is known as the wOBA cube as seen at

http://triplesevenproductions.com/daily-fantasy-sports/sabermetrics-dark-full-terrors/

Oblarg
9 years ago
Reply to  orval

Yep, this is essentially what I did (though my analysis was restricted to 2D).

dtpollittMember since 2016
9 years ago

Things I learned from this article and the Passan one:

1 – A visualized WAR tree is an awesome idea and something I can immediately see myself manipulating and showing non-statheads. Me want.

2 – What do I want WAR to be? I want to be able to compare player quality and performance on an equal playing field. Sometimes I want to compare same context, same team members, same pitchers / fielders, same field dimensions, same positions, same historical differences. I want to be able to compare, contrast, and argue about the quality of players and use WAR to help do that. I do not have an answer re: Either’s home run. That is probably not a home run in most parks and situations, like you said. So do we give Either credit? Or devalue Lester? What about having a “luck” factor? You can see catch or hit probability, but is it possible to measure the luck-ness of a play, or is that already within probability? I don’t know.

3 – Tom Tango isn’t his real name? I bought and read THE BOOK and that’s not his real name? Is Dayn Perry Tom Tango?

Brian ReinhartMember since 2016
9 years ago
Reply to  dtpollitt

His real name is Sam Samba.

mike sixelMember since 2016
9 years ago

I don’t know, a HR is a HR….just as a strikeout is a strikeout. How far do you take this? If 85% of the time an 85MPH FB right down the middle is hit, but in this case it is a strikeout, is there still a strikeout? What if an umpire misses a call that is correctly made 99% of the time? Does that get “reversed”? How would that even work if that mistake resulted in a walk or strikeout, since there are no future pitches? I’d think “true” outcomes should stay. Where do you draw the lines? 51%? 75%? Three sigma?

giantsprospectsMember since 2023
9 years ago
Reply to  mike sixel

Agree with this, it opens up a can of worms. But I would still like to see it, likelihood of outcomes is an interesting variable to add to the pot.

swingofthings
9 years ago
Reply to  mike sixel

Not really a fair analogy. In this case, Ethier hit a routine fly ball and the wind – a factor outside any player’s control – pushed it out. If a batter swings and misses at an 85 MPH FB right down the pipe, what factor outside of the players controlled that?

brad.vargas
9 years ago
Reply to  swingofthings

Mike isn’t talking about swinging and missing, he is referring to a batter that gets takes a ball out of the strike zone, but the umpire calls him out for strike 3. Take that a step further, how about a batter on a 0-1 count and he takes a ball outside, but it’s called a strike. Now he is down 0-2 and has to shorten his swing. That bad call directly impacts that player’s AB, but is not as obvious as it being on the third strike.

What about a batter that is 0-2 and the next pitch is a strike, but it’s called a ball. Then he drills the next pitch out of the park. His metrics will show the AB was very productive. When in fact it should be a K. Ultimately the result of the AB is the result of the AB. The anomalies even themselves out over the season/career. That is why you can’t take SSS seriously. Either in actual results or metrics.

GoNYGoNYGoGo
9 years ago

Dave,

Very interesting article. What would be nice to see is that rather than a new-WAR discussion would be using the new hit/out probability stats to measure clutch.

Ethier hit a HR in the play-offs, how clutch, despite the wind, etc. Next better hits a frozen rope down the line that a 3B catches. (Measured as un-clutch) Next one hits the ball 400 feet that I caught over the wall (also measured as un-clutch).
Pivoting the probability of hits (or non-hits) with a WPA measurement would be a great way for a Tom-Tango like person to look for a clutch factor.

Zach Walters Appreciation Guild
9 years ago

RE24 should be the basis for Bonds WAR, if nothing else.

Look at his WRAA! And then look at his RE24! The gap!

scotman144Member since 2016
9 years ago

Much like we currently have WAR (FIP based) and RA9 WAR (ERA based) for pitchers I’d be in favor of aWAR (actual WAR calculated by giving value to what happened per RE24 base/out states) and eWAR (expected WAR calculated by the expected run value of each event: e.g. 5% of a HR for Ethier in the scenario above).

Give me a leaderboard of who actually has impacted games the most (aWAR) and who could most reasonably be expected to do so going forward (eWAR) and we are all good.

Unrelated: Give me “bases probability” not “hit probability”. Expected SLG is a more useful number than expected AVG IMO.

dtpollittMember since 2016
9 years ago
Reply to  scotman144

I think this is a great idea, but sadly it does nothing to promote, as Passan says:

“Don’t be scary”

and

“What will make our version of WAR intriguing,” Willman said, “is the way we’re going to make it accessible.”

scotman144Member since 2016
9 years ago
Reply to  dtpollitt

If “what actually happened” and “what’s likely going forward” being two distinct concepts is “too scary” I give up on the population at large.

Call me skeptical that they’ll manage to blend the two intents /questions posed into one figure but I’m excited to see them try.

dtpollittMember since 2016
9 years ago
Reply to  scotman144

Yeah I agree with that. It IS exciting to see “barrels” discussed on MLB Network, though. Good first steps from BAM, I think.

Six Ten
9 years ago
Reply to  scotman144

Yeah, I like all of this.

scooter262
9 years ago

I like WAR to give credit/blame based on what actually happened. Teams should go with the most predictive stats they can get, but I think fans want more of a “value produced” stat. And I think WAR should stay as context-neutral as possible.

Jaack
9 years ago
Reply to  scooter262

I think this is a little contradictory. Context-neutral stats like wOBA and wRAA are inherently based on things that didn’t happen – a player is awarded ~1.2 runs for a double, no matter the context, but some doubles are worth substantially more than that, and others are worth substantially less.

Don’t get me wrong, linear weights is a really great model, but it gets treated as gospel truth a little too much I think – there’s plenty to be learned from moving the needle in both directions.

Personally, I’d love to see both a statcast-based WAR and a RE24 based WAR.

rlwhite
9 years ago
Reply to  scooter262

There are many fans that want the most predictive stats too. Fans love to predict: how their team will do, whether this or that FA will be worth a contract, how prospects will impact their team, etc. And I haven’t even mentioned fantasy baseball yet.

swingofthings
9 years ago
Reply to  scooter262

I want to use stats to get as close to measuring a player’s true talent as possible, which is really the most basic goal of measuring them at all. And predictive stats in the vein of FIP are more adept at that than value-produced stats like ERA.

bigriggs42
9 years ago

Hypothetically, what if Ethier knew how strong the wind was and intentionally tried to hit a fly ball? Is that still considered luck or did he have some influence over the outcome?

That makes me think of another scenario: what if a lefty purposely tries to hit the ball the opposite way due to a shift, hitting a weak ground ball to 3B that 99% of the time is an out? Is that not skill and an earned hit rather than luck? Now that I think about it I’d assume player positioning would be calculated with ground balls so maybe I answered my own question.

Damon G.
9 years ago
Reply to  bigriggs42

I came to post a question very similar to yours. The biggest problem I have with stripping away too much context is that ballplayers might be specifically *using* that context (perhaps even subconsciously), and so it’s unfair to dock them for it.

Your Ethier example is a good one. Another example is a pitcher who knows he has a great defense behind him, so he’s not always trying to strike out everybody. If Kevin Kiermaier makes a great catch, you might say the pitcher got lucky, but the pitcher *knows* Kevin Kiermaier can make such catches. If Kiermaier wasn’t there, the pitcher’s entire approach might be different. So is it fair to count this against the pitcher?

swingofthings
9 years ago
Reply to  Damon G.

I don’t think the Kiermaier example holds as much water. I’m fairly certain a pitcher is always trying to get routine-out contact or avoid contact altogether. I can’t prove this doesn’t happen any more than you can prove it does, but it’s hard for me to imagine a pitcher ever throws a pitch with the intention of allowing hard/line drive contact, or that they would ever consider a Kiermaier diving grab to stop a double a success for themselves.

Mostly I could just see this happening in hitters’ counts – on a 3-0 pitch, the pitcher might lob it in there, thinking that it’s better to leave it up to defense than the guarantee the batter a base. But should I don’t think this desperate plan makes the hard-hit gapper a win for the pitcher, even if Kiermaier runs it down.

Damon G.
9 years ago
Reply to  swingofthings

“it’s hard for me to imagine a pitcher ever throws a pitch with the intention of allowing hard/line drive contact”

Of course no pitcher would want this outcome, but they might pitch in such a way that it’s more likely to happen than it would be if Kiermaier wasn’t in center field. For example, a pitcher with a great defense might be more aggressive at staying in the zone. The advantage to this is not that they will give up hard hit balls that will be miraculously caught; it’s that they will be less likely to walk hitters. But the only reason the pitcher feels comfortable doing this is because *if* a ball is hit hard it’s more likely to be caught than it would be with an average defense behind him.

Do pitcher’s actually think this way? I don’t know. I don’t even know how one would find out, as the approach pitchers take might not be conscious strategy, but rather “feel of the game” honed by years of playing baseball.

My main point is that I think you run a big risk in stripping out too much context of devaluing pitchers who are good at using the context to their advantage.

vtsoxfan
9 years ago
Reply to  bigriggs42

Good comments. Dave made a similar, if slightly different, point a year or two ago about catcher framing. If we actually gave catchers direct credit for framing runs, we’d have to subtract those runs from the pitcher, since we’re saying that the catchers are more responsible for creating those runs than the pitcher is. But it’s likely that pitchers who have good framing catchers have learned to trust their catchers to steal them strikes and are thus changing their pitch mix to take advantage (largely by throwing lower in the zone). A true assessment of that pitcher’s skill wouldn’t simply turn those stolen strikes back into balls and punish the pitcher accordingly, since presumably the pitcher would have thrown different pitches if he’d had a worse catcher. It’s all about figuring how much players really leverage their context.

Arjon
9 years ago
Reply to  bigriggs42

Along similar lines, the thing that came to my mind was LHBs trying to hit oppo balls in the air at Fenway.

stonepie
9 years ago
Reply to  bigriggs42

i would imagine (or hope) that statcast adjusts for shifts and corrects for positional alignment. so if a guy bunts down the 3rd base line with no defender in sight, odds are thats a hit. he deserves credit for producing a state where a hit is likely to occur, just like a guy hitting a line drive in the gap. The issue is, both are still hits. a iso/slg/woba version of hit probability needs to be incorporated as well.

SamF
9 years ago

Why not have two versions at opposite ends of the spectrum? You guys already do something similar for pitching WAR, providing both RA9 and FIP based WAR.

Johnny Dickshot
9 years ago
Reply to  SamF

Yeah, it seems that there’s a pretty simple solution for the conundrum Dave posits.

Do one version of WAR where, e.g., Ethier gets credit for a HR and the pitcher gets credit for giving up a HR, and one version where Ethier gets credit for the expected run value of the ball he hit (which would be very low in this example) and the pitcher gets gets credit for expected run value of the weak fly ball that he induced.

Fans would clearly want both versions.

bartelsjason
9 years ago

On the other hand, Ethier and Lester would both be aware of the wind and fence, thus both them adjusting their approaches for this matters. Either knows its a short fence and that it wouldn’t take much for the ball to fly out of the park in that location due to the wind. A hitter should not be dinged too much for possibly just taking advantage of the elements presented to him, as he may choose differently in other elements.

GoatHerderMember since 2016
9 years ago
Reply to  bartelsjason

I was thinking the exact same thing. How does this system account for players adjusting to the factors outside their control.

Six Ten
9 years ago

One of the things that’s hard here is that WAR has been, properly speaking, a descriptive statistic rather than a predictive one, although we might think of it predictively when it comes to things like free agent signings, starters, Hall of Fame odds for active players, etc. The temptation to think of it predictively gets stronger as the precision of the data available gets stronger. Statcast-based WAR would raise that temptation exponentially, because what amazing data. But really, WAR is still trying to tell us what happened, not what will happen next time.

That swing from Ethier did not result in an ideal flight path, but you can’t describe the outcome of the at-bat as negative for the Dodgers. A version of WAR that calls that outcome negative begins to cross from the observable realm into the platonic realm of forms, and while that’s helpful for predicting I don’t think it helps us understand what Ethier contributed to his team winning or losing that game, that series. And if it doesn’t help us understand that, I’m not sure it’s doing its job.

I’m aware there’s a contradiction in my own thinking, since I also think it’s fine to call it worth 1.4 runs, even though it only added 1. But where I think I want to resolve that contradiction is here:

By assigning a win value to every play, you imply that every play is a discrete event. If every play is a discrete event, it can’t possibly be Ethier’s fault that someone failed to get on base in front of him. If, conversely, the outcome of only 1 run means his hit shouldn’t get as much credit as a grand slam would imply that plays are not discrete events. And if they are not discrete, then you can’t capture a WAR value on a play-by-play basis.

Which is to say: outcomes that are discrete to a given play seem like things that should factor into WAR. Outcomes that are not discrete to the play should not. Process inputs are valuable as long as they don’t lead to discounting play outcomes.

drew_willyMember since 2020
9 years ago
Reply to  Six Ten

You wrote “outcomes that are discrete to a given play seem like things that should factor into WAR. Outcomes that are not discrete to the play should not. Process inputs are valuable as long as they don’t lead to discounting play outcomes.”

I am inclined to agree, but only because of what I want out of WAR – a figure that tells me what, in theory, a player contributed to his team. I want to know what to give him credit or blame for. This though, leads me to have conflicted intuitions; I sort of want to agree with you here but sort of want to disagree. Take the Ethier example. Sure, we shouldn’t credit Ethier based on the base-state before him when he hits the HR; here, I think we agree, since you suggest the base-state isn’t his “fault”. But should we credit him for the HR at all considering that the wind wasn’t his “fault” either? I feel the pull towards saying “no, he should be blamed for a poorly hit ball, and boy did he get lucky.” And, mind you, I am a Dodgers fan who was deeply invested in the outcome of the game, in what “actually happened” (as many who poo-poo “non-descriptive” or “forward-looking” stats like to say).

You talk of what Ethier contributed in the game and the series – “I don’t think it helps us understand what Ethier contributed to his team winning or losing that game, that series. And if it doesn’t help us understand that, I’m not sure it’s doing its job.” – but does ignoring the wind-factor aid us in seeing what Ethier contributed in the game? Again, I am conflicted, but I sort of want to say “no, it is better to take into account the wind and chide Ethier for a poorly struck ball. He deserves blame, not praise.” I think crediting Ethier for the HR starts to seem similar to crediting batters for reaching base on egregious errors or blaming pitchers for unearned runs.

It isn’t so much about being “descriptive” vs. “forward-looking”, as it is about crediting players for all and only those things for which they are responsible. It just so happens that, often times, doing so is also a better predictor of future performance than is allowing outside forces (weather, park dimensions, other players’ contributions, etc.) to influence what we think of a given player.

John Autin
9 years ago

I think a Statcast WAR would be interesting and useful. But I would rather it not become the first-cited WAR. I think trying to judge every outcome on the basis of neutral conditions (like park and weather) is too granular.

We expect a pitcher to adjust his approach to the conditions, at least somewhat.

And batters who are well-suited to their home parks deserve full credit for out-performing others in the same conditions. Wade Boggs hit many soft flies to LF in Fenway, partly *because* he knew the park often rewarded such batted balls. So I don’t think Statcast WAR can yet fully fill the “descriptive” bucket.

CamdenWarehouseMember since 2025
9 years ago

I’d love to see a new metric that takes luck out of the picture entirely. I wouldn’t call it WAR though. If that is left as is and the new metric is brought into the picture, people can gravitate to the one they like.

I’m really confused by the third to last paragraph. The cited poll question doesn’t seem to back up that readers don’t want a WAR model that strips luck out. Isn’t whether or not the runner is driven in a matter of luck? And 95% don’t want WAR to differentiate. Are you suggesting that the results from this question mean readers want a triple to be a triple regardless of how likely the outfielder is to catch it? Without asking that particular question, I don’t think the results can be used for this purpose.

Rols1026
9 years ago

So glad someone else pointed this out because that reference confused me so much. I would love to hear if anyone has an answer as to why this was included in the article.

david k
9 years ago

I have a question about the attempt to quantify an outcome based on what would have NORMALLY happened vs. what ACTUALLY happened. Say you’re a pitcher on the Royals, where you know you have a big ballpark and you have outfielders that can chase down a lot of balls. Perhaps you pitch to that situation, and not even try to strike everyone out. So if you give up 7 batted balls during a particular game that all turn out to be outs, but in an “average” ballpark with “average” defenders, 6 of those would have been hits, should the pitcher be penalized for that? Maybe he would have pitched differently under different circumstances? So I think it’s almost impossible to achieve a perfect result in all of this.

Rols1026
9 years ago
Reply to  david k

That’s why this is so difficult. The current versions of WAR don’t consider luck enough while this just takes it too far. Although I’m not complaining, this is interesting and keeps everyone thinking.

Fidrych
9 years ago

No matter how granular a metric gets, there are still too many unknown variables to presume the value of an event in a vacuum. There’s too much context built into the reality of a baseball event, and trying to extract even most of the context from the situation is futile.

Contextual factors exist in degrees, and the bulk of those factors can’t be honed down to the suggested level of precision. For one thing, we would have to be mind readers and using reverse psychology to attempt to ascertain what Lester and Ethier were both factoring in. In other words, all we can do is guess on those counts. So our supposed precision is cancelled out by our greater guesswork. There is no way to fine-tune speculation to the degree which is being proposed here.

Another manifestation of this is our loose interpretation of what constitutes a player having “control” over a particular event. Again, the dynamic of control comes in varying degrees. All pitchers and all batters have varying levels of “control” (“influence” would be a more realistic term) over every pitch and every ball in play. At our available level of precision somewhere around 95% (we can’t even know with precision what that is), it would be silly to think we could deal in the realm of 0.1% individual instances of precision and have those somehow “correct” our built-in margin of error.

Also at issue: defining what is or isn’t luck, or to what extent. We think we’re getting nearer our goal, but we’re just finding interesting rabbit holes to explore. Complexity does not necessarily translate into greater accuracy.

Llewdor
9 years ago

What happens when players adjust their approach to the context? If there’s a wind coming in from left, does that encourage a lefty to pull the ball, even if he normally wouldn’t? Any system that ignored all context (and I can see that appeal of such a system) would necessarily assume that player performance also ignores context. But then we know it doesn’t, because players can intentionally hit fly balls (or, for a more obvious example, bunts are intentional).

So ignoring all context probably isn’t the way to go. I think we should give hitters credit for homeruns (except inside-the-park homeruns), because they took advantage of the playing field as it was at that moment to hit a homerun. And we should penalize pitchers for homeruns (except inside-the-park homeruns), because they gave the hitter a ball he was able to hit for a homerun.

We should be willing to defer to Statcast in pretty much all cases of a ball-in-play, however.

DMT
9 years ago

It’s kind of amusing to me that the trend is to strip context from the sport that has 30 different stadiums, unlike any other sport. Isn’t context fundamental to the sport?

swingofthings
9 years ago
Reply to  DMT

The goal isn’t to remove the effects of context from the game, it’s to learn how to account for it when trying to figure out how good players are.

For example, one park having an offensive environment as extreme as Coors field is perhaps a positive for the game – it creates interesting variety. That doesn’t mean we should just make no effort to figure out how good Arenado actually is, separate from the influence of Coors.

rlwhite
9 years ago

I think there is room for both a descriptive WAR as we have now and a predictive probabilistic model that tells us what a player’s physical abilities could contribute on the field. I just wouldn’t want them to both use the WAR name. The difference needs to be clearer than that.

rlwhite
9 years ago

It could revolutionize the FA market if we could quantify the impact of footspeed on OF defense, differentiate it from reads and routes, derive aging curves specifically for foot speed, and then plug it back into a WAR model to predict the decline in an OFer’s defense over a 5-10 year contract.

rlwhite
9 years ago

There should be totally different names for descriptive WAR vs predictive WAR, like the difference between OBP and wOBA. Not simply fWAR and sWAR. They should use the same scale but be clearly different.

Paul22
9 years ago

Well, you can always have both context neutral and context WAR. Dont have to choose. One is used for best player analysis, the other for MVP.

As with statcast, a MLB WAR makes you worry about transparency. Statcast as I know is available to only a selected favored few. While the public WAR is not very transparent when it comes to defensive runs saved, due to the metrics themselves over BIS licensing restrictions, thats a small part of WAR.

Then there is the potential conflict of interest. MLB controlling the measure of a players value, or at least fans perception of that which can influence teams decisions. MLBPA should be raising holy hell over this and insist they get a seat at the table. But Clark just wants to get along and wont risk leaving the middle ground

.

shoewizardMember since 2016
9 years ago

Exciting stuff. I can’t wait to see more. Regarding the “tension” that Dave is referring to, for me personally there is no tension at all. As a fan, who loves the action, the story, the narrative, I love WPA, because that number coincides with the narrative. As an analyst, I like everything that strips out luck and gets us closer to “true talent” evaluation. They are just not mutually exclusive for me. The grand failure is one of communication. We just need to be clear about what we are saying. I’ll hand out my MVP awards heavily influenced by what happened in high leverage situations and measured however, (WPA, Clutch, etc). I’ll hand out my player contract extensions, free agent contracts, and pursue trades based on the true talent evaluations.

bndj88Member since 2017
9 years ago
Reply to  shoewizard

Exactly my sentiments. You can have both peacefully co-exist, and look at them in the appropriate context.

Joshua Northey
9 years ago
Reply to  shoewizard

Yeah this is not a hard concept. On my men’s league hockey team we have a few subs who don’t play unless the regular players cannot make it. One of them is terrible, so we try to not have him play if at all possible.

He has played two games and have 2 goals and 2 assists, which is great. It doesn’t mean he is a good player, he is still terrible talent wise.

Dissociating results outcomes from process should not be conceptually difficult for people who read fangraphs.

Joshua Northey
9 years ago

I don’t want to make too wild a suggestion here, but why not just have a diverse messy family of figures some of which are more descriptive focused and some of which are more predictive focused. As long as the Dave Camerons, and Rob Neyers, and Jeff Sullivans of the world stay on message and stay consistent with their use and description of what a particular figure means, I don’t think there is a problem.

Most people are idiots and you are never going to reach them, stop trying to. The conversation you are interested in leading/participating in is with the top 20-25% of fans, and just stick to that. Don’t bend over backwards to have just 1 or 2 unitary frankenfigures to make things simpler for people who aren’t going to understand it anyway.

Don't talk to me in the Uber pool, I dont know you
9 years ago

I think when we talk about this kind of thing we forget that there are thinking human beings playing the game in the moment. When a player sees a short porch, or knows that the wind plays a certain way, they may decide to change their approach to aim for that easier homerun, because they don’t care if in another stadium there fly ball would be an out. Say you are a Yankee, you may change your whole hitting approach (Brian Mccann) to hit fly balls to right field. If those fly balls were hit at At&t they would be easy out, but you don’t care about At&t or any of the other 28 stadiums you don’t play in because in Yankee Stadium you hit a homerun. I don’t think we should negatively view a player who has realized the opportunity cost of pulling fly balls to the short porch is too high to just try and go for the best possible contact. This to me is true of many context neutral stats, as much as we would like to try an determine how good a player is if he played by himself on a neutral sight, this isn’t the case.

WARrior
9 years ago

I’m sure it must be obvious to others that if we go down the path of stripping as much context as possible from events, the only things that will ultimately matter are exit velocity and launch angle. And at that point, why have defensive players on the field? All you need is a pitcher throwing to a catcher, with the run value produced by every hitter determined by Statcast, and run values added up to determine who wins the game.