Yes, Hitter xStats Are Useful

Some of the most frustrating arguments involving baseball statistics revolve around the use of expected stats. Perhaps the most frequently cited of these metrics are Statcast’s xStats, which use Statcast data for hitters to estimate the batting average, on-base percentage, slugging percentage, and wOBA you’d “expect” a hitter to achieve. Investigating how predictive xStats are compared to their corresponding actual stats has been a common research exercise over the last few years. While it depends on the exact dataset used, xStats by themselves generally aren’t much better than the actual stats at predicting the next year’s actual stats. But that doesn’t mean we should simply discard expected stats when trying to evaluate players.
While I’m not going spend too much time talking about how predictive xStats are versus the actual ones, I do want to briefly touch on some of the existing work on the subject. Jonathan Judge at Baseball Prospectus examined many of the expected metrics back in 2018. He also spoke with MLBAM’s Tom Tango about the nature of expected stats and their usage:
Earlier this week, we reached out to BAM with our findings, asking if they had any comment.
MLBAM Senior Database Architect of Stats Tom Tango promptly responded, asking that we ensure we had the most recent version of the data, due to some recent changes being made. We refreshed our data sets, found some small changes, and retested. The results were the same.
Tango then stressed that the expected metrics were only ever intended to be descriptive, that they were not designed to be predictive, and that if they had been intended to be predictive, they could have been designed differently or other metrics could be used.
One of my colleagues, Jeff Zimmerman, wrote about xStats in the fantasy context in 2018. Justin Mason looked into the data in 2021 and found that xStats are less predictive than actual ones.
It’s always good to have the most up-to-date information, so let’s start there. I pulled every player with consecutive 200 PA seasons since 2015; there were just over 1,800 season-pairs. I then ran the r-squared (the coefficient of determination) for xBA, xOBP, xSLG, xwOBA and their observed counterparts in the first season, and compared it to the second season:
| Relationship | R-Squared |
|---|---|
| xBA to Next Year BA | 0.173 |
| BA to Next Year BA | 0.163 |
| xOBP to Next Year OBP | 0.236 |
| OBP to Next Year OBP | 0.210 |
| xSLG to Next Year SLG | 0.226 |
| SLG to Next Year SLG | 0.189 |
| xwOBA to Next Year wOBA | 0.221 |
| wOBA to Next Year wOBA | 0.179 |
One thing that’s worth noting with this data set, which comes right up to Monday morning, is that I’m getting slightly better correlations than others have gotten in the past. The cause of that is tricky to identify, though one possible explanation is that the change from Trackman to Hawk-Eye in 2020 has helped to improve these metrics.
Still, the relationship between the expected stats and the actual ones is only slightly stronger than the alternative. That doesn’t mean, however, that the expected stats don’t matter when making evaluations.
When constructing a model, a developer will engage in a process called “dimensionality reduction.” There are many methods for doing this, but the basic idea is to take a dataset and reduce the number of features while still preserving the validity of the model. One thing that all the methods share, however, is that they don’t simply throw out variables because they have similar or even lesser correlations with the dependent variable. Even a variable that performs worse than another can still contribute to making a model more accurate than it would be otherwise. This is not an uncommon occurrence.
Imagine you’re trying to model someone’s life expectancy. Age is an extremely important variable. But factors such as whether the person is a smoker, their socioeconomic status, and their health history are also variables that, if known, serve to make the model more accurate than simply using age alone. The key is determining whether those lesser variables are capturing some useful information that age alone is not. Let’s use a baseball example to demonstrate this.
Since the divisional era started in 1969, the r-squared for OBP vs. runs per game is 0.73, basically meaning that 73% of the observed variance in OBP explains the observed variance in runs per game. For SLG, that number is 0.82. Team OBP has a weaker relationship with runs per game than Team SLG, but using both makes the model far better. OPS’ r-squared is 0.905, and OBP*SLG is 0.911. Now take the last part of the triple-slash, batting average. The r-squared for BA and runs per game is 0.53, but in this case, it’s not adding information that OBP and SLG aren’t already capturing; the r-squared for a model of runs scored per game only improves to 0.914 when incorporating BA. Indeed, when OBP and SLG are used, BA is actually a very slight negative factor, because OBP/SLG combinations slightly underrate walks.
So the pertinent question isn’t whether xStats are better than actual stats at predicting future performance, but whether they improve our ability to predict future performance when used in conjunction with actual stats. Below are the RMSE (root-mean squared error) for the relevant stats:
| Stat | RMSE |
|---|---|
| BA to Next Year BA | 0.0347 |
| xBA to Next Year BA | 0.0317 |
| xBA and BA to Next Year BA | 0.0312 |
| OBP to Next Year OBP | 0.0370 |
| xOBP to Next Year OBP | 0.0350 |
| xOBP and OBP to Next Year OBP | 0.0345 |
| SLG to Next Year SLG | 0.0760 |
| xSLG to Next Year SLG | 0.0739 |
| xSLG and SLG to Next Year SLG | 0.0719 |
| wOBA to Next Year wOBA | 0.0403 |
| xwOBA to Next Year wOBA | 0.0385 |
| xWOBA and wOBA to Next Year wOBA | 0.0378 |
[Chart fixed, it was very late and I flipped the xStats and Actuals -DS]
Knowing the error ranges of these stats doesn’t directly tell a user how to treat this data. So rather than ask what the error of each stat is, I went back to the full dataset and instead asked what linear combination of the xStat and the actual stat have worked best:
| Stat | xStat | Actual Stat |
|---|---|---|
| Next Year BA | 73% | 27% |
| Next Year OBP | 70% | 30% |
| Next Year SLG | 59% | 41% |
| Next Year wOBA | 65% | 35% |
As a simple rule of thumb, you won’t do too badly if you simply regress xStats a third of the way towards the actual ones. But as you might have guessed, that changes depending on the player’s number of plate appearances.
For BA, if you only look at the players with at least 600 plate appearances in the first season, the ideal BA mix is 37% BA, 63% xBA. When you only look at the players with between 200 and 300 plate appearances, that becomes 10% BA and 90% xBA, a drastically different number. Naturally, this reflects the fact that the longer a player outperforms their xStats, the closer to the actual stats you expect them to be in the future. But let’s calculate that, too.
Adding multiple year inputs into the mix, I calculated the stabilization point for each of these four stats. This is the number of plate appearances at which the xStat and the actual stat have equal predictive power:
| Stat | Stabilization Point (PA) |
|---|---|
| BA | 1154 |
| OBP | 1007 |
| SLG | 607 |
| wOBA | 766 |
Sticking with using only xStats and actual ones to predict the future, you can approximate how much of the actual stat to use with the formula Actual % = PA / (PA + Stabilization Point).
If you’re trying to renovate your home, you can’t use a screwdriver for every task. But if you throw away your screwdriver because there’s a lot it can’t do, you’ll regret it when you encounter a screw. xStats aren’t a predictive model by themselves, but they can be a crucial part of a predictive model. The zStats used in ZiPS look at things like spray tendencies and speed to improve accuracy, but xStats are still a useful tool.
I’ll look at the pitcher side of the equation in a future piece.
Dan Szymborski is a senior writer for FanGraphs and the developer of the ZiPS projection system. He was a writer for ESPN.com from 2010-2018, a regular guest on a number of radio shows and podcasts, and a voting BBWAA member. He also maintains a terrible Twitter account at @DSzymborski.
Great article Dan, thank you.
Thank you Dan, very xCool!
This article makes me want to learn more about the concepts discussed so I can understand it better.
I would guess the reason they’re similar is because most hitters have xStats similar to their actual stats. Deviations from that are sometimes due to baserunning, sometimes due to shifting (less so than in the past) and sometimes due to luck. But over 150+ games you would expect these to either an average out or have a very small effect.
But it does seem to be useful when someone is dramatically underperforming or over performing their expected stats. Bobby Witt from earlier this year and Aaron Judge from 2021 are good examples where they were hitting the ball very well and getting very unlucky.
The tricky part is that sometimes doubles and triples and BABIP can be “earned” by expected stats and still not be easy to repeat. Sometimes a guy is just really hitting the snot out of the ball, so it’s earned, but it also doesn’t mean it’s going to continue. It’s just hard to keep hitting the ball that hard sometimes.
I did a quick data check for 2023 prior to July 1st for wOBA and xwOBA and compared to post July 1st to see if intraseason xwoBA was better or worse for predicting wOBA than wOBA (min 100 PAs before and after)
xwOBA to wOBA had an R2 of 0.14.
wOBA to wOBA had an R2 of 0.07.
Not as good as the season 1 to season 2 data and only for a year, but it seems like xwOBA may be able to detect a change in future wOBA a little quicker than wOBA. Not sure how useful this is outsode of the big changes.
I would love to hear someone take a shot at the discrepancy between Cody Bellinger’s xstats and his real stats. His AVG EV is in the 18th percentile, Hard Hit 9th, Barrel 25th, all way below average. Meanwhile, back at the ole ballpark, he’s shoving it in both real (4 WAR YTD) and fantasy baseball, currently the 9th ranked player overall in the Yahoo game.
Great article, and one I’ll likely continue to refer back to!
One question: This part “the r-squared for a model of runs scored per game only improves to 0.914 when incorporating BA. Indeed, when OBP and SLG are used, BA is actually a very slight negative factor, because OBP/SLG combinations slightly underrate walks.” is tough for me to follow.
The model of OPS had a value of .905, and OBP*SLG had a value of 0.911. Incorporating BA improved the model by 3 points to 0.914, so how was BA a negative factor? Did you mean to say BA/OBP and BA/SLG combinations slightly underrate walks, or am I not following this correctly?
Oh, I mean a negative not that it makes it slightly more accurate, but that once you have OBP/SLG, a higher BA is actually a *very* slight *bad* thing. If all you know about two teams is that one is a 230/320/420 team and the other is a 260/320/420 team, you’d actually expect the lower BA team to score a few more runs a year. Walks that advance runners are a bit more likely than productive outs that advance runners.
Understood!! I remember another article about that, which explained it quite well. Thanks Dan, I appreciate your work and the response 🙂
How much more predictive are zStats than xStats?
I talked about that more in a few places!
https://blogs.fangraphs.com/hitter-zstats-entering-the-homestretch-part-1-validation/
https://blogs.fangraphs.com/pitcher-zstats-entering-the-homestretch-part-1-validation/
One benefit I have is that I don’t need it to be a fairly simple concept that can easily be used by the public. I can make the model as robust as I can and it can adapt to changes.
Thanks!
Great article!
What is an intuitive explanation for why xStats have higher r squared to next year’s stats, but also higher RMSE? Usually, they would go the opposite direction.
Because the author is an idiot who flipped the numbers in his columns.
Fixing.
I’ve been waiting for this one for awhile. Thanks Dan!
What about pitcher list xwoba in relation to statcast xwoba predictiveness? Pitcherlists uses statcast data but also uses horizontal spray which statcast doesn’t.
Some guys with significant difference
Much higher pitcherlist xwoba:
Mike Ford
Isaac Parades
Royce Lewis
Wilmer Flores
Corey Seager
Much higher statcast xwoba:
Judge
Tatis
McCormick
Teoscar
Expected statistics are constructed to take three-dimensional actions and compress them into a two-dimensional system. It removes the aspect of the direction of the batted ball. I believe that, as a consequence, potentially useful information is lost. This information is found in the traditional statistics which, indirectly, includes information about direction.
This is demonstrated in the performance of certain players such as Isaac Paredes. By expected stats, last year Paredes was a sub-average batter with an xwOBA of 0.297. This year his xwOBA has barely improved, 0.317, which is the league average. But given the fact that he has been successful in pulling the ball repeatedly, his wOBA last year was an above-average value of 0.323 and this year, he is among the top 20 at 0.371.
I don’t know why direction is not included in calculating xBA values. The argument may be that direction is something batters have no control over. If this were true, there never would be a need to shift the position of fielders.
OPS or any other traditional stats also work well. xStats are absolutely not crucial to anything – they never have been and they never will be.