FIP vs. xwOBA for Assessing Pitcher Performance
At a basic level, nearly every piece at FanGraphs represents an attempt to answer a question. What is the value of an opt-out in a contract? Why do the Brewers continue to fare so poorly in the projected standings? How do people behave in the eighth inning of a spring-training game? Those were the questions asked, either explicitly or implicitly, by Jeff Sullivan, Jay Jaffe, and Meg Rowley just yesterday.
This piece also begins with question — probably one that has occurred to a number of readers. It concerns how we evaluate pitchers and how best to evaluate pitchers. I’ll present the question momentarily. First, a bit of background.
Fielding Independent Pitching, or FIP, is a well-known tool for estimating ERA. FIP attempts to isolate a pitcher’s contribution to run-prevention. It also serves as a better predictor of future ERA than ERA itself. The formula for FIP is elegant, including just three variables: strikeouts, walks, and homers. It does not include balls in play. That said, one would be mistaken for assuming that FIP excludes any kind of measurement for what happens when the bat hits the ball. Let this be a gentle reminder that home runs both (a) are a type of batted ball and (b) represent a major component of FIP. There is, in other words, some consideration of contact quality in FIP.
Expected wOBA, or xwOBA, is a newer metric, the product of Statcast data. xwOBA is calculated with run-value estimates derived from exit velocity and launch angle. Basically, xwOBA calculates the average run value of every batted ball for a hitter (or allowed by a pitcher), adds in the defense-independent numbers, and arrives as a wOBA-like figure. The advantage of xwOBA is that it removes the variance of batted-ball results and uses a “Platonic” value instead.
The introduction of Statcast’s batted-ball data is exciting and seems like it might help to better isolate a pitcher’s contributions. But does it? This is where I was compelled to ask my own, relatively simple question — namely, is xwOBA better for assessing pitcher performance than the more traditional FIP? What I found, however, is that the answer isn’t so simple.
The differences between FIP and xwOBA, as well as the similarities, deserve some exploration.
Both xwOBA and FIP include strikeouts and walks in their formula and both use runs as the basis for determining value. The metrics differ in two very meaningful ways. One difference is in presentation. FIP is made to look like ERA, probably the most widely used pitching statistic by most baseball fans. A good number for ERA is likewise a good number for FIP. On the other hand, xwOBA was created to look like wOBA, a popular sabermetric statistic that was made to look like on-base percentage. That difference doesn’t meaningfully change what the statistics do, but one is generally from a pitcher perspective (FIP) while the other is more easily associated with hitters (xwOBA).
The second difference comes in the batted balls measured. FIP uses only home runs. One could argue that FIP’s use of homers could be considered a proxy for all batted-ball contact, but I won’t be making that argument here. Alternatively, xwOBA — as I’ve said — uses the launch angle and exit velocity for all batted balls to provide a run value for those events. Both metrics convert the inputs to runs, FIP expressed as earned runs per nine innings and xwOBA in runs per plate appearance.
The stats work very similarly. To demonstrate this, consider the graph below, a plot of 72 pitchers who recorded at least 2,000 pitches in both the 2016 and 2017 seasons. The plot below shows xwOBA and FIP from last season with all xwOBA data from Baseball Savant.

The two produced similar numbers, although the similarities do not end there. The table below shows the correlation of the two metrics with ERA during those seasons.
| Metric | 2016 | 2017 |
|---|---|---|
| FIP | 0.56 | 0.64 |
| xwOBA | 0.63 | 0.66 |
The numbers were nearly identical last season, though FIP looks slightly off during the 2016 campaign. More on the 2016 data will come later, but first, the graph below shows those same 72 pitchers’ xwOBA in 2016 and 2017.

The relationship here is a pretty good one — much better than the r-squared for ERA among the same group, which comes to just .12.
Now compare those numbers to FIP from the last two seasons.

In terms of correlation, the r-squared is very similar to xwOBA’s, if a tiny bit behind. This, again, reinforces the similarities between the two statistics.
Now, let’s run another comparison testing their potential predictive capabilities, this time comparing 2016 xwOBA and FIP to 2017 ERA. First, xwOBA:

The relationship here isn’t the strongest one we’ve seen as we move further away from like groups, but there is a correlation between the two groups. Also, keep in mind: the r-squared is much higher than simply using ERA between the two seasons.
Now, let’s compare 2016 FIP to 2017 ERA.

Again, we receive almost identical correlations for FIP and xwOBA when compared to the following season’s ERA.
I said earlier I would get back to the 2016 numbers showing that xwOBA was more closely correlated to ERA than FIP. I broke down that season in half for the 72 pitchers and found something interesting — namely, that xwOBA had a stronger correlation between the first and second halves than FIP and a slightly higher correlation from first-half xwOBA to second-half ERA compared to first-half FIP and second-half ERA. The table below contains the results.
| 2016 | 2nd Half xwOBA | 2nd Half FIP | 2nd Half ERA |
|---|---|---|---|
| 1st Half ERA | — | — | 0.03 |
| 1st Half xwOBA | 0.26 | — | 0.16 |
| 1st Half FIP | — | 0.17 | 0.12 |
These numbers potentially indicate that, in smaller samples, xwOBA might stabilize more quickly and do a better job of predicting future ERA than FIP currently does. Unfortunately, this finding does not hold up to further scrutiny. Here are the same numbers as in the last table, except from 2017.
| 2017 | 2nd Half xwOBA | 2nd Half FIP | 2nd Half ERA |
|---|---|---|---|
| 1st Half ERA | — | — | 0.09 |
| 1st Half xwOBA | 0.24 | — | 0.12 |
| 1st Half FIP | — | 0.25 | 0.21 |
Compared to itself by halves, xwOBA did just as well in 2017 as it did in 2016, while FIP moved up to the same level after struggling to show a relationship in 2016. FIP came out better in its relationship to second-half ERA. Last year’s xwOBA relationship to second-half ERA didn’t quite measure up. It seems possible that 2016 was simply a very volatile year. It was the first full year with a juiced ball, and perhaps there was some added movement in the numbers. FIP is heavily reliant on homers as a measure of talent, so perhaps that helps explain the weaker results the season before.
We will need more years of data for further study, but thus far, I’m comfortable saying that xwOBA and FIP are pretty similar metrics. We have more data than we’ve ever had before about the quality of all batted balls, and there’s some hope this might lead us to a better, more confident metric as it relates to pitcher skill. From this data, that doesn’t appear to be the case. These results provide support — or perhaps just mimic — a similar study I ran back in August when I could not find a predictive skill in a pitcher’s quality of contact.
It’s certainly possible, maybe even likely, that a pitcher has control over the contact he gives up, but these numbers don’t support that finding with xwOBA. To be fair — if a metric needs to feel as though it is being treated fairly — FIP sets a fairly high bar when it comes to measuring pitchers, and being roughly as good as FIP is a victory in and of itself. So, good job xwOBA.
Craig Edwards can be found on twitter @craigjedwards.
while predictive value is important, there’s also explanatory power about past results.
in this case xwoba simply captures more information, namely the batted ball profile for a significant portion of balls in play. it’s useful in answering the question of ‘why is this pitcher over/under performing his FIP by so much.”
But the R^2 data shown above is explanatory value. Predictive value would compare, for example, 2017 projected ERA based on FIP with 2017 actual ERA, or 2017 projected FIP with 2017 actual FIP, etc.
xwOBA has no explanatory advantage over FIP from the data above.
In order to demonstrate that xwOBA explains why a pitcher over or under performed his FIP, a controllable skill for managing contact quality would have to prove statistically significant. I don’t believe such a showing has yet been made.
uh this article is comparing 2016 with 2017. it’s not past-looking.
you’ve got this entirely backwards.
Excellent post!
Aren’t these results a clear win for FIP? Since xwOBA essentially takes up where FIP leaves off ( also tracking K’s and BB’s, but then treating HR’s and BIP based on EV and LA) don’t similar season to season correlations and similar future season ERA predictive ability mean that xwOBA isn’t adding anything?
Simply put: If I can created a metric that is just as predictive as xwOBA by assuming that all non-HR BIP are the same (basically what FIP does) then is all that added input really of any real value?
FIP doesn’t necessarily assume that all non-HR BABIP is ‘the same,’ as the author noted, it can be seen instead to use HR rate as a proxy for BABIP distribution generally.
FIP assumes all non-hr BABIP are results of factors beyond the pitcher’s control and ignores them. Part of the beauty of FIP is it’s simplicity and through this article and TapeyBeercone’s comment (nice name btw), how well it performs.
A more relevant question might be: what is FIP adding that Steamer doesn’t already capture?
showing the R^2s are inconsistent over different samples doesn’t really show anything. Particularly under the assumption that the ball changed at some point in 2016 (especially if we assume it’s at the all star break). That would create a ton of noise in both 2016-to-2017 correlations and 2016 1st half to 2nd half correlations.
On top of this, none of this is particularly rigorous. Even an ANOVA can at least show the statistical significance of these predictors, you’re simply showing R^2s which aren’t really saying much.
Finally, when you’re doing 1st-half to 2nd-half correlations, you make no attempt to ensure each pitcher in the sample pitched enough in both halves. If there are a couple players who made 18 starts in the first half and 2 in the second half, it could severely mess up the sample. You really need to ensure that the samples being compared have more consistent variance.
The ball changed in 2015. In 2016, the ball was juiced the whole year. As for concerns on first and second half, of the 288, half seasons, all but 6 pitchers had at least 40 innings and every pitcher made at least 5 starts.
Nah, you are simply wrong about this whole thing. Below is the lg avg for xwOBA vs. wOBA during the Statcast era.
xwOBA – wOBA
2015: .309 – .315
2016: .316 – .322
2017: .311 – .327
Besides, your previous study on pitchers’ contact management was also just flat out wrong, because this contact management thing is more about the BIP type frequencies and less about the contact-authority suppression; every BIP type (FB, GB, PU), except for LD, exhibits strong year-to year correlation.
ISO:
2013: .143
2014: .135
1st Half 2015: .143
2nd Half 2015: .158
1st Half 2016: .163
2nd Half 2016: .161
1st Half 2017: .171
2nd Half: 2017: .171
xwOBA vs. wOBA on FB (20-50 degrees)
’15: .449 – .452
’16: .463 – .483
’17: .464- .506
A single predictor r-squared is literally just the correlation squared (hence the term r-squared). You’ve got a sample size of about 72 or so and a correlation of something like .47 with only one predictor, which is comically high for one predictor. That’s a t-value in excess of 4, with a p-value waaaay less than whatever threshold you want to reasonably set.
How would there be no relationship between any reasonable indicator and the statistic it was designed to be related to?
The question is: does it add anything that we don’t already know?
Since the correlation of the outcome variable with itself is far lower than the correlation between the indicator and the outcome variable, I would answer that it does add a lot we don’t know.
Well, FIP isn’t really one predictor, it’s walks, strikeouts, and home runs.
It’s a composite variable, entered into the regression equation as a single predictor for a single outcome. You could, in theory, enter them all in separately in a multiple regression framework.
xwOBA eliminate ballpark effect right ? Which FIP and ERA doesn’t do.
So maybe after consider ballpark effect, you would have different result?
I actually do not know if this is the case or not. I think xwOBA is assuming average outcomes regardless of park.
xwOBA is not adjusting for park, only factors by season:
Expected weighted on-base average (xwOBA) is formulated using exit velocity and launch angle, two metrics measured by Statcast.
In the same way that each batted ball is assigned a Hit Probability, every batted ball has been given a single, double, triple and home run probability based on the results of comparable batted balls — in terms of exit velocity and launch angle — since Statcast was implemented Major League wide in 2015.
All hit types are valued in the same fashion for xwOBA as they are in the formula for standard wOBA: (unintentional BB factor x unintentional BB + HBP factor x HBP + 1B factor x 1B + 2B factor x 2B + 3B factor x 3B + HR factor x HR)/(AB + unintentional BB + SF + HBP), where “factor” indicates the adjusted run expectancy of a batting event in the context of the season as a whole.
http://m.mlb.com/glossary/statcast/expected-woba
A batted ball with same launch angle/Exit Velocity should have same xwOBA regardless of park.
But the same batted ball can create different ERA/FIP result in different ballpark.
This is an awesome article, and I think this is the right question to be asking.
The fact that xwOBA and FIP perform similarly despite xwOBA having much greater inputs would seem to favor FIP, under a desire for greater parsimony. I don’t see it that way. xwOBA is an alternative measurement of the same thing, and as such should give us different answers at the individual level even if the variance explained is similar in the aggregate. Having two measures that are equally valuable in the aggregate but that give us different answers at the individual level is incredibly useful if you’re trying to triangulate between different sources.
I think the next step would be comparing pitchers where there is almost no difference, where FIP dramatically favors one, and where xwOBA dramatically favors another (if you’re looking for new article ideas, Craig, you could do about 20 articles on this topic, and that could be one of them). I suspect that scaling xwOBA to FIP would take longer than the actual comparison.
Ultimately, I bet you that if you did a confirmatory factor analysis on some subset of FIP-, xwOBA, SIERA, cFIP, some park-neutral version of RA9, etc you could predict a factor score that would have incredible predictive validity* (in less technical terms, triangulating between multiple similar-performing measures with different inputs almost certainly would yield better estimations of true performance).
.***
*There’s some overlap between those estimators in terms of inputs though, so either you’d want to exclude some or do some fancy correlated-errors modeling to pull that part out.
Yup, excellent article and some pretty strong comments IMO.
So basically, you’d like to see FG reinvent Steamer’s algo from scratch?
Thanks for an interesting article.
A quick comment on the methodology: To evaluate the performance of your models I think it would be better to predict the n+1 year ERA from the model fit from n year FIP or xwOBA to n year ERA, and then compare the variances in the observed versus the expected ERAs between models. Your question is how good are your two independent variables at predicting specific ERAs for pitchers, rather than how well correlated are the independent variables with ERA between years. As you have done it, it is possible that a model could have a higher correlation between years even if the actual prediction is worse (if, for example, the slope of the line or the intersect change between years).
I doubt it would change your results, but I think it would be statistically more correct.
cool, thanks
I think the next step for Statcast should be incorporating horizontal launch angle, which I’m sure is easy to determine. I feel like this would really hone down on the expected run values as this could differentiate between two balls hit at the same EV and LA, however one is straight at a left fielder while the other is in the left field gap. It would be pretty interesting!
According to a Perpetua article from January, Trackman is already measuring both the horizontal (angle off the bat) and spray (bearing) angle but just not releasing it to the public. However he has incorporated the ‘bearing’ (angle between home plate and the point at which the ball lands) into his xStats calculations (xstats.org). I believe he just derives it from the BIP data of where the ball lands. Where xwOBA is only using EV and LA, Perpetua’s xOBA is using “the vertical and horizontal launch angles, exit velocities, batted ball distances, game time temperature, and ball park are taken into account. All other factors are ignored. The angle and exit velocity information is fed through an algorithm that lumps together similarly hit balls and finds their average success rates. Game time temperature and ball park information are used to adjust the exit velocities.”
He dove into spinrates back in January in a really fantastic piece for anyone interested in the nuts and bolts of this stuff: https://www.fangraphs.com/fantasy/step-aside-statistics-it-is-physics-time/
Hey thanks for the comment, I was hoping they used some thing like that. It’s definitely way more accurate
Seems like an intermediate step could be reviewed – what is the relationship between FIP and a Statcast FIP utilizing xHR? Perhaps there is some “luck” involved with HR 🙂
JK on the luck thing but the analysis would interesting.
For example, I took a quick look at Perpetua’s xStats. I chose Lance Lynn. He gave up a lot of HR last year. His scFIP runs a half a run lower than FIP because HR=27 and xHR=23.2.
I believe his scFIP calc includes horizontal angle as well as EV and LA. I also have not seen his detailed calculations. His sheet was easy to grab for a comparison.
The market seems to value xwoba more than FIP based on recent contracts.
Three examples that quickly come to mind are Marco Estrada, Sabathia, and Lynn.
That may be true but it doesn’t matter to the analyses at hand. Apples should be compared to apples before they’re compared to apples and oranges.
The most obvious followup question is of course how does xFIP compare to FIP and xwOBA? We all know that the HR/FB rates are just about as noisy as ERA itself, so why are we comparing a statistic like FIP with such a noisy component instead of xFIP, designed to remove that noisiness? And IIRC, doesn’t xFIP predict ERA better than FIP anyways?
Thought about including xFIP. It’s very good at predicting itself for hopefully obvious reasons, and it is about as good as FIP and xwOBA with future ERA, but it is generally not as good with simultaneous ERA so it doesn’t do a good job of explaining what has happened/what is happening now. Given its lack of advantages, it is a pretty big disadvantage.
This has been done quite often in the past.
https://www.fangraphs.com/tht/the-great-run-estimator-shootout-part-2/
https://www.fangraphs.com/tht/should-we-be-using-era-estimators/
https://www.fangraphs.com/tht/predictive-fip/
The truth is, none of them are designed to predict. The people most serious about prediction (GMs, handicappers, and DFS players) are using far more accurate techniques which, one created, are no more cumbersome to generate that FIP is.
I’ve been using this equation to figure an xwOBA-FIP:
((((xwOBA * 1.33)*TBF))+(3*BB)-(2.25*HBP)-(2*K))/(IP+3.125)
It’s not as predictive as DRA but it seems to be as good as FIP. I don’t have the numbers with me because I’m at work, but if people think this isn’t laughable I’ll try to find the R-squared numbers on my home CPU.
xwOBA is clearly superior because it just doesn’t rely on one arguably arbitrary type of hit (the home run) but looks at all the batted ball data
Is there a way to marry metrics like that into a single exploratory metric?
@SteamerPro 🙂
No one actually uses FIP for prediction. If you want to evaluate predictive compare to Steamer’s projected ERA.