The Case Against a Case Against FIP

© Steven Bisig-USA TODAY Sports

At FanGraphs, our headline WAR number for pitchers is based on FIP. Because of that, and because people enjoy debating and arguing, there’s a yearly refrain that you’ve probably heard. “FanGraphs pitching WAR only considers (X)% of what a pitcher does, how can that be used for value?” No one would dispute that year-one FIP does a better job of estimating year-two ERA than ERA does – or at least, not many people would – but the discussion around whether FIP does a good job of assigning year-one value is alive and well.

One reason for this view is pretty obvious. FIP considers home runs, strikeouts, walks, and hit batters to estimate pitcher production on an ERA scale. Our WAR does some fancy stuff in the background – it treats infield fly balls, which virtually never fall for hits, as strikeouts, and it adjusts for park and league. In the end, though, it’s estimating pitcher value using just three (well, actually four — HBPs always draw the short straw) outcomes. There are a lot of other outcomes in baseball!

In 2021, roughly 39% of plate appearances ended in a homer, strikeout, walk, hit batter, or infield pop up. One thing you could think, in recognition of that fact, is that FIP-based WAR doesn’t consider enough of a pitcher’s production. You wouldn’t use 40% of a hitter’s plate appearances to calculate their WAR, so why do it for pitchers? But that doesn’t actually make sense, as David Appelman pointed out to me recently. Assuming “average results on balls in play” is actually going to be pretty close for every pitcher, by definition.

How so? Consider it this way. Let’s say you’re a pitcher who for whatever reason allows a .350 BABIP. That’s awful! That would be one of the 10 highest single-season BABIPs allowed since 2015 (minimum 100 innings pitched). Let’s get a little more specific and say that you, the reader, are Brady Singer, with a .350 BABIP in 2021. Oof, sorry Brady (but thanks for reading).

FIP assigns Singer league average results on balls in play. Our modified FIP-WAR does a little better by including infield fly balls. That’s not how it broke down for him in real life, though. How poor of an estimate was FIP’s “Singer will have average BABIP” assumption? It’s actually pretty easy to do the math.

After excluding infield fly balls, Singer allowed an fBABIP (I’m just calling it that for the duration of this article, but that’s not a real term) of .359. The league as a whole allowed an fBABIP of .301. Over the 368 balls in play that Singer allowed (again, excluding infield fly balls), that’s a difference of 21 hits. That’s a lot of hits – and it’s also just 3.6% of the batters Singer faced in 2021.

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

Put another way, FIP’s estimation of Singer includes 96.4% of the batters he faced. That’s less than 100%, but if your only criteria for which statistic you want to use is what percentage of batters faced it “considers,” FIP is going to be pretty close to 100% across the board, not the 39% that it takes directly from player stats. Averages are powerful that way, and the population of pitchers who throw 100 innings in the major leagues has a near-perfect normal distribution of fBABIP:

The league average fBABIP over that time is .305, with a standard deviation of roughly 25 points. 67.8% of pitcher BABIPs are within one standard deviation of average; 96% are within two standard deviations. There’s no skew. A statistics textbook should use this as a case study.

If you think of FIP as certain outcomes (the three true outcomes plus pop ups) with a BABIP assumption, you’re basically correct. As it turns out, that BABIP assumption is really close to capturing all the balls in play. For 68% of pitchers, FIP correctly pegs at least 98.5% of their batters faced. For 95% of pitchers, FIP gets 96.5% or more correct. It’s even more accurate than that in our WAR calculations, because we include park and league adjustments, which I haven’t done here and which causes some of the variation. The argument that FIP is ignoring 60% of what a pitcher does isn’t right – the assumption that the pitcher is average is, in itself, a pretty good description of that last 60%.

You might reasonably say that BABIP isn’t the only thing that matters when it comes to balls in play. If a pitcher is giving up rocketed line drives left and right, the doubles will mount, and doubles are more damaging than singles despite looking the same in BABIP. To account for this, I calculated wOBABIP – wOBA on balls in play, although excluding infield fly balls again – for the same cohort of pitchers.

This gets away from the “but-FIP-is-40%-of-a-pitcher” argument that kicked the article off, but I think it’s still instructive. FIP does a pretty good job of estimating this part of pitching too. For every pitcher with 100 innings pitched in a single season since 2015, I calculated the difference between that season’s actual wOBABIP and the league average wOBABIP for that year. I converted that into runs, then converted those runs into runs allowed per nine innings.

The average absolute difference between the runs a pitcher would allow with league average results on balls in play and what they would allow with their actual results on balls in play is 0.43 runs per nine innings. Half the pitchers in baseball had a gap of less than 0.35 runs per nine innings. Only 7% of pitchers had a gap of one run or more. You can see the output of both my BABIP and wOBABIP calculations here. Again, I didn’t consider park adjustments — the actual random variation will be smaller.

Why not use the actual wOBABIP to improve FIP, then? It’s quite unclear how much of this variation is something that a pitcher accomplishes via their own skill and how much comes down to either the breaks of the game or defense. Per Statcast’s OAA, the average variation due to defense from pitchers who allowed at least 250 balls in play in a season between 2016 and ’21 was 0.18 runs per nine innings. It would be strange to credit or debit pitchers for that in a value metric – the fielders are the ones responsible for those runs in either direction.

That doesn’t take shifts into account, or whether the wind was blowing in or out that day, or even which of the two baseballs they happened to throw. How much a pitcher does to influence their results on balls in play is an ongoing debate, but even if it weren’t, FIP captures a good chunk of those balls in play by assuming league average results.

Does that mean FanGraphs’ formulation of FIP is a flawless, unassailably perfect way to assign value to pitchers? Definitely not. Whether you want to isolate pitcher value or isolate “what happened on the field,” there are key shortcomings of using either FIP or RA9-WAR.

Both have a strained relationship with defense, and particularly with how a pitcher’s incentives change based on the defense behind them, but the possible shortcomings don’t stop there. FIP ignores sequencing – a pitcher who walks the bases loaded and then gets three outs without allowing a run might not succeed in the long run, but in that inning, he certainly didn’t allow any runs. Maybe he even got those outs in ways that didn’t tax his defense – routine grounders and lazy fly balls. Sequencing has a lot to say about how many runs a pitcher allows in a given season, even if pitchers mostly don’t show any ability to control sequencing over time.

Likewise, ERA- and RA-based WAR have faults. ERA is extremely strange – an official scorer’s decision doesn’t mean squat for how a pitcher performs, but it changes how many earned runs they allow. Even beyond that, should we give different credit to a pitcher who allows a bases-loaded screaming liner that the center fielder corrals, or to one that finds a hole in the outfield? What if the ball is hit to the exact same place and the defender just got a good or bad jump?

While we’re there, what about what happens after a pitcher leaves? RA9-based WAR will tend to flatter pitchers with better bullpens. If one pitcher leaves with the bases loaded and gets three runs for his trouble while another leaves in the same spot and gets a zero, did we learn anything about the difference in their skill? Clearly not, but the RA9 tally would tell you the two performed differently.

I won’t act like I’ve solved the debate or that one WAR is clearly superior to the other. But the complaint that FIP only considers a minority of what a pitcher does doesn’t track to me. On the low end, it’s giving an accurate accounting of 95% of the batters a pitcher faces. Those last 5% can have an outsized influence on how many runs a pitcher allows – a few extra singles here and there or a poorly-timed double can have outsized effects on a pitcher’s line.

Some of that probably accrues to the pitcher, some to the defense, and some to luck. How to assign credit and blame among those three things is the central disagreement in how to assign WAR to pitchers. FIP goes to one extreme; RA9 WAR unadjusted for defense goes to the other. I don’t know the right answer to this disagreement. But I do know that “FIP only considers 40% of balls in play” doesn’t give FIP enough credit.





Ben is a writer at FanGraphs. He can be found on Bluesky @benclemens.

119 Comments
Oldest
Newest Most Voted
Inline Feedbacks
View all comments
Kevbot034
4 years ago

Funny enough, Singer also has a 350 BABIP again and I am curious suddenly about what drives that? FIP usually tends to penalize pitch to contact types, doesn’t it? But it seems to still like Singer specifically despite him allowing a billion baserunners with a stupid high BABIP.

Ivan_GrushenkoMember since 2016
4 years ago
Reply to  Kevbot034

XFIP, SIERA and xERA are all closer to FIP than to ERA. This seems to support the hypothesis that BABIP is not Singer’s fault very much unless he’s a terrible fielder. This might or might not mean that FIP is enough and those other things don’t add much.

Also since 2020 Royals have been the 3rd best defensive team by Def but with only average range. Maybe this means they aren’t positioning well behind Singer. I don’t know

Last edited 4 years ago by Ivan_Grushenko
RonnieDobbs
4 years ago
Reply to  Ivan_Grushenko

That is unthinkable that pitcher defense would play any role at all in BABIP. Do you know how few balls are fielded by a pitcher? The expectation is that he never makes a single play. If the pitcher put his hands in his pockets, the rest of the infield would be OK with that. On the contrary, BABIP is about two things – execution and luck.

Defensive metrics are worth absolute zero.Someday everyone will agree with this statement when all of the things that we currently use are thrown out or overhauled.

FrancoeursteinMember since 2025
4 years ago
Reply to  RonnieDobbs

I have always informally considered controlling the running game as pitcher defense. I think that was one of the many reasons Julio Teheran was able to run such a high FIP-ERA for many years

CC AFCMember since 2016
4 years ago
Reply to  Kevbot034

Depends on how you define “pitch to contact types.” If you just mean “guys who have low strikeout rates,” then no. To some extent, some pitchers have batted ball profiles that can lead to sustainably low or high babips, but they’re not necessarily “pitch to contact types,” see e.g. Scherzer, Max.

As regards Singer specifically, it’s been less than 6 innings this year. Not enough to draw any conclusions.

Six Ten
4 years ago
Reply to  Kevbot034

Singer isn’t really a pitch to contact pitcher. He only has two pitches and one of them is a sinker, which traditionally suggests a pitch to contact approach. But out of the 129 pitchers who threw 100+ innings last year he had the 53rd most K/9 and 18th most BB/9. So not blowing guys away, but that’s a lot of PAs ending without contact too. And he was pretty decent at avoiding home runs too.

Sinkers are good at preventing fly balls, and they are perfectly fine on balls in play as long as they don’t become fly balls. But if the batter elevates one, it gets tagged. That means home runs! Unless, of course, you play in a park with a gigantic outfield. Then it’s all about what happens to that fly ball in play. High BABIP: bad! Low BABIP: good! Well, FIP doesn’t like worrying about BABIP.

Which is to say: if you throw a sinker and you play in a really large stadium, it seems like FIP might tend to give you too much credit rather than too little. And that’s Singer.

dl80Member since 2026
4 years ago
Reply to  Kevbot034

But he’s not really. He’s averaged over 9 K/9 the two years he’s had the inflated BABIP (slightly above league average). And he’s given up fewer line drives than average, also.

The one thing that stands out is that he gets a lot of groundballs, which are more likely to become hits.

But all of their infielders are above average defensively this year, so it’s not like he’s being victimized by bad defense. And it was mostly true for last year, also.

So it’s either random bad luck or the team does a terrible job shifting.

Ivan_GrushenkoMember since 2016
4 years ago
Reply to  dl80

If the problem is Royal Shifting he’d be a good trade target for a better shifting team

Jason BMember since 2017
4 years ago
Reply to  Ivan_Grushenko

I asked Dennis Miller for his opinion. He said “I haven’t seen this kind of Royal Shifting since the Dutch royal family fled to Canada, Chachi.”

TheGarrettCooperFanClub
4 years ago
Reply to  dl80

The Royals shift the least in all of baseball (17.3%)! Perhaps that is most of the issue.

Six Ten
4 years ago

It’s possible shifting is the issue, but if there’s a non-luck explanation I think it’s more likely the type of contact Singer allowed. In 2021 he allowed a ton of medium contact and quite a lot of line drives. Medium line drives are far and away the contact type with the highest BABIP. The Royals’ shift rate in 2021 was 27.3%, 19th rather than 30th.

I still think there’s probably a lot of luck involved. But only having two pitches makes the contact quality explanation pretty enticing.

Ivan_GrushenkoMember since 2016
4 years ago
Reply to  Six Ten

The medium contact and line drives should be in xERA no?

RonnieDobbs
4 years ago
Reply to  Kevbot034

BABIP punishes people with poor command and/or bad luck. The guys that throw hard but don’t have to execute very well to succeed get pummeled by BABIP. I think you have it completely backwards. People with command and poise are better at managing BABIP. Missing bats does not lead to batted ball events. Few things are more detrimental to building a better understanding of than creating false generalizations.

Kevbot034
4 years ago
Reply to  RonnieDobbs

No, I don’t think I really have it backward, pitch to contact guys tend to not have great FIP, or at least outperform their FIP, because they don’t K many. Anyone walking a ton is going to be detrimental, regardless of if they are strikeout artists or pitch to contact guys. Command and poise relies on your defense – something FIP intentionally ignores. I’m also confused how I would be creating a false generalization anyway, when it was written as a question.

Ivan_GrushenkoMember since 2016
4 years ago

The alternatives include xFIP, SIERA and others that attempt to include what happens when bat hits fair ball. Also ERA or RA might be a better indicator of performance with a multi year period. The debate isn’t FIP right or wrong. It’s FIP vs alternatives

Trevor May Care Attitude
4 years ago
Reply to  Ivan_Grushenko

I always think of Ricky Nolasco in these situations. While rare, there are guys who consistently underperform (and overperform) their FIPs/xFIPs over multi-year samples. Ricky racked up nearly 1900 IP with a FIP over 1/2 a run lower than his ERA.

Ivan_GrushenkoMember since 2016
4 years ago

The flip side of Nolasco is Mariano Rivera whose career ERA- and FIP- were 49 and 62

Seattle Homer
4 years ago
Reply to  Ivan_Grushenko

Pardon the cliché, but Mariano Rivera is the exception that proves the rule. The assumption for FIP is that the outcomes on balls put in play will tend to even out for all pitchers. And for most pitchers that’s true. But Mariano might be the single best pitcher in recent-ish history at forcing bad contact on balls in play. Sooo much weak contact and sooo many broken bats.

Related, I love the story of the (Twins?) presenting him with a chair made entirely of broken bats during his farewell tour.

BenZobrist4MVP
4 years ago
Reply to  Ivan_Grushenko

Rivera was much better by ERA- than FIP-. Oh, and he still has the best FIP- in history of pitchers with 800+ IP.

sadtromboneMember since 2020
4 years ago
Reply to  Ivan_Grushenko

The problem with ERA is that you need to be able to observe a pitcher with different defenses behind them. But because multi-year contracts or many years of team control is the norm, it’s hard to get that. And when they do, often teams will change the pitch mix.

sadtromboneMember since 2020
4 years ago

People who think that ERA is more useful than FIP in assessing pitcher value don’t understand how correlated errors work. Because pitchers play in front of roughly the same defense every night, it means that we can’t ever really parse what’s due to the pitching and what’s due to the defense. Attempts to correct for defense using regression like pitching bWAR are flawed because we don’t have a counterfactual, so it tends to overfit and assign too much credit for defense instead of too little.

The reason why FIP does a better job of predicting future ERA than ERA itself is because ERA is an inferior measure for capturing pitcher value. FIP has predictive validity, and ERA doesn’t.

Now, I’m of the mind that SIERA is definitely a superior measure to both, and that xERA has a lot of promise. I think xERA might undervalue ground-ball pitchers, but the concept of xERA is exactly what we need. I think it’s worth a conversation about whether there should be an xERA-based WAR, maybe an xERA/FIP hybrid or something. Might make comparing historical data difficult but if RA9-WAR gets its own tab it seems like xERA-WAR could very get one as well. Although in practice, xERA and FIP tend to wind up in similar places because true outcomes actually do drive a dramatic amount of pitcher value.

Ivan_GrushenkoMember since 2016
4 years ago
Reply to  sadtrombone

In cases where there’s a persistent unidirectional difference in Pitcher BABIP vs average over multiple years with different defenses behind him there’s a reasonable argument that the particular pitcher has some significant influence on BABIP

sadtromboneMember since 2020
4 years ago
Reply to  Ivan_Grushenko

In that case, I think it’s worth a deeper dive into defenses and batted ball metrics (beyond just xERA). This is more or less the reason why I think that xERA might undervalue ground ball pitchers–I started noticing that very groundball-heavy pitchers were outperforming their xERA. Even more than outperforming their FIP.

Last edited 4 years ago by sadtrombone
Ivan_GrushenkoMember since 2016
4 years ago
Reply to  sadtrombone

Then SIERA should help

Joe Joe
4 years ago
Reply to  sadtrombone

xERA’s problem is that it is scaled from xwOBA. xERA, IMO, needs to have pitcher specific linear weights for batted ball events (though would need to be heavily regressed early in season or pitcher’s career). Extreme flyball and extreme groundball pitchers likely create a very different context for different events than the average pitcher.

George ResorMember since 2016
4 years ago
Reply to  Joe Joe

I think you could use something like base runs with the expected outcomes instead of custom linear weights

Left of Centerfield
4 years ago
Reply to  sadtrombone

“Attempts to correct for defense using regression like pitching bWAR are flawed because we don’t have a counterfactual, so it tends to overfit and assign too much credit for defense instead of too little.”

And FIP goes too far the other way by claiming to completely ignore defense. Except you can’t really ignore defense which is why the name for FIP (Fielding Independent Pitching) is a misnomer. You can’t really sepárate defense from pitching (which Ben alludes to in his article).

Just one example: If you’re pitching behind a poor defense, you have to work harder to get outs. And that can lead to frustration, tiredness, etc. which results in extra walks and homeruns relative to a pitcher pitching behind a good defense.

Also, I never understood the argument that FIP has value because it has predictive validity for ERA while at the same time claiming that ERA is a flawed measure. The ability to predict a flawed measure is hardly a ringing endorsement of FIP.

OtisMember since 2019
4 years ago

ERA is an outcome; i.e., runs are scored, and attributed to the pitcher because he’s the one on the mound (minus the ones directly attributed to errors).

FIP is an attempt to assess how much of that outcome is attributable to the pitcher, stripping out luck and defense as much as possible.

Luck and defense will still happen, obviously, but the idea is that you will have a clearer understanding of what the pitcher himself will be contributing to the equation and future results (such as ERA) going forward.

Spahn_and_Sain
4 years ago
Reply to  Otis

(…but plus the ones where he’s not the one on the mound anymore, but he was the one on the mound when the runners reached base).
With baseball stats, there’s always a caveat or two.

joe_schlabotnik
4 years ago
Reply to  sadtrombone

I don’t think anyone is out here thinking that ERA has any predictive power. What I think team wheeler of the burnes/wheeler battle thinks the predictive power doesn’t matter when you’re looking at a full season. ERA tells you what happened, even if it was luck driven. The same way someone can run a higher wOBA than their xwOBA. If at the end of the year those figures are high, i just think the luck is a cool feature of it. Sometimes lets just allow baseball to have some magic

NATS FanMember since 2018
4 years ago

Excellent work! I’m with FIP!

soddingjunkmailMember since 2016
4 years ago

The trick here is that we’re trying to assess two things that are separate but frequently conflated.

  1. What is the value of person’s contribution?
  2. What is the underlying or fundamental quality of person?

These things are similar enough that they’re frequently the same (or close enough), but sometimes (and particularly over samples that are not large enough) weird things happen.

The difference becomes clear to me when I try to apply them to different situations. For example, I care very much about #1 and not much at all about #2 when handing out MVP awards and the like. But if I’m offering a contract, I care not at all about #1 and very much about #2.

So, to bring this full circle to the author’s article: Is Fangraphs’ FIP formulation appropriate?

Depends on what you’re using it for!

sadtromboneMember since 2020
4 years ago

There is such a thing as playing above one’s true talent level, but if that’s the case you would also usually see that in FIP. The former things you’re talking about are covered by SIERA, FIP, xERA, and ERA (in the last case, badly). In the latter case you’d be looking at ZiPs or Steamer; ZiPs and Steamer rely on many of the same inputs as FIP / xERA / SIERA but not entirely, the weights are different, and there’s a lot of regression to past experience and the age curve.

soddingjunkmailMember since 2016
4 years ago
Reply to  sadtrombone

Sure. And that’s the point I’m trying to make. FIP is the right tool to pull out of the toolbox sometimes, but depending on what you’re trying to do other tools may be more appropriate.

Where I see problems on the internets is when people are imprecise about the question they’re attempting to answer and/or use metrics(WAR) as a bludgeon to answer too many questions. (like my example above)

sadtromboneMember since 2020
4 years ago

I must have misunderstood what you were saying, then. I agree with that more general point.

DLHughey
4 years ago

This distinction is what I think the article misses, and it turns this article into something that’s 10+ years late to the discussion, content on creating the strawman of “it’s better than ERA/RA,” but not really focusing on justifying why, and more importantly, when it makes sense to use FIP.

As a stat that is intended to be descriptive of talent but functions better as a predictive stat of outcomes, it’s just weird to use it to quantify the value of a player’s past contribution.

Ivan_GrushenkoMember since 2016
4 years ago
Reply to  DLHughey

It’s not meant to be a predictor nor measure of talent. It’s meant to be an evaluator of performance isolating controllable events. The questions are whether there are better ways of doing that, and whether this is an objective worthy of headlines

Last edited 4 years ago by Ivan_Grushenko
joe_schlabotnik
4 years ago

i completely agree!

Brad JohnsonMember
4 years ago

I still don’t see why we don’t use some variation of SIERA-WAR instead. At least offer it in the same table as RA9-WAR. FIP misses the mark in one direction. ERA misses it in the other. FIP is probably marginally better overall unless we’re making specific statements about what actually happened in one year. SIERA also misses the mark (in the same direction as FIP), but it’s a LOT closer to it (the mark that is).

The other choice is to make it clearer that WAR for pitchers is exceptionally imprecise. And, ideally, reinforce how you quickly evaluate a pitcher without referencing WAR. Dave Cameron always used to say a full season WAR is basically +/- 0.5 wins. For pitchers, it might be more like +/- 1.5 wins. (I don’t know, I only guestimate. I let others do the math).

Ivan_GrushenkoMember since 2016
4 years ago
Reply to  Brad Johnson

How do we know that SIERA is “closer to the mark”?

AustinMember since 2019
4 years ago

With FanGraphs adopting OAA as part of their WAR calculation, I have started to wonder if xERA might find its way into pitcher WAR.

Domingo AyalaMember since 2024
4 years ago

In terms of the value of what actually happened, wouldn’t bWAR be a better indicator? FIP is only showing true talent, but if you got good results, that value still occurred.

CC AFCMember since 2016
4 years ago
Reply to  Domingo Ayala

Not be overly semantic, but it depends how you define “value.” In my view, bWAR or a RA9 based war would be more accurately called descriptive of what happened. If you define “value” as results, then sure.

If you define “value” as meaning the specific influence of one player, i.e. the pitcher, then I don’t think so. In that case, bWAR depends on a lot more factors outside of the pitcher’s personal control than a FIP based WAR. The FIP based WAR does a better job isolating the pitcher’s contribution to the runs being scored, imo, so in that sense, I think it is better at determining “value” as defined by one player’s contribution.

Ivan_GrushenkoMember since 2016
4 years ago
Reply to  Domingo Ayala

To go further WPA shows how much you actually affected winning, even closer to “value of what actually happened”.

sadtromboneMember since 2020
4 years ago
Reply to  Ivan_Grushenko

WPA only really makes sense as a game and team-level measure. It’s literally defined as measuring excitement in the game.

Ivan_GrushenkoMember since 2016
4 years ago
Reply to  sadtrombone

It’s defined as increasing chance to win not excitement

hazelrah
4 years ago
Reply to  Domingo Ayala

I agree that we want WAR to be a quick number for how good someone performed in a year. Don’t lower a pitcher’s WAR because they were lucky… The team still got the benefits and we’re not measuring “true skill” with WAR!

Hitter WAR is based on what actually happened but pitcher is based on “what should have happened” and that’s a mistake in my opinion.

dodgerbleu
4 years ago
Reply to  hazelrah

Huge mistake.

Domingo AyalaMember since 2024
4 years ago
Reply to  hazelrah

This is exactly it.

Ryan DCMember since 2016
4 years ago
Reply to  Domingo Ayala

No, that is not it at all and I don’t quite get how people keep making the same mistake after all this time. FIP DOES measure what happened, it just measures a subset of what happened rather than every single event—but the entire point of the article you are commenting on and thus hopefully read is that the subset of events it measures captures enough data to tell us not only what happened but also what is likely to happen going forward, and does so more effectively than the criteria used by, for example, ERA.

dodgerbleu
4 years ago
Reply to  Ryan DC

Nope

Ryan DCMember since 2016
4 years ago
Reply to  dodgerbleu

Yes, FIP measures home runs that actually happened, strikeouts that actually happened, and walks that actually happened (and hit by pitches that actually happened).

hazelrah
4 years ago
Reply to  Ryan DC

We’re saying WAR should be a measure of what has happened and not try to predict what will happen going forward. The same as it is for hitters

2wins87
4 years ago

FIP ends up being really accurate because the X% of balls in play actually correlate really well with what is included in FIP (Ks and HRs mostly). Because of this, it might still be the most accurate predictor we have in aggregate over short samples, since the attempts to adjust for defense might always introduce too much noise to be as predictive over a season or less. However, if this noise is itself uncorrelated then an RA9 based approach can still be more predictive with a large enough sample.

The more nuanced problem is how large the sample needs to be. I think it is uncontroversial that there are guys who consistently get weaker or harder contact on average than would be expected from a simple correlation to Ks and HRs. Anecdotally, I think it takes at least 3 or 4 seasons to really be able to know this though, and that’s for the most extreme cases. For a guy where the effect is small but real, we’re probably never even going to know it.

Given a 20 year career for an extreme case like Mariano Rivera though, the 35% difference between RA9-WAR and FIP-WAR is real and something we should be aware of at the very least.

tung_twista
4 years ago
Reply to  2wins87

>FIP ends up being really accurate because the X% of balls in play actually correlate really well with what is included in FIP (Ks and HRs mostly).

This is a key point that took me a long time to grasp.
The condition for FIP being a useful measure is NOT
‘all pitchers induce similar level of contact’ (as many people might naively think)
but rather
‘all pitchers with similar level of FIP induce similar level of contact’
Not exactly surprising that pitchers with high K% and low HR%, BB% generally induce weaker contact than pitchers with low K% and high HR%, BB%.

sadtromboneMember since 2020
4 years ago
Reply to  tung_twista

This, along with a few other comments on this article, are among the best I’ve seen here in a long time. Well done, this is an incredibly subtle point but so important.

2wins87
4 years ago
Reply to  tung_twista

Yes, I do think it’s important to note though that there are clear examples that show that not all of the variance is random for the level of contact of guys in a certain bucket of FIP. I think an xERA based WAR could outperform FIP, but even that might not overperform in a one year sample. A lot of variance in the level of contact is just random noise, and if this noise is much greater than the skill variance then it could still take a lot longer than one year for the population variance to swamp the sample variance.

g4Member since 2020
4 years ago

I can see that FG commenters are well beyond me in terms of stat theory, but I very much appreciated this article. Over the years the oft-repeated 40% criticism had burrowed its way into my subconscious so it’s nice to learn that the other 60% isn’t outright ignored by FIP, which at least mitigates its biggest (surface) flaw.

One thing I question — and encourage readers to explain — is whether FIP fares better/poorer than other measures over far less than 100 IP, which encompasses full seasons for the majority of pitchers these days. I understand SSS hinders any metric, but is FIP’s underlying assumption on BIP comparatively more vulnerable when it comes to measuring relief pitchers? When I see a FIP cited 11 innings into April, I can’t help but wonder if an admittedly abominable but basic stat like WHIP doesn’t offer more clarity on how a guy is throwing the ball. (I know WHIP wouldn’t account for, like, all 3 hits yielded thus far being homers, but it’s also bloody unlikely such extreme results will continue.)

When citing FIP/SIERA/ERA+/etc in routine baseball writing, is there a consensus on the most useful stat to illustrate short-term performance? Or is sample size pretty much irrelevant when stacked against the more fundamental disagreement as to the degree to which pitchers can contribute to outcomes of BIP?

Thanks!

dl80Member since 2026
4 years ago
Reply to  g4

The problem with WHIP is that it doesn’t, as you noted, differentiate between different kinds of hits. It also inaccurately values low-walk, high hit guys over high walk, low hit guys, since it is putting a walk and a homer both in as the same value.

That said, FIP or xFIP should be fine for most pitchers even with small samples. It will get some wrong, but probably fewer than just WHIP alone.

As an example, Zach Greinke and Kyle Wright have nearly identical WHIPs, but Wright’s FIP is 2 runs lower because Greinke isn’t striking anyone out and is instead relying on getting bad contact. That could continue (it has been for years), but Wright is more likely to be better going forward.

2wins87
4 years ago
Reply to  g4

The simplicity of FIP is exactly what makes it perform better for smaller samples than other more complex metrics. Basically every variable you add to a model has some random noise associated with it. Even if this variable is measuring some skill that is not completely correlated to the other variables you are already using, generally it will take a larger sample to to find the signal from the noise.

As I’ve mentioned in other comments, because quality of contact is already highly correlated to FIP, the added utility of including some measure of contact is relatively small, even if it is a real improvement. We should expect to need a larger sample for this improvement to actually be realized since we have also added a bunch of random noise.

weinventyou
4 years ago

improved pitching WAR = (fWAR + bWAR) / 2

Left of Centerfield
4 years ago
Reply to  weinventyou

That’s what Tom Tango said years ago. Just average fWAR and bWAR and be done with it.

Pirates HurdlesMember since 2024
4 years ago

This alludes to the often stated”just average out the projections” that we hear every winter. I’ve yet to come across any evidence that doing so is a good idea statistically. I’d have the same question for Tango, if the average fo the two correlates better than one or the other with future performance?

sadtromboneMember since 2020
4 years ago

Often, multiple indicators of the same construct mean that you can parse out error associated with each one. But to do that, you need more than two, and there are more sophisticated measures than just a simple average.

weinventyou
4 years ago

It may or may not be the best way to go about it, but it certainly is one of fastest and easiest. Also, I wish I could somehow work “you know, I recently had an idea that was very similar to what Tom Tango said years ago” into conversations, but somehow the opportunity never comes up!

hazelrah
4 years ago
Reply to  weinventyou

I’d like to see RA9 or ERA for most innings, with FIP or league average scoring rates for innings where a pitcher was relieved

MorboTheAnnihilator
4 years ago

I like to paraphrase Winston Churchill statement about Democracy when I think about FIP, “FIP is the worst predictive measure of pitching performance – except for all the others that have been tried.”

hazelrah
4 years ago

WAR should be about what actually happened though, not predicting true talent

MorboTheAnnihilator
4 years ago
Reply to  hazelrah

I understand the sentiment but don’t necessarily agree with it. WAR can have different purposes. I’m far more interested in a version of war that has predictive power because a version with predictive power is less likely to credit a player for good or bad luck than one based on what happened. In my opinion WAR should try to determine the actual talent of a player and not how lucky or unlucky they are in a given year. Now if you argue luck is a talent then so long as it’s consistent then WAR a talent based war should capture that.

hazelrah
4 years ago

Agree to disagree I guess. I think most fans (and sports media people) looking at WAR use it to see how good a season someone had, not how good they should be or will be in the future. And for hitters that’s basically what it is, but for pitchers it isnt

NATS FanMember since 2018
4 years ago
Reply to  hazelrah

Your wrong. FIP determines how good a pitcher was despite the quality of defense. That’s far more accurate than a measure that allows a good or bad defense mess with the results.

nomarwomacks
4 years ago

The point that FIP only impacts the PA*BABIP-lgBABIP is a clever point I’d hadn’t thought of before and really does a great job of making me grasp just why FIP/DIPS stuff works.

However, “It’s conceptually obvious FIP is fatally flawed” is really impossible to get around and that’s at the heart of the “FanGraphs pitching WAR only considers (X)% of what a pitcher does, how can that be used for value?” objection Clemens is referencing. The amazing thing is more that “let’s completely ignore anything the defense has a role in generating” works as well as it does.

It’s just obviously wrong to assume Home Runs allowed (and, for fWAR, IFFP%) is the only measurement of soft/hard contact just can’t be right. I know Ben addresses this in his article (in a genuinely interesting and informative way) but you really just can’t get beyond the fact that there obviously is some better way at grasping this even if we currently don’t have access to it. FIP may be a good proxy for “true deserved xwOBA allowed” but it’s in no way conceptually measuring the BIP half of that variable.

Charles BalterMember since 2019
4 years ago

It’s definitely a useful statistic. From my vantage point, the blind spot for hardcore adherents to modern statistics and metrics is not realizing that statistics like WAR or FIP, as much of an effort as people make to make things scientific and unbiased, reflect (on some level) a human being’s OPINION of what should or shouldn’t have taken place.

If we’re criticizing traditional statistics on pitching and defense, we say, “That error charged to that person shouldn’t have happened. It should have been ruled a hit, and the pitcher should have been charged for those runs.”

Advanced metrics can have biases built-in, as well. When judging defensive metrics, just like the traditional official scorer, we’re judging what we feel should have happened. A ball leaks through the infield untouched for a hit. UZR or DRS aims to determine whether that play should have been made, and assigns the player a defensive value based on that.

I’m not sure there’s one perfect statistic for evaluating pitching. I think you have to look at a lot of different statistics, as well as watch the pitcher perform. Two Hall of Fame pitchers, Mariano Rivera and Greg Maddux, routinely outperformed their FIP. If you didn’t watch them pitch, you might say that they pitched in good luck. If you did watch them pitch, you understand why they were more successful than some numbers would suggest.

Contrarily, if you watched A.J. Burnett pitch, you know why his ERA was often higher than his FIP.

It’s really hard to find the perfect metric/statistic. That’s why we have so damn many of them.

Adam Daily
4 years ago

People will always inherently trust fewer events that “actually happened” even though they are very often poorer estimates of future performance.

LTGMember since 2020
4 years ago

This is all great detailed work. But I worry that someone might read it and think that FIP is the best ERA-estimator out there, when that is a very tendentious position. At this point we should be studying the differences among FIP, SIERA, dERA, xERA, and whatever other estimators there are that I can’t recall or haven’t heard of yet.

sadtromboneMember since 2020
4 years ago
Reply to  LTG

A while ago people were talking about DRA and DRA- at Baseball Prospectus. There are a couple other xERA-type estimators like FRA and pCRA. I’m not sure dERA has much of a role anywhere.

Radhames Liz
4 years ago

More articles like this please! Even if I do prefer SIERA and xERA, this is a fascinating debate.

Timmeh49Member since 2017
4 years ago

A request: would it be possible for Fangraphs to include “WAR-FIP” (i.e. the FIP flavor that FG uses in their WAR calculation — the one that counts IFFB as strikeouts) in the “Value” pane for player stats? Or at least make it available via Custom Tables?

jakesingi
4 years ago

Great article… never thought of it this way until now.

RonnieDobbs
4 years ago

I will take WHIP over any other pitching metric. WHIP and OPS changed things a lot for the better. Here we are decades later trying to find the next worthwhile metric – I am not holding my breath.

Sultan of Say
4 years ago

Are we at a point where we can use Statcast data? Each batted ball has a probability of being an out. If we simply use that, doesn’t it isolate each event to the pitcher?

dodgerbleu
4 years ago
Reply to  Sultan of Say

Statcast data doesn’t reset until the end of year, I believe. So this year, most everyone is underperforming their Statcast data because Statcast is using 2021 data (from the 2021 ball(s)). Given that it’s not a real time stat (I believe), it can’t really be used exclusively.

Left of Centerfield
4 years ago

Ben is a bit dismissive about using information like doubles and triples to improve FIP. I think he’s wrong about that.

For one thing, home runs are relatively rare events and any sort of rare event has a large error around it. Obviously, we know the exact number of home runs someone gives up. But two pitchers could pitch exactly the same and give up a very different number of home runs based on factors outside of their control, such as the swing decisions of the batters that they happened to face. Now the error around home runs is certainly reduced over the course of a pitcher’s career. But for shorter periods, such as a single season, the error is still quite large.

All of that is exacerbated by the fact that home runs are weighted so much heavier in the FIP formula than the other components (13 for home runs, 3 for walks, 2 for Ks). So adding more information regarding hard-hit balls (such as doubles and triples) into FIP would certainly help.

And then there’s this. In his rookie season, Shane Bieber had a 3.23 FIP. In his second season, it was 3.32, basically the same. His ERA though dropped from 4.55 to 3.28. And there’s an easy explanation for why that happened. In his rookie season, he gave up a double or triple every 11.5 PAs. His second season, it was down to one every 22.6 PAs. There’s little doubt that Bieber pitched much better his second season. But FIP misses that improvement because it fails to include doubles and triples. (his FIP was clearly too low in his rookie season and probably about right in his second season). This is obviously just one example, I’m sure there are a lot more.

I honestly don’t understand the rationale for excluding what could be very valuable information in FIP. Why not try to make the measure better than it is?

sadtromboneMember since 2020
4 years ago

This is essentially the theory behind DRA at Baseball Prospectus. It’s pretty complicated, and I would be awfully worried about model overfitting as a result. And it’s also very much a black box, with a whole lot of complicated variables in it that are not at all the same across pitchers. So it also tells us less about what types of things make someone successful, which is a useful goal but maybe not for all purposes.

The good news is that it is pretty sophisticated modeling, and it seems more accurate than FIP (and way more accurate than straight ERA). So it deserves a longer look.

Ivan_GrushenkoMember since 2016
4 years ago

How is 2B and 3B better than just measuring really hard hit balls?

Left of Centerfield
4 years ago
Reply to  Ivan_Grushenko

I went with doubles and triples because they fit better with the framework of FIP which relies on outcome measures (HR; BB; Ks) rather than process measures.

Mahoney
4 years ago

I have no problem with any of the variations of WAR, or other metrics, as long as they are at least somewhat transparent about their composition, so we can debate their respective virtues (one thing I don’t particularly like about DRA).

If I had to decide, I’d use RE24 as the basis for runs above/below average for both offense and pitching + defense before applying the positional and replacement level adjustments. RE24 allows for the possibility for situational skill in how each inning’s PA-level events sequence into run expectancy, so I think this fits in somewhere. The challenge would be to find a way to back out the value of the defense at a play-by-play level to get to the pitcher’s contribution to RE24 prevention – it seems like it should be doable, but would require a lot of play-level data tabulation that needs to fit to an acceptable defensive metric (most of which are not publicly available at a play-by-play level).

2skupz
4 years ago
Reply to  Mahoney

I think this implies an ability to control when a double happens, which I’m not a huge believer of. But, interesting thought. I’d say that’s a stat that should exist!

Last edited 4 years ago by 2skupz
Antonio BananasMember since 2026
4 years ago

Why not just use batted ball information? Some hard hit balls are fly outs, some are home runs. A 330 foot line drive is a home run in Yankee Stadium, a 405 foot shot to center is a flyout in Citi Field.

Ivan_GrushenkoMember since 2016
4 years ago

Is xERA more serially correlated and/or predictive than FIP or other estimators?

sadtromboneMember since 2020
4 years ago
Reply to  Ivan_Grushenko

It’s not correlated with the pitchers’ defense, if that’s what you mean.

Using xERA (or other batted ball information) is where I think we should probably end up. You want a method that it is independent of fielding, and batted ball information is pitching independent. It doesn’t matter what the defenders are behind the pitcher, where they are positioned, etc. The bat comes off the ball and you get everything before the defense steps in.

That said, I’m not sure xERA or xwOBA is mature enough to serve in this role yet.

CeetarMember since 2020
4 years ago

Personally it just bothers me that it’s called “Fielding independent” but your FIP goes down if you get a groundout to second.

NATS FanMember since 2018
4 years ago
Reply to  Ceetar

but it doesn’t. FIP assumes all defense dependent plays are made at the league average rate. It does not adjust for individual defensive plays. in other words, it does not matter if an individual play results in a groundout to second, a missed catch, or an E4. FIP will be the same cause the average result is used.

Last edited 4 years ago by NATS Fan
Mahoney
4 years ago
Reply to  NATS Fan

Technically, Ceetar is right….outs of any kind, including balls in play, help increase the denominator of FIP, which is innings pitched. So if you have two pitchers with the same BB%, K%, and HR% allowed, the one with the lower BABIP will have a slightly lower FIP as well.

Which, at the end of the day, strengthens Ben’s argument a bit more.

CeetarMember since 2020
4 years ago
Reply to  Mahoney

Yeah, the “we’re basically assuming that other stuff is all random/rounded out/negligible” helps me square it in my head a little better. I understood what it was doing, and obviously I believe the numbers when they say it’s a decent predictor of next year ERA, I just couldn’t quite shake the denominator thing,and always felt like it was missing something as a result. It probably is, but this helps me believe that it’s probably something somewhat negligible.

D-WizMember since 2019
4 years ago

That’s about as cogent an argument for FIP/against RA9 as I’ve ever heard. Great stuff!

themastiff
4 years ago

This is a compelling defense of FIP vs other box score derived stats. How should we think about newer stats such as xERA, xBA which rely on pitch tracking data from Statcast? Presumably they really do “consider” 100% of the action and also do a good job of isolating pitching from defensive performance.

NATS FanMember since 2018
4 years ago

Right now the Nats starting pitchers have the following FIPS and ERAs:
fip era
Corbin 3.31 7.16
Fedde 4.39 4.68,
Sanchez 3.74 6.75
Gray 3.94 3.12
Adon 5.58, 7.38 he will be replaced if/when Strassburg is back
Rodgers 5.42 4.41 just moved to bull pen

Im really hoping FIP turns out to be a great predictor of future performance because we have some decent FIPs and dreadful ERAs.

Last edited 4 years ago by NATS Fan
sadtromboneMember since 2020
4 years ago
Reply to  NATS Fan

Corbin is in a weird spot because he’s been hit hard but very few balls have gone out of the park. So his xERA is probably a better indicator of what’s going on, which pegs him more at about a fifth starter. It’s entirely possible he’s just played some really good offenses so far but it’s also possible he’s not locating. The ERA is bad, but I think he’ll be playable going forward.

This is probably also true of Aaron Sanchez.

That said, if you want predictions of the future, your best bet is to take a look at ZiPs rest of season projections. Although all of them are off in terms of absolute magnitude of runs allowed / created because of the new ball, since it’s happening to everyone many stats normalized to league average will be fine.

rjm311Member since 2022
4 years ago

Would you still recommend doing this for the career stats? I think of a player like Nolan Ryan where his FIP for his career was 0.22 lower than his ERA over more than 5k IP. Is evaluating him using FIP still preferable?

2skupz
4 years ago
Reply to  rjm311

I think the point is that 180 innings or so is almost always more noise than signal. Famoulsy Maddux has at least one two years stretch where his BABIP was like .100 different. I think when you have Ryan innings you could say ERA (or, more accurately probably runs/9) are a better indicator of talent.

Morris GreenbergMember since 2020
4 years ago

One thing I think should be added to FIP is pitcher pickoffs. It’s probably a small effect, but there is plenty of evidence that pitchers have a large effect on stolen bases and the difference between getting people out a portion of the time and someone turning a single or walk into being on second seems like a signal worth measuring that is for the most part independent of other fielders.

2skupz
4 years ago

That isn’t wrong, but, I do love the simplicity of being able to take 4 stats and combine them with integer math to get a reasonable guess at FIP. Enjoy the simplicity. SIERA may be a better candidate for using pickoffs. Though again, you aren’t getting that much closer and it’s a lot harder to do.

szakylMember since 2024
4 years ago

How correlated are double and triple rates to HR rates? It would make sense that pitchers who give up more HR would also have elevated 2B and 3B rates. IFF they are mostly correlated, is it reasonable to believe that the HRs included in the FIP calculation could include a greater coefficient to account for this?

Last edited 4 years ago by szakyl
ReynoldMember since 2024
4 years ago

Thank you. Cogently argued.

Xerostomia
4 years ago

I think I followed the argument. At the same time if WAR is based on FIP, then players like Brad Zeigler perhaps have been unfairly punished?

evo34Member since 2023
4 years ago

I think FIP is fine for what it is, but people constantly misuse it to predict future ERA. Specifically, FIP does not account fully for park factors (HRs are noisy), defense (by definition), past level of competition, etc.

Yes, it’s better than ERA for prediction, but if you really want an ERA prediction, use the BAT or Steamer Ros projections. If FIP was truly a better predictor, do you think Carty and Cross would bother with their respective algorithms?

(No, this is not a strawman, as both authors and commenters on this site are regularly looking for differences between FIP and ERA and attributing it all to variance, as if FIP represents a pitcher’s true ERA and we have no way of knowing more than what FIP suggests).

Last edited 4 years ago by evo34
TommyfastballMember since 2016
4 years ago

What the heck did I just read? That assuming average results is somehow different than ignoring balls in play? Love Ben, but ths was a waste of my time.

hughduffy
4 years ago

I think Ben gets at the core of the issue when he writes, “For 95% of pitchers, FIP gets 96.5% or more correct.” For the bulk of pitchers, it doesn’t make that much of a difference. 

That also means that for 5% of pitchers, it makes a real difference.

My problem with using FIP for WAR are those outliers, mostly that 2.5% of pitchers at the top whose contributions aren’t properly measured. We know the value of soft contact, of limiting exit velocity, limiting hard hit balls, and inducing ground balls. If they have it, good pitchers use good defense, because it allows them to be more efficient and stay in games longer.

So if we’re talking about past value and trying to measure pitcher contributions, pitchers knowing how to induce the contact they want to get outs, while difficult to measure, should show up. There’s a reason why Zack Greinke has averaged a -0.583 RA/9 Differential from 2015-2021. He’s good at reading hitter swings and getting them to hit balls to his fielders.

These outlier seasons aren’t easily replicated year-to-year for good reasons: defenses change, pitchers get older, change teams, and the margin for error is very small when you pitch to contact. But pitchers can read swings, can see what the hitters are trying to do, and pitch them so the hitters do what the pitchers want them to do.

Is it skill? Breaks of the game? Defense? There’s definitely some skill involved, especially since some pitchers do outperform their FIP for extended periods of time. Getting hitters to hit balls directly to your fielders is a skill, one that Greg Maddux did masterfully during his 4-consecutive Cy Young years.

That’s a big part of why I don’t think FIP-based WAR is good for awards season. We are looking at the top pitchers and FIP-based WAR consistently fails to measure the total contribution of pitchers.

In 2021, Brandon Woodruff had a -0.764 RA/9 Differential. Walker Buehler had a -0.733 RA/9 Differential. Max Scherzer had a -0.663 RA/9 Differential. Zach Wheeler had a -0.140 RA/9 Differential. Corbin Burnes had a 0.127 RA/9 Differential. Who wins the NL Cy Young? The guy with the fewest innings and best three true outcomes numbers.

Is it defense? Brandon Woodruff’s RA/9 differential was 0.891 lower with the same defense. The breaks of the game? Burnes’s HR/FB ratio was 6.1%, the lowest in the major leagues among qualified pitchers, that’s a break for you. Both Burnes’ and Wheeler’s RA/9 Differential were within the average due to defense. Woodruff’s, Buehler’s, and Scherzer’s RA/9 Differentials were way outside the average, and the difference was large enough that it was not just random chance, not just the breaks of the game, and not just variations in defense.

The real problem with FIP based WAR is in many ways the same problem with FIP: the outliers.

AaronSabin
4 years ago

The argument against FIP is even more simple than you said, Ben. FIP actually rewards pitchers like Bill Singer for giving up hits! Because it is calculated on a per-inning basis, the more hard-hit balls a pitcher gives up, the more likely they are to become hits, and the more chances the pitcher has to strike the next guy out. If it were calculated on a per PA basis, it would be neutral to contact.

szakylMember since 2024
4 years ago
Reply to  AaronSabin

Pitchers also have more opportunities to give up a walk, hit-by-pitch, or home run because they’re not out of the inning yet.

dukewinslowMember since 2020
4 years ago

I’m not sure what the outcome variable you would want is (ERA? Eh), but you could just throw the inputs for FIP and…. Everything else in a model that predicts ERA and do dimension reduction on the “everything else” (two step LASSO IMO) to see what the marginal improvement to prediction is.

Angelsjunky
4 years ago

When considering FIP, I always think of Tom Glavine and Javier Vazquez. Glavine was, in some ways, the third wheel to Maddux and Smoltz; he didn’t have Maddux’s mastery of the zone and his repertoire, nor Smoltz’s blazing stuff, but he knew how to pitch. He was a classic “pitchability” guy, as expressed by his 3.54 ERA to 3.95 FIP. Vazquez had great stuff and was a very good pitcher, but was never as good as he “should” have been, as you can see with his 4.22 ERA and 3.91 FIP.

Now if we use FIP, they’re basically equal – and no one in their right mind would have said Vazquez was an equal pitcher to Glavine (or slightly better, actually) – despite the impressive peripherals.

Now you could argue that Glavine is in the Fame of Fame because he won 305 games, while Vazquez won only 165. Certainly, if Vazquez had won 250+, he’d at least be in the conversation. But I’m talking about their “pound-per-pound” value, and WAR likes Vazquez more in terms WAR per IP, and FIP says they’re equal – and I don’t think that adequately represents their comparative quality as pitchers.

FIP doesn’t account well for intangible qualities – or what some call pitchability: for a pitcher’s knack for getting out of jams or knowing how to pitch to certain batters or even, I hate to say it, the ability to establish a reputation or cozy up to umpires so they might treat you more favorably. ERA does – not intentionally, but by default – because it represents actual results.

So for me, ERA is a better indicator of overall quality of performance, while FIP is useful as a predictor of future performance.

Ryan DCMember since 2016
4 years ago

I feel like I’m going to go insane: FIP MEASURES WHAT ACTUALLY HAPPENED! THE WALKS HAPPENED! THE STRIKEOUTS HAPPENED! THE HOME RUNS HAPPENED! THE HIT-BY-PITCHES HAPPENED! THEY ALL ACTUALLY HAPPENED! FIP DOES NOT MEASURE ANYTHING THAT DID NOT ACTUALLY HAPPEN! ALL THE INPUTS HAPPENED!! AHHHHHHH!!!!

Domingo AyalaMember since 2024
4 years ago
Reply to  Ryan DC

Nobody is questioning what it includes. It’s what it excludes. I’m not saying FIP is right or wrong, but your comment is not what anyone is arguing.

Angelsjunky
4 years ago
Reply to  Domingo Ayala

Yes, exactly this.

Domingo AyalaMember since 2024
4 years ago

What would happen if we calculated hitting WAR by their three true outcomes?

Angelsjunky
4 years ago
Reply to  Domingo Ayala

We’d get a very inaccurate picture of a hitter’s value.