On the Consistency of ERA

We know that ERA isn’t a perfect indicator of a pitcher’s talent level. It depends a lot on the defense behind the pitcher in question. It depends a lot on luck in getting balls in play to fall where the fielders are. It depends a lot on luck in getting fly balls to land in front of the fence. It depends a lot on luck in sequencing — getting hits and walks at times where it doesn’t hurt too much.

That’s why we have DIPS. Stats like FIP, xFIP, SIERA, my recent SERA, and Jonathan Judge’s even more recent cFIP all attempt to more accurately measure a pitcher’s talent by stripping those things out. But what if there was an easy way to figure out how much ERA actually can vary? How likely a pitcher’s ERA was? What the spread of possible outcomes is? The aforementioned ERA estimators do not address that issue. They can tell you what the pitcher’s ERA should have been with all the luck taken away (or at least what they think the ERA should have been), but they can’t answer any of the questions I just posed.

But SERA, which simulates ERA instead of using a formula, can be modified a little bit to help us out. If we, instead of simulating hundreds of thousands of innings at once, break the simulation up into 50- or 100- or 200-inning parts, we can find the distribution of outcomes for that pitcher. This is what I think the real advantage of SERA is. It’s a decent ERA estimator — not quite as predictive as xFIP — but it’s biggest asset is the ability to tell us about the variability in a pitcher’s ERA.

With a new script, we can now set the IP to a lower number, iterate that hundreds or thousands of times, and see the distribution of outcomes. Generally the distribution is not Normal, but luckily most of the time the distribution is around the same. It usually looks like this (the dotted line is the mean):

ERA_variability_density

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

 

Of course, the spread changes based on the inputs (this was made using league average as the inputs), the innings per simulation, and even the number of iterations, but it’s usually pretty similar. The 10th percentile is generally between 1.05 and 1.2 standard deviations from the mean, and it is usually skewed slightly right (which is pretty intriguing; I don’t know why that is — but my guess is that really bad values are easier to get than really good values).

For an average pitcher, the 10th percentile is just under 3.00, the 25th percentile is just over 3.25, the median is about 3.63, the mean is about 3.66, the 75th percentile is just about 4.00, and the 90th percentile is just under 4.40. These are kind of like what percentile projections for Pecota do — estimate what the best- and worst-case scenarios are. I will later combine this with pSERA to look at each pitcher’s uncertainty, as well as plain mean projection, is for next year.

The next step is trying to figure out what makes a pitcher’s ERA more stable or unstable. Intuitively, you (or at least I) would think that strikeouts and walks both work to decrease volatility because they take out batted ball luck. However, that is not the case:

K and BB Heatmap

 

The StdDev is the standard deviations of the 1000 simulations of 180 innings each I ran using the given K% and BB% (with batted ball inputs all set to league average of 44.8 GB%, 34.4 FB%, 20.8 LD%, 9.6 IFFB%, and 9.5 HR/FB%). You can see that as K% increases and as BB% decreases, the standard deviation gets lower. That means that yes, a higher strikeout rate does work to decrease the variability in ERA, but a higher walk rate does not. This is probably because the sequencing luck involved with putting more runners on base has a greater effect then the batted ball luck involved with allowing more contact.

So, then, visiting the K-BB% leaderboards should give us a good sense of whose ERA last year was a pretty good indicator of their true talent. The higher the K-BB%, the lower variability there is in that pitcher’s ERA.

What about batted balls? Which batted ball types make ERA the most unstable? Here’s a chart similar to the one above, using the same numbers of 1000 iterations of 180 IP, showing GB% and FB% and the resulting standard deviation of the ERAs. LD% isn’t shown on the graph, because graphing four variables is hard, but it’s just 100-GB%-FB%; it’s lower towards the top-right and higher towards the bottom-left.

BIP Heatmap

 

The LD% here is sometimes negative and sometimes crazy high, which is obviously unrealistic. But the overall trend is consistent throughout: the higher the GB% and FB% — and in turn, the lower the LD% — the more consistent and less variable the ERA is. It’s the same thing as strikeouts and walks. Since line drives turn into hits more often, allowing more of them means more runners on base and the ERA is more dependent on sequencing luck.

But LD% is usually unreliable and doesn’t carry over very much year-to-year. GB% and FB% are much more stable, so we should look at those to see which is more important to reduce variability in ERA — that will tell us more about pitchers’ future ability to maintain a consistent ERA.

GBFB

 

Another clear trend. This one shows that a higher GB% and a lower FB% decreases the variability. I had expected that a higher FB% would lead to a more stable ERA, since more fly balls gives IFFB% and HR/FB% a better chance to normalize and be closer to the league average. But that isn’t the case. When you think more about it, it makes sense, too.

A pitcher with a 50% fly ball rate (which is very high) would allow about 285 fly balls over 800 batters, assuming average K%, BB%, and HBP%. Using sampling techniques, we can figure out that one standard deviation of HR/FB% would be roughly 1.8 percentage points. For a pitcher with a 25% FB% (which is pretty low), the standard deviation is about 2.5 percentage points. That’s a huge gap in FB%, but not a huge difference in the variation of HR/FB rate. (All this is assuming that pitchers do not have control over their HR/FB and IFFB rates, which isn’t totally true, but is an assumption that holds well enough to be able to generalize safely.)

Much more important than allowing HR/FB and IFFB rates to stabilize is preventing big hits like home runs and doubles altogether, since they put more runners on and create more variability in sequencing. Ground balls do that very well; almost no ground ball ever ends up as something other than an out or a single, save the occasional one of these. Fly balls, on the other hand, end up in extra-base hits quite often (over 20% of outfield fly balls go for extra bases).

In a nutshell, I think that the most important takeaway is this: good pitching = more stable ERA. A better pitcher will have an ERA that is more indicative of their actual skill, while a worse pitcher who puts more runners on will have an ERA that could be very different from what their actual talent level is. More runners on leads to more uncertainty. The very close correlation between K%, BB%, GB%, FB%, and LD% to the ERA variance tells us that there is a very tangible effect of those things. A higher strikeout and ground ball rate and a lower walk, fly ball and line drive rate not only lead to a lower ERA, but also to a much more consistent one.

Next week, I’ll take a look at pSERA for next year and use it to find how much the ERA of each pitcher can be expected to vary.





Jonah is a baseball analyst and Red Sox fan. He would like it if you followed him on Twitter @japemstein, but can't really do anything about it if you don't.

21 Comments
Oldest
Newest Most Voted
Justin
11 years ago

“Intuitively, you (or at least I) would think that strikeouts and walks both work to decrease volatility because they take out batted ball luck.”

I don’t think that’s true for a per inning type estimator like ERA. Pitchers still have to get three outs per inning, so walks shouldn’t have much effect on how much contact is allowed per inning. We measure hitters on a per PA basis, so for them walks will take out batted ball luck.

Justin
11 years ago
Reply to  Justin

I should have added that I really enjoyed the article and look forward to more.

BipMember since 2016
11 years ago
Reply to  Jonah Pemstein

Guys who walk a lot of batters should probably see slightly fewer balls in play per inning because more guys on base means more of their outs will come from double plays and guys making outs on the bases, as a portion of their total outs made, compared to a similar pitcher who walks fewer batters. However, that effect is likely very small.

mario mendoza
11 years ago

nerdgasm

wildcard09
11 years ago

This combined with the cFIP article you linked to, and the pitch sequencing article from last week is all excellent stuff. I’m fascinated by pitching valuations and predictors, and the work that is being done recently is just amazing. Everybody keep doing a great job and thank you all for this.

jim fetterolf
11 years ago

” It depends a lot on luck in getting balls in play to fall where the fielders are.”

Beyond luck it also depends on the pitcher hitting the spot with a quality pitch. The defense is set and defenders lean expecting a certain location of a certain pitch to a certain batter. A pitcher who misses his target or throws a hanging slider or fastball up and middle will tend to have much more bad luck than a pitcher who can hit the target.

jim fetterolf
11 years ago
Reply to  Jonah Pemstein

Unfortunately the metrics aren’t available to quantify it, so it gets down to eyeballs. Maybe in a few more years advanced data generators will get a handle on it.

BipMember since 2016
11 years ago
Reply to  jim fetterolf

I’m not sure that this is any more than a tiny effect. Fielders are not usually running all over the field in between pitches. Even if it was advantageous to do so due to changes in the expected hit location density map for each pitch type, you would then have to consider how defensive positioning could tip off the batter to what pitch is coming.

I think that if we get more in depth batted ball data, we’d probably find that pitchers who consistently hit the corners just get weaker contact on average, and that explains most of the rest of BABIP that isn’t explained by luck.

jim fetterolf
11 years ago
Reply to  Bip

Shifts are an example of positioning for pitches to certain hitters and the results are non-negligible.

Jim Price
11 years ago
Reply to  jim fetterolf

Defensive shifts are more about the hitter. Teams don’t put on a shift for different pitches. Maybe an infielder will move a step if he knows the pitch but that’s about it. I would guess most of the fielder’s do not know what pitch is coming any better than the hitter does.

Cats
11 years ago

That script is just 790 if statements. I don’t even..

Robert Hombre
11 years ago

URL is beautiful. The article is beautiful. You, Jonah – you are beautiful.

In the mind, anyway. You might look like that jackanapes Cistulli for all I know.

Not a Statistician
11 years ago

Wait, isn’t this just because pitchers with better K/BB and higher groundball rates have lower ERAs? Isn’t this a statistical inevitability?

BipMember since 2016
11 years ago

This is about the expected variation in ERA, not the expected value. We know that a guy can pitch like a 3.00 ERA pitcher, but end up with a 2.50 or a 3.50 ERA due to good or bad luck. The question the article is trying to answer is whether we can predict the variability. And, in fact, a guy with a high K rate, low BB rate and high GB rate will be more likely to have an ERA closer to his skill level than a guy with the opposite set of peripherals. The latter guy will likely have a higher ERA, because he’s worse in the areas that matter, but he’s also more likely to have a ERA that deviates widely from his skill level – in either the positive or the negative direction.

DNA+
11 years ago

I suspect the skew is because the scale is truncated at zero. The most dominantly pitched innings (three strikeouts) are no better in terms of outcomes than a less dominant inning (two strikeouts, and a pop out), or even sometimes a poorly pitched inning (walk, walk, pop out, pop pout, walk, pop out). On the other hand, there is no theoretical limit to how many runs you can give up in an inning.

Bottom line
11 years ago

Even though ERA may be biased, it is based on real data; i.e. actual performance during a season, albeit somewhat skewed by the vagaries of errors. So if I am a GM and want the most consistent pitcher then I am interested in the pitchers with the highest K%-BB% and the highest GB%.

Perhaps a test of this theory would be to plot change in ERA over two seasons vs. GB% + (K%-BB%) in season one.