Are Pitching Projections Better Than ERA Estimators?

ERA estimators estimate how well a pitcher pitched in the present, and pitcher projections estimate how well a pitcher is expected to pitch in the future. Naturally, we’d expect projections to more-accurately predict pitchers’ future performances, since that’s what they’re designed to do. But it appears that ERA estimators can figure future performance quite well — and SIERA, in particular, has actually done a better job projecting pitcher performance that than traditional projections.

Projecting pitchers is harder than projecting hitters. Not only do pitchers’ skill levels change often, but simply estimating a pitcher’s skill at any given time is challenging. Differentiating their performances from that of their fielders’ — and removing luck — are difficult tasks necessary to isolate pitchers’ true talent.

Testing an ERA estimator’s ability to predict future ERA is the most common method of assessing its reliability, because if it is similar to a pitcher’s future ERA, then it is probably picking up a pitcher’s true skill level. The most common metrics for testing an ERA estimator are correlation with future ERA or the Root Mean Square Error (RMSE) with future ERA. The difference between using correlation and RMSE is that the correlation studies how much two numbers move together, while RMSE studies how close they are.

Say you were a general manager a year ago, and you saw two pitchers on the trading block who were coming off sub-3.00 ERAs. Their names were Jaime Garcia and Mat Latos. You knew both were due for a reversion to the mean. If you wanted to target the superior pitcher, you would want the ERA estimator with the superior correlation with future ERA. But if you were more interested in pinpointing each of their future ERAs to determine how competitive your team would be if you got either, the ERA Estimator with the superior RMSE should be your focus. In other words, the pitching metric with a best correlation with future ERA ranks them correctly, while the pitching metric with the best RMSE is better at predicting their performance in an absolute sense (rather than relative to other pitchers).

But you might wonder: why not just use a future ERA projection? The implicit assumption is that if you only want to predict future performance, you might as well factor in several years of data, aging, park effects and changes in skill level that a more sophisticated system would recognize. But I would strongly argue that doing a projection requires two steps:

1) Figuring out how well a pitcher actually pitched
2) Figuring out how this will change

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

If you don’t do the first step well, what can you expect from the second step?

Consider the following test for 2011 that shows the correlation of future ERA with several ERA estimators and projection systems. ERA estimators are presented in red and pitcher projections are presented in blue.

Estimator(N=258 pitchers)
Correlation of 2010 Statistic with 2011 ERA*
SIERA .480
ZiPS .470
xFIP .438
ERA .424
Marcel .420
PECOTA .413
FIP .402
tERA .396
Oliver .371

*For all tables, I use only pitchers with 40 IP in both years.

SIERA actually tops the three most commonly cited projection systems in correlation with future ERA. In fact, you would have done a better job just looking at park-neutral SIERAs in 2010 than all four projection systems’ park-adjusted ERAs. Interestingly, 2010 ERA itself didn’t fare that badly when ranking pitchers’ 2011 ERAs — and it actually topped Marcel, PECOTA and Oliver.

What if we look at RMSE? This will penalize ERA estimators that let luck play too large of a role, even if the estimators appropriately rank pitchers. But here, we again see that SIERA tops the pack, and xFIP actually beats all three projection systems — plus PECOTA and Oliver — rather handily.

Estimator(N=258 pitchers) RMSE of 2010 Statistic with 2011 ERA
SIERA 1.048
xFIP 1.069
ZiPS 1.071
Marcel 1.076
FIP 1.132
PECOTA 1.155
Oliver 1.168
tERA 1.171
ERA 1.221

Just to make sure that this discovery wasn’t a 2011 quirk, I looked at the previous six years. Note that PECOTA has changed architects several times, but this will aggregate their performances. Also note that I left out Oliver because it wasn’t available for all six years.

Estimator(N=1,576 pitchers) Correlation of Statistic with Next Year’s ERA(2006-2011)
SIERA .428
ZiPS .402
PECOTA .394
tERA .384
Marcel .377
xFIP .375
FIP .355
ERA .328

Over time, it appears that SIERA is the best at ranking pitchers — though tERA does edge Marcel. ZiPS is best among projections, though PECOTA’s strong 2007 and 2008 performances kept the aggregate score close.

Estimator(N=1,576 pitchers) RMSE of Statistic with Next Year’s ERA(2006-2011)
SIERA 1.126
Marcel 1.132
PECOTA 1.141
ZiPS 1.143
xFIP 1.148
FIP 1.212
tERA 1.236
ERA 1.387

The lesson learned from the last RMSE table is that, over time, projections do a better job being closer to future ERA than ERA estimators — other than SIERA. But SIERA remains on top. It’s worth noting that the reason Marcel tops PECOTA and ZiPS here is that Marcel has a far lower standard deviation (.527 for Marcel, versus .652 for PECOTA and .721 for ZiPS). It might not rank pitchers as effectively, but it predicts their performances better simply by assuming that all pitchers are more average than they appear. When Ubaldo Jimenez or Javier Vazquez fall back to earth, Marcel catches them; but Marcel also expected Tim Lincecum and Cliff Lee to regress significantly, as well, while ZiPS took a stronger stand on both. This made ZiPS better at correlation and Marcel better at RMSE.

What’s particularly interesting about this discovery is that ERA estimators should have far inferior correlations and RMSEs. Projection systems use several years of performance to understand a pitcher’s true talent level. ERA estimators only use one year of data. Ryan Vogelsong had a very low 2.71 ERA in 2011, but his SIERA was only 3.97. That’s well and good, but years of ERAs and SIERAs both hovering around 5.00 suggest an even harder crash in 2012 — unbeknownst to Vogelsong’s 2011 SIERA. Using more data and more highly regressed data should significantly benefit projection systems. But that extra information isn’t enough to overtake the benefit of SIERA.

The reason this is happening is because we still aren’t incorporating the lessons of ERA estimators properly into developing projections. ERA estimators give a truer estimate of how well the pitcher actually pitched. If we continue to ignore the interplay between different pitching skills and their effect on runs prevention, we will fall short in our ERA projection. First we need to understand the information contained in strikeout rate as it pertains to BABIP, HR/FB and situational pitching. We need to know what information is and isn’t contained in batted ball data. And we must comprehend how all of these statistics combine to affect run prevention. Once we know all of these, then we can understand how well a pitcher is pitching at the moment — and this can be carried forward to predict how well a player might pitch in the future.

When predicting the future, using several years of SIERA will do better than one year, and adding park effects and aging to previous years’ SIERAs will also help. Most importantly, regressing SIERAs towards the mean is necessary if we’re interested in approximate talent level, rather than just rankings. After all, the gap in RMSE is small between SIERA and projection systems, but SIERA’s correlation advantage is far larger. At this stage, using SIERA as a jumping off point appears to be the best method to project pitcher performance.





Matt writes for FanGraphs and The Hardball Times, and models arbitration salaries for MLB Trade Rumors. Follow him on Twitter @Matt_Swa.

70 Comments
Oldest
Newest Most Voted
Yirmiyahu
14 years ago

I like the idea of using defense-independent pitching info in the projection systems (Marcel, PECOTA, ZiPS). But pitching performance seems so volatile, I wonder if using just the prior year’s data might be more accurate than using multiple years of data anyway.

It would be relatively easy to make a Marcel projection using SIERA instead of ERA to test this.

Vin
14 years ago

Always great to see Matt Swartz contributing to the site. Another very informative piece.

jcxy
14 years ago

dumb question: what is the value of predicting ERA year over year as opposed to FIP or xFIP? in other words, ERA says less about a pitcher’s true talent level than FIP/xFIP, right?

what am i missing?

Nik
14 years ago
Reply to  jcxy

I suppose using ERA is just a way to neutralize the guys that over and under-perform their projected FIP/xFIP.

Baltar
14 years ago
Reply to  Matt Swartz

FIP is not an estimator, though it may be used as such, just as ERA can.
It is a measure of what actually happened independently of defense.
Thus, using FIP as the target stat for pitching estimator’s would be an improvement over using ERA.

jcxy
14 years ago
Reply to  Matt Swartz

ok, that makes sense. thanks!

Lewie Pollis
14 years ago
Reply to  jcxy

FIP and xFIP are estimators of ERA, so theoretically in the long run ERA = xFIP/SIERA/whatever you use to describe true-talent level. Projecting FIP instead of ERA would be like predicting what political candidates’ Intrade odds will be instead of who will win the election.

Awesome article, by the way.

Paul
14 years ago
Reply to  Lewie Pollis

Great analogy. Agree that this is a great article and much needed.

John R. Mayne
14 years ago
Reply to  Lewie Pollis

This is the comment of the day.

This is a fascinating finding. I wonder what it is that the projection systems are doing wrong. (Except for Marcel; I know exactly how Marcel works.)

There are a couple of cheap and easy things to improve our projection from year-1 SIERA (such as velocity, as Matt pointed out at THT in “You Shall Know Our Velocity,” which was kind to one of my articles) and that should extend SIERA’s lead.

Matt, or any other qualified person: Is there a good reason for the correlation scores to drop so significantly when going from the one-year sample to the five-year sample?

Travis L
14 years ago
Reply to  jcxy

Fantasy purposes as well.

Dekker
14 years ago

SIERA and xFIP are likely better than projection systems because they are better suited to remove biases like extreme park factors or fielding quality.

Yirmiyahu
14 years ago
Reply to  Matt Swartz

It’s more of a test. The reason we know that FIP/xFIP/SIERA tell us more about a pitcher’s true talent is because they are better at predicting future ERA.

Yirmiyahu
14 years ago
Reply to  Matt Swartz

Oops. I apparently replied to the wrong comment. That was meant to go to the guy asking why we cared about ERA.

But, speaking of park factors…. Someone in the fangraphs forums was asking a good question about SIERA’s park adjustment. How does that work? Are the underlying variables (K%, BB%, GB%, FB%) adjusted by park? Or does the park adjustment come after the main calculations?

David AppelmanFanGraphs Staff
14 years ago
Reply to  Matt Swartz

As we’re calculating it, there are no park adjustments since it’s really just: SO, BB, GB, FB. I think when Matt was testing it, adding park factors made pretty much no difference.

Yirmiyahu
14 years ago
Reply to  Matt Swartz

Good to know, guys. Thanks.

Are the park factors for K/BB/GB/FB available anywhere? I know they’re small, but I’m curious about the effects of certain parks.

Matt H
14 years ago

I’m sorry, I know you explained it like 50 times, but I’m still a little confused about the difference between correlation and RMSE. If I understand it correctly, correlation doesn’t care about how far apart the estimator/projection and next year’s ERA are, but how consistent that difference is among all pitchers. RMSE, on the other hand, simply measures how close the estimator/projection is to the next year’s ERA. So, the projection system could systematically project ERAs to be a full point lower than they end up being, and have a perfect correlation while having a bad RMSE. Correct?

Anon
14 years ago

Any thought of using SIERA for WAR calculations rather than FIP?

Yirmiyahu
14 years ago
Reply to  Anon

The Davids (Appleman and Cameron) are the guys to ask, not Matt Swartz. I think the answer would be “no,” with the usual reason that WAR is not supposed to be predictive- – it’s supposed to reflect what actually happened.

Now, I’m personally not content with that answer. By using FIP, you’re including one luck-based variable (HR/FB%), excluding other luck-based variables (BABIP, LOB%), and excluding some skill-based variables (GB%, SB/CS).

FIP is nice insofar as it’s simple and clean. The Three True Outcomes (BB, K, HR) are the pitcher’s responsibility; everything else is attributed to team defense.

Dekker
14 years ago
Reply to  Yirmiyahu

fWAR properly takes in account ballpark factors to largely neutralize the HR/FB luck.

Joe
14 years ago
Reply to  Yirmiyahu

Dekker…. no it doesn’t. Yes. it takes into account park factor but it doesn’t take into account the variation (luck) factor.

You are assuming the majority of year to year HR/FB ratios for a pitcher is based on the park and that is clearly not the case – just look at a player who has played in the same place for year and ho much the HR/FB% fluctuates year to year.

As an example Roy Halladay had a 5.1% HR/FB ratio last year, the previous year (in the same park) he had a 11.3% ratio.

Jack Nugent
14 years ago

Chris Volstad’s 2011 SIERA?– 3.84. Carlos Zambrano?– 4.46.

I’d be more confident Volstad could post that sort of ERA if I didn’t think the Cubs’ infield defense might be just marginally better this year than it was last. Volstad and his career 50+ GB% better hope Starlin Castro gets his act together at SS pretty quick…

Mike PodhorzerMember
14 years ago

My projection method involves me projecting all the underlying components and then spitting out various ERA projections using the ERA estimators. That way I can capture the guys who have shown the true skill of keeping a low BABIP, while also not using actual past ERAs to project future ERA.

Jeff K
14 years ago
Reply to  Mike Podhorzer

Is there evidence for true skill for low BABIP other than a high flyball rate? Nearly all of the top 30 qualified starters (since 2004) with the lowest BABIP are flyball pitchers.

Two of the exceptions, Price and Niemann (with only average GB rates in the 43-44% at that) benefit from the Rays strong infield defense.

Zambrano is another exception. His high IFFB/FB ratio is likely a contributing factor. So that is some evidence of true skill for him, but IFFB/FB doesn’t seem to correlate extreme stronly with BABIP.

johnnycuffMember since 2017
14 years ago

i did a correlation comparison a few weeks ago using pitcher seasons (min 100IP for starters, 40IP for relievers; N=328) from 2009 and 2010, comparing them with their results in the following year. here’s what i came up with:

zips: 0.298833053
marcel: 0.309433008
rotochamp: 0.473114922
fans: 0.486526149

the numbers themselves aren’t as important as the overall trend. obviously the fans don’t adhere to one specific methodology but perhaps rotochamp’s methods could be enlightening?

johnnycuffMember since 2017
14 years ago
Reply to  johnnycuff

(it should be obvious, but i’ll add just to be clear that this is correlation of ERA)

Yirmiyahu
14 years ago
Reply to  Matt Swartz

Wait. If its correlation (rather than RMSE), wouldn’t changes in run environment not matter?

johnnycuffMember since 2017
14 years ago
Reply to  Matt Swartz

perusing the rotochamp site yields these:

All of our predictions use mathematical algorithms that look at key player performance indicators based on historical performace. We generally look at these indicators over a 3-year period with the most recent history weighed the most.

on pitchers specifically:

We discount the traditional metrics like ERA when predicting 2011 performance and use the more reliable metrics like FIP and xFIP to generate pitching projections.

not a ton of information there. my guess is that their projections were better geared to the lower run environment because they don’t take the previous year’s ERA into account in the first place.

i went back and looked at the data. turns out the rotochamp projections i have only coincide with 2011 since fangraphs only began carrying them last year. i split the ZIPS data out into 2010 and 2011 results and got the following correlations:

ZIPS (2010): 0.315901648
ZIPS (2011): 0.270765573
ROTOCHAMP (2011): 0.473114922

and i did the RMSE:

ZIPS (2010): 1.231978994
ZIPS (2011): 1.195200293
ROTOCHAMP (2011): 0.894394372

looks like ZIPS was hurt by the lower run environment in 2011 in the correlation, but it was still vastly outperformed by rotochamp in both seasons.

caveat: each set had >200 pitchers but not the identical number since ZIPS forecasts more players than rotochamp does. this could throw the numbers off a bit.

i’d be happy to send you my data if you’d like or you can just get it from fangraphs, since that’s where it all came from.

Jordan
14 years ago

Does a 3 year weighted average of SIERA’s do a better job than last year’s SIERA?

Paul
14 years ago

Excellent, clearly designed study that importantly used valid methodology (i.e., some desired result did not impact study design).

The second-to-last paragraph is the best thing I’ve ever read on this site.

Really fantastic work, Matt.

studes
14 years ago

Matt, I need to read this a bit more, but I don’t think your assessment of projection systems is correct. For instance, Oliver does take the “interplay between different pitching skills and their effect on runs prevention” into account (I believe), though perhaps not the way you would. For instance, take a read of Brian Cartwright’s article in this year’s THT Annual. The results may not yet bear it out, but the work is happening.

studes
14 years ago
Reply to  Matt Swartz

Thanks. I agree about the linear weights approach.

Joe Peta
14 years ago

Matt,

As always great to see your work, regardless of site. Consistent with everything you’ve ever posted (readers new to Matt’s work should peruse BP archives for his initial work on SIERA, home field advantage, Cole Hamels, and even his electrifying performance on BP Idol) it’s great to know you’ll follow-up frequently in the comments section.

Given that a few questions:

1) I’ve heard Nate Silver (a bit testily, I think) take pains to point out PECOTA is an algorithm. (Therefore, not a formula?). From his perspective, how is that different from xFIP, SIERA, etc and is that a possible defense against your findings?

2) SIERA is scales to average ERA, correct? Will the weightings change this year as a result? Is it a moving average scale or simply last year’s ERA, etc? When will you publish new factors and if so do you ever go back and change old weightings if say, 2011 was a lower run environment than expected or does, for example, Ryan Dempster’s 2009 SIERA stay fixed?

3). Why not scale SIERA to RA? Or, if a team really wanted to know how many runs allowed a pitcher might give up in the future, should it just take all pitcher’s ERA and multiply it by the league average multiple (I think it’s about 1.09%)?

Jesse
14 years ago

Did you weight the innings counts at all? What happens if you don’t include starters/relievers?

AustinRHL
14 years ago
Reply to  Jesse

I would like to know this, too. It would also be nice to see actual data on using a three-year weighted mean SIERA to project future ERA.

The results of this study were quite shocking to me, but I’ll be convinced (and very impressed) once this particular methodological aspect is explored.

Jesse
14 years ago
Reply to  Matt Swartz

This is very interesting analysis. I’m a big fan of siera and think that as you say its really important to figure out ways of using siera type thinking to modify the data that goes into projections. I still very much believe in the core of the Nate silver Pecota model (kind of hard to get a handle on what it’s become) for projections.

I think there could be a plausible explanation of the greater problems projections had with relievers. I’m guessing that a lot of the variance in this sample (and the efficiency of ERA estimators) is the tendency of regression to mean. When you examine a “season” as a data point, relievers suffer from way more sample size issues and as such their unweighted tendency of regression to mean outweighs that of starters. The difference between the projections and the estimators is the introduction of comparables to assessment of true talent level to find some sort of progression over a career. That should be more accurate for starters than relievers, particularly if your assessment of true talent is not as robust as the new era estimators. \

oooh crap just scrolled down and saw the bztips discussion. Well… glad to see it.

Newcomer
14 years ago

I’d like to see a further study of this taking a look at pitchers in different aging buckets. As MGL keeps arguing at The Book Blog, pitcher aging is considerably different than hitter aging. By looking at different age groups, we might find certain parts of the aging curve where the aging adjustment of projection systems is counterproductive.

As one possibility, there is some evidence to suggest that pitchers typically begin their decline in the early 20s, at least in performance measured by rate stats (workloads are usually increasing). Perhaps projection systems are applying a positive aging adjustment to young pitchers, expecting improvement, while the pitchers are on the aggregate declining. That could explain why an accurate estimator that applies no aging adjustment might be a stronger predictor. SIERA could be a strong enough measure of ability where the improper aging hurts the projections more than the larger sample of data helps them.

Excellent analysis. I hope someone can take a deeper look.

bztips
14 years ago

Jesse hinted at an issue — starters vs. relievers — that has not received enough attention. I’ve computed my own SIERA regressions using Retrosheet data from 2003-2010, and found consistently that you get substantially different coefficient estimates (and significance levels) if you estimate separate regressions for starters and relievers. Note that this is very different than simply tacking on a reliever adjustment value after the fact (which is how the current form of SIERA works).

I’m a big fan of Matt’s overall approach, but I honestly think he hasn’t done enough statistical testing of the actual regression results (which is not the same thing as testing how well SIERA predicts next year’s ERA); this has led him to overemphasize and over-interpret the importance and meaning of the squared and interaction terms in his equation.

bztips
14 years ago
Reply to  Matt Swartz

Thanks Matt for the prompt reply.

The confidence intervals probably do overlap.

But even without splitting the sample between starters/relievers, most of the time I found that the interaction terms were completely insignificant, and many times the squared terms were as well. I never could duplicate your results, although I came “relatively” close.

Part of this is undoubtedly due to the difference between Retrosheet and BIS; I guess I really should download the data from FG and re-do the whole thing, but I’m not really looking forward to that!

A few really basic questions:
Which years did you actually include in your estimation?
Did you weight your observations (by IP I assume?)
Exclude observations with <40 IP?
Does BB include or exclude intentional walks? HBP?
Do flyballs in your netGB number include both outfield FB and popups?
I assume you estimated with just a single constant term? (as opposed to a year dummy for each year?
Dependent variable was same-year ERA — ballpark-adjusted or no?

Thanks again. YDM.

BoSoxFan
14 years ago

In the article where Colin Wyers criticized SIERA (which by the way I thought was a really lousily written article, people don’t use career SIERA, and it was shown to be better for 1 season which most people use ERA estimators for) it showed that plain ERA was better over a larger sample, so maybe, using the reliability or something, we could find the percent that we usually regress, and use SIERA for that instead. That would probably give the best estimate.

Also minorleaguecentral.com carries minor league SIERA for anyone who wants to know.

Joe Peta
14 years ago

Matt,

Do you have a feel, or have you done a study to determine at what point mid-season SIERA is more predictive of rest-of-season ERA than say, last year’s SIERA? For instance at Memorial Day this year how much weight on current year SIERA would you assign vs. last year’s SIERA as a predictor of rest-of-year SIERA?

I remember last May, SIERA favorite Matt Garza did not maintain his newly found level of strikeout rates but on the other hand, Charlie Morton really had transformed himself into a ground ball pitcher.

Thanks for your thoughts.

chuckb
14 years ago

Where did you get the idea that a year ago, Jaime Garcia was on the trading block? That absolutely is not true. I’m not sure it’s particularly relevant to the article but Garcia was never on the trading block. I’m not sure Latos was a year ago either — perhaps I’m wrong there — but I’m certain the Cards did not consider trading Garcia last year.

Zach K
14 years ago

Matt,

Did you use 2012 SIERA to project 2006-2011 ERA? And is SIERA based on a regression including data from 2006-2011? If yes to both, then we would naturally expect SIERA to project 2006-2011 better than other systems, since it was built from the data contained therein.

I would suggest that the study is only reliable if we use SIERA as it would have been calculated in 2005 to project 2006-2011.

Best,
Zach

Zach K
14 years ago
Reply to  Matt Swartz

Ah I see. That’s very interesting, since it implies there is no “True SIERA.” Thanks for your response.

Joe Peta
14 years ago
Reply to  Matt Swartz

Based on the quickly stabilizing variables, this means that SIERA stabilizes fairly quickly as well, right? How quickly do you think you can rely on a current season’s SIERA to be more predictive than the prior year’s SIERA?

bztips
14 years ago

In case Matt (or anyone else) is still interested, I’ve taken Matt’s specs that he listed above and tried to duplicate his SIERA equation as close as possible, using all the same variables and for the same time period 2002-2010. I couldn’t duplicate his results exactly, because:
a) I used Retrosheet data and he used BIS data
b) I did my own park-adjusted ERA calculations

But I came fairly close to his published results. Then I split the sample and ran separate regressions for:
–starters only
–relievers only
–American League only
–National League only

Here’s what I found:
The coefficients on K’s, BB’s and netGB’s are pretty stable and (usually) statistically significant across the entire sample and all sub-samples; this is good.

The squared term on strikeouts is very unstable and highly insignificant in 3 of the 4 split samples (all except NL); this is contrary to what Matt suggested would happen.

The squared term on walks is unstable and highly insignificant, even in the initial unsplit sample.

The squared term on netGB’s is stable and significant – nice!

Among the interaction terms, only the K/netGB interaction is fairly stable, though not very significant statistically. The other two interactions are very unstable and very insignificant.

I got adj-R^2 of between .36 and .40 for all runs except the relievers sample, which came in at under .28; clearly there are some real differences in how well a SIERA-type regression tells us what’s going on with relievers as compared to starters. This was my concern from the beginning.

As a quick follow-on, I re-ran all 5 regressions excluding those terms that were consistently unstable/insignificant. This left me with the following set of 5 variables: Ks, BBs, netGBs, netGB^2, and K/netGB interaction. This yielded very similar results to the initial set of results that included all those other terms; in fact, I got slightly HIGHER adj-R^2 in most cases, meaning that the excluded variables really don’t belong in the equations.

So I’m left to conclude that Matt really may be over-interpreting the meaning and importance of all the non-linear terms in his equation. The overall approach is still a great way to try to go beyond the oft-repeated claim that “strikeouts and walks are the only things that we can measure a pitcher by”; but a somewhat simplified version would probably do just as well in predicting out-of-sample ERAs.

bztips
14 years ago

Matt, yes I understand the marginal impacts implied by the functional form of your equation. I suppose it’s a matter of perspective — while you may be comfortable “knowing” that there are complicated and indirect effects between Ks, BBs and GBs regardless of how well they show up in your equation, my perspective is that the data may not support what you think you know.

I doubt that the differences I see when I break the dataset into starters and relievers are due mostly to just sampling error — I have 1745 observations for starters and 1494 for relievers. Obviously there is some impact from the fact that the avg number of innings for starters in my sample is around 138, while the avg of relievers is 61. But still, the idea that all of the insignificant terms I see in the reliever equation are just due to sample size issues doesn’t make sense. Here are the t-statistics I got for the reliever equation:

SO/PA -3.41
(SO/PA)^2 -0.17
BB/PA 0.97
(BB/PA)^2 0.43
nGB/PA -0.92
(nGB/PA)^2 -2.23
(SO*BB)/PA 0.99
(SO*nGB)/PA 0.79
(BB*nGB)/PA 0.37

And again, the R^2 was around 0.28. What this says to me is that there’s something else beyond linear or nonlinear Ks, BBs and nGBs, and their interactions with each other, that explain reliever performance; as opposed to your suggestion that we KNOW these things are tied to performance and it’s just small sample sizes or bad Retrosheet data that’s preventing us from seeing it.

As to your specific questions:
The BB^2 term in the starter equation was 41.48 with a t-stat of 1.68; for relievers it was 10.24 with a t-stat of 0.43.
The SO^2 term in the starter equation was 9.11 with a t-stat of 1.26; for relievers it was -1.00 with a t-stat of -0.17.
So for starters I’m not too concerned about the significance levels, because the point estimates are similar enough to what you found for the entire sample. But not so for the relievers — there’s got to be something else going on that is just not being captured well by the variables in your equation IMO.

Thanks again for continuing this discussion; the more we can learn from the data, the better off we’ll all be.

bztips
14 years ago

OK, I know we’re beating a dead horse, but here’s (hopefully) my last shot at it.

You can’t have it both ways:
1) claiming that coefficients matter but the large standard errors don’t when it comes to deciding which explanatory variables belong, and
2) using these same standard errors to argue that the coefficients from the reliever sub-sample aren’t really all that different from the coefficients from the full sample.

If in fact there’s all this noise due to their smaller number of innings pitched, then doesn’t it make more sense to just exclude relievers from the analysis altogether and apply your conclusions just to starters?? Instead, you’re making an assertion (and that’s all it is) that even though we can’t see it in the data, we should just ASSUME that relievers’ performance can be explained by the same model that explains starters’ performance.

slash12Member since 2021
14 years ago

very nice article, I’ve been wondering how these things stacked up for a while, but haven’t taken the time to do the work, thanks!

Larry
13 years ago

You wrote “When predicting the future, using several years of SIERA will do better than one year, and adding park effects and aging to previous years’ SIERAs will also help”

How do you add park effects?

Matt Swartz
13 years ago
Reply to  Larry

Just multiply by them. So, like if a pitcher’s home park has a 1.20 park factor, then him pitching half his games on the road would give him a 1.10 park factor for his team. So say this pitcher’s SIERA is 4.00, then his expected ERA since he’s pitching half his games in a high run environment would be 4.40.

lschechter
13 years ago

Matt,

I’d like to have a more private conversation about some of this stuff, and I have a few general questions about fangraphs, too. Can I e-mail you?