Adam Wainwright, Run Clusterer
On Monday night, I was watching the Cardinals battle the Royals when I heard something that stopped me in my tracks. As Adam Wainwright labored in the sixth inning — two runs in and runners on the corners with two outs — the Cardinals announcers mentioned one of Wainwright’s greatest strengths — in their minds, at least. “That’s something that Adam Wainwright is really good at, is not compounding the inning… going back and getting the next guy.” I’ve been a Cardinals fan my whole life — and to that tidbit, I said, “Huh?”
It was, in truth, something I’d never thought about. Are some pitchers better than others at turning off the tap, amping up their performance when they need it and keeping crooked numbers from getting even crooked-er? My saber sense was tingling — something about this didn’t sound quite right. But of course, these spots are exactly where if a pitcher could bear down more than expected, it would make the most difference. I decided I’d try to find out how real this effect was.
Defining what I was looking for turned out to be a difficult. What, exactly, does “not compounding the inning” mean? The announcers seemed to think it meant that Wainwright pitched better after runs were in, or at least pitched the same while most pitchers in baseball got worse. Either way, the general idea was that his ones and twos turned into threes and fours less often than average.
One possible reaction to that might be “So?” His ERA is his ERA, regardless of whether it comes via a three-run spurt and eight zeros over nine innings, or three one-run frames and six zeros. To that I say: reasonable point. There are still reasons to care, though. For one, if a pitcher were actually prone to clustering, they’d tend to underperform their FIP over time. One of the reasons home runs are so bad is because they always result in runs, whereas other hits can be scattered around in otherwise dry innings without damage. A cluster-prone pitcher wouldn’t have that advantage; when you give up baserunners in bunches, a single and a home run become much closer in value.
In the same way, a pitcher who was prone to lots of singleton runs allowed but then mysteriously got better after letting one in would beat his FIP over a long time horizon. Base/out states tend to be more dangerous after a run has scored, naturally enough. Getting better then, or not getting worse while most pitchers do, would be quite the superpower.
Of course, that’s not necessarily the right way to think about it. There’s a simpler way to take this statement. Maybe Wainwright simply has fewer crooked numbers, as a proportion of the runs he gives up, than the average pitcher does. There doesn’t need to be a provable reason why that should happen, or even a real advantage to displaying that behavior. Maybe Wainwright simply allows runs differently.
I wasn’t exactly sure how to approach this exact problem, so I decided to define terms very narrowly and answer some problem, rather than spending forever thinking of how to do it. Is it the right problem? You tell me. Here’s what I did, though: I looked at each inning that a pitcher both started and finished and grouped them by how many runs were allowed.
Why exclude innings where they were pulled? Because we can’t know what would have happened. What relievers do with the scraps they have to pick up doesn’t necessarily mean much about the pitcher who left the mound. We could assign some run value based on the base/out state when the pitcher left — but that wouldn’t do what we want. We’re looking for a place where pitchers behave differently than a naive run expectation.
Here, for example, is Wainwright’s runs allowed distribution across every inning he has both started and finished in his career:

That’s a broad picture, but it gives you a general sense of the shape of the innings he allows. About 60% of the time that he allows at least one run, it’s only one. Put another way, if you know only that Wainwright allowed at least one run, it will be exactly one 60% of the time. If you know that he allowed at least two runs, it will be exactly two runs 65% of the time. If you know that he allowed at least three runs, it will be exactly three 68% of the time.
Working out what “average” is in this statistic is tricky. In 2020, for example, pitchers as a whole check in at a 60.1% rate of one-run innings out of all the innings in which they allowed a run, almost exactly identical to Wainwright for his career. But this statistic isn’t talent-level agnostic; the better the pitcher, the higher their proportion of one-run innings should be. Jacob deGrom checks in at 63%, Clayton Kershaw at 68%. Among pitchers who have completed 30 innings this year, the slope comes out thusly: for every point of ERA below average, pitchers have a four percentage point higher rate of holding their opponents to just one run.
Some era adjustment is necessary, because the offensive environment has changed, and that itself could change the rate of one spots as compared to other run-scoring innings. I looked at every complete inning since 2005 and found a rate of … 60.1%. Okay, so maybe we don’t need to adjust for the run-scoring environment.
Over his career, Wainwright has an ERA 0.8 runs better than league average. We’d naively expect him to have a distribution of 63.4% one-run innings out of all of his run-scoring innings. By this metric, he actually allows more big innings, relative to his overall skill level, than the average pitcher!
There are plenty of problems with this way of looking at things. An inning with a single and a homer, for example, counts as a “big inning” even though the pitcher never had a chance to bear down after allowing the first run. Maybe the better question is what wOBA a pitcher allows after a run has scored, or their strikeout rate, or something along those lines. But it’s my article, and I like this formulation for its simplicity, so we’re sticking with it.
There’s another question worth answering here: Wainwright doesn’t seem to have this skill, but does anyone? Is it actually a skill, or something that happens randomly to pitchers? To test this, I looked at every pitcher who started and finished at least 100 innings in 2016 and 2017 combined. I assigned each of them a projected one-run percentage (there has to be a better name for this, I just can’t think of one) based on their ERA.
Let’s use Chris Sale as an example. He threw a whopping 441 innings over those two years. When he did allow a run, he allowed a lot; he allowed one run 51 times, two runs 20 times, and three or more 16 times. That works out to a 58.6% one-run percentage. He had a 3.12 ERA over those two years, though, far better than average, which means we’d predict him to have a one-run percentage of 67.4% (with huge error bars, to be fair). Thus, we assign him a “one-run error” of -8.8%, the difference between our goofy prediction and his actual rate. I did the same thing for every pitcher.
Next, I did the same for 2018-2019. Let’s look at Sale again. In 2018 and 2019, he had a 3.21 ERA (consistent!) and a 54.1% one-run percentage. We would have predicted him to have a 64.4% one-run percentage, which means he was again an outlier. Sale seems to have fewer one-run innings, as a percentage of his scoring innings, than the average pitcher of his caliber, not that there are many pitchers of Sale’s caliber.
From here, I took every pitcher with 100 innings in each of my two time periods and divided them based on how much they beat or missed my prediction in 2016-2017. Those groups look like so:
| Quartile | 16-17 Prediction | 16-17 Actual | 16-17 Error | 16-17 ERA |
|---|---|---|---|---|
| 1 | 62.7% | 53.4% | -9.4% | 3.97 |
| 2 | 61.2% | 58.1% | -3.1% | 4.25 |
| 3 | 61.4% | 62.0% | 0.6% | 4.21 |
| 4 | 62.6% | 70.2% | 7.6% | 4.00 |
Sale would be in that first group, the one with a much lower actual one-run rate than you’d expect from their ERA’s. It’s not clear whether there’s any sample bias — the pitchers with the least and greatest errors are the two groups with the lowest ERAs, and good luck explaining that. Let’s see how each group did in 2018-2019:
| Quartile | 18-19 Pred | 18-19 Actual | 18-19 Error | 18-19 ERA |
|---|---|---|---|---|
| 1 | 61.6% | 61.6% | 0.0% | 3.82 |
| 2 | 59.0% | 58.6% | -0.3% | 4.39 |
| 3 | 59.0% | 59.0% | 0.0% | 4.39 |
| 4 | 60.5% | 61.5% | 1.0% | 4.06 |
Well, so much for that. Pitchers who are at one extreme or the other for two years (quadrants one and four) had almost exactly equal one-run rates in the next two years. The middle half was again the worst group by ERA, which makes sense — it’s the same pitchers who were the worst in the first two years. But they, too, had aggregate one-run percentages almost exactly on top of a naive prediction.
In the end, I’m not sure what to say about this analysis. Maybe the way I framed it was wrong. Maybe it’s a real skill that only shows up intermittently, or there are so few players who actually possess it that grouping into quartiles obscures their skill. Maybe — and I’d say this is most likely — it’s nearly impossible to perceive the difference between 60% and 63% of your run-having innings be one-run jobs. If that’s the case, it wouldn’t be surprising that your brain links a positive trait — keeping the inning manageable — with a pitcher awash in other positive traits, and Wainwright certainly fits that bill.
Great players are — well, they’re great! Adam Wainwright qualifies as one of those over the course of his career. Be careful about haphazardly ascribing traits that go beyond what you can see in the numbers, though. That’s a great way to end up believing something that simply doesn’t appear to be the case. Being great at measurable things is impressive enough without the need to ascribe bonus intangible excellence.
Ben is a writer at FanGraphs. He can be found on Bluesky @benclemens.
I haven’t read the full thing yet since I’m leaving for work but check out Yimi Garcia’s 2019 stat lines. He had the lowest BAA while also allowing basically every hit to be a home run. He’s the kind of player that you expected to surrender a single run and that’s it. Maybe that profile in a starter than reliever is what you’re looking for.
Baseball announcers are always spouting non-sense. They have to talk continuously for so long it’s inevitable.
Pitchers talking about hitting and hitters talking about pitching.
A slightly related point is that the interpretation of most baseball metrics is under the implicit assumption that P(x|S) = P(x) where S is situation specific knowledge, x is an outcome of interest, and P is a probability. There is no reason for one to expect that P(x|S) = P(x) actually holds other than it makes life very simple.
I think that this really shortchanges sabermetric analysis as a whole. The Book, for example, spent a great deal of time doing the exact opposite of what you’re talking about. “Old school” baseball thinking was rife with assumptions that P(x|S) != P(x); clutch, ERA, pinch hitting specialists, platoon this-and-that, and so on. The Book looked through those examples and found compelling evidence that P(x|S) = P(x).
FIP, for example, isn’t a statistic because we say it is with no thought about whether it’s actually true. It’s incontrovertibly true that you’ll do a better job of guessing next year’s ERA (which is essentially a summation of outcomes given different situation specific knowledge) if you simply use FIP (which gets rid of S entirely by using regression weighting for K, BB, and HR).
I’ll agree with you that public analysis these days probably errs on the side of removing context too often — see all of Russell Carleton’s work for an example of someone who is really good at putting context back in where it matters. It’s generally easier, now that we have building blocks like FIP and wOBA, to ignore context than to dig in and use it. But I’d hardly argue that there’s *no* reason to expect that P(x|s) = P(x). There are plenty of reasons, even if people are too quick to generalize “there is no clutch” into “context never matters.”
Thank you for the literature review on this subject, I appreciate some historical perspective and the additional references.
I do not think that FIP rids the conditioning on S entirely. Pitchers may have different FIP in different situations (later in the game, change teams from year to year, injuries to competing lineups etc). I think the prediction problem is more complicated than posited here. It is absolutely true that you can improve upon guessing ERA by using FIP (or a general regression model with each component of FIP estimated separately). However, this is not removing S from the equation, it is just saying that the interactions between S and other terms of a predictive model for ERA exhibit less variation when FIP is used. For an oversimplified mathematical presentation of ideas, let’s posit two general additive error predictive models for ERA
model 1: ERA_new = f(FIP, S_new) + error
model 2: ERA_new = g(ERA, S_new) + error
where S_new is new situational information for the season under study, ERA_new is the ERA that you want to predict and both ERA and FIP are the previous seasonal values. Technically speaking ERA = ERA(S_old) and FIP = FIP(S_old) where S_old is the situational information of the previous season. We’ll assume that the error distribution is centered and exhibits roughly the same variation for the two models (perhaps this is a reach).
We agree that the signal strength for the measurable terms is higher for model 1 than model 2 under this array of assumptions, but this does not mean that S is removed from the equation entirely. I wonder how this prediction setup would do when S_old = S_new perfectly! This is a theoretical abstraction that can never be fulfilled, but I would guess that FIP and ERA would yield the same predictive power in this setting (up to error).
As a general statement, it should not be expected that P(x) = P(x|S). However, I can see that assuming so may not produce inferential problems in some settings.
I really don’t suspect that it would change your conclusion at all (and it’s also probably a more difficult analysis), but I think the approach you mentioned of dividing the inning into before and after allowing a run and comparing the pitcher’s performance in those two parts might be a better way to capture what the broadcasters seemed to be suggesting – that, once an inning starts going downhill, Wainwright is better than average at halting that “momentum” and not allowing (m)any more runs. So allowing either a two-run double or a solo homer would theoretically kick this ability into action in the same way. Again I am very skeptical that this is a real skill, but it would be cool if it was!
I don’t see why it couldn’t be real – many SP reserve to go deeper into games, from Verlander throwing harder in the 9th than the 1st to any number of pitchers not using their 3rd pitch until the second time through the order – though it does seem to be an exception rather than a rule. That said, aside from the solo HR, I would think that if a pitcher doesn’t turn it up a notch as soon as they get in trouble to prevent that first run from scoring at all, it doesn’t make sense. And by that theory, they’re already pulling out all the stops by the time the first run scores, so the comment kind of falls apart.
Also, we have this data in a sense with things like LOB% and pitcher performance with runners on base. I think we already measure this, just a bit more rationally.
I fear that the vast majority of what you are looking at/analyzing is how long the pitcher’s leash is. This is more a reflection on the coach decisions than pitcher behavior, since allowing 4 runs and getting pulled does not enter the data set but allowing 4 runs and finishing, does.
So… the announcer’s comment came after 2 runs had scored and there were runners at the corners?
Then… you define a “big inning” as one where more than one run is allowed?
That’s an abuse of logic.
This is something I wonder about all the time regarding hitting. Is it more useful to have a guy who’s consistently 1/3 every night or a guy who goes hitless one week and looks like an MVP the next?
Rather than just looking at runs per inning, I think you’d want to get more granular and look at how many hits/walks are allowed per inning and go from there. Are pitchers who scatter hits better than ones that mow everybody down until they get blown up ?
I don’t really know enough about the study of statistics to formulate the question better, but I’m sure this is a question of the “clumpiness” of the distribution somehow.