Corbin Carroll Is Even Better Than Advertised

Not every out is created equal. Take this fly out from Corbin Carroll, for example:
A lot of things can happen when you make an out with the bases loaded. You could strike out, leaving every runner in place. You could hit into a double play, an inning-ending one in this case. You could ground out some other way, or hit an infield fly. But Carroll’s here was the most valuable imaginable; with one out, he advanced every single runner, including the runner who scored from third.
Mathematically speaking, you can think of it this way. The average out that took place with the bases loaded and one out lowered the team’s run expectancy by a massive 0.61 runs in 2024. That’s because tons of these outs were either strikeouts (bad, runner on third doesn’t score) or double plays (bad, inning ends). But Carroll’s fly out was far better than that. It actually increased the run expectancy by a hair; driving the lead runner home and moving the trail runners up a base is exquisitely valuable.
That’s not the only way this could have gone. Consider a similar situation, a groundout from Aaron Judge:
Like Carroll, Judge batted with a runner on third and fewer than two outs. In this situation, the average out is bad, lowering run expectancy by 0.514 runs. But Judge’s was obviously worse. It cost the Yankees all the expected runs they had left in the inning, naturally, which added up to just a bit more than 1.15.
One thing you could say is that both players made outs on this play, and that since players don’t control when and how their outs come, we should just give them all the same value. That’s a reasonable stance to take. FanGraphs’ marquee offensive stat, wRC+, treats all outs as identical. It’s not just us, though. Batting average, on-base percentage, slugging percentage, wOBA, you name it; “an out is an out is an out” is a core part of the way it works. Notwithstanding minor exceptions for “intentional outs” like sacrifices, pretty much everyone treats outs as interchangeable.
I wanted to dig a little deeper than that, though, so I started making some assumptions and doing some math. First, I looked at how costly an out was in every base/out state in 2024. Then, I broke outs into four categories: balls caught in the air, grounders that didn’t result in double plays, double play grounders, and strikeouts. For each base/out state, I calculated how much better or worse each out type was compared to the “average” out in that base/out state.
Let’s take another example, because I’m talking about a lot of numbers here and it’s easy to get lost. With the bases loaded and nobody out, the average out lowered run scoring expectancy by 0.565 runs in 2024. The average strikeout in that situation lowered run scoring expectancy by even more, 0.765 runs. This makes good sense to me; the “average” out scores a run fairly often, and a strikeout doesn’t. On the other hand, a fly out with the bases loaded lowers run expectation by only 0.494 runs. It’s slightly better than the average out – there are sacrifice flies to think of.
Continuing on, a non-double-play grounder is a great outcome, lowering run expectancy by a mere 0.164 runs. Since we know only one out was made, a groundball is amazing. Maybe the runner scored, maybe there was an error, maybe a couple runners advanced to boot. On the other hand, a double play is a disaster, lowering run expectancy by 1.091 runs. Sometimes that’s the runner at the plate plus the lead runner, sometimes it’s a classic 6-4-3 with a run scoring, but on average, hitting into a double play smarts.
You might wonder why I grouped all fly balls together but kept double plays separate from all other grounders. It’s a matter of accounting. Imagine hitting a medium-depth fly ball that the right fielder easily corrals. With a fast runner on third, that’s probably a sacrifice fly. With a slow runner, it’s probably just an out. But the batter doesn’t have control over that, and we already give baserunners credit, so giving the batter points only if a sacrifice fly is completed feels wrong. Instead, I merely assigned the average value across all fly balls, with baserunners getting their credit added or subtracted for their contributions separately.
Double plays aren’t like that. The batter exerts a ton of control over them. Take Carroll, for example. He’s no stranger to grounders, but he’s also hilariously fast. Like, turn-out-the-lights-and-be-in-bed-before-it-gets-dark fast. Carroll hit into just three double plays last year. Manny Machado, famously not fast, hit into 25. You also can’t hit into a double play if you don’t hit the ball on the ground, so fly ball artists do better than grounder-heavy types. The point is, hitting into a double play really is something you can put on the batter, so I do. Everything else, I merely point to the average outcome.
If you’re either very bored or reasonably good at computer programming, you can go through every single out that major leaguers made in 2024 and note how much better or worse they were than the average out in a given base/out situation. That’s more or less what I was doing above, but to be slightly more specific: A fly out with the bases loaded and one out is 0.224 runs better in expectation than an average out, so I credited Carroll with those extra runs. A double play grounder with one out and runners on first and third is 0.708 runs worse than an average out, so I docked Judge that amount. I did this for every single play in the 2024 season and then just hit sum.
My selection of Carroll to lead this column off wasn’t just random. His outs were a lot better than the average out made in the situations that he batted, to the tune of a whopping 8.5 runs in aggregate. Only one other player, Jackson Merrill, even topped six runs. This is a zero-sum game: For every out with better-than-average results, there’s an offset somewhere. It’s also a low-volatility game; 80% of players fell between -2 runs and 2 runs across the entire year.
Carroll’s game is made to break the mold of “average out,” though. That blazing speed means he almost never hits into double plays, which is a huge part of his score. He rarely strikes out – strikeouts are better than double plays but worse than everything else. He hits a decent number of grounders, too; they just never become double plays. That means he’s frequently advancing baserunners but almost never creating extra outs.
Merrill, the other standout of the productive out, succeeds for similar reasons. He rarely strikes out. He’s fast. He hit into only two double plays all year, albeit with fewer groundballs thanks to his batted ball mix. But it’s a similar formula: Put the ball in play but avoid the worst type of ball in play.
Those are reasonably large adjustments. Eight runs is the better part of a win; if you were to credit Carroll for his good outs, his value last season would’ve been closer to 5 WAR than the 4 WAR we marked him down for. I think there’s a good argument that he shouldn’t get all the credit: He doesn’t control who’s on base and how many outs there are when he bats, for example, and he had a ton of runners on third base when he batted with less than two outs — 36 to be exact, 20 more than Steven Kwan, to name a hitter with a similar batted ball distribution. But I feel comfortable saying that Carroll’s batted ball distribution and speed out of the box meant that his outs were less damaging than average, and in a way that none of our current hitting stats tabulate.
The other side of the coin? It’s Judge, of course. Judge cost the Yankees a whopping 8.8 runs relative to average with his deleterious outs in 2024, one of only three hitters who lost six or more runs in this accounting (Machado and strikeout king Tyler O’Neill were the others). Honestly, it wasn’t the strikeouts so much as the fact that Judge hit into a whopping 22 double plays.
Some of that was out of his control, because he led the majors in double play opportunities. Batting after Juan Soto will do that. But in 200 opportunities, he hit into 22 double plays. José Ramírez was second in the majors in double play opportunities – and he hit into only nine out of his 155 chances. Bryan Reynolds hit into only seven in 148 bites at the apple. Heck, Soto himself had 138 chances to hit into a double play and ended up with only 10 of them. Judge’s performance in double play situations was bad, and it happened frequently.
As in Carroll’s case, I’m not sure how much of the blame from those -8.8 runs truly belongs to Judge. A lot of it is situational, and those situations tend not to repeat from one year to the next. It’s also not a big deal in the grand scheme of things; if you insisted on counting every drop, Judge would have produced about 10.3 WAR last year. (This would have made for a much more competitive MVP race between him and Bobby Witt Jr., but even with this adjustment, nobody else would’ve been at their level.)
Still, these run adjustments are real. Different ways of making outs have always been treated mostly the same (batting average never gave you an extra out for a double play), but they’re slightly different in practice. You can’t consider Carroll’s offensive prowess without noting his ability to advance his team’s agenda, at least a little bit, even when he fails. Likewise, Judge’s season inarguably involved some painful failures that weren’t completely accounted for in his batting line.
Once again: Most of these come out in the wash. It’s likely that Luis Arraez’s -1.4 runs, Shohei Ohtani’s +2.6 runs, and Soto’s 2.4 runs are all just noise that will cancel out in the long run. But I like calculating things, and I particularly like hunting for hidden value. Avoiding double plays while still advancing runners is unquestionably valuable – and now, that value is at least a little bit more out in the open.
Ben is a writer at FanGraphs. He can be found on Bluesky @benclemens.
Wow, Ben. Outstanding!
Among the uncontrollables are haw pitchers react. Do they react by throwing harder and getting less sink? Or do they take a deep breath and throw their best pitch? Jim Palmer was asked what do you throw when it’s 3-2 against a good hitter and there are runners in scoring position? He said “a fastball.” It’s my best pitch.
How pitchers react
It might all come out in a wash for most players, but 10 years from now I will also not be surprised if Shohei somehow had positive run values in this metric every single year. How? I don’t know. He just will.
I think this is super-useful, thanks. One of the things that’s always bothered meis that a lot of the things batters are taught from childhood about “good outs” (ie, hitting the ball to the right side with a runner on second) aren’t really captured the right way by stats. This really helps get to that. Wither bunts? Is there any better way to capture their relative value than we currently have?
I though GIDP was in the Fangraphs BsR, it says so in the glossary entry for BsR. So players do get some credit for this.
Yes I’m curious how much more information this granular calculation adds to the baserunning value that’s already captured. I guess this is answered in the article–Judge gets about a win less, Carroll gets about a win more–but I’d love a bit more mention of how/why that isn’t adequately represented in BsR already.
That glossary entry needs to be updated – we switched to using Statcast’s xBR plus wSB for baserunning in 2024, which now excludes GIDP. We did that to keep using a modern advanced baserunning statistic, but that excludes ‘baserunning’ value added while batting, such as double plays.
That makes sense. Thanks for the explanation.
This is the stat for the old timers who grumble that some guys “have it” because they “play the right way.” Those hustling guys who don’t strike out much and are over to first base in a flash are absolutely the guys who shine on this. Although it usually doesn’t add up to that much it absolutely does for some players!
I am one of those old timers who strongly believes this is a big part of the game and the reason stats simply can’t grasp the entirety of the game. The claim that “an out is an out is an out” is a fallacy and a major one. I mentioned that I was quite pleased when Tyler O’Neill did not return to the Red Sox for 2025. It was because of these little nuances that can turn a game around. Only 61 RBI’s on 31 HR’s means something. He drove in just 20 runs on anything but a HR. It says that his outs were not productive outs and that he left lots of runners stranded. I grew up with a game in which getting a runner from 2nd to 3rd with no outs was “getting the job done” as Mel Allen or Vin Scully would be sure to mention. It is things like the different value of outs that cannot be quantified that keep me entranced by the vagaries of baseball.
It’s worth sorting things into “things that are just noise” and “things we haven’t measured properly yet.” This one was the latter. But now we have measured it properly! I did not expect it to be this big, but that’s why we do it.
Since the bulk of this value seems to be wrapped up in double plays, the takeaway I get is: “don’t hit the ball on the ground unless you are fast enough to beat out the double play, and maybe try to make a little more contact as long as it doesn’t cause you to hit a grounder”.
Oh, if it were only that easy!!
Judge’s GIDP distribution is really interesting – he hit into 11 from season start to May 9. His to-date season OPS at that point was .855 — he was starting to come out of that early season slump. There was also a stretch in August where he hit into 4 over 9 games.
Stupid question: Were all outs made with the bases empty assigned the same value? If a player reached on an error — whether after a K, grounder, or flyball — it wouldn’t register in the database as an “out,” correct?
That is correct – the Run Expectancy doesn’t care how we got to there being runners on the corners and one out.
Not quite, though this could be done a few ways. Our calculation of wOBA bundles errors in with outs (hitters don’t get any credit for reaching on error in wRC+ or wOBA as calculated here). Statcast’s version credits the error roughly the same as a single. I threw in errors here because for our statistics, they count as an out, so the ‘average’ groundout has some chance of being an error. You could calculate the stats differently, with errors included in the ‘reach base’ buckets and outs being only ones that ended in an out, and then the bases empty situations would be exactly identical.
So if I’m understanding right, the claim isn’t exactly that Carroll is a great situational hitter. The claim isn’t that he’s very good at adjusting his approach to the situation, so that if he does make an out with runners on base at least it’ll be a productive out. Right? Rather, the claim is just that he’s the sort of hitter who makes the sorts of outs that just so happen to be productive when there are runners on base? That is, he always avoids striking out (regardless of the situation), and he’s always too fast to double up (regardless of the situation), and those skills just so happen to be quite useful with runners on?
If that’s right, I would be quite interested in an analysis that went the next step, and asked whether certain hitters are good at *changing* their approach with runners on, so as to avoid unproductive outs. I’d also be interested in whether making this kind of adjustment with runners on is worth the (presumed) cost to power.
Seems the claim is that he’s a decent situational hitter who is fast enough to avoid GIDPs, which tracks pretty well with him being the most valuable baserunner in the league the last two years
I don’t think this article makes any claims about Carroll’s situational hitting (changing his approach based on the context).
It’s interesting to me that Arraez didn’t show up as a Carroll-like outlier in this analysis. I take it that merely avoiding Ks isn’t sufficient on its own, you also need to be fast enough to avoid GIDP? Or was it more just comparative lack of opportunities with runners on?
My guess too. He beats the ball into the ground a lot and is slow af
I am all the way here for the RE24 Agenda.
Long may it reign – I’d love a write up using it as a lens for reliever value!
I’ve been advocating for this ever since I heard of this stat. My dad and I used to have conversations about how every single legacy pitching metric was a terrible way to rate relievers: Wins, Saves, ERA (better than the first two, but still not very good for this), etc. We thought there should be some kind of points system that rewards pitchers for coming into a game in a jam, and getting out of it, and on the flip side, being docked for creating that jam. Then I was introduced to Fangraphs and saw RE24, and it was like a halo of light enveloped me. This is exactly what we were looking for! But nobody uses it. Yet.
I see someone has already mentioned Judge’s DP distribution, but it’s worth noting that 10 of his 22 GIDPs for the year came before the end of April, which was also a period when he really wasn’t an exceptional hitter.
Once he was locked in and lifting the ball, he became something akin to a right-handed Babe Ruth and also someone who didn’t ground into too many double plays.
Judge is so overrated!
This is fantastic research, and while I feel like there could be a lot of followup, this is likely to be the best article I read all 2025 (unless there are more articles about good and bad bunts).
I have often wondered that when a team outscores/underscores its expected runs if out quality not being properly valued is a strong reason.
Define OOPSY yesterday please.
This is a great article. I’d love to see the stats for each team as well and a table of the Top 100 or so players. I think this is on to something important.
I’ve been fascinated by – but lack the skills or time to do any research and analysis myself – the idea of productive outs for 20 years. I wonder if teams that overachieve their expected win totals or talent (say 1969 Mets, 2012 Orioles, etc.) have good results for productive outs.
Ben has pulled back the veil a little and I’m excited to see what else he has in store in this area, which I think is more important than it seems to be on the surface – or, maybe it isn’t. It will be just nice to see what happens with a closer look and a wider time span.
Since you were wondering if those numbers were just “noise” and may possibly not be consistent from year to year, is it possible for you to re-crunch the numbers going back maybe 5 years to see how the data fluctuates for each player? Of course, we only have two years of data for Carroll, but we do have more then 5 years for Judge, Ohtani, Machado, etc.