An Early, Nerdy Look At The Challenge System

Troy Taormina-Imagn Images

In the new season’s early going, the challenge system has been all the rage across the majors. If you don’t believe me, you can read ESPN’s coverage of it, or The Athletic’s, or MLB.com’s, or … well, you get the idea. The coverage has been extensive and positive, and I couldn’t agree with its enthusiasm more. I love the new system, and I’m also really excited to think about challenges in general. There are so many fun angles to consider. So here’s the math nerd’s take on what challenges have looked like so far, and what I’m most interested to learn about them moving forward.

How I’m Thinking About Challenges
Every time a strike or ball is called, there’s an opportunity for a challenge, at least so long as the relevant team has one remaining. That makes it easy to measure the prospective value of a challenge on any given pitch: It’s worth however much flipping the result of that particular pitch would change the game situation in the challenging player’s favor. All we have to do is figure out how many runs were likely to score in the inning in each case and compare the two.

That sounds hard, but it actually isn’t so bad. All you have to do is construct an RE288 matrix, which measures how many runs have scored, on average, after the game reaches a given combination of outs, baserunners, balls, and strikes. For example, over the past 10 years of major league play, teams have scored 1.02 runs per inning after a batter reaches a 1-0 count with a runner on third and one out, but using our matrix, we can work out all of the possible run expectancies a batter could reach in that plate appearance:

Run Expectancy, Runner on Third, One Out
Strikes
Balls 0 1 2
0 0.98 0.90 0.81
1 1.02 0.94 0.81
2 1.09 0.99 0.88
3 1.19 1.10 0.96
MLB, 2016-present

Now let’s imagine a catcher weighing whether to challenge a called ball. To determine the value of successful challenge, we can calculate the change in run expectancy of a ball versus a strike. Of course, the situation matters. We’re most interested in those counts where the outcome would result in a walk or a strikeout, as it’s the difference between first and third with one out, or a man on third with two outs. Let’s take a look:

Run Value of a Successful Challenge, Runner on Third, One Out
Strikes
Balls 0 1 2
0 0.12 0.13 0.43
1 0.15 0.18 0.50
2 0.20 0.22 0.58
3 0.12 0.26 0.84

Those are listed in run values, and they’re are big numbers. Flipping a 3-2 pitch from a walk to a strikeout is worth a whopping 0.84 runs; it’s the difference between a jam and a comfortable inning. Flipping the first pitch of an at-bat is much less impactful. And that’s in an important spot, with a runner on third and fewer than two outs. Next, imagine that our batter hits a sacrifice fly, leaving the bases empty with two outs. The challenge values for each pitch in the next batter’s time at the plate are far lower:

Run Value of a Successful Challenge, None On, Two Out
Strikes
Balls 0 1 2
0 0.03 0.04 0.08
1 0.04 0.04 0.09
2 0.06 0.07 0.13
3 0.07 0.10 0.23

Flipping a walk to a strikeout still matters, but almost everything else is low value. Even if you don’t do the math, you know this intuitively. If an umpire misses a call deep in the count with a runner on third and one out, it stings. It feels like it could be a key turning point in the game. If a bases-empty, two-out count gets to 1-0 when it could have been 0-1, it hardly matters.

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

Runs, not flipped calls, are the end goal of the challenge system. You could win 20 challenges in an 0-0 count with two outs and the bases empty and still help your team less than winning a single challenge in a 3-2 count with a runner on third and one out. I’m quite confident that this is the right way to think about it. Tom Tango’s extensive overview of the challenge system uses the same methodology. It’s what I came up with prior to consulting other folks at the site, and it’s also what they came up with before they talked to me. That’s a good sign that runs are the correct currency for challenge value.

Why not use win probability and leverage index? Because you don’t get to take your challenges home with you. If you’re trailing 7-1 in the eighth, leverage index would tell you there’s basically no play that’s really worth challenging. The win probability gains aren’t that high when you’re almost sure to lose anyway. That isn’t right, though. All you can do is use the challenges you have to add as many runs of value per game as possible. A run is a run is a run. That doesn’t change depending on whether you’re behind or ahead, or what inning it is. If you’re trying to measure how much challenges have impacted a team in retrospect, win probability is a reasonable measure. If you’re trying to determine how players should behave, run value is the way to go.

Catchers Are Better At Challenging Than Hitters or Pitchers
With that run value framework in mind, we can take a look at all the challenges that have occurred so far (well, through March 30) and note some broad patterns. Through Monday’s action, teams have issued 227 challenges, 124 of which have been successful, good for a 54% conversion rate. The defense has challenged more frequently, and it’s mostly been catchers; 118 of 124 defensive challenges have come from behind the plate. Hitters have challenged only 103 times.

Defenders have been successful on 57.6% of their challenges, while hitters have been successful on 50.5% of theirs. This tracks with the spring training and 2025 minor league data. It’s just easier for catchers to judge the strike zone. They get to look at the ball directly as it comes in, rather than catching a side-on glance at it, and they’re inches away from the plate instead of 60 feet.

It’s worth noting that 57.6% is not 100%. Catchers have a better idea where the ball crossed the plate, but they clearly don’t have a perfect picture. I think that major leaguers could have already told you this. They’re never totally sure whether or not a pitch was a strike until they go back to the dugout and take a look at their iPad. Why don’t players challenge every bad call? Because they don’t always know whether or not a call is bad in the moment.

Batters Understand Run Leverage
I wanted to understand what makes a player more likely to challenge, so I dug into the data more deeply. I took the population of challengeable pitches in the 2026 season (again, through March 30) and noted each one’s potential challenge value. I split the pitches into three buckets, each with a roughly equal total challengeable gain. For example, the 5,984 lowest-leverage challenge opportunities carried a total challengeable gain of 485 runs, while the next 2,571 challenge opportunities were worth up to 524 runs. The 890 highest-importance challenge opportunities were worth up to 441 runs on their own. (If you’re wondering why the buckets aren’t perfectly equal, it’s because I can’t slice things up infinitely; tons of pitches have the exact same value thanks to the way RE288 works.)

If players are behaving optimally, you’d expect to see a lot more challenges in the highest-value bucket. And great news: That’s exactly what’s going on so far. Batters challenge more frequently, even with a lower accuracy rate, in more important situations. They’re less accurate in those important situations, which is rational. They’re less accurate because they’re challenging more. And they’re challenging more because the value of succeeding is quite high:

Batter Challenges By Run Leverage
Run Leverage Challenges Opportunities Challenge Rate Success Rate Runs Gained Per Challenge Runs Gained
Low 49 2074 2.4% 55.1% 0.05 2.5
Medium 32 689 4.6% 43.8% 0.09 2.8
High 22 216 10.2% 50.0% 0.26 5.8

Now, this data suggests that hitters don’t have a great idea, in the aggregate, of whether a pitch is in the strike zone. In low-leverage situations, you want to be very sure of yourself before challenging. You get unlimited correct challenges, but only two wrong ones. That means that optimal behavior early in the count or with the bases empty is something like “challenge if you know it’s a ball, otherwise save it for a bigger spot later.” And yet hitters have been successful barely more than half the time. Yikes!

That’s not to say that there are no gimme challenges. Elly De La Cruz has only challenged one pitch all year. It came in a 1-0 count with a runner on first and two outs, squarely in the “low run leverage” bucket. That’s the kind of challenge you only make if you’re sure – but De La Cruz was, and he was right.

Of course, there are reasons to challenge even when there aren’t many (or any) runs on the line. The lowest-value challenge by runs added came on a borderline pitch to Spencer Torkelson in an 0-0, bases empty, two out situation. That sounds bad. But it was the ninth inning and Detroit had two challenges remaining. Spend ‘em if you’ve got ‘em in that situation.

Most low-run-leverage challenges are bad, though. That’s how you end up with such a low success rate, and so few runs added per challenge. Denzel Clarke challenged this one. Evan Carter challenged this one, and even made an “I’m not sure” face as he did it. Vladimir Guerrero Jr. burned a challenge with two out and no one on in the first inning. These guys are all undoubtedly being told by their teams to only challenge if they’re sure. They challenged anyway, in a situation where the gain was so small that they had to be certain they were right to have it make sense. It’s just hard to know where the ball crossed the plate!

Catchers Understand Run Leverage, Too
Catchers face a different value proposition than hitters. The population of pitches they technically can challenge includes a lot that they definitely won’t – curveballs in the dirt, fastballs to the backstop, and other miscellaneous non-competitive pitches. Hitters swing at most of the pitches they’re sure are in the strike zone, so the population of pitches they might challenge is disproportionately full of borderline calls relative to catchers. In other words, you can’t compare each side’s challenge rates one to one. But while the denominator is different, catchers behave similarly to hitters, challenging much more frequently in important situations:

Catcher (And Pitcher) Challenges By Run Leverage
Run Leverage Challenges Opportunities Challenge Rate Success Rate Runs Gained Per Challenge Runs Gained
Low 59 3910 1.50% 62.7% 0.06 3.4
Medium 43 1882 2.30% 58.1% 0.13 5.6
High 22 674 3.30% 45.4% 0.22 4.8
Catchers represent all but six of the challenges here.

We already know that catchers do better than batters overall. They challenge more low-importance pitches than batters and succeed at a higher rate. They also challenge more frequently when winning a challenge is more valuable, despite it coming with a lower success rate. They still aren’t perfect, of course – again, calling balls and strikes is very hard.

Catchers have challenged in the lowest-importance spot – bases empty, two outs, 0-0 count – five times already. They’ve only been successful 60% of the time. Yainer Diaz missed one in the first inning, and it wasn’t even close. Jonah Heim challenged a pitch that was more than two inches low, perhaps having fooled himself with his own frame job. This shot of him watching the challenge result live is extremely relatable:

So catchers still have some work to do. But it’s clear that they’re generally thinking about this the right way. Challenge rates go up and challenge success rates go down as the rewards for a successful challenge increase. That’s how it should be. In other words, catchers are rational actors, even if they aren’t perfect arbiters of the zone.

Low-Importance Challenges Are Still Happening Too Frequently
The exact math behind a breakeven challenge success rate is tricky, and I’m not confident that I’ve solved it yet. You have to consider your subjective confidence that you know whether the pitch was a ball or a strike, the reward for a successful challenge, and also how many more opportunities to challenge you might have in the game. Given a limited number of incorrect challenges but an infinite number of successful ones, your certainty changes the likelihood of you paying any cost whatsoever for your challenge. I can give you an approximation, though, and I think it has some clear takeaways.

Imagine two hypothetical situations. In one, a catcher has a 50/50 shot at flipping a walk to a strikeout. Since this is hypothetical, let’s say it’s worth 0.7 runs to do so. In the second scenario, the catcher is 80% sure that they can win a low-value challenge, worth 0.1 runs if successful. Here we’ll say that the value of one unspent challenge is 0.1 runs, the average run value gained per challenge issued so far. These numbers are all roughly representative of real-world run values, though the 80% challenge success rate is probably optimistic given what we’ve seen so far.

The math here is pretty easy. Before the challenge in the first scenario, the catcher’s team had 0.1 runs worth of unspent challenges. Fifty percent of the time, they get 0.7 runs and retain a challenge for a total of 0.8 runs of value. The other 50% of the time, they lose and end up with 0 runs of challenge value. The net value, then, is the difference between 0.4, the expected value of challenging, and 0.1, the expected value they held before challenging. That challenge is worth 0.3 runs in expectation, in other words.

Now let’s ask another question: What’s it worth to a catcher to challenge those 80%, low-importance situations over and over until they miss? We can run the math on that too. Twenty percent of the time, they get nothing. Another 16% of the time, they win one challenge and then lose the next. You can keep going down the line like that, adding 0.1 runs for every successful challenge and solving the entire equation. I went as far as calculating the odds of the catcher hitting 25 challenges in a row. Sum up the expected value of the ones they win, and you get that an 80% success rate in low-leverage circumstances, repeated infinitely, is worth 0.39 runs. Given that the initial unspent challenge was worth 0.1 runs, that’s a gain of 0.29 runs of value. In other words, it’s equally profitable to challenge low-importance calls, even with 80% certainty, as it is to challenge a pure coin flip in an important spot.

Of course, I’m probably overestimating the value of that 80% strategy. No one gets 25 chances to challenge in a game. If you limit the catcher to five challenges and assume that they’ll just pocket the unspent challenge for later if they hit all five, that strategy is worth 0.19 runs in expectation. It’s just really hard to be certain enough of a low-importance challenge to offset the value of having one left when you really need it.

You might notice that both strategies have positive expected value. Why not do both, then? That certainly seems reasonable to me. If you have a catcher who can overturn calls with an 80% certainty rate, he should probably do that until he loses a challenge. But the risk of not having a challenge remaining for the high-leverage spots is real. In finance, we used to call this picking up pennies in front of a steamroller. The odds are good – but the rewards aren’t enough to justify the risk.

We Don’t Know Who’s Good Yet
It’s going to take a while to figure out who’s actually good at challenging. There aren’t that many observations, and not every observation has equal value because of the differing run values for different challenges. Winning the most challenges isn’t inherently great. Neither is having the best challenge winning percentage. Teams also seem to be changing their behavior on the fly. It’s going to take a long time to weed out the best from the worst with so much variance. To make matters even more complicated, there’s definitely value in not challenging at times.

That said, we can say who has accrued the most value from challenges so far. It’s Eugenio Suárez. His two challenges – maybe you’ve seen them – were worth a combined 1.73 runs of added expected value. Kyle Schwarber comes in second with 0.71 runs added. On the catching side, Edgar Quero has added the most value, but he’s used a ton of challenges to do so and hasn’t been all that successful on them. Interestingly, pretty much all of the catchers who have accrued the most value have also missed a fairly low-importance challenge already:

Top Catcher Challenge Run Values
Player Successful Challenge Value Chall Overturn Runs/Challenge Best Success Worst Failure
Edgar Quero 1.3 9 4 0.14 0.46 0.08
Salvador Perez 1.0 5 4 0.21 0.50 0.11
Samuel Basallo 0.8 3 2 0.27 0.44 0.21
Nick Fortes 0.8 5 3 0.16 0.51 0.07
Patrick Bailey 0.8 5 3 0.15 0.51 0.10

It’s Hard To Measure Certainty
As Tango’s research into challenges shows, “just challenge the ones that are obviously wrong” does not describe the reality on the ground. Catchers in spring training challenged just 35% of pitches that were three or more inches inside the zone and were called balls. Those are obvious strikes; they’re a strike by more than a baseball width. These are free! If you challenge them, you get the strike and don’t lose a challenge. And yet catchers, even in a training environment where they were surely encouraged to experiment with the new system, just weren’t sure enough. Heck, they only challenged 70% of pitches two-plus inches into the zone in full counts.

In other words, you can’t just look at where the pitch ends up and say that every single bad call will get cleaned up. No one actually knows where the ball crossed the plate when they’re challenging. No one knows the exact physical location of the strike zone, either. The zone is a theoretical box, not a physical one, and pitches are moving so quickly and so much that batters and catchers are inferring their trajectory rather than perceiving the ball continuously throughout its flight. Sometimes, a catcher is sure the ball is way outside and doesn’t challenge, but in actuality, the ball nicked the zone. Sometimes a hitter is convinced a pitch was in the zone when it was actually four inches low; the opposite happens too. “Sure thing challenges” are sometimes actually pitches that shouldn’t be challenged.

Challenges Work In Two Main Ways
First, they correct some egregious calls early in games. Now, not every egregious call will get corrected; that’s just not how this works. But umpires have called strikes on 31 pitches that crossed the plate 2.5 or more inches away from the zone this year, and batters have successfully challenged nine of them. Similarly, umpires have called a ball on 11 pitches that were in the regulation zone by at least 2.5 inches; defenders have challenged six of them, prevailing each time. The league has cut egregiously wrong calls by somewhere between a third and a half.

Another thing challenges do? Make sure that a lot more of the highest-importance calls are right. So far this year, 624 pitches have been called a ball when the difference between a ball and a strike is worth a third of a run or more. Twenty-six of those calls were wrong, but 10 got challenged and corrected. Batters haven’t done as well, what with not having as good of a sense of the strike zone and all, but 207 pitches have been called a strike when the difference between a strike and a ball is worth a third of a run or more. Thirty-eight of those calls were wrong, and hitters challenged and overturned 11 of those, meaningfully reducing the percentage of important calls that get missed.

That’s great! Egregious misses and high-importance misses are the two times I’m most interested in having a robot ump correct the record. We don’t know a lot about challenge skill on an individual level yet, and I’m not even ready to say anything about which teams are doing the best. Again, this is very noisy data. But challenges are reducing the number of incorrect calls, and they’re doing so in a predictable and desirable way.

I’m excited to continue learning more about this system. My chief takeaways, though, are that it’s doing what I hoped it would, and that fans seem to love it so far. There are tons of fun research questions to consider. When catchers bat, are they better at challenging? How much more should star hitters challenge, and how much more should catchers challenge against star hitters? Does umpire identity change team challenge behavior? What’s an optimal challenge strategy? How does it change based on your roster? Who’s the best at it? Who’s the worst at it?

I don’t know the answers to any of those questions yet. But I hope to find out in the future. And in the meantime, what’s not to like? Calls are more accurate. Both the in-stadium and on-TV experience of a challenged pitch have been rousing successes so far. Players seem to be enjoying themselves. I’m not surprised that the challenge system is a success, because it worked in the minors and copies a system that worked well in tennis. I’m happy that we have it, though, and I think that in short order, everyone will wonder why we didn’t always let players do this.





Ben is a writer at FanGraphs. He can be found on Bluesky @benclemens.

50 Comments
Oldest
Newest Most Voted
MichaelMember since 2017
4 months ago

Fabulous article Ben.

LouisMember since 2024
4 months ago

Seems to me it’s going to be an interesting NL MVP race between Joey Wiemer and
Sal Stewart. Some risk that Aaron Judge and Cal Raleigh get cut or at least moved to the bottom of the lineup. C’mon, it’s only five days into the season. This is very thoughtful and detailed analysis. But this data is likely going to be very different a month from now. (E.g.: once the Yankees prohibit Jazz Chisolm from challenging, the hitter success rate will go up.)
One thing this analysis confirms and I don’t think will change with more data: there are still a lot of wrong calls. The ABS challenge system, by punishing for incorrect challenges, assures the perpetuation of some incorrect calls. Why anyone wants a system like that, I don’t know. Take the strategy out, take the umpires out. If, as the current system provides, the umpires are sometimes wrong but the computer is always right, why don’t we go with always right all the time?

HappyFunBallMember since 2019
4 months ago
Reply to  Louis

Because when a fully automated strike zone has been tested in MiLB it was not always right. Presumably the ABS challenge system, as implemented, will also not always be right. But on a limited basis those wrong calls are not only less frequent but perhaps even less likely.

Also, no one wants every borderline or disagreeable call to be challenged out of habit or spite. You can’t just allow for an unlimited number.

quincy0191Member since 2020
4 months ago
Reply to  HappyFunBall

“Right” is a matter of subjectivity. The strike zone is made up, which means it’s whatever we say it is, which means if we say it’s what the robot calls, then that’s what it is. The full-ABS system provides consistency, which I think is really what everyone wants.

You could also define the zone as “what the umpire says it is” and then guess what, umpires have 100% accuracy. The problem is that isn’t what we have defined it as – there’s a rulebook – and the inconsistent nature of human judgment means that definition implies the zone moves, which seems bad (but again isn’t inherently wrong).

Moreover, even if we assume the ABS zone isn’t always right – however that would happen – a full-ABS zone would be more right than the umpires, because that’s what’s being used to check the umpires. If challenge accuracy was 100%, then we would essentially be in a full-ABS world – whenever the ump and ABS disagree, ABS wins. So that’s the same as just always using ABS, and since it is theoretically possible that players could be 100% correct on challenges, there’s some world where challenge ABS is functionally the same as full ABS. Which sort of makes it weird that we aren’t just doing that, though presumably the drama and the “human element” folks want the inconsistency.

Bottom line, challenge ABS opens a door that I don’t think the league will actually walk through, but it does mean people can peek through, What’s on the other side is definitionally better because the league has already said it is, that’s why it’s allowed to overrule the umpires.

sandwiches4everMember since 2019
4 months ago
Reply to  quincy0191

“Right” is a matter of subjectivity.

It ceases to be subjective when we codify it in the rulebook. Whether it is an artificial construct or not, the strike zone has an actual definition. There does exist for any given pitch, a correct (or “right”) classification of the pitch as a strike or a ball.

If you’re watching or playing baseball, you understand the pragmatic limitations on correctness. For almost all of baseball’s history up until the last X years, it just wasn’t technically feasible to have anything but a human arbiter classify the pitch. Such standards have evolved over time — technological advances have moved toward standardization.

The practical problem for full ABS is time. Even though the challenge system is fast, there is a delay. It is much easier to have the computer system clean up any aberrations with post-processing when you’re not expecting a “real-time” answer. Even with gobs of processing power and working memory, there are still limits on how fast you can get a correct answer when the “problem” is calculating where in 3D space an object traveled by correlating visual inputs from multiple camera sources with noisy backgrounds.

What’s on the other side is definitionally better because the league has already said it is, that’s why it’s allowed to overrule the umpires.

It’s not definitionally better if the constraints of reality and expectation would reduce the quality of the output given by the computer system.

Now as far as the reductionist argument that challenge ABS is functionally equivalent to full ABS goes: while there does exist a world where players would correctly challenge with 100% accuracy (and 100% sensitivity — meaning they wouldn’t miss any opportunities to challenge), it doesn’t mean that such a world is feasible, per se.

Just because there is some world where these things hold doesn’t mean that this is that world. Baseball as a game would look very different if players had that level of ability to discern the position of the baseball.

Last edited 4 months ago by sandwiches4ever
soddingjunkmailMember since 2016
4 months ago
Reply to  HappyFunBall

>Because when a fully automated strike zone has been tested in MiLB it was not always right. 

Presumably it’s right more often than the umpires though, correct? (Or else why use it to overrule the umps?)

Generally, I’m not in favor of of things that interrupt the flow of gameplay or defer excitement. In a vacuum I’m in favor of getting more calls right, but I don’t like doing so at the expense of entertainment value. (For example, I didn’t like tennis’ challenge system, but I love the automated line calls.)

I’d like to see baseball go the same route – automated balls and strikes that keep fans in the moment.

HappyFunBallMember since 2019
4 months ago

As I understand it, the full robo-ump zone was not only less accurate than the challenge system … as another poster in this thread mentioned the difference in a couple of seconds of processing time is enormous … but it also brought into play some aspects of a 3-D zone that hadn’t really been considered. Namely pitches that are outside of the zone when they cross the front of the plate but curve or dip into the back corners of the zone.

The ABS challenge zone, you see, is not a 3D zone. It is a plane located at the midpoint of home plate.

Now, while it may be desirable (eventually) to train both batters and pitchers to more fully understand the full depth of the strike zone, if you try to do that as part of the process of acclimating everyone to the notion of robo-umping in general, you’re going to have way more controversy than MLB is interested in taking on.

What MLB wants, today, is the same strike zone but called with more accuracy.

lukeMember since 2024
4 months ago
Reply to  HappyFunBall

Why would/should full ABS have to use the 3D zone? It could use the same plane as the challenge system. You’re comflating two separate issues.

Last edited 4 months ago by luke
HappyFunBallMember since 2019
4 months ago
Reply to  luke

I’m not conflating issues. I’m pointing out that the two versions tested in MiLB are significantly different systems. ABS challenges are not simply snapshots of a larger fully automated model.

The fully automated system tested in MiLB did use a model of the strike zone that covered the entirety of the plate in three dimensions. As I understand it, it was also a fixed zone regardless of the batter’s height or stance. And here’s the kicker: The fully automated version did not actually measure strikes directly. Rather, it measured the trajectory of the pitch and calculated strikes based on whether or not the ball passed through the location box of the strike zone.

The ABS challenge system uses a flat plane, the upper and lower bounds of which are fixed to the player’s measured height. Furthermore, it only observes the moment that the ball crosses that plane.

katmanisaliveMember since 2024
4 months ago
Reply to  HappyFunBall

I will say that just because they chose to test something that was objectively not the rulebook strike zone does not make full ABS not the answer.

They chose to test an obviously bad version of ABS technology. One could argue they did that specifically to avoid having to go to full ABS.

I personally am a bigger fan of the challenge system because who really cares what the call is on borderline pitches. Does increasing accuracy so every pitch that touches the zone with 1/1000th of an inch accuracy improve the game? Unlikely.

But let’s not pretend their test of a bastardized form of ABS collected any meaningful or useful data.

HappyFunBallMember since 2019
4 months ago
Reply to  katmanisalive

One could also argue that they attempted complete solutions that just aren’t ready for prime time yet.

The robots are coming for all of our jobs. Some sooner than others. Umpires are a particular subset of the working class for whom there is actually significant public sentiment behind putting them all out of work. Poor them.

Shirtless George Brett
4 months ago
Reply to  HappyFunBall

Yeah the 3D strikezone was basically too accurate. Pitches that wound up in the dirt would be called a strike because it clipped like a sliver of the very front of the zone. Same with like a high curveball that just nicks the back of the zone. And god forbid a knuckleballer showed up lol. Technically correct but visually it just looked wrong. Which is the opposite of what ABS was designed to address. They found a 2D zone set in the middle best mimicked how pitches are actually called.

Side note, I pointed this very problem out probably a decade ago when we first started talking about Robo umps. MLB could have saved alot of money by just reading fangraphs comments 😁

darrenasuMember since 2025
4 months ago

Totally agree. The entertainment value of ABS has been a rousing success!

carterMember since 2020
4 months ago
Reply to  darrenasu

If I did like it, or if I didn’t… I’m fine with it now. Ballparks go crazy for it, it’s entertaining! Are teams using them right? No. Did the mlb say they have been doing it wrong today, yes! But all things considered it’s been entertaining for sure

Tigers SuperfanMember since 2024
4 months ago

Fantastic article as always. Does Baseball Savant have markings for challenges on their pitch by pitch data. I only see in the “des” column which is only marked on pitches that end at bats. Curious for advice obtaining the data to do further research.

TangotigerMember since 2016
4 months ago

Check the Search tool on Savant

bubblesMember since 2024
4 months ago

Long term I think teams will reduce offensive challenges due to lower success rates and defensive ones by the catchers will go up success rate wise and volume wise. Catchers have a volume advantage in that they see so many more pitches and can challenge more to keep improving their judgement overall.

With the higher defensive success rate and volume I expect, it will move the needle slightly further away from more offense which the league wants.

g4Member since 2020
4 months ago
Reply to  bubbles

I think a big advantage catchers/defenses have as it stands is simply the definition of a strike itself versus how umps have defined it for years. When I see these challenges won because a sliver of the ball ticked a corner of the defined zone — including pitches that are moving away from the zone post-snapshot — I can’t help but think, damn, that pitch was totally unhittable and the ump was totally justified, in spirit if not letter of the law, in calling it a ball. There’s so many of those types of reversals that catchers are very likely to win … and I kinda wish they wouldn’t.

If MLB wants more offense, and challenges take away the up-to-now commonplace blurring of the zone, it can do so by officially shrinking the zone, which IMO could be as simple as revising the rulebook text from “any part of the ball” to read “half the ball”.

Any part of the ball through any part of the 3D dimensions made for a perfectly practical zone when pitchers topped out at 89 MPH. But at 99 MPH? Tightening up the target area is a sensible offset.

si.or.noMember since 2017
4 months ago
Reply to  g4

revising the rulebook text from “any part of the ball” to read “half the ball”…Any part of the ball through any part of the 3D dimensions

Well, the challenge system is a 2d plane. So idk.

1 . Would the proposed change (half a ball, 3d zone) actually shrink the zone compared to the current 2d plane?

2 . That would make for a visually challenging ID (like, what if the nfl rule were “half the runners foot must be out of bounds” — and then extrapolate that to a 3d box).

Last edited 4 months ago by si.or.no
g4Member since 2020
4 months ago
Reply to  si.or.no

I worded that poorly. Umpires would continue to follow a rulebook definition of a 3D strike zone, and that definition would be tweaked to require “the majority” of the ball to pass through the zone at any point. My feeling is that this is already what many of the umps with tighter zones have been doing in practice for years.

Then, when the 2D plane is used for challenges, it would be exactly like it is now except a larger proportion of the ball would be required to fit within the box to qualify as a strike.

This isn’t really an ABS complaint. It’s a by-the-book strike zone size complaint. But by taking away an ump’s avenue to subjectively squeeze the zone, the presence of ABS could lead to its expansion (in practice).

bosoxforlifeMember since 2016
4 months ago

The one thing I was certain about was that the ABS challenge system would be very popular. I saw three games in Worcester in 2024 and it was clear that the crowd was completely into it and responded enthusiastically when the Hawkeye image came up. Beyond that, my impression has been that the players have not figured out the value of each situation as Ben clearly points out. Wasting challenges in very low leverage situations has been far too common. I was under the impression that Hawkeye calls were instantaneous and accurate. The umpire was beeped and made the call much like as if he had made the call himself. Speed and accuracy are both required and any delay is probably not acceptable. The quality of some of the umpires has also been exposed to the public, not just us nerds here on Fangraphs, and just a couple of more challenges would be better. All in all, I love it!

cashgod27Member since 2024
4 months ago

I think challenge data is going to be more valuable on the team level than the player level. I’d imagine every team has built virtually identical run expectancy matrices that show when players should challenge, but some teams are going to be better at communicating that to the players.

So far, the Yankees have been the best team at challenging both on offense and defense. That’s a system.

sadtromboneMember since 2020
4 months ago

After reading the comments here I am excited to see a deep debate over the exact meaning of “truth” in relation to strike zone calls, whether the umpire or ABS system is the proper authority, and how these decisions should be litigated and / or overturned. And I am looking forward to all of those debates occurring in the 20 seconds between each pitch.

George ResorMember since 2016
4 months ago

It’s probably to early to tell, but does having a call overturned have an impact on the umpire, do they call more strike after getting a ball overturned or vise versa

Sporter's Five HorsesMember since 2014
4 months ago

Love this. I agree runs not win expectancy are the right metric.

But when it comes to opportunity cost of a failed challenge, I think cost should be measured in win expectancy. Tactically, I have no clue as to how or if that can be captured well.

amiller78Member since 2025
4 months ago

my strategy would be to challenge every call by CB Bucknor

TKDCMember since 2016
4 months ago
Reply to  amiller78

He’s not even the worst so far this season but all over the internet he’s the only umpire you hear about. Bucknor was wrong on 6/8. Chad Whitson was wrong on 7/7. Did anyone hear about that? Curious! Could be a coincidence, I guess.

I’m not defending Bucknor. He’s never been a great umpire really in any way. The hyper focus on him across the baseball world, with plenty of others also worthy of criticism, should make a thinking person a little uneasy.

kramericaindustries
4 months ago
Reply to  TKDC

Is Whitson worse than Bucknor or were the Yankees and Giants just better at challenging? Remember the Red Sox ran out of challenges after like three innings, so for six innings beyond that any of Bucknor’s incorrect calls were beyond reproach.

Bucknor: https://pbs.twimg.com/media/HEleFC1WIAAPMQo?format=jpg&name=4096×4096

Whitson: https://pbs.twimg.com/media/HElfIe5WMAA8jYb?format=jpg&name=4096×4096

Whitson called a better game. The challenges were just well spotted.

TKDCMember since 2016
4 months ago

Ok, was Bucknor’s game actually an extreme outlier for overall ump score? Ron Kulpa, who also has a long history of less than ideal umpiring, had the same accuracy and 1% better consistency than Bucknor. Nary a peep.

ScottyBMember since 2017
4 months ago
Reply to  TKDC

Bucknor was wrong on 20 out of 80 pitches in the shadow zone in that game!!! It wasn’t just the challenged ones he messed up on

amiller78Member since 2025
4 months ago
Reply to  TKDC

clearly I meant he was the only one worthy of criticism

Mr. RedlegsMember since 2026
4 months ago

Ben, would it be possible to track umpire success rate post challenge vs. pre? I’m curious how much a hurt ego will impact the umpires after they’re continually shown up in front of thousands of people.

ScottyBMember since 2017
4 months ago

Suarez gained 1.73 in expected additional runs in his two challenges in the one at bat (Bucknor fail!!!!), but ended up grounding out in his at bat, leading to 0 actual additional runs. Yours is the correct analysis, but, man, it’s hard to average these things out.

One issue in terms of catchers challenging perhaps incorrectly is that the ABS strike zone does not account for batter stance, etc., whereas 15 years of muscle memory and thought process accounts for it. Over time, I suspect catchers will get more and more accurate.

Kellen VossMember since 2026
4 months ago

As a Tigers fan, Javy Baez should be banned from doing this at all costs.

LenFuegoMember since 2025
4 months ago

Personally, I am looking forward to the umpire shaming … of course, with Angel Hernandez out of the league, it won’t be quite as fun.

bosoxforlifeMember since 2016
4 months ago
Reply to  LenFuego

Bucknor, who should never be behind the plate again after Saturday’s fiasco, took one directly in the mask today and was forced to leave the game. This is an opportunity for the league to say he was disabled by the incident and gently send him into retirement.

kramericaindustries
4 months ago

I’d be curious about parsing this further. You mention an example of Torkelson essentially giving a “fuck it, nothing left to lose” challenge, which can skew data when in that kind of circumstance. It’s to be expected the odds of success will be lower when you’re throwing caution to the wind and just hoping for something because you can’t roll those challenges over to the next game. How many “garbage time” challenges (for lack of a better phrase) are there and how do the numbers look when those are filtered out?

Chris PMember since 2025
4 months ago

A great read. I really wish I could focus my more technical writing work like this – such an underappreciated (and underrepresented) quality!

South DetroitMember since 2020
4 months ago

Fantastic article. I am intrigued by the strategy teams will incorporate and this provides some early insights. I’ll also be interested in discovering who benefits the most and that will be answered in time.

Also wonder if this will be modified in future seasons to allow teams to reset challenges in extra innings, or preserve for last out of game regardless of prior challenges, etc.

Jason ScottMember since 2024
4 months ago

I am enjoying ABS but would recommend 2 tweaks: allow two batter and two defensive challenges (4 total), and allow every team a single “extra” challenge if needed in the 9th. Another idea might be to allow more or unlimited challenges on at-bat ending 3-2 calls, where accuracy has the most impact.

Sporter's Five HorsesMember since 2014
4 months ago
Reply to  Jason Scott

I agree on 1 extra 9th inning challenge! I think it should follow same logic for existing challenges….aka, once incorrectly used, it’s burned for 9th, and, extras.

But I do think there’s a reasonable fear of end-games crawling like the NFL running out the clock, which is why I’d be wary about rules on unlimited challenges for 3-2 counts, or even bumping 2 challenges to 4.

I’m sure there will be tweaks after a season or two of live data.

Jason BMember since 2017
4 months ago
Reply to  Jason Scott

Another idea might be to allow more or unlimited challenges on at-bat ending 3-2 calls, where accuracy has the most impact.

This is a terrible idea – every single 3-2 pitch would be challenged because why not, if they’re free and unlimited?

Having a low limit like it is now (even if it is tweaked in some way) is absolutely the way to go – over time we’ll see which teams/catchers/batters are better at it and, as importantly, know when to use their challenges. Things that make in-game strategy more important and reward good decision-making, I’m all for.

Last edited 4 months ago by Jason B
chewbaccaMember since 2025
4 months ago

Great article! Thank you…

Michael CecchiniMember since 2026
4 months ago

Great article as usual, though in a very unscientific poll of 5 fans, I found 3 who either dislike the challenge system or feel “meh” about it (including me). Yes, everyone loves to roast bad umpires but the extent of hate I’ve seen has confirmed in short order why the ump union preferred full abs. We’ve also seen players pilloried for “burning” a challenge, and while less common now, that is sure to increase later in the season and in the playoffs.
More broadly, there’s a certain absurdity to making players of a game also double as its officiants—while still enduring loads of missed calls.

samathMember since 2025
4 months ago

This runs-based analysis is clean, but has some pretty obvious flaws. The value of retaining your challenge(s) is the entire counterweight against challenging, and the fact that that changes over the course of the game is huge, and needs to be actually modeled. Otherwise, the fact that teams are challenging much more later in the game (even if it isn’t close) looks odd. In the limit, if a game ends on a called strike three, the batter should basically just always challenge, even if it doesn’t seem to be close; the value of retaining the challenge is zero if the game is over. (Same goes for walk-off walks, but those are much rarer.)

NATS FanMember since 2018
4 months ago

My Nats were strong in spring, but have sucked so far this season. They keep blowing calls in low leverage situations early in the game. then when they could use a call late, they don’t have them. Very frustrating!

viceroyMember since 2020
4 months ago

Thanks for getting a jump on the analysis of the challenge. Really adds a deep layer of strategy to the game that, if we see it long enough, I’m sure teams will find a way to “Rays” it via some unusual strategies. If catchers have already added half a run in a week, then perhaps players could earn a full win’s worth of WAR or more.

The tricky thing with the challenge system is that it seems far more strategic in timing than advanced stats aim to capture. Protection in a lineup has been debated and generally not shown to exist (perhaps because pitchers can’t just execute on demand). However, a challenge seems to be a much more calculated decision, so a challenge on a ball to a batter with two outs before Judge comes up seems far more than valuable than a two out challenge with Austin Hedges on deck. Will be interesting to see all the year end analysis of this thing

Roger McDowell Hot Foot
4 months ago

One thing not mentioned here that I didn’t anticipate going in, but have seen in practice already, is that (pitcher and?) catcher challenges can be used to clean up framing failures. Francisco Alvarez had a successful ball-to-strike challenge a few days ago that was a borderline cross-up, or at least a pitch that didn’t remotely hit the target, where he was moving his glove all the way across the plate and so basically fooled the ump. A reverse-frame, basically. Keith Hernandez in the booth called it out as “a strike you would never get called before ABS.”

I’m not sure how to quantify this but it seems important to keep an eye on. Framing still matters a lot, because there are too many borderline pitches for ABS to affect them much in the aggregate, but ABS can also buy a catcher out of framing in decisive spots.

mike sixelMember since 2016
4 months ago

Nine overturned calls in the KC MN game. Fans have, apparently, been correct in saying umpires aren’t all that good at their job.

mgwalker
4 months ago
Reply to  mike sixel

Define ‘good’. They are the best (humans) in the world at what they do.