Never Use an ABS Challenge in This One Weird Count

If you clicked through to read this post, you’ve probably visited the ABS Challenge Leaderboard on Baseball Savant at some point this season. While you were there, you may have sorted by Won% to see which players have been the most successful with their challenges. And if you, like me, are a bit of hater, you also reverse sorted to see which players are now considered a fire hazard because of how rapidly they burn through challenges. In that case, you know that James Wood has won just 20% of his 15 challenges, that Josh Naylor owns a 25% success rate on 12 challenges, and that Jazz Chisholm Jr. has a 27% hit rate on his 15 challenges. Players this bad at picking their spots probably shouldn’t be allowed to challenge at all, right? Well the truth is, those samples probably aren’t large enough to definitively signal an inability to consistently win ABS challenges. Or maybe they are large enough, but it’s tough to say for sure because the ABS challenge system hasn’t been in place long enough to generate the volume of data needed to determine an appropriate sample size.
But even if there were absolute certainty about which players lack the eye for challenging ball/strike calls, sitting a player down and telling him he’s not allowed to challenge anymore because he sucks at it isn’t exactly the best strategy. It runs the risk of damaging the relationship between the player and the team and it shuts down the opportunity to improve with additional reps. And let’s say that player is in the box for a pitch that absolutely should be challenged — given the short window to challenge following a call, a batter paralyzed by self-doubt or concern over potential reprimand is set up to fail. It’s also much easier to communicate and get buy-in on a single, team-wide philosophy than it is to devise a bunch of player-specific exceptions to the rule.
The good news is that there’s a straightforward method for eliminating many of the most infuriating failed challenges, a method independent of any given player’s ability to judge whether a pitch was in the zone. Because there’s more to challenging than assessing whether a ball/strike call is correct and then assigning a level of certainty to that assessment. If you’ve ever watched a batter on your favorite team spend a challenge on an 0-0 count in the first inning, you know that it’s vexing on multiple levels. Even a successful challenge in that scenario doesn’t offer a significant swing in advantage, since it’s just flipping an 0-1 count to a 1-0 count (a swing in run expectancy of about a tenth of a run, depending on the base-out state). And to make things even more maddening, it also tightens the calculus around future challenges, since an additional failed challenge risks leaving the team unable to act on a potential missed call in a late-and-close situation.
Every failed challenge stems from a fundamental skill issue in reading the location of the pitch and judging the likelihood of an overturned call, but in some cases, the pain of failure is compounded by a lack of situational awareness. Before the ball leaves the pitcher’s hand, players should consider whether even a successful challenge is likely to have a meaningful impact given the current context of the game. Implementing an overarching strategy around which pitches are worth challenging from a situational perspective would limit the likelihood of a failed challenge in the early innings having a detrimental effect later on. Instead, those unsuccessful challenges would be concentrated in the high risk/high reward scenarios where failure is more acceptable because of the increased benefit associated with success.
But before we can determine which pitches merit the use of a challenge, we need to know the situations where a flipped ball/strike call is the most likely to influence the outcome of the game. Fortunately, we have a few ways to measure the magnitude of the impact of a successful challenge, such as Leverage Index, Win Probability Added, and run expectancy. When Ben Clemens wrote up his initial takeaways on the challenge system at the start of April, he advocated for using RE288 (the version of run expectancy that includes the pre-pitch count in its calculation) when evaluating the optimal usage of a team’s allotted challenges. I mostly agree with the logic he presented to justify that decision, but I do think there’s one component of Leverage Index that shouldn’t be completely cast aside.
RE288 does exactly what its name suggests, which is use historical averages to estimate the number of runs expected to score following each of the 288 distinct combinations of count, outs, and runners on base. Leverage Index works similarly, but its unit of output isn’t runs, but rather a rating of each situation’s importance relative to the outcome of the game. Leverage Index can be adapted to include the current count, but the standard version is calculated based on outs, runners on base, inning, and score differential, and it’s score differential and inning that distinguish Leverage Index from RE288.
Score differential we can ignore. Ben made a solid argument for discarding the score as a variable in the context of ABS challenges, noting that challenges are a finite resource within a single game. They don’t rollover from one game to the next, so there’s no benefit to sitting on them even if the game is a blowout.
Inning, on the other hand, is still worth keeping an eye on. Since teams are allotted only two incorrect challenges, the amount of game still left to play is a relevant component of the risk calculus around challenging a call. The primary consequence of a failed challenge is the potential inability to challenge at an important moment later in the game, but the likelihood of being caught without a challenge diminishes as the game progresses.
I’ll get to how exactly I want to incorporate inning into challenge strategy in a bit, but first let’s return to the notion of setting up situational guardrails to ensure challenges are consistently used on the pitches that matter most. We can isolate those pitches by using RE288 to compare the expected run scoring if the umpire’s call stands against the expected run scoring if the call is overturned.
As an example, consider the challenge initiated by Orioles catcher Samuel Basallo in the top of the ninth inning of a road game against the Mariners last Thursday night. Batting with two outs, a runner on second, and one team challenge remaining, Basallo tapped his helmet following a called strike in a 3-0 count. A called strike in that situation has an RE288 of 0.39, whereas a ball and the resulting walk would have increased the RE288 for the remainder of the inning to 0.44, an improvement of 0.05 runs (a relatively modest shift within the scope of possible increases in run expectancy). The call was confirmed, leaving the Orioles without challenges for the remainder of the game. But given that there was just one out remaining and Baltimore trailed by three runs, the Orioles were unlikely to see another borderline call. Then again, this game featured nine total ABS challenges and seven overturned calls, so another missed call by home plate umpire Tyler Jones wasn’t completely out of the question.
With a framework for estimating the potential run-scoring benefit associated with challenging any given pitch in hand, we can now bucket pitches based on the magnitude of the change in RE288. In his piece, Ben referred to the bump in expected runs scored as run leverage, so for the sake of consistency and because I find the term pretty apt, I’ll do the same here. I also used a similar methodology to his when defining the upper and lower bounds for the run leverage buckets. The low-leverage bucket encompasses all changes in RE288 less than 0.13, the medium-leverage bucket contains the values from 0.13 to 0.33 (inclusive), and the high-leverage bucket holds the pitches with a change in run expectancy over 0.33.
The table below shows the distribution of pitches thrown in low, medium, and high run leverage situations this season, the distribution of pitches challenged in each of those situations, and the challenge and success rates for each classification:
| Run Leverage | Opp. | Challenges | Challenge Rate | Success Rate | % of Opp. | % of Challenges | RV Per Challenge | RV Per Overturn |
|---|---|---|---|---|---|---|---|---|
| Low | 111,240 | 2,579 | 2.3% | 56.8% | 65.7% | 54.4% | 0.04 | 0.07 |
| Med | 42,674 | 1,449 | 3.4% | 50.7% | 25.2% | 30.6% | 0.11 | 0.21 |
| High | 15,012 | 710 | 4.6% | 45.9% | 9.1% | 15.0% | 0.23 | 0.49 |
Though high run leverage pitches make up less than 10% of all pitches thrown, they comprise 15% of all pitches challenged, which is to be expected given the higher payout associated with a successful challenge on those pitches. On the flip side, low run leverage pitches make up less than 55% of challenged pitches despite representing just over 65% of overall pitches. Again, this makes logical sense, given the diminished reward for successful challenges in this bucket.
When measuring across the total population of players, it seems the general strategy of challenging fewer low run leverage pitches and more high run leverage pitches is already in effect. However, zooming in reveals a wide swath of individual players who have yet to get with the program. Below is a rundown of the worst offenders, which is to say, the batters with at least five challenges on low run leverage pitches and a success rate on those challenges that sits at or below 40%:
| Player | Total Challenges | Success Rate | Low Leverage Challenges | Low Leverage Success Rate |
|---|---|---|---|---|
| Gary Sánchez | 22 | 40.1% | 17 | 35.3% |
| James Wood | 15 | 20.0% | 11 | 27.3% |
| Jazz Chisholm Jr. | 15 | 26.7% | 10 | 30.0% |
| Marcell Ozuna | 14 | 42.9% | 9 | 33.3% |
| Ceddanne Rafaela | 14 | 28.6% | 8 | 12.5% |
| Josh Naylor | 12 | 25.0% | 8 | 25.0% |
| Ronald Acuña Jr. | 16 | 43.8% | 8 | 37.5% |
| Chase DeLauter | 11 | 27.3% | 8 | 37.5% |
| Willi Castro | 11 | 27.3% | 8 | 37.5% |
| Rhys Hoskins | 10 | 50.0% | 8 | 37.5% |
| Andrés Giménez | 9 | 11.1% | 7 | 14.3% |
| Mauricio Dubón | 10 | 30.0% | 7 | 28.6% |
| Jonathan Aranda | 13 | 30.8% | 6 | 16.7% |
| Junior Caminero | 9 | 33.3% | 6 | 16.7% |
| Brett Baty | 6 | 16.7% | 6 | 16.7% |
| Jesús Sánchez | 9 | 33.3% | 6 | 33.3% |
| Adolis García | 9 | 44.4% | 6 | 33.3% |
| Ben Rice | 9 | 33.3% | 5 | 20.0% |
| Luis García Jr. | 8 | 25.0% | 5 | 20.0% |
| TJ Rumfield | 8 | 37.5% | 5 | 20.0% |
| Nick Gonzales | 7 | 28.6% | 5 | 20.0% |
| Miguel Vargas | 10 | 60.0% | 5 | 40.0% |
| Alex Call | 8 | 50.0% | 5 | 40.0% |
| Carlos Cortes | 8 | 50.0% | 5 | 40.0% |
| Spencer Torkelson | 7 | 42.9% | 5 | 40.0% |
| Michael Harris II | 6 | 33.3% | 5 | 40.0% |
This pattern of behavior is most apparent among batters. Catchers who challenge low run leverage pitches at an above-average clip do so with enough success that it isn’t a glaring issue. And of course, pitchers aren’t challenging frequently enough in any situation to really play a role in this conversation.
Going back to the Savant leaderboard referenced above, nearly all of the players who are in the bottom 10 of Won% appear in the table above. (Only two of Matt Chapman’s 10 challenges have been in low run leverage situations, while Randy Arozarena was one low leverage challenge shy of qualifying, though he has a 0.0% success rate in those situations). Removing these players’ low run leverage challenges doesn’t necessarily improve their overall success rate, but it does reduce the damage done by burning challenges in minimally impactful situations.
More broadly speaking, if players stopped challenging in low run leverage situations, how much value would be lost? Could that value be made up by re-allocating lost challenges to medium and high run leverage situations? Does it make sense to cut out all low run leverage challenges, or is there a way to be more purposeful about it?
Starting with the last question, there is something to be said for making sure there are challenges available for borderline calls in high run leverage situations. But teams wouldn’t need to behave all that conservatively with their challenges to stay in the clear on that front. High run leverage pitches make up under 10% of all offerings, and only around 4% of all high run leverage pitches have been challenged so far this year. Based on some quick back-of-the-envelope math, that’s maybe two pitches per game. Moreover, high run leverage pitches are fairly evenly distributed throughout the game, as are missed calls by umpires, so there’s no need to play it safe in anticipation of a run on challenges in the eighth and ninth innings.
Adding another useful data point for pacing challenges, Baseball Savant classifies each taken pitch as either “reasonable” to challenge or not. Reasonable is defined as meeting at least one of the following criteria:
- The umpire’s call on the pitch was actually incorrect.
- The pitch was located within three inches of the edge of the zone and an overturned call would lead to a swing in run expectancy of at least 0.3 runs. That is, the potential benefit of getting the call overturned justifies challenging, even if the pitch only meets the broadest definition of “borderline.”
- The expected challenge rate of the pitch is at least 20%.
Through Saturday’s games, the league is averaging just over three reasonably challengeable pitches per team-game. Assuming those pitches, like missed calls, are distributed evenly throughout the game, teams wanting to make sure they have challenges available for such pitches need to either get one of their first two challenges right or pass on low run leverage challenge opportunities early in the game. And this is where the current inning becomes a useful variable to challenge strategy. Because fully opting out of all low run leverage challenges would be leaving runs on the table. There aren’t enough challenge-worthy medium and high run leverage pitches to spend challenges on. Too many challenges would go unused, and not using a valuable resource is just silly.
To determine the most reasonable approach for trimming low run leverage challenges, I divided the nine regulation innings of a baseball game into three-inning chunks and measured the run value gained on successful challenges in each block of innings, broken down into high, medium and low run leverage buckets. I did not attempt to quantify value lost on unsuccessful challenges; that’s a topic for another day. Here’s what I found:
| Innings | 1 – 3 | 4 – 6 | 7 – 9 |
|---|---|---|---|
| # of Low-Leverage Challenges | 708 | 841 | 992 |
| Total Low Leverage Run Value | 32.7 | 36.7 | 35.2 |
| Low Leverage Success Rate | 65.4% | 59.3% | 49.0% |
| # of Med-Leverage Challenges | 392 | 470 | 563 |
| Total Med Leverage Run Value | 43.9 | 52.2 | 57.6 |
| Med Leverage Success Rate | 53.3% | 52.5% | 47.8% |
| # of High-Leverage Challenges | 196 | 226 | 264 |
| Total High Leverage Run Value | 47.6 | 51.3 | 56.7 |
| High Leverage Success Rate | 49.8% | 45.1% | 44.7% |
Unsurprisingly, low run leverage challenges in the first three innings yielded the lowest total run value, and though low run leverage challenges offered slightly less value on a per-challenge basis during innings four through nine, the difference was nominal. If teams were to sacrifice low run leverage challenges during the first three innings of play, would they be able to make up that value elsewhere? Mathematically, yes. League-wide, they would re-gain 245 failed challenges. Use those in high leverage situations instead, and that value could be recouped and then some. But that requires finding an additional 140ish high run leverage pitches worth challenging, which would be a 20% increase in total challenges for high run leverage situations. That’s a pretty big ask.
But as alluded to above, it’s probably unnecessary to swear off all low run leverage challenges in those first three innings. Maybe just the lowest of the low run leverage situations would do. How would that work in practice, though? It’s not as though players (or anyone else!) would be willing to memorize the RE288 table. Thankfully, it turns out several of the lowest-leverage situations have characteristics that are easy enough to distill down into digestible chunks one could actually commit to memory.
Here’s the simple summary of situations that basically never merit using a challenge during the first three innings of a game:
- Any 2-1 count. It doesn’t matter how many outs there are, and it doesn’t matter if there are runners on base. Batters in a 2-1 count should keep their hands as far away from their helmets as possible.
- With the bases empty, any count that’s made up entirely of zeros and ones (0-0, 0-1, 1-0, 1-1), regardless of the number of outs.
- Any 0-0 count with fewer than two baserunners. Again, regardless of the number of outs.
- With two outs, no RISP, and fewer than two strikes.
Stop challenging in these scenarios and teams are all but guaranteed (depending on specific personnel) to lose fewer challenges, increase their odds of still having a challenge for a critical moment late in the game, and see better overall returns in the instances where they do opt to challenge.
Convincing players to challenge less because their failed challenges are hurting the team is difficult and uncomfortable. But this approach is about convincing players to challenge less because eliminating challenges in these four very specific situations provides an extra layer of cunning gamesmanship. It’s a message some need to hear more than others, but it’s also one that will appeal to anyone with an unrelenting need to exploit whatever competitive edge they can find. And last time I checked, that’s most professional baseball players.
Kiri lives in the PNW while contributing part-time to FanGraphs and working full-time as a data scientist. She spent 5 years working as an analyst for multiple MLB organizations. You can find her on Bluesky @kirio.bsky.social.
It certainly seems weird, or at least counter-intuitive, that 2-1 would be the lowest leverage count, lower than both 3-1 and 1-1. Can you offer any intuition guidance there?
3-1 is one pitch away from a walk and 1-1 is one pitch from a bad batting count. 2-1 is one pitch from one pitch away if that makes sense (I think)
This doesn’t really help my intuition. What makes 1-2 such a bad batting count that it’s much more important to avoid than 2-2? What limits the offsetting value of 3-1 as a “good batting count”? How is 0-0 not the most “one pitch from one pitch away” count of them all?
I’d be curious too. In 2026, the league has a .580 OPS after a 2-2 count and a 1.054 OPS after a 3-1 count.
Here’s OPS after every count in 2026:
1-0 .816
2-0 .970
3-0 1.221
0-1 .601
1-1 .668
2-1 .795
3-1 1.054
0-2 .449
1-2 .484
2-2 .580
Last year was pretty much the same
To be thorough, the OPS after 0-0 (ie, every PA) is .720
I also can’t imagine using it with 2 outs and nobody on in the early innings regardless of how many strikes there are.
Those match my intuitions on how much better hitters do with 3-1 vs 2-2 counts. Other than 2-strike counts, 2-1 has the largest swing in run OPS, so something isn’t adding up for me.
Yeah, that stood out to me too. Before you showed this table I thought “surely the difference between a 3-1 count and a 2-2 count is bigger than the difference between a 1-0 count and an 0-1 count.” Looks like my intuition was right.
You left off 3-2
I did!
After 3-2, it’s a .780 OPS
What’s the opinion on whether the challenge system affects actual hitting? The basis of the question is that hitting is incredibly hard and requires lots of mental power as well as physical tools. You need to review the count, history of the at bat, history of that game’s at bats, maybe history against that pitcher, etc. There’s a lot to think about and that’s not including challenges.
Does the ability to challenge impact actual hitting performance if you include in it pre-swing thoughts? I would want to make the ability to challenge as simple and minimal as possible so as to minimize the impact to the actual at bat. Maybe that can be accomplished by the above? But maybe the above is already too complicated?
Is this a crazy opinion?
If you’re crazy then you have company because I, too, have presumed that the subconscious aspects of this newfound freedom must affect some batters more than others. Granted, challenging is still very new right now; over time, the gap between most affected and least affected is likely to narrow. Certainly this would be nigh impossible to measure, even via qualitative research methods, as many batters may not be fully cognizant of their own wandering minds.
In the sample scenario Kiri cited of Sam Basallo’s PA, I wonder if with 2 outs in the 9th and a challenge remaining, Basallo went up with a firm intention to use it, and thus took the first 3 balls waiting for that strike-challenge opportunity. A good strategy IF the pitcher is wild, but was his mindset equally prepared to blast a first-pitch cookie fastball into the seats? Or did the carrot of aggressive challenging increase his passivity early in the count? Some intrepid reporter like David Laurila might consider asking players these types of questions after the fact.
Though the lingering wounds of rampant, forbidden digital transmission of stolen signs might make this a sensitive sell, I wonder how long it will be before a team assigns a coach to a dugout traffic light, communicating — before the pitch — if a batter or catcher is allowed to challenge the ensuing pitch via red/yellow/green situations.
This tactic could be fun or annoying, but at the very least it’ll provide fodder for an amusing Davy article someday.
The Brewers tried exactly that in Spring Training! The league shut them down.
https://www.mlb.com/news/brewers-using-green-index-cards-in-phase-two-of-abs-plan
uh, I didn’t hear about that; thank you for sharing. Clearly the league is concerned about blurring the optics of pitch stealing because I don’t see how signaling challenge status is any less appropriate than a coach signing a bunt, steal or take directive to the hitter.
Yeah, perhaps the comparison with your proposal is not apt, because there was always the chance that a Brewer could catch a glimpse of the card after the pitch was thrown and before deciding to challenge, which would be forbidden.
The main problem with any system is that the information has to live inside the batter’s own head, alongside all the other important information about the pitcher’s arsenal, what kind of swing the situation calls for, and so on. Challenges are so rare it’s not clear it’s worth that mental capacity.
I can’t see how the league would know or be able to prevent the 3rd base coach from relaying a sign prepitch just like they do for any situation.
Didn’t the Brewers test this and the league said no?
Anon- what is the OPS for a 3-2 count?
Using Anon’s numbers, here is the difference in OPS for a strike call or a ball call at each count.
strike ball
0-0 -0.119 0.096
1-0 -0.148 0.154
2-0 -0.175 0.251
3-0 -0.167 -0.221
0-1 -0.152 0.067
1-1 -0.184 0.127
2-1 -0.215 0.259
3-1 ? -0.054
0-2 -0.449 0.035
1-2 -0.484 0.096
2-2 -0.580 ?
Based on this simple table, the biggest leverage would be 2 strike counts that are called strike 3 since a strikeout has an OPS of 0. A walk has an OPS of 1, and a successful challenge on 3-1 and 3-0 counts somewhat counterintuitively reduces OPS (but guarantees a base runner)- making it situationally valuable, but not always.
I can’t possibly be the first one who has thought of it this way, but it’s amusing to me that in a 3-0 or 3-1 count (and maybe 3-2?) a walk is actually worse than the average outcome.
I’ve always been surprised more hitters don’t swing 3-0 given you know you are getting a fastball down the middle. This year, when a batter puts a 3-0 pitch in play they are hitting .368 with an .846 SLG%, a number that is likely to go up by the end of the year since that doesn’t include a full summer of warm weather hitting (in 2024 it was .426/.890).
Sorry about that.
After 3-2, OPS Is .780
A walk doesn’t have an OPS of 1. It has an undefined 0/0 slugging percentage, so the OPS is also undefined.
Technical Man gets technical
OPS is imperfect as a metric, as it underweights the relative value of OBP. I think taking the walk and thus eliminating the possibility of making an out is more valuable than the added xSLG. Another way to look at this is run expectancy, since now the next batter has an 0-0 count with an added baserunner. Of course, that complicates everything, since you have to account for the number of outs, baserunners, etc. But my gut says taking the walk is going to almost always be better than allowing a strike.
OPS splits are also a LOT easier to find than wOBA splits and close enough to be relevant.
BTW, bold call on “taking a walk is going to almost always be better than allowing a strike.”
Agreed, but I think it breaks down when you’re talking 100% chance of not making an out (taking the walk) vs the increased slugging % (a lot of the high OPS on 3-1 and 3-2 counts is that the walk happens anyway).
I am stunned by how many wasted challenges I see. I am writing this at 8:45PM after watching Hunter Goodman of the Rockies challenge a 3-1 pitch to Wilyer Abreu in the top of the 1st with 2 outs and the bases empty. Goodman got lucky by a tenth of an inch. Abreu walked on the next pitch anyway. To challenge a close pitch in that ultra-low leverage situation was insane. What are the coaches doing if they are not helping the players from making these egregious mistakes?
Tangentially/ABS related, I wonder if someone at FG could look into whether an umpire’s ball/strike accuracy increases, decreases, or stays the same after a team is out of challenges. I’ve anecdotally noticed some effect there, that umps seem to be less accurate when they know that one or both teams is out of challenges.
It was driven home to me in the Nats game tonight. In the bottom of the 8th they loaded the bases with one out but Curtis Mead was called out on a pitch that wasn’t a strike, but the Nats had frittered away their challenges (as a team, they are so, so, so bad at challenges) and so had no recourse. That pitch was at least close. Unlike the pitch where the umpire called Dylan Crews out on a 3-2, two out pitch that was nowhere near the zone. Heck, I’d have been mad if Crews had swung at it. That’s a swing from +1 run and still batting to out of the inning. Lucky for the Nats they won anyway.
I don’t know how to edit my own post here… so separate post…
Updated with the 3-2 count OPS from Anon.
strike ball
0-0 -0.119 0.096
1-0 -0.148 0.154
2-0 -0.175 0.251
3-0 -0.167 -0.221
0-1 -0.152 0.067
1-1 -0.184 0.127
2-1 -0.215 0.259
3-1 -0.274 -0.054
0-2 -0.449 0.035
1-2 -0.484 0.096
2-2 -0.58 0.2
3-2 -0.78 0.22
Seems to me that the most value is in 2-2 and 3-2 counts. And it would be situationally most valuable when runners are in scoring position and 0 outs. Should be combined with run expectancy tables to identify optimal times to challenge. Seems very solvable with the right input data.
You can only edit posts for a very short amount of time, like 15 minutes.
I usually love your work, and maybe my problem here is with the headline, but I labored through this long article only to find one sentence related to the headline, and no description of why not to challenge in 2-1 counts, except that you better not.
I probably missed something, but if so, at least folks can have a laugh at my expense.