JABO: The Best and Worst Managers at Challenging
After the conclusion of yet another great regular season of baseball, we can now start in on the exciting year in review retrospectives. The stats are in, the playoffs are scheduled, and we can look back on the totality of 2015’s regular season with a full sample of what worked and what didn’t for players and teams. The first day after the end of the season can be disappointing for fans who don’t get to root for their favorite club during October, but it’s also our first chance to draw that all-important end line that frames this year along with all the others that have come before it.
The same goes for evaluating managers. We have an idea of the managers who have done a good job; many of them still have games to play. We also have an idea of those who haven’t done a good job — even up to the point of knowing who has a chance of being fired. The drawback, unfortunately, is that we still don’t have a great way of truly evaluating managers, and so we look to the small amount of data that we do have when we try to gauge their performance. Dave Cameron talked about this in the context of filling out his NL Manager of the Year ballot last year — here’s a paragraph from that piece relevant to what we’re discussing today:
“Evaluating player performance is tricky enough even with all the amount of information we have about their performance; with managers, we’re basically just guessing. We can speculate about things that we think matter, but we don’t really have much objective data to support these thoughts.”
Dave’s right — we have very little data, and the data that we do have isn’t terribly useful for evaluation. That being said, there is one newer area of data with respect to managers that I find interesting, and it lends itself to not only understanding an aspect of performance, but also — in this year’s case — serves as a window into the operating style of particular managers.
That data is the result of the fairly new system of manager replay review, and this season of baseball has produced some very interesting results. We already had a post earlier in the season on manager challenges, looking specifically at Kevin Cash, and his rather “unique” style of challenging (not waiting for any sort of video consultation from his coaches/advisors before popping out of the dugout to signal for an official review). That post theorized on a way to rank managers on their challenge ability; this post will go a step further in refining an attempt to do that.
We’ll do this a few ways. First, we’ll start by looking at successful challenges. This data comes from Baseball Reference, with the site originally ranking managers by challenge success rate, or the percentage of the time managers were correct out of the total number of times they challenged. That actually isn’t the best way of looking at this data, however, as an absolute value of how many successes during the season is probably more valuable: since there is no penalty for losing a challenge (other than the loss of a potential opportunity to use the challenge in the future), there shouldn’t be any penalty in our ranks for not succeeding with a review.
Think about it another way: challenging unsuccessfully at some point during a game is always better than ending a game without having challenged, as a manager has given their team at least a chance to improve upon their possibility of winning the game (however small that chance might be).
With that said, let’s take a look at the number of successes by manager, along with the total number of challenges that they’ve initiated (represented by the dot). Teams that have had multiple managers in 2015 have been grouped together:
Owen Watson writes for FanGraphs and The Hardball Times. Follow him on Twitter @ohwatson.

The 2nd and 3rd graphs made me feel like I was looking at performance comparisons in a review at AnandTech. To clarify, this isn’t a complaint.
Keep up the good work!
Interesting read, but a few things if I may. First, while the ‘dot’ in the first bar graph is the total number just using or seeing that number in parenthesis would be nice. I understand percentage does not tell the full story, so would support your reasoning behind using total won challenges. That being said, I’d like to compare their success rates with total numbers. Also what’s the actual correlation between the total number and the high leverage? Did you explain that and I misread it?
Nice to see the numbers back up my gut reaction that Hurdle (2nd overall) is doing a very good job with his challenges.
Yet another thing Matt Williams is terrible at!
Though, to be fair, the replay system seemed highly questionable. It really did feel like the Mets had infiltrated the replay room in New York on occasion.
And gNats fans wonder why they’re so ridiculed…
Great article. My additional question is, are the high-leverage measures skewed by winning early? In other words, will a team like this year’s Brewers or A’s find themselves later in the season low on “important” moments, but teams in the thick of a playoff race, like the Cubs, have lots more, thus affecting these metrics?
This is game leverage, so a 1-0 game in the 7th counts the same whether it’s April or September, Cubs or A’s.
That said teams that play more close games have more high leverage situations throughout the season. But I expect the effect is small — this isn’t the NFL where the worst teams are out of it at halftime fairly often.
Why did you choose to use Leverage * successes rather than sum(WPA)? The latter seems more direct, and is in units of wins.
Seems like a better final graph would be to evaluate each manager based on WPA. I would think you could do that by looking at each play and seeing the impact that the overturned call had on the win probability of the manager’s team. Would also combine total successes and “clutchness.”
Would also be interested to see a similar analysis for last year to determine if it’s a repeatable skill.
Except that hopefully the whole stupid challenge system will be scrapped next year in favor of the umps just quickly reviewing all close plays.
Quickly reviewing plays? Have you ever seen an official quickly review any play? In any sport? Especially hit by pitches? And for the record, don’t spend 10 minutes looking to see if a player was hit by pitch on video. Just roll up his sleeve and see if there’s a red Mark/swelling. Common sense guys.
Yes. Goal line technology in soccer and tennis has been implemented extremely well and quickly. Officials are notified goal/not goal or in/out, and everyone moves on. The entire process takes less than 10 seconds.
MLB has chosen to implement a slow, time-consuming approach, just as the NFL, NHL and NBA have too. These breaks generate additional ad revenue time, the difference with these leagues is they are not suffering from the perception that games take too long.
Regarding the last graph where you put things together, does that assume their success ratio is evenly distributed among the various leverages?
It looks like Matheny was successful 50% of the time (16 out of 32). At the extreme end, maybe his 16 successes were in the 16 highest-leverage situations. Or at the other end, maybe he was successful on the 16 most meaningless challenges. Is that taken into account? Or are you looking at it as if he’s 50% successful in each situation?
Enjoyable read.
I wonder to what degree this is basically random and small sample size. Matt Williams had a high success rate and an average number of successes last year. Lloyd McClendon was middle of the pack last year — and anecdotally seemed to be poor with challenges — and very close to top in success and % this year.
Aside from Cash, I imagine most teams have the same strategy for challenges incorporating the views of the player involved, the video guy, and the manager watching the replay on the Jumbotron. I suspect some managers are more/less willing to challenge low leverage situations in early innings and likewise there may be different approaches to high leverage, low chance of success situations. I don’t recall many stories about managers wasting challenges early or failing to challenge a blown call that impacted the game.
I.e., if what are the top guys doing right and the bottom guys doing wrong?
I had the same reaction. I think the data is interesting, and I’m glad Owen wrote the piece, but I’m not convinced it tells us much of anything. I just don’t see much spread in terms of challenge strategy – it seems a byproduct of opportunity as much as anything else.
I mean, are Maddon, Hurdle, and Yost judiciously holding onto their challenges until just the right moment, whereas Ausmus, Bannister, and Williams are blowing theirs early? Are good managers adept at spotting calls that are ripe to be overturned while bad managers let them slide? I doubt it. My guess is that Maddon just happened to have a lot of calls go against his team at key moments (discernibly incorrect calls – a key point), whereas most of the other guys just weren’t in that position as frequently.
At the very least I’d be convinced this is more signal than noise if these numbers held up year to year. But I suspect ‘clutch challenges’ are more akin to clutch hitting. That is, managers do make clutch challenges, but it’s not a detectable skill that distinguishes them year-in-year-out.
This is interesting, I did a quick comparison of the Review Score in the final graph with each team’s Luck (Wins – Pythagorean Wins) and there is a definite positive relationship, though the R^2 is only 0.08. It would be interesting to see this analysis done with several years of data
Not surprised to see Farrell near the bottom percentage-wise. Some of the calls he challenged this year made me wonder if he was actually at the game. It seemed like Lovullo had a little better luck, and would like to see what his numbers were.
I think Clint Hurdle success may have something to do with the speed of his outfielders on the base paths. Most of his challenges involved bang-bang plays at first with Marty and Polanco beating out a lot of ground balls for singles. And slides at second. Kang has a weird way of applying tag that basically block the umpire’s view from the tag.
Huh? I would guess his success is due to having a very good staff member behind the scenes watching replays quickly and accurately.
Few things.
1) This is a measurement of the success of the video review person as much as the manager
2) The % success rate needs to be factored in somewhere, I don’t agree that is a nonfactor. Some managers will not challenge early in a game if something is borderline. And the loss of a challenge early is real and there is no way of quantifying the opportunity cost of no other challenges without going through game by game when there was a failed challenge.
3) There are some throwaway challenges late in the game when you are at the point of ‘might as well challenge it as the game is over anyway (or nearly over)’ Are these impacting the clutch index even though they are certainly losing challenges (is it helping push up the average index?). Is the combo clutch+success chart only on successful challenges? If not, why use average of all challenges and not simply the successful ones?
Perhaps certain managers are just victims of more egregiously bad calls more often than others, so their win % will just naturally be higher through no skill of their own. And maybe that’s just due to bad luck, or umpires that hate certain teams/managers more (just kidding on that last one)