PitchingBot and Stuff+ Pitch Modeling Is Now on FanGraphs!
I’m happy to announce that we now have two pitch quality models, PitchingBot and Stuff+, available for your perusal on the FanGraphs Leaderboards.
PitchingBot is the brainchild of Cameron Grove. We worked with Cameron to be able to run and maintain his model in-house at FanGraphs; he has since joined an MLB front office. You can read all about PitchingBot in the FanGraphs Library here.
In short, PitchingBot takes inputs such as pitcher handedness, batter handedness, strike zone height, count, velocity, spin rate, movement, release point, extension, and location to determine the quality of a pitch, as well as its possible outcomes. Those outcomes are then aggregated and normalized on a 20-80 scouting scale, which is what is displayed on the leaderboards.
There are three models used in PitchingBot: Overall (botOvr in the leaderboards – stuff and location), Stuff (botStf – stuff only) and Command (botCmd – location only). These are available for a pitcher’s entire combined arsenal and for each individual pitch type.
There’s also botxRV100, which is the expected run value per 100 pitches, and botERA, which is that same run value converted into an easy to understand ERA scale.
Eno Sarris and Max Bay created Pitching+, with inspiration from work by Ethan Moore, Harry Pavlidis, and Jeremy Greenhouse, among others. Eno and Owen McGrattan currently maintain and work to improve the model, with engineering support from Matt Dennewitz. Stuff+ has been written about extensively at The Athletic. Additionally, Stuff+ and its corresponding models — Pitching+ and Location+ — are detailed in the FanGraphs Library here.
Stuff+ only looks at the physical characteristics of a pitch, including but not limited to: release point, velocity, vertical and horizontal movement, and spin rate. Generally, the model aims to capture the “nastiest” pitches in baseball. Location+ is a count- and pitch type-adjusted judge of a pitcher’s ability to put pitches in the right place. It ignores a pitch’s physical characteristics and looks at count, pitch type and location. The overall model, Pitching+, is not just a weighted average of Stuff+ and Location+ across a pitcher’s arsenal. Rather, it is a third model that uses the physical characteristics, location, and count of each pitch to try to judge the overall quality of the pitcher’s process. Batter handedness is also included in Pitching+, capturing platoon splits on pitch movements and locations.
Stuff+, Location+, and Pitching+ are all on the familiar “+” scale (like wRC+), with 100 being average.
It’s worth noting that all pitch modeling data can be rolled up to the team level and can also be broken down by custom date ranges.
David Appelman is the creator of FanGraphs.
Best website news since Goldstein’s departure
Weird dig.
Love it!!!
Yay! More tools to play with. Me likey
Awesome stuff! Will it eventually find it’s way to the player pages as well?
I’m wondering about this as well. Would love to be able to add some of this to my custom dashboard on the player pages
Yes, also hoping for this.
Something cool to play around with in this (as mentioned by Eno on the podcast) is that you can show that BABIP correlates with Stuff+, meaning that pitchers actually have some control on their BABIP. Take that, FIP stans!
I’m pretty sure its been known for a long time that pitchers have some influence over BABIP. It was one of the earliest counter points when Voros McCracken unleased DIPS on the world. Its just not a large enough effect to be super meaningful.
I’d say it isn’t super meaningful for most pitchers. For the Chris Young’s of the world, it is super meaningful that they can control contact significantly more than the average MLB pitcher. The Dodgers pick up about 4 wins a year lately controlling contact.
Even for someone like Young there isnt a ton of correlation from year to year in BABIP which suggests there is still quite a bit of luck involved. In the years where he had at least 100 IP his BABIP was:
.291
.226
.241
.254
.287
.238
.209
He was able to keep it below the league avg which is not nothing, but that’s still quite a bit of variance for something that is supposedly “controllable”.
Its probably also not a coincidence that he was putting up lower BABIP’s when he was pitching in San Diego, Seattle and KC (3 pretty pitcher friendly parks.) than he was in Texas (which is not a pitcher friendly park)
In our projections https://theathletic.com/4206001/2023/02/21/starting-pitching-ranks-sarris-2023/ stuff plus had a major effect on Babip, and our projections have a larger Babip spread than any other, just saying.
WOAH I was not expecting this. This is an awesome addition!
This is really awesome. Great addition!
Super excited about this! What years are data available for?
This is awesome! I’ve been looking around for Stuff+ style models for a while. I love that it’s here!
Awesome! Will this be available on the Graphs section of player pages? Preferably in scatter plot by game/month, rather than line graphs.
This leaderboard gives you a sense of how insane Jacob deGrom’s stuff is right now. From 2020-2022, deGrom has the #2 fastball, #1 slider, and #1 changeup. The slider and changeup are at something like 150, which is insane. The PitchingBot model is a little lower on the fastball (grading it “only” a 66) but hangs a 73 and 75 on the slider and changeup.
Also, for those of you who like to argue about Julio Urias, you have a new favorite stat to argue with (depending on your position). PitchingBot thinks that Urias is the second-best pitcher in the major leagues from 2020-2022. Stuff+ is much more modest, putting him in something like an 8-way tie for third place.
Yippy skippy, riveting
The leaderboard is really pretty neat. Although I don’t know what to make of it since the order isn’t what you’d expect either from other methods of ranking pitchers either through systems or people. Which makes me wonder what’s the missing ingredient here. Is it pitch sequencing? Is it just plain luck factors? Something else?
Now I’m curious about things like what’s keeping Kevin Gausman from being a Cy Young contender every year, or why Jose Urquidy isn’t the best Astros starter?
Yeah, I’m a bit curious, As a Cleveland fan, I first looked at their starters and was a wee bit surprised to see Bieber, McKenzie, and Quantrill ranked 36th, 37th, and 44th out of 45 starters last year. Quantrill I get but Bieber and McKenzie? Guess it means that Austin Hedges was the true AL MVP last year!
I mean the fact that they’re all on the same team suggests to me that defense is a pretty big part of that! Also framing and game calling, but the Guardians were either 5th or 6th in OAA.
True, but FIP has them 7th, 25th, and 34th. That’s a pretty big disconnect between two measures that should be defense neutral/independent.
The Guardians were wondering that too, which might be why they hired the guy that developed PitchingBot!
These are fantastic additions to the site. Well Done!
Explain this to me like I’m my dad.
Very pumped for those Stuff+ numbers to finally be public after reading Eno’s work at The Athletic the past few years.
Alright, you win
I’ll pay for a subscription now
Am I reading the leaderboards correctly that the Pitching+ leader is only at 111, meaning that the very best pitcher in baseball is only 11% better than league average? I’m sure I’m missing something here.
It’s centered on the per pitch level, so when you aggregate to pitchers it’s a different spread. You can see the spread here: https://library.fangraphs.com/pitching/stuff-location-and-pitching-primer/ 111 is two plus SD above average
As another commenter asked, will you please explain this to me like I’m my dad?
I think the idea is that the very best pitcher in baseball’s average pitch is 11% better than a league-average pitch (of the same type), averaged across all the pitches he throws. Turns out being that much better than average on every single pitch you throw is really hard.
Fantastic “get”!
Why does stuff+ think Patino is great, but every other metric think he’s below average?
It’s rating the quality of his pitches, not the quality of the results
This is great! Will the data be updated daily in season? Or will it be done in dumps like the defensive stats?