Pitch-Framing Data Is Going Insane
The season’s complete, which means the numbers are official. This is convenient for a writer, because it means there shouldn’t be any issues anymore with comparing 2017 to another full season in the past. A full season is a full season. So how about a quick full-season review of the pitch-framing data? There’s something interesting going on. Something dramatic, something that shakes the foundation of the numbers themselves. I have the graphs to prove it.
The most advanced pitch-framing information is available at Baseball Prospectus. It’s long been the gold standard, and so many hundreds of hours have gone into generating the results that get published on the sortable leaderboards. There are two framing metrics of note, for catchers and for entire teams. One is just framing runs above average, which is self-explanatory. The other is CSAA, or called strikes above average. This is directly related to framing runs above average, but it’s expressed as a rate stat. I think that’s all you need to know. In this post, I’m going to use them both.
There exist 10 years of detailed information, based on the pitch-tracking technology that’s been in existence. This is a plot of standard deviations, year to year, over the course of the decade. This is on the team level, using runs above average. This is just an examination of each year’s spread.
There’s currently less spread than there was in the past. There hasn’t been any real change since 2015, but there’s still less spread than in, say, 2008. This is presumably related to a point I’ve made before: More teams than ever are aware of the value of good framing. So more teams are prioritizing it, which makes it harder to stand out. When you raise the floor, you narrow the distribution. There are still differences between the best teams and the worst, but the gaps are somewhat smaller. Neat stuff.
But that’s not really what I want to show you here. Instead, I’d like to call your attention to something taking place with individual catchers. I gathered data for every catcher since 2008 who’s had at least 2,000 framing opportunities in consecutive seasons. Here’s how the numbers held up, in terms of CSAA, between 2013 – 2016.
You see that pretty strong, linear relationship. You want that relationship to exist; that way you can have the confidence you’re measuring something real and sustainable. That plot comes with an R^2 value of 0.49. The slope is 0.69. A good framer in one year was likely to be a good framer in the next year, and the opposite was also true.
Moving on now, here’s the same plot, except for only 2016 – 2017.
Look, there’s still some relationship. This isn’t the picture of randomness. And yet, the R^2 value is 0.20. The slope is just 0.42, which is down 40% from in the earlier plot. Something just happened. Or, something is continuing to happen. I think this last plot drives it home. Here are all the year-to-year R^2 values, in terms of CSAA.
The sustainability is eroding. Which means the predictability is eroding. Sure, there was a little spike a year ago, but the last three years have the three weakest relationships in the sample. And, actually, the last four years have the four weakest relationships in the sample. Which pitch-framing performance was first being measured, one of the things that made it so exciting was that the numbers held up so well, year to year. It was essentially proof of signal. The method of measurement hasn’t meaningfully changed ever since. It’s the same system. If anything, it’s more advanced now than ever. It’s had the benefit of time. But the year-to-year relationships are disintegrating. A good framer in 2016 was still likely to look like a good framer in 2017, but that couldn’t be said with very much confidence. The data is getting increasingly random.
Welington Castillo is currently a free agent. He just spent the year with the Orioles. When he became an Oriole, he had the record of being a below-average framer. Last year, he performed like an above-average framer. Ditto J.T. Realmuto. Ditto Stephen Vogt. Buster Posey, meanwhile, got a lot worse, and so did, say, Tony Wolters. I have a sample of 393 catcher season pairs. Of the 24 biggest year-to-year changes in CSAA, seven of them just happened between 2016 and 2017. I don’t even know how to explain what’s happened with Chris Iannetta.
Of those 393 catcher season pairs, Iannetta is responsible for nine of them. But of the seven largest year-to-year changes, whether for better or worse, Iannetta’s been responsible for three of them. All three are from the most recent seasons. From 2014 to 2015, Iannetta got dramatically better. From 2015 to 2016, he got dramatically worse. And from 2016 to 2017, he got dramatically better again. Iannetta might be the current face of pitch-framing uncertainty. Or maybe it’s Jonathan Lucroy, who just keeps on declining. I don’t know. Things are just weird.
There are a few possible explanations. One, it’s all a blip. I don’t know. Maybe. Numbers do funny things sometimes. Two, there’s something wrong with the actual data, which might be related to the recent switch-over from PITCHf/x to Trackman/Statcast. That wouldn’t explain what already seemed like a trend before 2017. And all this information comes from Pitch Info, which takes deliberate care to make all necessary adjustments and corrections.
And three, get used to this. This could be the new normal, the consequence of more teams caring about how their catchers catch. Maybe framing is easier to teach and learn than we thought. Maybe more catchers than ever know what they’re supposed to do, and so the baseline for everyone is so high that random volatility plays a larger-than-ever role. Even if everyone were exactly the same, there would still be variation, because the baseball season isn’t infinitely long. If every team, for example, were a true-talent .500 ballclub, a season would still end up with 90-win teams and 90-loss teams. You could detect this volatility because, in subsequent years, you’d observe further randomness. Performances wouldn’t correlate so well year to year. That’s what we’re seeing with framers.
In a sense, you could see this all coming. It’s been possible to forecast, and I’ve written about this on multiple occasions. But it’s still pretty stark to look at that 2016 – 2017 R^2. A far weaker relationship than ever. Seemingly far more randomness than ever. Is Welington Castillo actually a good pitch-framer now? I never would’ve believed it, but the data says what it says. I don’t know what to think about Castillo, and I don’t know what to think about a lot of different guys. Pitch-framing has entered a strange new era. An era in which it still matters, but an era in which it’s not easy to tell who’ll actually stay good at it. It’s difficult to justify a heavy investment in something that now comes with such a high degree of uncertainty. But that same uncertainty is what the league has to reckon with.
Jeff made Lookout Landing a thing, but he does not still write there about the Mariners. He does write here, sometimes about the Mariners, but usually not.




Did BaseballProspectus make any alterations or updates to how they measure it? As awesome as BP can be, sometimes I worry some of their stats may suffer from overfitting.
Also, and I think BP does try to compensate for this, are the pitchers’ contributions maybe being undervalued in someway? Wilder pitchers with more movement on their pitches perhaps are affecting framing numbers more than is currently accounted for, maybe.
Perhaps the curveball revolution and the death of the fastball?
Start making more money weekly… This is a valuable part time work for everyone… The best part ,work from comfort of your house and get paid from $100-$2k each week …•••••••??
???USA~JOB-START
But what about the umpires? Framing amounts to fooling the ump (I’d argue that it’s therefore a form of mild cheating) – and presumably they don’t like to be fooled, now that they know it’s happening, and try not to be. Some framing is probably easier to detect; some catchers (Lucroy) may draw special attention.
Some of the glove movement by the catchers is so obvious that it would seem to be self-defeating. The umpires should be able to detect even relatively small movement and that should tend to make them call the pitch a ball.
I remember Jason Kendall writing in his book that any movements to pull a pitch back to the zone were more likely to get a strike called a ball than the reverse. We see it working in games sometimes, but maybe it just stands out more than clean, quiet receiving.
To this point – I don’t recall the name of the software, but some teams (think I saw it associated specifically with the Twins) are using a program that tracks catch probability by distance from the pocket, to help train their catchers to move the glove less. Yasmani Grandal does well with this though he ends up with a league-leading number of PBs, with all of his clanks, as a result.
Talking out of me bum, but I always wondered if umps were harder on the notorious good framers in an effort to get better grades.
Here is Jeff making a real thought out mathematical based post and ol’ majnun is coming along and lobbing maybe-bombs, weak!
Possible, but for an umpire concerned about his grades, the rational approach would be to try getting every call right regardless of who’s catching.
As an umpire, if I see a borderline strike from Perez, I give it to him, but if I “see” a borderline strike from Posey, then I do Bayesian updating and am more reluctant to call it.
That would actually be a rational approach.
I doubt this is happening though because I don’t think umpires are this sophisticated.
Umps do this for borderline strikes based on count whether they realize it or not.
I should note that I’ve also calculated the yearly standard deviations in individual-catcher CSAA, with a minimum of 2,000 framing opportunities. The three lowest standard deviations are from the last three years, with the most recent season having the least spread of all. Helps to support the idea of the rising floor.
It’s all Wieters. 4th most framing opportunities and ranks 103 out of 110 in framing runs with a -13.6 in 2017 (-4.6 in 2016). He single handedly deskilled framing.
Gee. Is it possible that catcher framing metrics are actually highly dependent on the quality of pitchers that said catcher is receiving for?? Pitchers that are consistently and constantly around the zone get more calls on pitches that are slightly off the plate than pitchers who are always all over the place. Pitches that hit their spot an inch or two off the plate are more likely to get strike calls than pitches that are actually in the zone but make the catcher reach due to the pitcher missing his spot completely… regardless of who’s catching!
Its not rocket science. Stop making it rocket science.
Practically everybody believes that the quality of pitchers affect catcher framing.
However, most catchers receive balls from similar pitchers from year to year.
If you could show that there were uncharacteristically larger changes in catcher/pitcher combination for the past several years, you might be on to something here, but it seems unlikely.
An appropriate question might be: Does BP control for pitchers in their model?
The corresponding answer would be: Yes, we do. You can even look up pitcher CSAA on the site if you were so inclined.
What is the control? When I consider the pitcher’s part I look at how far the catcher’s glove moves to receive the pitch. More motion likely leads to less free strikes. Glove on the black that doesn’t move probably gets more calls than glove in the middle that moves eight inches right. That’s on the pitcher.
People can give a much more sophisticated answer than me but I believe they use mixed models to try to give credit to the ump, batter, pitcher, and catcher.
Perhaps a dumb question, but how are ‘framed’ strikes at the top and the bottom of the zone accounted for? Do we have data on the height of the top and bottom of the strike zone from year to year? Have those changed?
Jeff, can you report the average number of opportunities in your samples for each year, as well as the r-squared? I’ll show you something cool you can do.
2008-2009: 5687, 5640
2009-2010: 5679, 5171
2010-2011: 5005, 5248
2011-2012: 5643, 5182
2012-2013: 5197, 5378
2013-2014: 5142, 5357
2014-2015: 5508, 5059
2015-2016: 5171, 5075
2016-2017: 4917, 5113
Oh, and:
0.72
0.72
0.65
0.82
0.69
0.60
0.36
0.56
0.20
Thank you, excellent.
Year2Year r Opps RegressionAmount
2008-2009 0.85 5,664 1,011
2009-2010 0.85 5,425 968
2010-2011 0.81 5,127 1,232
2011-2012 0.91 5,413 565
2012-2013 0.83 5,288 1,078
2013-2014 0.77 5,250 1,528
2014-2015 0.60 5,284 3,522
2015-2016 0.75 5,123 1,723
2016-2017 0.45 5,015 6,199
1. We convert r-squared to r.
2. We have the average number of opportunities each season pair.
3. We figure the “regression amount” as (1-r)/r * Opps
That figure is the most important figure: it tells us how many opportunities until what we observe is half-real and half-luck.
For something like a batter’s K, that number is around 100 PA. For OBP and wOBA, it’s 300 PA (though even more recently). For BABIP, it’s 4000 BIP.
Here we see that in the 2016-17 season pairs, we need an enormous amount of opportunities in order to find the talent, compared to what we used to be able to find.
Why is that? When EVERYONE has a talent, it’s hard to find it. So, for something like wOBA, it’s hard to find, because you basically NEED to be a half-decent hitter to make it in MLB (yes, even guys that seem “horrible”). But for something like K/PA, you don’t need to have that talent, since you can get by doing something else good.
Here, we have something fascinating, where up to a few years ago, it didn’t seem like teams knew to target framers, so, you had all kinds of catcher framing ability. But, it seems that that tide might be turning, exactly as Jeff is pointing out. And we see that here quite strikingly. When everyone has the talent, it’s harder to find that talent.
And hence, we need to regress the observations a great deal in order to best estimate it.
I have an idea. Let’s create an expansion team, the Cleveland Arachnids, and fill the roster with zero-talent ham sandwiches (I and my 68 mph fastball volunteer as tribute). By setting a true floor for talent and greatly expanding the league’s range of talent, would that lessen the need for regression? Or is there an underlying assumption of a normal distribution of talent, and skewing the distribution left would only make regression less effective?
Tango, the CSAA numbers are already shrunk toward the mean on an annual basis as a function of the multilevel model, and most of these involve an enormous number of opportunities to boot. I’m not sure how much more regression there is to be done, in a “most likely contribution” sense.
Yes, the year to year correlation is based on observed samples and not on the regressed values, thereby invalidating my technique (in this case).
This would be like me running a correlation of Marcel 2016 to Marcel 2017.
To that end therefore, simply showing the distribution in CSAA is sufficient to show how much spread in talent there is each year. (If I understand how Jonathan has constructed the metric.)
I’m late to the discussion but I don’t think that the regressions in the model take it all the way to the “talent” level. I think they are something like I do in the current version of UZR. I am trying to remove some of the measurement error by I’m still trying to capture performance and not talent. For example, a mediocre defensive talent can have a spectacular year in the field, for whatever reasons.
If we find that a player had supposedly a spectacular year on defense it is always part measurement error (the imperfect batted ball data makes it look like a lot of plays were harder than they were) and part actual spectacular play. A players prior seasons can help us out separating the two.
That’s the kind of regression I THINK Jonathan et. al. does.
So, Tango, in order to drill all the way down to talent, I think more regression needs to be done. I THINK we can use your same method to come up with that regression amount (y-t-y correlations). Because there is already some regression built in to the seasonal numbers, the y-t-y “r’s” will be higher and thus the “regression to talent” will be less as we would expect.
This whole conversation should be an article.
Off topic, but I’d hoped that some of the Orioles disastrous pitching results this year could be blamed on Castillo. Turns out maybe not so much…
Maybe umpires are aware of the data and are now less likely to be susceptible to presentation.
This has to be because of the strike zone changes, right? Just from observation it seems like they’re not calling the low strike nearly as much, which would obviously hurt certain catchers and help others.
Jon Roegele looked at the strike zone mid-season (https://www.fangraphs.com/tht/midseason-2017-strike-zone-review/) and found that it had shrunk somewhat. He also noted that data from Atlanta’s new park looked screwy, and excluded it. It seems likely there are multiple things going on here: the compressed range, a change in measurement technology, a change in the strike zone, a new park, ….
Two variables would be pitchers and umps. Pitchers who miss spots cause the mitt to move. Umpires who think they are being fooled will adjust.
Another component that might be worth digging into this offseason: Have hitters become better at identifying and reacting to pitches that are easier to frame?
For instance, the word is out now that taking a low pitch with Lucroy behind the plate is as good as taking it down the middle, so hitters work to foul off those pitches more than they otherwise would. And even if they swing and miss, Lucroy still gets the strike, though not the pitch framing credit. The net effect might be that pitchers still profit with him behind the plate, but Lucroy’s gaming numbers become worse.
Just a thought.
*framing numbers
I watched a lot of Lucroy last year as a Rangers fan, and I almost think he may have had some sort of wrist/forearm injury that went undisclosed.
He seemed incapable of ‘sticking’ pitches lower in the zone this year—he would receive the pitch, momentum would carry his wrist down and then he would sort of loop it back up to where he originally caught it.
Could also explain his lack of pop with the bat this season as well.
It’s interesting that you use you use Beef Wellington as an example of the ongoing randomness, because his progression as a framer is linear. Before anyone was paying attention to framing Beef was considered a good glove behind the plate. As soon as framing was a real thing, Beef was considered one of the worst gloves. As Beef doesn’t return my calls, I’m unable to verify that he has made framing a top priority, but over the last 3 years his framing progressed in a linear fashion to where it is today.
This brings up the point of consciousness and framing. Who is trying to frame and who is not? The Observer effect at work. As soon as you observe a phenomena, you change it.
It’s pretty obvious that when we changed data source couple of yrs ago then all of sudden correlations changed.
How much relationship might there be to the effect of wear and tear on framing ability? Injuries, tired catchers’ legs late in the season, I wonder what the impact on receiving is.
This one aspect of the data revolution boils my blood. We are commoditizing a skill that is disturbingly close to cheating. Seems to me that if we discovered certain catchers getting an advantage this big via, say, hypnotizing the ump before key at-bats, we would convene an emergency council to decide how to STOP THIS UNETHICAL MANIPULATION.
I attended a talk by Commissioner Manfred on Monday. When asked about the prospect of electronic ball-strike calls within 5 years, he reasserted that the technology is not ready to make those calls fast enough. (He also suggested that the power of making those calls is largely what gives the plate ump the status that discourages players from constantly challenging him — a laughably circular argument, but hey, you don’t get where Manfred is without a little sophistry.)
There was a Q&A afterward, but I didn’t get my question in:
“What active steps is MLB taking to help immunize umpires to this voodoo garbage?”
I feel like that if you were alive when someone started throwing curveballs or sliders you would’ve thought that was cheating
That is a really inane interpretation of what I said.
I have two questions:
1: Do you recognize that this has been going on as long as baseball has been played, and cathers have been praised for their ‘receiving skills’, but only recently did someone figure out how to try to estimate it’s value?
and
2: Do you recognize that ‘framing’ is not exclusively stealing ‘miscalls’ out of the zone, but also improving the odds that borderline strikes are called correctly? There are corners and edges of the strike zone that are not always called correctly and pitch framing is as much about getting the right call on a borderline strike as it is about getting the wrong call on a borderline ball.
Austin Barnes lead MLB in called strikes above average at 0.0035. He caught 3.5 more strikes than average per 100 pitches. No determination is made about how many of those balls would have been called correctly; some of them would have been wrongly called balls on the edges of the strike zone, especially the top corners.
If this were a plague on baseball, it’s one that has existed for the entirety of the game’s history without anyone trying to address it. All BP ever did was quantify its value.
I feel like some minor part of it could be that umpires are more aware of the great pitch framers and are maybe giving them less calls. Seems kinda stupid but it could help explain it somewhat. The umps could be thinking, it looked borderline to me and I know this guy is a great framer so maybe he’s trying to fool me, I’m going to call it a ball.
I now see this point has already been made, and the shame is setting in… lol
Prospectus’ numbers are the best PUBLIC numbers, for sure, but that doesn’t mean that they’re accurate, or that they’re properly distributing the value between the pitcher, catcher, and umpire. I mean, they basically ignore the umpire and they undervalue the pitcher’s contribution because they have no way of tracking a pitcher’s command – how far the pitch misses where the catcher is set up- which is obviously an influence on called strikes. Baseball Info Solutions at least has commandf/x and can include pitcher command in their calculations, and consequently the value of a called strike is somewhat more believably spread out between pitcher, catcher, and umpire (the umpire also has more value in the DRS metric, as opposed to being basically ignored by Prospectus). I would say DRS has a much more reliable framing metric – on top of a much more reliable all-around catcher defense metric.
Fangraphs uses the old catcher DRS that doesn’t include framing, but if you check baseball-reference you’ll notice that a catcher’s DRS has a different value there, because B-ref has been listing the updated DRS that includes framing. So for a bit of a sanity check of Prospectus’ numbers you can look at a catcher’s DRS on fangraphs (for instance, Tyler Flowers is -9 DRS there) and then check B-ref (11 DRS) and infer from the difference that DRS’ framing has Flowers for +20 runs against the +25 on Prospectus. James McCann is -15 runs on Prospectus but only -2 by DRS. Lucroy is -17.7 on Prospectus and -10 by DRS. Austin Hedges is +20 on Prospectus, +13 by DRS. The numbers on Prospectus tend to place an inflated portion of the run value on the catcher. Just something to think about.
I have no real answers, but here are some brainstorming ideas:
– More relief pitcher innings means catchers are less able to predict where, when, and how each pitch will move and where they need the glove to be because of a lack of familiarity.
– More velocity (from starters and relievers) means it’s increasingly difficult for catchers to get their glove to the right location before the ball gets there, and thus increasing glove movement after/during the catch.
– The changing ball makes it increasingly difficult for catchers to predict where the ball will end up.
– The changing ball somehow bounces around more in the glove, requiring catchers to try to move the glove farther so that they are catching everything deep in the pocket instead of the edges.
– More high-velo relievers with shaky control means the ball is less likely to hit the target perfectly.
Thoughts?
Any reason to believe the increased rate of fastballs up in the zone would cause more randomness? It would be very baseball if an unintended consequence of the approach to attacking all the good low-ball hitters led to a decreased consistency in pitch framing. Catchers are probably more well-versed at framing low pitches since that’s where pitchers have been taught to throw! Could be that teams adopting the approach more have seen more variance..
Also, was there a concerted effor to raise the zone? I think there was but then it went away?
You’ve probably thought about all of this and even looked into it, but yeah, I like to comment. Fangraphs is fun.
Don’t know if you know this but the y-t-y correlation is directly proportional to the spread of talent in the league. If that spread decrease, which is has, then the y-t-y “r” will decrease proportionally as well. If we get to the point where everyone has the same talent (that never happens of course), the the y-t-y “r” (or any time period to any other time period) will be zero. Those correlations (their magnitude) are a function of exactly 3 things: One, the amount of randomness in the measurement (basically measurement error) – for framing it’s probably pretty close to a binomial variance since we’re using strike zone data. Two, the sample size underlying each period in your regression (the “x” values and “y” values), in this case, one season each (although of varying numbers of called pitched and games played). Three, which is the important thing here, the variance of true talent. The higher the variance, everything else being equal, the higher the “r”. The lower the variance, the less the “r”. That’s why you’re getting the results you’re getting. It would be a shock if the “r” hasn’t gone substantially down as the variance of talent has gone down.
That being said, you should not be seeing any more fluctuations in framing numbers in 2016 or 2017 than you did in earlier years. If you are that’s either a random fluke or a change in the measurement procedure. The fluctuations among players has nothing to do with the variance in talent. If you underlying sample sizes are lower (are catchers playing fewer games, or getting less called pitches?), then of course you’ll see in increase in variance among individual players.
Finally, is it harder to project players when those “r”‘s go down? Depends on what you mean by easier or harder and it depends on why the “r”‘s are lower. If they’re lower because your sample sizes are lower than your projections have a larger error bar. If they’re lower because the variance in the population has decreased then the projections are actually MORE accurate. Imagine that everyone had the same framing talent (zero by definition). The y-t-y “r” would be zero of course. What would the projections look like? Everyone would be regressed 100% toward the mean and have a zero projection and those projections would be perfect! If we had a very small variance in the population, such that everyone were around zero, but not quite, it would be very hard to distinguish one talent from another without an enormous sample size, and we would still project everyone to around zero. Is that considered accurate, to call everyone a zero if we knew that true talents ranges only from -1 to +1? I don’t know. Depends on how you define accurate.
Whether umpires are being “less fooled” (I doubt that to any significant degree) or there is more selection going on by teams for good framers (I think that is definitely true to a significant degree) doesn’t matter. Either way results in less talent variance. Whether it’s the umpires or the pool of players, they are indistinguishable. That’s because good framing IS, by definition, fooling umpires. If teams only use good framers or we go to computer strike zones, we’ll end up in the same place. No framing. When we say, “there is no (or little) talent” in something (like DIPS) we ALWAYS mean, “there is no (or little) variance in talent in the major leagues. We don’t mean NO TALENT. (What does “no talent” mean anyway? – talent is always relative.)
This might be something that Mike Fast or Jeff has addressed elsewhere, but what is the consensus view on the best way to integrate the totality of measurable defensive skills at catcher to provide an integrated view of runs saved at the position? Is this the aim of DRS? Is framing relatively more important than throwing/stealing, and more important than blocking? Is there any way at all of figuring out who is good at sequencing, tunneling, etc.?