Batted Balls: It’s All About Location, Location, Location

BABIP is a really hard thing to predict for pitchers. There have been plenty of attempts, sure, but nothing all that conclusive — probably because pitchers have a negligible amount of control over it. So naturally, when I found something that I thought might be able to model and estimate pitcher BABIP to a high degree of accuracy, I was very excited.

My original idea was to figure out the BABIP — as well as other batted ball stats — of individual pitches from details about the pitch itself. Velocity, movement, sequencing, and a multitude of other factors that are within the pitcher’s control play into the likelihood that a pitch will fall for a hit (even if to a very small degree). But much more than all of those, pitch location seems to be the most important factor (as well as one of the easiest to measure).

I got impressively meaningful results by plotting BABIP, GB%, FB%, wOBA on batted balls, and other stats based on horizontal and vertical location of the pitch. So I came up with models to find the probability that any batted ball would fall for a hit with the only inputs being the horizontal and vertical location (the models worked very well). I even gave different pitch types different models, since there were differences between, for example, fastballs and breaking balls. I found the “expected” BABIP of each of each pitcher’s pitches, and then I found the average of all of those expected BABIPs — theoretically, this should be the BABIP that the pitcher should have allowed.

There was absolutely no correlation to actual BABIP allowed, and there was even less correlation between one year’s “expected” BABIP and next year’s BABIP. My idea didn’t work at all. I was pretty crushed, since I thought that I’d be able to find something revolutionary with this new method. But I did find some pretty interesting things along the way, which is why I’m writing this article.

First, some details about my methodology. The way I did this was to group all pitches of a certain vertical or horizontal location together and calculate the BABIP, wOBABIP, GB%, or what have you of all the pitches in each bucket (which had widths of 0.1 feet); I then came up with an equation to model those stats based on the results. So for each pitch, you can roughly expect what the wOBABIP will be solely from its location — average all of a pitcher’s expected pitch wOBABIPs and you have his expected wOBABIP. This is a different methodology from most other ways of modeling BABIP and other similar stats, which look holistically at a pitcher’s whole season (or some other chunk of time) and draw conclusions from that.

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

My data was all from pitchf/x, for which we have information going back to 2007. Errors and fielder’s choices, unfortunately, had to be excluded from all calculations as the MLBAM play description strings — which contain the batted ball type — don’t mention what kind of batted ball it was on a reached-on-error or fielder’s choice. Instead, they say something like “Player X reaches on fielding error by Player Y,” whereas a normal play description string would say something along the lines of “Player X singles on a ground ball to center fielder Y”. Additionally, all bunts were excluded.

You may also notice that popup percentages (PU%) are very high compared to the FanGraphs numbers, which is because MLBAM data has a more liberal definition of popups. (An aside: I tend to prefer popup rate — that is, popups divided by all balls in play — to infield fly ball rate (IFFB%), which is popups divided by just fly balls, as it is more stable and there is really no similarity between a popup and an outfield fly ball. For that reason, I am also going to use OFFB% — outfield fly balls divided by balls in play –instead of FB%. OFFB% + PU% = normal FB%. I refer to OFFB% as FB% throughout this post.) HR/FB% numbers may also seem high because I am defining that as home runs per outfield fly ball instead of the regular FanGraphs definition, which is home runs per all fly balls (including popups).

So let’s now take a look at the results. First, we’ll look at just the vertical location of pitches and how that relates to various batted ball stats. How about BABIP to start off?

BABIP

What you’re looking at here is a plot with BABIP plotted on the x-axis and distance from the top of the strike zone — which is located at y=0 — on the y-axis; negative numbers are below the top of the zone and positive numbers are above. The reason I measured the vertical distance in feet above/below the top of the strike zone is because different batters are different heights, and measuring absolute height would be less accurate. I set the axes the way I didto give a better visual representation of what we’re trying to see. (The same plot, just with horizontal distance instead of vertical, is flipped — that I will show later.) The size of the circles represents how many pitches there were in that location in the dataset (the scale is on the right).

There’s a pretty clear relationship between vertical location and BABIP! Nothing groundbreaking, it’s all intuitive and expectable (pitches in the middle of the zone fall for hits more often). I was certainly surprised by the strength of the relationship, although I suppose you shouldn’t be, since I just told you in previous paragraphs that there was a strong relationship. Let’s now look at wOBABIP (which, for the sake of ease and typing, I will call just wOBA throughout this post. No other type of wOBA will be mentioned).

wOBA

 

 

This graph is nearly identical to the BABIP one except for a different scale and a more drastic drop as the pitch gets lower in the zone. Next, let’s look at different batted ball stats – GB%, FB%, LD%, and some others. FB% and GB% come first:

FB

GB

Holy tight relationship! This is much better than BABIP and wOBA. Again, the findings are not so surprising — lower in the zone, you get more grounders; higher in the zone, more fly balls — but the strength of these relationships (albeit the fact that they are not linear) is again surprising.

PU

Another tight relationship! This one was even more encouraging to me when I saw it than the previous two were, because popups are one of the driving factors behind BABIP and have much more influence over it than either outfield fly balls or ground balls do.

HRFB

LD

These last two are a little weaker, as the circles don’t fit as tightly on the line. That’s to be expected, though: LD% and HR/FB% are notoriously unstable year-to-year, so it would follow that it’s harder to estimate them from secondary factors. However, there is still a clear pattern, and it’s a closer relationship than what might have been expected. This is more encouraging stuff.

Now onto the horizontal distance. As I mentioned before, these graphs have the axes switched from what they were in the previous ones, so realize that when you’re looking at them. The distance here, too, is adjusted, so righties and lefties are on the same scale — a positive value is always a pitch farther outside, and a negative value is always a pitch farther inside — so it’s essentially looking at this from a righty’s perspective. x=0 is the center of the plate; the edges of the strike zone are at x=±.708333.

BABIP wOBA

Much like vertical location, BABIP and wOBA follow similar patterns to each other, and of course, a pitch in the middle of the zone is more likely to fall for a hit — and more likely to fall for an extra-base hit — than a pitch on the edge.

FBGB

And as with vertical distance, there is an extremely close relationship between horizontal distance and both GB% and FB%…

PU

… as well as with PU%. This was another good sign, and it looked to me at the time like we should be able to predict at least PU%, GB%, and FB% pretty well if nothing else.

HRFBLD

And HR/FB% and LD% also show a fairly nice relationship here. This was all good.

But these don’t tell us too much by themselves. They’re interesting, but not too surprising or applicable. To get a better sense for what location does to BABIP and other stats, we need to look at both kinds of distance on the same graph. Since plotting in three dimensions is hard, the way I will display these is with heatmap-style graphs. The axis scales are read the same as the ones in the graph above, only now they are on the same plot; the color of the boxes shows how high or low the statistic in question is at that spot. The white dotted line is a general representation of the strike zone; the zone obviously changes based on the height of the hitter, but this is the average one.

BABIP

Meh. Actually, there’s not so much of a relationship for us to see, although the model I came up with had a pretty high correlation (north of a .8 r^2). wOBABIP shows a much clearer graph:

wOBA

There we go! There’s a really visible pattern now. A line extending from the upper-outside corner to the lower-inside corner seems to be where hitters do the most damage — interestingly and maybe not coincidentally coinciding with the line for effective velocity. If you look closely back at the BABIP graph, you can see the same line, only less pronounced. The message is simple: in general, hitters do better with pitches in the middle or low and inside.

FB GB

These last two are, I think, more fascinating visually than analytically. The smoothness of the graphs here for OFFB% and GB% could have been anticipated because of how high the correlation is for each of vertical and horizontal location.

HRFB

Nothing here at all. Oh well. The graphs plotting HR/FB% against only one type of location showed promise, but putting the two together shows none.

LD

Line drives follow a similar pattern to that of BABIP and wOBA if you look closely enough. It’s weaker than those two, but it’s there. Which makes sense: line drives are by far the best type of batted ball for hitters, as they land for hits the most.

PU

And, lastly, popups. This one, in the same vein as OFFB% and GB%, follows an extremely tight relationship in each of the individual location types and then also when combining the two. And since popups are the best results a pitcher can hope for if the ball is to be put in play, maybe pitchers should start throwing high and inside more.

How useful is this information? Not terribly useful, unfortunately. There are a few reasons for why the models from this don’t explain pitcher BABIP (or wOBABIP) at all.

First, these graphs are made with samples of thousands of pitches for every location bucket. If the sample is decreased to just a few dozen pitches at most, the variability shoots up and the model becomes worthless.

Second, these graphs are generalized to the population of all hitters. Every hitter is different, and pitchers have to attack them differently. Pitching only in spots where you can expect a low BABIP from the average hitter — like low and outside, for example — might work against some hitters, but not others, and pitching according to the hitter you’re facing is more important than pitching according to what the league average is.

And third, even though a very clear pattern exists for BABIP and wOBA, that pattern covers nearly the entire strike zone, save only some of the corners, and that is where the vast majority of balls in play are hit from. The expected wOBABIP for pitchers isn’t going to vary all that much, really.

I did, however, enjoy making and looking at these graphs. They provide a concrete, quantitative look at how the location of a pitch matters to the ball’s journey off the bat.





Jonah is a baseball analyst and Red Sox fan. He would like it if you followed him on Twitter @japemstein, but can't really do anything about it if you don't.

29 Comments
Oldest
Newest Most Voted
Robin
11 years ago

You forget to say “Batman” at the end

channelclemente
11 years ago
Reply to  Jonah Pemstein

Gotham’s favorite son. Beautiful work.

Adam
11 years ago

Fascinating graphs, and very well presented!

I was surprised at the higher pop-up percentages for outside pitches than inside pitches. The jammed pop-up to the pitcher/middle infielder is a staple of pitchers everywhere, and yet for whatever reason I’m struggling to call to mind what it looks like to pop up on a pitch away. Is that just me?

mattdecap
11 years ago
Reply to  Adam

The heatmaps are all from the righty perspective, so that is high and in, not high and away.

novaether
11 years ago

This is really cool stuff. I’m also saddened that there was no real correlation or predictive power that came out of this, though. Perhaps a similar analysis for batters might be more fruitful.

One quick question – is the “Distance from Center of Strike Zone vs. PU%” graph backwards? It looks like it’s saying that outside pitches generate more pop-ups, when you might think that the high and inside fastball would make pitching inside more favorable for fastballs.

I’d also be curious to see big differences for lefty and righty pitchers, as well as pitches to lefty and righty batters.

Andrew
11 years ago

Following up on Adam’s comment, on the PU by inside/outside, the chart is clearly HEAVILY weighted to the outside, but then on the 2D plot is shows weighted to the inside.

I suspect one of your charts there is corrupted with bad application of data.

BM
11 years ago
Reply to  Andrew

I too am wondering if this is the case, but then also wouldn’t you probably expect that the BABIP chart is also backwards? After all the author comes right out and says that PU% is the most important for BABIP, so you would expect the high pop-up area to be the low BABIP area.

Since the charts seem to agree, the implication is that either both are off or neither one is. At first glance it seems like a sign error on the x-axis to me, and that makes intuitive sense because sign errors are fairly common. Let me know if I’m wrong.

BM
11 years ago

With caveats about the possible error, this does conform to the “effective velocity” hypothesis. Of course the whole thing falls apart if, when you add walks back in, you’re allowing an OBP over the league average to go along with all those fly balls that generate an ISO over the league average (here’s looking at you, Trevor Bauer). If only pitching exclusively to low/outside and up/inside was easy.

ncgostl
11 years ago

Very cool stuff. Thanks for writing it up.

When you introduce the horizontal location of the pitch, you write ” The distance here, too, is adjusted, so righties and lefties are on the same scale — a positive value is always a pitch farther outside, and a negative value is always a pitch farther outside.”

Is the positive value outside and negative value inside?

Adam
11 years ago
Reply to  Jonah Pemstein

But now it says that a negative value is outside and a positive value is inside, which makes no sense given that the label on the x-axis says “distance outside from the center of the strike zone”…

jdbolickMember since 2024
11 years ago

Awesome work. Kudos.

NATS Fan
11 years ago

Perhaps BABIP is as much of a function of team defense as it is the pitch/pitcher. Perhaps adding in defender locations and ranges would make this all clearer.

Damaso
11 years ago

maybe BABIP is effected most of all by quality of scouting reports on hitters and effectively exploiting hitters’ weaknesses consistently?

how hard would it be to cross reference pitch selection for each hitter with those individual hitters’ heat maps?

Brian
11 years ago

This is fantastic work. It’s obvious, in the sense that it is intuitive, but that doesn’t mean that it’s not also great. It is worth proving things that we already know but have not gone through rigorous testing of.

As for what this data can be used for … I mean, this certainly implies that the ability to control BABIP exists. If a pitcher can live in those areas where there are low BABIPS, high and low in the zone, or away from hitters where GB% is high, they by definition will give themselves a chance to beat luck.

I can’t wait to see how this data is used in the future.

obsessivegiantscompulsive
11 years ago

Nice article and nice charts!

Perhaps you can fix this but it has been bothering me lately that the mantra is drilled into us that pitchers have negligible effect on BABIP. And yet, there are pitchers who defy the hegemony that is DIPS.

Tom Tippett showed this long ago that there are pitchers who succeed in spite of what DIPS says, and he showed a variety of types, like crafty lefties and others, who succeed in spite of DIPS pronouncement. Matt Cain, until his downturn that started with his Perfect Game, had a pronounced much lower BABIP than the .300 mean everyone is suppose to regress to, for example. And there have been others like Barry Zito, who passed the 7 season threshold for data necessary to statistically show a pronounced lower BABIP than the mean.

Why don’t anyone do something like you did here, taking the classes of, say, Tippett’s categories, and seeing how these pitcher defy DIPS? (well, except for knuckleballers, we know how they do it) See how crafty lefties do their thing. See how Matt Cain does his thing. Fangraphs famously had a series of articles discussing this Cain conundrum, and published a study showing that it was not just Cain allowing below 10% HR/FB but the entire pitching staff, maybe that would be a nice thing to study with your analysis methodology here, all the Giants pitchers and how they suppress HR/FB. Wouldn’t it be much more useful and interesting to find out how pitchers are able to defy DIPS expectations?

And the thing is, Mike Fast of the Houston Astros but prior to that prominent saber, said in an interview that BABIP is not a mean in the minors, that pitchers in the minors do have control (or perhaps better to say, lack of control) over batted balls, but that the ones who make the majors reach a certain level of skill of preventing hits. http://whattheheckbobby.blogspot.com/2012/12/an-interview-with-astros-analyst-mike.html

Here’s a couple of other interesting articles:
http://baseballanalysts.com/archives/2009/07/can_pitchers_co.php
http://www.crawfishboxes.com/2014/9/15/6143809/sabermetrics-looking-for-weak-contact

So, to me, there seems to be some sort of understanding that while most pitchers can’t control BABIP, there are some who can, and yet it seems like the vast majority of our best saber brain power is devoted to understanding something that yields no advantage to either the pitcher or hitter, DIPS, and yet there are clearly pitchers who break the mode and people shrug their shoulders and basically say that they are the exceptions that prove the rule.

Perhaps it is a matter of most pitchers not knowing what the few know. Just like the ones who learn how to prevent hits well enough to stick in majors are separated from the ones who don’t in the minors, maybe the exceptions in the majors know something that the rest don’t know, or rather, they happen to learn at some point the key to doing what they do, without knowing that they were learning the key. What if the exceptions are analyzed to figure out how they do it, and then those key lessons are taught to other pitchers?

High and TIght
11 years ago

From a pitching perspective, it would seem that pitchers should really develop a consistent high and tight pitch that is in the strike zone. If that were coupled with a low and outside pitch, I think that batters would be consistently flummoxed.

highrent
11 years ago

I think you have something here. You said you looked at movement. What about adding movement or even velocity. I would imagine there isn’t a super strong correlation with position otherwise pitchers would be always pitching to the same spots as you said. But adding movement to it may yield something. Ultimately I think its the whole package. however this is also the reason weak contact is so hard to consistently achieve because its a myriad of factors not just where you throw it. Its not an easy thing to quantify. Keep at it.

Corey
11 years ago

This seems obvious, but I didn’t see anything in the article to suggest you did. Did you control for batter handedness? If you didn’t lefties following the same pattern as righties would corrupt much of your data, which means some of those correlations to pitch location and outcome are even stronger. That stronger relationship between pitch type and outcome could give you a better predictive relationship to the pitcher’s actual babip. You probably did, but I thought I should mention it since it didn’t seem clear to me that you did.

Corey
11 years ago
Reply to  Jonah Pemstein

I figured you must have, I just didn’t see it. Nice work, sorry it wasn’t a little more fruitful, I’m very intrigued by babips.

Craig Wall
11 years ago

Awesome work, Jonah. It sure underscores the premium value of command in any pitcher. Most pitchers aren’t trying to throw “meatballs” down the middle but they miss!

We’ve seen correlations of velo and FIP. Perhaps it’s time to see correlations between command and FIP. I concede that command inside the strike zone and it’s correlational value makes more a much more complex study.

pft
11 years ago

“probably because pitchers have a negligible amount of control over it”

I still don’t agree with this. Pitchers can control the quality of contact with location, stuff and pitch sequencing. How hard a batted ball is hit goes a long way to determining if the ball is a hit. Maybe the new data coming out of statquest will help as the batted ball types classified by interns is likely riddled with inconsistency

glib
11 years ago

excellent work, congratulations.

Brad McKay
11 years ago

Major contribution here. Really appreciate it.