What Hard-Hit Rate Means for Batters
Recently, one of the hot topics in baseball statistics has been the appearance of a measurement for hard-hit balls: here at FanGraphs, we added hard-hit rate to our leaderboards before this season, adding along with it a wealth of opportunities for analysis. An issue with any new statistic is that it can be cited without fully knowing its true use or impacts, and so hard-hit rate has been making the rounds in player analysis, generally cited in respect to how well or how poorly they have been performing.
For hitters, it might go without saying that hitting the ball harder is generally a good thing: the aim of hitting, in a certain sense, would seem to be to hit the ball as hard as possible as often as you can (except in the cases of bunting or other situational circumstances). However, it hasn’t been clear yet how hitting the ball hard impacts other rate and counting statistics, and that seems to be a hole in our understanding of a statistic that is undergoing a moment in the spotlight.
The aim today is, at the very least, to explore how hard-hit rate impacts a few of those stats, as well as to begin a conversation that more astute statistical minds may be able to take to deeper and exciting places. There are a couple levels to this piece today, but there are surely many more that I have not reached: I don’t intend to make hard conclusions, but rather to explore and provide a well-intentioned foray into the data. With that said, onward.
To begin, we should remind ourselves of some research that has already been done on these digital pages: year-to-year correlation of hard-hit rate among pitchers and batters was the subject of this piece. A brief summary of the results from that: hard-hit rate seems to be a skill for batters, but not so much for pitchers. That’s good news for us today, as we’ll strictly be looking at hitting statistics. Tony Blegino’s treatises on batted-ball data should also be seen as a preface.
The Results
Today we’ll be looking at all of the hard-hit data we have before this year: 2002-2014. One thing that should be noted: hard-hit rate was graded visually prior to 2010, whereas afterward it was graded by batted-ball type, hangtime, and distance. Further studies might look into the possibility of changes between pre-2010 data and post-2010 data, but I’m currently assuming they’re about on the same level. The sample size we’ll be looking at is qualified batters for each season. Because of the large sample (almost 2,000 data points), we should understand that our p-values associated with these models will be very low. We should primarily look to R-squared values, listed on the charts, to analyze the strength of the correlations.
We’ll start with HR/FB%. Let’s take a look at the relationship between hard-hit rate and the percentage of fly balls that go for home runs:
There is certainly something here, even though our r-squared shows that using hard-hit rate as a predictive statistic for increased home run/fly ball rate could be problematic. Another interesting thing to note: we could fit an exponential regression onto this dataset to slightly increase our predictive ability, but that also brings with it a few other issues — I went with a simple model instead. In case you were wondering, those three outliers at the top of the scatter are, from left to right, Jim Thome (2002) and Ryan Howard (2006 & 2007).
Next up we have ISO. Let’s check out the scatter of ISO vs. hard-hit rate:
Again, we have a significant relationship here, this time between an increased hard-hit rate and increased ISO. However, we can also see that it is again going to be problematic to use the relationship as a predictive tool, which comes up often when looking at one baseball statistic in relation to another. Just like it is difficult to use a singular statistic as a way of measuring a player’s entire worth (the reason we have metrics), there are simply too many variables beyond how hard a hitter hits a ball to adopt it as a concrete way of predicting outcomes in relation to another statistic. That doesn’t mean we can’t identify relationships that exist, however, as one seems to here.
Now let’s try a more traditional statistic, and one that is at least related to ISO: slugging. Does increased hard-hit rate correlate to increased slugging?
The explained variance is lower than ISO, as we might expect, as slugging is a noisier statistic than ISO. Still, there’s a relationship, and this gives us some hope that we might find at least some correlation to hard-hit rate for other traditional statistics in further studies.
Finally, we’ll look at a catch all, and the correlation that might be the most interesting to us when evaluating overall offensive performance in relation to hard-hit rate: wRC+. Does hitting the ball hard more often lead to higher overall offensive performance?
Once again we find that there is a relationship, but only just under 40% of the variance is explained by our model. Given the large sample and type of data, expecting very high R-squared values is probably not the best hope for us, and after staring at this for many hours, I’m happier and happier with a marginal victory between hard-hit rate and a metric that measures a larger idea of performance.
A final one, mostly for the sake of doing so: this pertains to the belief that hard-hit rate might mean that hitters are showing a propensity toward having more line drives in their batted-ball profile. Is that true? Let’s take a look:
Someone took a shotgun to the chart. Hitters with under 15% hard-hit rates can post line-drive rates of 26%, just as hitters with 45% hard-hit rates can post line drive rates of 17%. There seems to be effectively zero predictive value to using hard-hit rate in this way. Although it might seem obvious, hard-hit rate should not be confused with line-drive rate in any way, as ground balls and fly balls can be hit hard, just as line drives can be hit softly.
A Final Issue
One point that has to be included when looking into the data we have is the high variability in league averages between different years of hard-hit rate. We can see jumps of almost 6 percentage points in back-to-back years, making me wonder whether a league adjustment could be beneficial due to the possibility of outside influences. Take a look at the difference in the jumps in hard-hit rate between seasons for our data set compared to line-drive rate (with scale adjusted for proper comparison):
Line-drive rate, much like other rate stats we use, is fairly stable between years; hard-hit rate does not seem to share that trait (at least in the data we have). With that in mind, I’ve performed a league adjustment for each season of both hard-hit rate and the statistics/metrics we’ve compared it to (it does not include park adjustments), to effectively create “Hard-Hit Rate+”, “ISO+”, “HR/FB%+”, etc. for each player’s individual seasons.
This, though it may be quick and dirty, puts each season’s performance in the frame of the rest of the league for that particular year. I then reran our regression models for each comparison with this data to find out if our correlations would get stronger. In the interests of saving space and not including each scatter plot again, I’ve created a table with each R-squared value of both the non-adjusted and adjusted data sets. Here are the findings:
| Non-Adjusted R-squared | League Adjusted R-squared | |
|---|---|---|
| HR/FB% | .46 | .64 |
| ISO | .49 | .68 |
| Slugging | .42 | .60 |
| wRC+ | .38 | .48 |
Here we see much stronger correlations, with the percent of the variance in our data explained rising by significant margins. We can still debate the level at which we want to accept our R-squared values as significant, but smoothing out some of those large fluctuations inherent in year-to-year hard-hit rate data seemed to provide us with better final correlations. A lot of baseball data is inherently noisy, with large and random variance: this makes it fundamentally difficult to apply statistical rules to it that we might use for data sets in other fields. All in all, I’m actually surprised at the moderate strength of how hard-hit data correlates to other statistics, given the large sample size.
There is certainly room for more study with respect to how hard-hit rate influences other aspects of the offensive game. I hope this preliminary foray — this balestra, if you will — has provided at the very least food for thought and discussion. If I have made any statistical errors in my analysis, I apologize, and know that it was with the purest of intentions that I set out to look into this topic. Given the overlaps and noise inherent to comparing sets of baseball statistics and data, we’ve found some interesting and meaningful correlations here. Preliminarily, we know that hard-hit rate’s impacts may be what we expected, and perhaps what we didn’t as well.
Owen Watson writes for FanGraphs and The Hardball Times. Follow him on Twitter @ohwatson.






hard hit and wOBA?
For wOBA, looks like the non-league adjusted correlation strength would be R2=.35, with league adjustment bumping R2 up to .47.
Any discussion on why a nonlinear fit looks like it might clearly be superior.
Yes, I touched on it a few times here in the comments.
here is a link to the r values for more stats: http://s8.postimg.org/pvo1chk05/Screen_Shot_2015_06_03_at_11_31_09_AM.png
The pre v. post 2010 stuff would be interesting. From a cursory look at the split season leader board, a few things pop out. David Ortiz has several top seasons both pre and post 2010. Miguel Cabrera, on the other hand, has 4 seasons in the top 30, all from post 2010.
I had been playing around with the 2015 batted-ball velocity data, and using this with batted-ball type to estimate batting runs resulted in an R-squared just under 40% as well.
I think the key goes back to Tony B’s first article on this. For both fly balls and line drives, if you graph results vs. batted-ball velocity there is a large section where hitting the ball harder produces worse results – namely balls that would have fallen in front of an outfielder will instead carry further, increasing the chance for an OF putout.
This area might be the biggest component of variance in batted-ball “luck”, where hitters lose out from squaring up better on the ball.
I am not surprised by the lack of correlation from hard hit and line drive percentage. If you are swinging hard every time you are going to miss sometimes.
A lot of those charts look like they could have second order relationships worth looking at.
hard hit% and BABIP?
I wondered about this. Turned out it looked the same as the LD% plot, just a total mess.
I was hoping to see that too!
Here it is! A little better than the LD% scatter, but still not much there.
Wow I expected that to have some correlation…
I wouldn’t expect much correlation between FB or LD BABIP and hard hit rate, but you’d think looking at BABIP off GB only to hardhit rate would give a pretty strong correlation.
But maybe not. Maybe all the balls poked the other way through the gap (especially on shifts) erode that.
Thank you, Owen.
I wonder what the result would be if you filtered out hitters with high ground ball rates. By removing them, would you get more of a correlation between the hard-hit rate and actual results, or a higher BABIP?
Is it just me, or do the first three graphs (HR/GB%, ISO%, SLG%) look like a linear fit isn’t ideal? The data seems to follow a more parabolic approach (low curvature, but still parabolic).
I had the same thought. It seems like there is small exponential growth going on. It seems a bit weird, but I think it may have to do with skill versus plate approach. Bad hitters can swing super hard every at bat and have a great hard% while sacrificing contact. But a hitter with a lot of talent and high bat speed is simply going to have a larger hard% than the rest without having a poor approach.
I’d be interested in looking at Hard% data along with contact rates or strikeout rates, so we can weed out the guys who are simply swinging hard at every pitch but who are not very skilled hitters.
I did fit an exponential line onto the HR/FB% correlation when running it, as I said above, which decreased variance by a very tiny amount (about .01). It also ballooned the standard error, so I ended up going with the simple, linear model.
At the very top end of the correlations (high HH%/high dependent variable), I think you’re right, but the bulk of the data seems to follow a linear trend.
It’s a good article, and upon reading your fit comments, it’s hard to argue, but the residuals just look wrong for that to be simple linear. If you segment the data and think two fits, then of course you have to suspect two models, right.
I’m curious how changes in hard hit rates for individual players correlate to these statistics. I high wOBA hitters tend to generally have higher bat speeds hit the ball harder which may explain a lot of the correlation we are seeing. But I wonder if we see large increases or decreases in an individuals hard hit percentage if that will be more or less predictive of wOBA.
Since LD% and HardHit% seem to be independent, does using both of them go further in explaining the hitting stats mentioned in the article? Especially LD% and Adjusted HardHit%.
I have a couple questions regarding the soft/med/hard hit rates. Do those numbers include bunts? Is there a total # of soft/med/hard hit balls instead of %? I’d like to look at these on a per AB ratio and would like to incorporate K’s or “no hit” AB’s.
For instance, before you incorporate K’s, Stanton has the highest Hard Hit % (49.2%) and Posey (34%) is about 50th, after I figure in K’s or “no Hits” (based on BA minus K’s to get approximate #’s) then Stanton’s Hard Hit % per AB drops almost 18 pts to 31.6% (12th) and Posey, who rarely K’s, only drops about 3.5 pts to 30.7% (16th). The highest Hard Hit per AB is actually Tulowitzki at 35.62%.
While Stanton’s original #’s makes it seem like he hits the ball hard half of the time, its actually less than a third. And if you don’t put a ball in play you have 0% chance of getting on base (minus odd plays, BB and HBP).
I believe they do include bunts. As far as Ks, the soft/med/hard only measures balls in play, so strikeouts are excluded. Doing a hard hit rate per PA would be a great idea!
I suspect the relationships would look different if you included every swing in the sample. I think you would see stronger relationships between hard hit % and wOBA/wRC+ if all swings are included.
Thanks Owen! So I would need to subtract bunts from AB’s
I’d be curious to see how predictive hard hit rate is of say ISO+ in the following year. So, if we had a set of batters with equal ISO+ in Y1, and split them based on their hard hit rate, would the batters with the higher hard hit rate have higher ISO+ in Y2?
Are infield flies included in hard hit %?
Most of them would probably fall in the soft or med category, but theoretically a batter could have a hard hit infield fly, if its apex was extremely high. Just a hunch, though.
So perhaps this correlates with Swinging Strike Percentage, and is indicative of a hitter’s approach? So the Steven Souzas of the world are just fine whiffing on pitches as long as they make hard contact, and the Ichiros are more about placing the bat on the ball at any cost. And then of course most hitters doing something in between. I wonder if certain hitting approaches are more optimal depending on the pitching style.
I hate to throw a wet rag on this, but I think all this analysis is pretty much a waste of time, given that the very important vertical launch angle is missing. Batted ball speed alone produces a very incomplete picture of the batted ball. To complete the picture, the vertical launch angle is also needed. There was a lot of good analysis done using the speed/angle combination for the April 2009 HITf/x data,including BABIP, wOBA, HR probability, SLG, etc. I personally learned a lot from those analyses. To learn more from StatCast will require getting the angle information. At least, that is my point of view.
Well, this article wasn’t really analyzing speed/angle combination. It was more looking at hard-hit rates by themselves. The author came to the same conclusion you must have; that hard-hit ball data, by itself, provides some information but not enough to really make any reasonable conclusions re: how well that particular batter hit.
There’s room to look into vertical angles and their impact on the stats Owen looked at here, but that wasn’t the focus of this piece. That subject is for another time.
It says that the hard/medium/soft categories have been determined by a combination of batted-ball type, hang time, and distance since 2010, although it doesn’t give the cutoffs for each category. I’m thinking you could back into the vertical launch angle (roughly) from the hang time and distance, making some simplifying assumptions on spin, wind speed, etc.
I’d love to know more about those analyses from the April 2009 HITf/x data. Are those in THT or Baseball Prospectus?
Here are some:
http://www.hardballtimes.com/using-hitf-x-to-measure-skill/
http://www.baseballprospectus.com/article.php?articleid=15532
http://www.baseballprospectus.com/article.php?articleid=15562
Also, here is a link to a talk I gave at the Saberseminar:
http://baseball.physics.illinois.edu/ppt/Saberseminar2013-v2.pptx
Let me take back my “waste of time” comment, for which I apologize. That is much too harsh. However, if we ever get the full batted ball data from StatCast (batted ball speed, launch angles, hang time, landing point, batted ball spin), the kind of analysis discussed in this article (and in many others that have appeared recently) can be done much, much better (and similar to the kind of analysis in the links I provided).
Thanks Dr. Nathan – some very cool stuff there!
I think we’re in a bit of an awkward phase between when contact quality information was not widely available to the public and the current situation, where there’s easily accessible bits of that info on Fangraphs, Baseball Savant, etc. So I think articles like Owen’s can serve a role in showing what some of the “partial” metrics like hard-hit % can and cannot do. But, I agree we need to gear up for the type of analysis that can leverage all the great new data that’s starting to accumulate on batted balls.
This is great – nice work. I’d expect the relationship between hard hit% and some of the outcomes you’ve modeled to be non-linear (and all the models you’ve chosen are linear). This is because I suspect (but obviously don’t “know”) that there is a threshold at which hard hit% makes a huge and positive difference in power/other outcomes much as there would be for soft % (though negative). I don’t know where this threshold might be, but testing out some other models (exponential, Weibull, Poissin, Cox, or maybe even a Bayesian model) might yield some more valuable insight.
As I said in a previous comment, I did mess around with an exponential model, and it decreased variance very slightly (at least for HR/FB%). It also greatly increased standard error (the reason why I stuck with linear), mostly because I think the relationship is linear for the bulk of the data, up until the confluence of high values of HH% and the dependent variable, when it seems to go a little exponential.
Sure, but there are many other options to model that relationship, all of which have tradeoffs in terms of accuracy/precision compared to a linear one. Just a thought for future research.
Definitely! Certainly something I’ll keep in mind if/when I revisit this. Thanks!
Do hitters show any tendencies to hit flyballs hard and groundballs not so hard? (or vice versa) I’d think a flyball puller would show that.
If so, it might take out some noise to plot HR% versus FlyHard%, since hard groundballs won’t be home runs.
Good timing, I took a look earlier at Hard Hit% and Contact rate for those with 100 or more PA:
R^2 values aren’t as high as you’re getting with these other looks, but I’m also only using data from 2015.
I think you have to keep hard hits separated by LD, FB and GB since they will all have different outcomes, and am still not very clear if the classification of hard/medium and soft is reasonable or not. Breaking it down by batted ball velocity should remove the subjective and arbitrary nature of the classifications and allow more detailed analysis, but MLB has not released anything in a usable format, and I am not sure they will
tz asked for links to analysis of HITf/x data. Here are some:
Peter Jensen: http://www.hardballtimes.com/using-hitf-x-to-measure-skill/
Mike Fast: http://www.baseballprospectus.com/article.php?articleid=15532
Mike Fast: http://www.baseballprospectus.com/article.php?articleid=15562
And here is a link to a talk I gave on the subject:
http://baseball.physics.illinois.edu/ppt/Saberseminar2013-v2.pptx
The recent articles by Tony Blengino seem to be based on HITf/x data that he has access to.
Let me take back my “waste of time” comment, for which I apologize. It is much too harsh. Nevertheless, StatCast gives us the ability to do this type of analysis (as well as other analyses that have appeared recently) much, much better (like the ones I have linked to). The data include not only batted ball speed but also launch angles, hang time, landing location, and spin. Hopefully the data will become public.
Thanks, Alan! Agreed on all counts. Fingers crossed on us getting more in-depth data soon.
There has to be significant classification bias introduced by these type of statistics. What exactly defines a line drive or a hard hit ball? How consistent are the people that collect this data? I’d love to see some of these correlations but using batted ball velocity instead.
This is exactly what I was thinking. What is the rationale for classifying soft/medium/hard instead of using mean velocity?
It’s amazing how baseball can keep on getting more and more statistics to analyze deeper and deeper.
I’m interested why you chose to do the analysis with wRC+. Of the stats you choose to analyze, I expected wRC+ to have the worst correlation as it takes into account BB and HBP in calculating it. Could you use a modified wRC+ that takes out BB and HBP inputs?
Who is the outlier at 40% HR/FB in the first graph?
I don’t understand any of this, but I know that the Royals hit the ball hard last year and made the WS, best team in baseball IMO
If you do a multivar regression of AVG HR and FB distance combined with Hard Hit% to ISO or wOBA, you get an even higher corr (.65 for wOBA).
Here is a link to the leaders for expected wOBA based off of Avg HR and FB distance and Hard Hit%: http://s10.postimg.org/skm1nlg4p/Screen_Shot_2015_06_03_at_11_37_11_AM.png
Spelled Blengino wrong. I know, who cares about spelling when you have 47 pretty graphs.