Confessions of a Baseball Analytics Writer

© Steven Branscombe-USA TODAY Sports

Jack Leiter will always have a special place in my heart. The Rangers’ top pitching prospect was the subject of the very first article I wrote for FanGraphs, which talked about, among other things, the unbelievable carry on his fastball and how it could lead him to big-league success. But we haven’t checked in on Leiter in a while, and well, his Double-A numbers have been ghastly: a 6.24 ERA in 53.1 innings pitched has somewhat muted the hype surrounding the righty. Though it doesn’t really change our outlook on Leiter, it’s still unsettling to see.

Part of that has been his inability to throw strikes, as Leiter is issuing well over five walks per nine innings. But more importantly, Leiter has lost a significant amount of his signature fastball ride in pro ball. Statcast data was available for this year’s Futures Game, during which Leiter’s dozen or so fastballs averaged 16.1 inches of vertical break – a far cry from the 19.9 inches I calculated in that debut article using TrackMan data. It could be a small sample quirk, and yet, the general industry consensus is that Leiter’s fastball is no longer transcendent. That’s a genuine problem.

What might the reason be? Maybe Vanderbilt’s TrackMan device wasn’t properly calibrated (as suggested by Mason McRae), leading to imprecise readings. But if that’s true (and maybe it isn’t), how could we verify it? What I came up with this: Using velocity, spin rate, and spin axis data from the 2021 NCAA Division-I baseball season, I built a model that estimates the vertical break of four-seam fastballs from righty pitchers. Once completed, I grouped the data by the pitcher’s team and looked at which schools over- or under-shot the model. Those with the largest residuals, in theory, are prime suspects for having miscalibrated TrackMan devices.

We have some evidence here. Among the schools with at least 2,000 righty fastballs in the database, Vanderbilt ranks ninth out of 48 in the average difference between actual and expected vertical break. As for Leiter himself? Across a not-so-small sample of 721 heaters, he generated 2.5 inches of extra ride over expected, which puts him squarely outside the confidence interval. It could also be that Leiter just isn’t throwing his fastball like he used to, but it does seem like TrackMan data had a hand in sweetening his statistical profile.

Even with diminished ride, Leiter’s fastball is still a plus pitch, and on the whole, he’s still one heck of a pitching prospect. But even in an era of sophisticated data, inaccuracies can be surprisingly common. TrackMan devices are operated and maintained by humans, after all, and to err is human. While having the requisite data remains incredibly helpful, a healthy dose of skepticism – and subsequent adjustments, such as removing outliers – goes a long way in making the most of it.

***

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

Making sure we aren’t being misled by the data is one thing. Deciding how to represent and communicate it is another. Lately, I’ve been writing a lot of articles about pitching, and a few of the comments expressed confusion over how pitch movement is indicated. As if baseball isn’t complicated enough, there is indeed more than one way to accomplish a seemingly simple task.

Because life is short and precious, here are the Cliff Notes. My preference is what’s known as “short-form” movement, or the expression of pitch movement relative to a pitch with zero spin-induced movement. Fastballs “rise” relative to that designated origin point, while breaking balls drop instead. Short-form movement reflects how hitters actually perceive pitches, as the illusion of rise is what beguiles them into swinging under high-spin heaters. It also creates a clear distinction between pitch types and how they behave, preventing us from mistaking changeups for sliders, for example. Short-form movement is what you’ll see on Baseball Prospectus (including Brooks Baseball) and this very site.

Then there’s “long-form” movement, which reflects how pitches move in real life. Fastballs still drop, but much less compared to breaking balls. This is what you’ll find over at Baseball Savant. I assume folks get confused because popular sites are using different methods of representing pitch movement, which is beyond understandable. But wait, there are even two types of short-form movement! The first, which comes courtesy of PITCHf/x, is measured 40 feet from home plate. The second, which comes courtesy of Statcast, is based on the entire flight path: 60.5 feet, minus the pitcher’s extension. They’re functionally the same, but one produces higher movement numbers than the other. More to the point, it makes our head hurt.

Life would be much easier if we could all agree on a single measurement, but given the sport we’ve chosen to arduously follow – how can you not be pedantic about baseball? – that’s probably not happening anytime soon. It’s not just pitch movement that’s drowning in semantics: Baseball’s trendiest breaking pitch is widely known as a “sweeper,” but in Yankee-land, it’s better known as a “whirly.” Spin efficiency (Rapsodo) is active spin (Baseball Savant), but some analysts take offense to the former, which implies the higher the efficiency, the better. Meanwhile, spin direction and spin axis are two entirely different things, but that’s scantly explained, so even smart writers will end up using them interchangeably.

Admittedly, I’m also part of the problem. On occasion, I’ll flip-flop between short- and long-term movement depending on what’s more convenient, in addition to omitting explanations that I assume just aren’t necessary. The truth is, there might be thousands of fans who aren’t as well-versed in baseball analytics as you think. It’s our responsibility, then, to make sure they’re accounted for.

***

As a FanGraphs contributor, there’s a certain amount of pressure to get things right, given the site’s reputation and amount of traffic. It doesn’t dawn on me like it used to, thankfully, but it’s still there in the back of my mind. Not that it’s a major issue – if you care about what you do, I think feeling at least a bit ashamed of a notable mistake is inevitable.

But you learn not to let those moments get a hold of you. You also learn that they present great opportunities to improve as a writer and an analyst. Earlier this month, I wrote about this season’s most and least consistent hitters, as determined by a series of calculations that I sufficiently explained and justified… or so I thought. Much to my dismay, someone in the comments pointed out that I had failed to normalize the hitters’ standard deviations in wRC+ based on their mean wRC+. Not doing so created a positive relationship between the two variables, from which many of the article’s conclusions were drawn. Ouch.

After review, I realized that, yes, I had made a fairly big mistake. There’s not much use in starting over with a new article, but I can make up for it here. First, below are the most consistent hitters, as of that writing, according to the normalized standard deviation in wRC+ (that’s regular standard deviation divided by mean wRC+, aka the coefficient of variation):

The Kings of Consistency, Revisited
Hitter Normalized Std. Dev. Mean wRC+
Patrick Wisdom 0.27 130.7
Pete Alonso 0.27 158.3
Ian Happ 0.34 127.5
Wilmer Flores 0.36 108.7
Adolis García 0.40 156.9

Next, here are the least consistent hitters:

The Finicky Bunch, Revisited
Hitter Normalized Std. Dev. Mean wRC+
Adam Duvall 1.42 54.1
Myles Straw 1.36 52.3
Javier Báez 1.33 72.0
Jorge Mateo 1.19 52.0
Owen Miller 1.10 98.3

There is some overlap: Alonso, Flores, and Wisdom are still in the top three in terms of consistency, and Miller remains mysteriously mercurial. Based on how many of the consistent hitters from last time have stuck around, much of where the normalization has played a role is in distinguishing actual streakiness from mere variance. Indeed, you’ll see that the most inconsistent list is no longer a list of the greatest hitters, which in retrospect didn’t make a whole lot of sense.

Still, adjusted standard deviation has a moderate correlation with overall wRC+, which suggests that good hitters really do tend to produce through streaks of brilliance. What Alonso and Co. are accomplishing remains special, albeit to a lesser extent. The correlation between standard deviation and strikeout rate is no longer nonexistent, but it’s weak enough to the point where it doesn’t warrant discussion. Case in point: Wisdom and Duvall, who occupy opposite ends of the consistency spectrum, are number one and three in strikeout rate respectively.

The takeaways aren’t dramatically different, but the names sure are. I’m disappointed for not having been more vigilant about how I presented the data before filing the article, but what’s done is done, and there’s this little follow-up to address what went wrong. While it would have been easier to ignore it altogether, I owe it to whoever is reading my work to be honest and self-reflective. After all, nobody wants to follow an analyst who pretends they’re right all the time.





Justin is an undergraduate student at Washington University in St. Louis studying statistics and writing.

12 Comments
Oldest
Newest Most Voted
vslykeMember since 2020
4 years ago

Thanks for this piece, I plan on referring back to the discussion of the different ways to measure movement frequently.

Anon21Member since 2018
4 years ago

how can you not be pedantic about baseball?

Excellent.

I’ll admit, my eyes glaze over whenever I start reading about pitching analytics. It just seems too technical to grasp, so I’ve never really tried to grasp it.

Paul-SF
4 years ago

“Making sure we aren’t being mislead by the data is one thing.” Making sure we aren’t being misled by homonyms is another.

FrancoeursteinMember since 2025
4 years ago
Reply to  Paul-SF

You dirty dog

mikejuntMember
4 years ago

This is a great piece and it reminds me (as a number of things do) of my discontent with the direction of defensive metrics, on FG and in general. FG has switched to StatCast OAA, which is useful in the context of ‘which player gets to more balls’, and not useful in the context of ‘which teams make outs more efficiently’. I think using OAA in player-war is acceptable (though it has flaws), but it’s unfortunate that it’s involved in team WAR totals and frequently used in our analysis of team defense. A team can have individually talented defenders and not have an especially effective defense. It can have mostly average or middling defenders and have a very effective defense. Because OAA is neutral about the question of ‘where did the fielder originally start’, it penalizes players who are positioned to make difficult outs into easy plays, and it can even overly reward an extremely talented defender who is routinely positioned badly; this 2nd player is provided with a lot more ‘difficult’ plays, and because of their talent converts more of them than their peers – accumulating a high OAA. But they may still make less total outs on those balls than a better positioned, average defender.

This always stands out to me about the Dodgers, who are of course the team I follow most closely. Since the start of 2017, the Dodgers as a team have allowed an opposing BABIP of .271. That’s a crazy number, with 2nd place coming in at .280. And nearly all the teams at .285 and below are known for their analytical bent and not necessarily for the talent of their defenders (Astros, Rays, Athletics). Yet they routinely rate out as average or middling by OAA, because the individual defenders are not that talented. Whatever they are doing isn’t captured by the metrics, and its making a collection of mostly average to questionable defenders into the most efficient defense in baseball – for half a decade!

But we often easily shift our analysis of teams to doing analysis of individual players, and so it’s far more common to see the Dodgers (and other teams on that list) described as having indifferent or average defenses, when they are in fact the most efficient defenses in baseball. We have disconnected value from the results by focusing tightly on the apples to apples comparison between individuals, and in doing so we’ve lost any sort of context for the overall outcome as it relates to actual hits and runs.

We do some of this with pitching (via FIP), but there we do it *because we know FIP is more predictive than all but the most complex batted-ball alternatives*. We know that we’re sacrificing information, but we’re getting better forecasting for the exchange. OAA doesn’t do that for defense, and we’re worse off in analyzing team defense via OAA than we were using DRS and arguably even UZR despite it’s massive failings in shifted scenarios and their prevalence.

Last edited 4 years ago by mikejunt
Anon21Member since 2018
4 years ago
Reply to  mikejunt

This is a very Bill Jamesian critique of letting models override reality. (I mean that as a compliment.)

mikejuntMember
4 years ago
Reply to  Anon21

Its not that though, we’re measuring reality, but we’re skipping a step. OAA does a great job of measuring the differences in actual player talent and skill given the same opportunities.

But we then skip to giving those plays a run value without accounting for whether or not they had more or less easy/hard plays than usual because of where they started, and the fact that teams clearly do better than others in terms of where they locate people.

In the context of evaluating whether Carlos Correa or Corey Seager will turn more balls into outs for your team, OAA is useful. For evaluating whether the Dodgers have a better defense than the Rockies, it is highly misleading. This is expected, because OAA isn’t *intended* to measure that. But it’s what we have so it’s what we’re using it for.

The contrast I would draw is with DRS, which has introduced both a Shift Runs and a Plus/Minus Runs stat that exist *only at the team level* and ensure that the teams overall DRS mostly aligns with their actual BIP supression abilities. OAA needs a ‘location runs’ element in addition to its range, success rate, etc measurements in order to be useful for evaluation at a team level. These decisions aren’t made by individuals in modern baseball, and it *shouldn’t* be credited to the players – but it’s undoubtedly effecting the game overall.

And it’s especially egregious for this particular website, where we have chosen to use FIP as the primary input for pitcher value. I personally like FIP a lot, because it is the most future-predictive, simple pitching statistic. But choosing to use it means you are giving pitchers *zero credit* for what happens on balls in play. This means you need to give *all the credit* to their defenders. As we are right now, there’s a huge swath of value that we’re simply not trying to measure at all, because we’re dependent on MLB-owned StatCast metrics. These metrics are very focused on individual player evaluation only, and conspicuously avoid measuring team performance – IMO, this is intentional, because MLB is trying to avoid having outside sources ‘spoil’ the R&D investments of teams. If we look historically, this exact thing happened 12 or so years ago with catcher framing: Several teams were ahead of the curve and recognized this was valuable, and they invested in figuring out how to measure it and taking advantage of this knowledge that their competition lacked. Then a 3rd party media site (Baseball Prospectus) figured it out and broke it and it very rapidly became an industry consensus and the R&D investment made by those initial teams was completely wasted. By controlling not just the feed of Statcast data but how it is analyzed in public and provided to 3rd party sites, MLB is able to ensure it goes in a direction that is interesting for lay fans and useful during contract negotiations but also avoids stepping on the toes of teams spending 8 or 9 figures on baseball R&D.

There was a time in baseball when, if you read the right people on the right websites, and because many front offices were resistant to change and new ideas, that you could be as informed or better informed than people working in baseball operations about players and how they could be expected to perform. That isn’t true anymore, and it’s rapidly becoming less true because of the proprietary nature of Statcast data and the ways that Statcast analytics avoid stepping on the toes of team R&D.

Anyway, I thank you for your intended compliment but also I decline a bit. James has made these kinds of critiques of WAR but I think that his critique is primarily driven by James having a significantly different interest in what advanced analysis is for. Bill James is fundamentally a backwards-looking analyst. He wants to compare yesterdays’ great player to Babe Ruth, contextualize performance and help identify the greatest performances of all time. These are necessarily results-based (who won and didn’t win is fixed, and the result should tell you that the teams that won the most were the best and the people who contributed the most on those teams were the most valuable).

Modern analytics very quickly moved towards predictiveness. The first question I ask, and probably almost everyone here asks, when they see an abnormally good or bad performance is “What’s causing it? *Is it going to last?*”. An individual great game or a nice streak is interesting, but what we’re really constantly asking is ‘What do I expect from this player in the future? how does this recent information change that expectation?”

Bill James isn’t interested in that and he’s not interested in measurement systems that are focused on it. Context-neutral statistics are extremely useful, indeed, they are required, for predictability. They’re also useless for James’ interest in historical comparison and recognition, because whether past events are significant or not is determined in large part by their results. This drives James’ modern critique of “models over reality”, but the reality that matters is that James and almost all of the rest of the baseball analytics community (including teams) are interested in different questions. He wants to compare people to hall of famers, and we want to know if this guy whose played like a star for 3 months is going to be this way in the future or if this is a one-off career year for him. The inputs we’ve chosen (like FIP!) are future-focused. James doesn’t care about that, because he’s interested in observing baseball, not in making better decisions about who you do and don’t want on your team.

Last edited 4 years ago by mikejunt
D-WizMember since 2019
4 years ago

Great stuff, glad you revisited that consistency piece!

Brad Lipton
4 years ago

I am sort of struck by this sentence at the end: “While it would have been easier to ignore it altogether, I owe it to whoever is read my work to be honest and self-reflective.”

I am not talking about the obvious typo in “read” — we all have submitted things with this type of error.

I am talking about the sort of “acclaim” (or a similar word — I can’t get to it) that the author is expecting because he didn’t ignore the error and moved on…but acknowledged it, was contrite and corrected it.

You mention “to err is human” — that is known. But it shouldn’t be a noteworthy achievement for someone to fix their errors and not take the “easier” way out.

Or maybe it is noteworthy…

Ukranian to Vietnamese to French is back
4 years ago
Reply to  Brad Lipton

And he himself said: “Even if it were legally binding, I would read my work, but usually I am self-reflective. »

I don’t say akin to amnesty in “read” – we put the word “no” for that type of amnesty.

“We can’t do anything about it,” said Henderson. ovo sav but, chips se chupa.

You despise the “conditional” ones. But I wouldn’t be mistaken for a saleswoman for someone who has the right to reach a goal, who would come out as a “legislator” no matter what.

And yes, he is not considered a ruler…

LaBellaVitaMember since 2018
4 years ago

There is so much in this piece that it would require a very long paper with research to comment on it. So let me limit myself to the concern of “misinterpreting the data” and argue that the correction to the “consistent hitters” study is an example. 

To identify consistent versus inconsistent hitter, I think your original strategy makes sense, i.e., the use of standard deviation of wRC+ as a measure of consistency. However, maybe to the chagrin of too many readers, batters who tend to average more than 100 tend to have a higher standard deviation. Why is this wrong? It isn’t. It is just how we interpret the results. Does that make these players more inconsistent than batters who are in the 50-80 wRC+ range? Yes. The better players have games of several hits and games of nothing but outs, including Ks. Weaker players tend not to vary as much – more outs and very few great games.

To correct what you believe is a mismeasurement, you define consistency using a relative standard deviation. Not surprisingly, by dividing by larger numbers, better players end up on the lower end of inconsistency. But is this actually a more accurate way of mathematically defining consistency? I would argue no for the reason I stated above – better players have wider variance in their performance. Relative standard deviation doesn’t help us identify the players who are streaky (Brandon Lowe) versus those who are consistent (Brett Phillips). In fact, by dividing by the wRC+ value, one can end up with the absurd situation where a player generates a negative relative standard deviation. 

The reason we are in this situation is because of the definition of the metric. wRC+ has a floor. And many players hit that floor each game. Technically, it has a ceiling, but nobody reaches it so that spot is essentially equivalent to infinity. Thus it is a distribution that is not Gaussian / normal, but similar to average speed, which makes using a normalized standard deviation questionable, and it possesses a zero mark which has moved to the right. Thus the possibility of a negative relative standard deviation.

I would argue your original definition works just fine if you make sure it is read properly. It is okay to compare two players who of essentially equivalent productivity, i.e. Pete Alonso vs. Giancarlo Stanton, or César Hernández vs. Alec Bohm. But trying to compare across all players leads to the misinterpretation that we should all be concerned about. 

Last edited 4 years ago by LaBellaVita
ajamespowellMember since 2019
4 years ago

Appreciate the reflection on your writing in the article. I’m just a reader who has fun reading baseball articles and learning more about the game. Your reflection is how I constantly feel about my understanding of the game. Just keep learning. Thanks for your writing.