April Hitting Stats Mean Nothing… Except When They Kinda Do
As part of my exhausting shtick, I like to respond “April!” to questions in my chats involving player performances in the season’s early going. This is effective shorthand when someone wants to know if, say, George Springer is a bust because he’s put up a .480 OPS in his first two weeks in the majors. It’s also dead wrong. April stats, in their proper context, are meaningful.
“But Dan, a few weeks of baseball is a tiny sample!” That’s correct, but you have to take into consideration the underlying reasons projections can prove to be inaccurate. It’s not just that things change, though they do — pitcher X learns a sweet knuckle-curve or batter Y realizes that not hitting everything into the ground might be good — it’s that it’s challenging to gauge where players stand in the first place. Players’ stats themselves aren’t even perfect at this. Tim Anderson hit .322 in 2020, but that doesn’t actually mean his mean batting average projection should have been .322. We don’t actually know if a theoretical player was “truly” a .322 hitter, a .312 hitter who got lucky, a very unlucky .342 hitter, or a .252 hitter who made a deal with a supernatural or extraterrestrial entity. A .300 hitter isn’t observed, they’re inferred.
The way most, if not all, in-season projections (or any projections, really) function is by applying what we call Bayesian inference. We won’t get into a full-blown math class, but in essence, it simply means that we update our hypotheses to take new data into account. And for players, data comes in all the time: every pitch or swing of the bat is new information about a player. It’s valuable information, too, as only the last handful of seasons have much predictive value and recent performance is the most useful.
We only have about two weeks of baseball in the hopper, but projections for some players have already shifted by considerable margins. I’m using ZiPS here — it’s very convenient to use the projection system I own and run — but this will likely hold true for other projections that update in-season, though the magnitude of shift may vary:
| Name | RoS BA | Ros OBP | RoS SLG | RoS OPS | Preseason OPS | Difference |
|---|---|---|---|---|---|---|
| Byron Buxton | .268 | .312 | .532 | .844 | .796 | 0.048 |
| Yermín Mercedes | .260 | .318 | .437 | .756 | .712 | 0.044 |
| J.D. Martinez | .274 | .347 | .516 | .863 | .826 | 0.037 |
| Tyler Naquin | .254 | .295 | .446 | .741 | .709 | 0.032 |
| Chris Owings | .258 | .311 | .441 | .751 | .720 | 0.031 |
| Jonathan Lucroy | .258 | .318 | .382 | .699 | .673 | 0.026 |
| Cedric Mullins II | .256 | .307 | .404 | .711 | .686 | 0.025 |
| Nelson Cruz | .284 | .365 | .567 | .933 | .911 | 0.022 |
| Jared Walsh | .250 | .311 | .497 | .809 | .787 | 0.022 |
| Omar Narváez | .254 | .344 | .414 | .758 | .736 | 0.022 |
| Pablo Sandoval | .239 | .295 | .404 | .700 | .678 | 0.022 |
| Phillip Evans | .263 | .331 | .411 | .741 | .720 | 0.021 |
| Tucker Barnhart | .244 | .324 | .395 | .719 | .699 | 0.020 |
| Akil Baddoo | .210 | .277 | .353 | .630 | .610 | 0.020 |
| Roberto Pérez | .205 | .296 | .356 | .652 | .633 | 0.019 |
| Wilson Ramos | .270 | .323 | .421 | .744 | .725 | 0.019 |
| Evan Longoria | .251 | .303 | .408 | .712 | .693 | 0.019 |
| Zach McKinstry | .231 | .291 | .387 | .678 | .659 | 0.019 |
| Mike Trout | .289 | .428 | .610 | 1.038 | 1.020 | 0.018 |
| Will Smith | .243 | .342 | .494 | .836 | .818 | 0.018 |
| Name | RoS BA | Ros OBP | RoS SLG | RoS OPS | Preseason OPS | Difference |
|---|---|---|---|---|---|---|
| Marcell Ozuna | .280 | .350 | .517 | .867 | .891 | -0.024 |
| Miguel Sanó | .225 | .319 | .510 | .829 | .852 | -0.023 |
| Martín Maldonado | .206 | .286 | .347 | .633 | .656 | -0.023 |
| Rowdy Tellez | .253 | .317 | .484 | .801 | .824 | -0.023 |
| Tommy Pham | .251 | .346 | .402 | .748 | .769 | -0.021 |
| Joey Votto | .236 | .338 | .372 | .710 | .731 | -0.021 |
| Jorge Polanco | .272 | .326 | .427 | .754 | .773 | -0.019 |
| Joc Pederson | .235 | .323 | .455 | .778 | .797 | -0.019 |
| Matt Chapman | .240 | .318 | .507 | .825 | .843 | -0.018 |
| Roman Quinn | .216 | .279 | .337 | .616 | .634 | -0.018 |
| Ozzie Albies | .280 | .328 | .496 | .824 | .841 | -0.017 |
| Sean Murphy | .234 | .304 | .416 | .720 | .737 | -0.017 |
| C.J. Cron | .269 | .340 | .526 | .866 | .883 | -0.017 |
| Elvis Andrus | .243 | .285 | .363 | .648 | .665 | -0.017 |
| Mitch Moreland | .231 | .305 | .418 | .723 | .740 | -0.017 |
| Anthony Alford | .211 | .280 | .350 | .630 | .647 | -0.017 |
| Gleyber Torres | .281 | .362 | .513 | .874 | .890 | -0.016 |
| Aaron Hicks | .233 | .355 | .419 | .774 | .790 | -0.016 |
| Alejandro Kirk | .254 | .323 | .417 | .740 | .756 | -0.016 |
| Jake Cave | .251 | .310 | .437 | .747 | .763 | -0.016 |
Probably the most fun name in the improvement category is Yermín Mercedes of the Chicago White Sox. A solid hitter in the minors but one with a poor defensive reputation behind the plate, Mercedes wasn’t considered a top prospect (he was 15th on Eric Longenhagen’s White Sox list this year). But the injury to Eloy Jiménez opened up an opportunity for someone to seize the DH job, and Mercedes has. Starting seven of the team’s nine games at DH, he’s gotten off to a blistering start, going 15-for-28 with five extra-base hits. It’s easy for people to dismiss comic book numbers, and a 298 wRC+ definitely fits into that category. But it’s those outlandish performances that are most likely to suggest a performance improvement. A .300 batter hitting .320 for a week seems “normal,” but it’s also less likely to represent a change in abilities than the same player hitting .800 for a week.
In this case, ZiPS and Steamer both give Mercedes a wRC+ boost of 13 points based on just these seven games of data. It sounds like a lot, but that’s a product of our uncertainty in making projections. Coming into the season, ZiPS was only 50% sure that Mercedes would have a wRC+ between 77 and 102. Think of a hurricane being tracked in the Caribbean, where even the slightest change in direction or timing can be the difference between the storm running smack-dab into the Atlantic coast or churning (mostly) harmlessly out into empty waters. In Mercedes’ case, with only a single plate appearance a year ago, his seven games in 2020 make up roughly 10% of his rest-of-season projection.
While the weights ZiPS uses are derived from baseball history, sometimes recent examples are the most effective illustrations. So let’s go back to 2019 — the last season of normal length — and look at how the players with the biggest changes in their projections at this point in that season fared the rest of the year. I looked at the preseason OPS projections and the RoS OPS projections through April 8, 2019 and compared them to the actual results for the 233 players who qualified for our splits leaderboards. From April 9 to the end of the season, using the updated OPS projection rather than the preseason one improved the average absolute error from 68 points of OPS to 62 points. The in-season model “expects” to come closer to the actual result than the preseason projection about 57% of the time. In 2019, that was 55.8% (130 of 233).
But the larger question is: does the model actually work well for the outliers? Is it too generous to players like Mercedes and too miserly to players like Marcell Ozuna? Again, let’s look back at 2019’s largest projection changes at a similar point in the season.
| Player | April 8 OPS | Preaseason | RoS Projection | Difference | Actual RoS |
|---|---|---|---|---|---|
| Tim Beckham | 1.314 | .679 | .722 | 0.043 | .665 |
| Cody Bellinger | 1.468 | .889 | .923 | 0.033 | .995 |
| Austin Barnes | 1.328 | .699 | .732 | 0.033 | .559 |
| Alex Avila | 1.324 | .688 | .720 | 0.031 | .716 |
| Kolten Wong | 1.235 | .732 | .762 | 0.030 | .750 |
| Anthony Rendon | 1.412 | .843 | .869 | 0.026 | .983 |
| Domingo Santana | 1.071 | .759 | .783 | 0.024 | .732 |
| Rhys Hoskins | 1.446 | .851 | .875 | 0.024 | .783 |
| Trey Mancini | 1.252 | .762 | .785 | 0.023 | .874 |
| Josh Phegley | .942 | .642 | .665 | 0.023 | .674 |
| Mike Trout | 1.508 | 1.023 | 1.045 | 0.022 | 1.052 |
| Tim Anderson | 1.326 | .685 | .706 | 0.021 | .837 |
| Christian Yelich | 1.340 | .901 | .922 | 0.021 | 1.078 |
| Clint Frazier | 1.292 | .767 | .787 | 0.020 | .756 |
| Dansby Swanson | 1.169 | .699 | .719 | 0.020 | .719 |
| Martin Prado | 1.141 | .647 | .667 | 0.020 | .504 |
| Brian Goodwin | 1.092 | .664 | .684 | 0.020 | .778 |
| Jorge Polanco | 1.116 | .717 | .737 | 0.019 | .826 |
| Adam Jones | 1.105 | .742 | .761 | 0.019 | .690 |
| Edwin Encarnacion | 1.142 | .782 | .801 | 0.019 | .850 |
| Player | April 8 OPS | Preaseason | RoS Projection | Difference | Actual RoS |
|---|---|---|---|---|---|
| Jesse Winker | .157 | .800 | .773 | -0.028 | .881 |
| Jurickson Profar | .350 | .726 | .700 | -0.026 | .754 |
| Mike Zunino | .198 | .682 | .658 | -0.024 | .586 |
| Chris Davis | .125 | .673 | .650 | -0.024 | .649 |
| Roberto Perez | .249 | .600 | .577 | -0.022 | .798 |
| Yasiel Puig | .354 | .825 | .803 | -0.022 | .809 |
| Garrett Hampson | .151 | .752 | .730 | -0.022 | .736 |
| Jesus Aguilar | .399 | .825 | .805 | -0.020 | .748 |
| Eduardo Escobar | .439 | .798 | .779 | -0.019 | .857 |
| Josh Donaldson | .554 | .879 | .861 | -0.019 | .923 |
| Kevin Pillar | .303 | .700 | .682 | -0.018 | .745 |
| Jackie Bradley Jr. | .416 | .754 | .736 | -0.018 | .764 |
| Ian Desmond | .369 | .725 | .708 | -0.018 | .826 |
| Grayson Greiner | .321 | .602 | .585 | -0.017 | .594 |
| Danny Jansen | .379 | .717 | .700 | -0.017 | .662 |
| Brandon Nimmo | .395 | .761 | .745 | -0.017 | .847 |
| Yolmer Sanchez | .122 | .678 | .661 | -0.017 | .664 |
| Brian Dozier | .340 | .787 | .772 | -0.016 | .801 |
| Yadier Molina | .458 | .709 | .693 | -0.015 | .737 |
| Brian Anderson | .422 | .743 | .729 | -0.014 | .846 |
In this case, ZiPS RoS did a little better than expected with the overachievers (closer on 13 of 20) and worse than expected with the underachievers (seven of 20). Rather than hitting on the 22-23 players the model expected, it only hit on 20 of the 40 outliers. However, if we expand our search to the top 20 positive and negative outliers from 2004-19, the RoS model was closer to what actually happened for 362 of the 640 players (56.6%). In other words, yes, these big leaps forward and backward for individual players have real meaning.
Yermín Mercedes will not finish the season hitting .500. He’s unlikely to get MVP votes or make people forget about Eloy Jiménez. But if you’re not a little more optimistic about him than you were two weeks ago, even after just seven games, you’re doing it wrong.
Dan Szymborski is a senior writer for FanGraphs and the developer of the ZiPS projection system. He was a writer for ESPN.com from 2010-2018, a regular guest on a number of radio shows and podcasts, and a voting BBWAA member. He also maintains a terrible Twitter account at @DSzymborski.
Ronald Acuña’s start is incredible but extremely out of line with his past (far fewer walks, far fewer K’s, ridiculous BABIP and ISO).
What’s his change, about 10 ops points?
If you click “Projections” at the top of a player’s dashboard it will put in the pre-season projection. For Acuna it’s 17 points of OPS and 4 points of wOBA
What do you mean, unlikely to make people forget about…who was that again? Let’s talk about Yermin Mercedes.
JD Martinez has 5 home runs in 38 PAs, which averages out to about 13 homers per 100 PAs. That means that we can expect that over the next 550 PAs he will hit approximately 71 home runs, for a total of roughly 76 over the course of the year.
How’s that for some fancy projectin’!
There a few great ones here that I find plausible, and actually a little conservative.
– You almost couldn’t ask for better peripherals than Phil Evans. Look at Baddoo, who everyone is going nutty about, and you see the regression monster staring back at you, but Evans looks pretty legit, yet he’s waaaay under the radar.
– I just have to think that Zips hasn’t caught up with 2020 Jared Walsh being legitimately different from pre-2020 Jared Walsh. Everything I’ve seen suggests it should lean towards the former.
– Whatever aging factor Zips uses needs to be discounted for Cruz.
Excellent points. I am often amused at how peripherals like BABIP, exit velocity, swing-and-miss rate, BB rate, K rate matter … until they don’t. Case in point with Evans v. Badoo.
I’m guessing the big factor here regarding Baddoo vs Evans is age and experience. Evans is 28 and Baddoo is 22. Evans has played three seasons at AAA while Baddoo hasn’t played above A ball. Evans may have really taken a step forward or it may just be a hot couple of weeks. Baddoo’s peripherals are worse but he’s doing it while very young and with little prior experience. There’s much more room to grow for him.
The age factor is critically important when projecting how each will do over the next five years but not so much with how each will do the remainder of the season, and which of those is more important depends upon one’s perspective, and if you are a fantasy baseball player, the type of league in which you compete. Baddoo may have a higher ceiling in coming years but also carries more risk he won’t sustain his early performance this season — If you believe in the unnamed peripherals.
The thing about very old players is that projection systems are giving you averages and at those ages there’s always a substantial probability that the player just never plays again. A player like Cruz will not experience gradual decline from a 130 wrc to 120 to 105 to 100. He will be close to what he is now and then one year he will show up in spring and just no longer be able to catch up to a fastball or suffer an injury from which he never recovers and be out of the league more or less instantly.
The degree to which the careers of age 40+ major leaguers, even great ones, rarely resemble a slow decline and much more often represent a 3-8 week sample of suddenly being below replacement level and then.beimf out of the league, is very consistently underestimated in the FG comments section.
The most likely way that Nelson Cruz’s mlb career ends is suddenly, whether via injury or sudden ineffectiveness. A median projection therefore will always have to include the 10-20% or whatever that it happens this year.
Unfortunately I can’t agree with you. Three of the greatest of all time who played into their 40’s all exhibited the classic gradual collapse. Stan Musial hit under .300 4 times in his last 5 years, Hank Aaron went .268-.234-.229 in his last 3 seasons and the less said about Willie Mays, the greatest all-around player I ever watched play, the better. Nellie Cruz isn’t quite a unicorn, David Ortiz’s last season was super, but he is the exception, and to think the industry thought the Orioles were crazy when they paid him $8M in 2014, then it was Seattle and then Seattle and Seattle and Seattle and Minnesota and Minnesota and Minnesota. He has proven a lot of people wrong and I hope he has another great year and rides out on his Boomstick.
My observation has been that the great player often has a pretty gradual decline from peak to still All-Star level and then, boom! They immediately drop to another level where they are average-ish which lasts for 3-4 years and then like mikejunt said, they drop to being unplayable and are out of baseball in short order. Musial follows that to a T – still great at age 37, averageish from 38-40, a nice bounceback year at age 41 and then unplayable at 42. Aaron had a shorter averageish period but he was great at 39, average at 40 and unplayable at 41-42. Those guys get to hang on for a few years because of name recognition but not everyone will get that
Aside from JakeDubois’s point–which I agree with–I think the real difference is that ZiPs absolutely hated Baddoo and thought he would fail spectacularly. A .610 OPS is really bad. In 2019, not a single qualified batter had an OPS that low. After changing the projection to a .630…it is STILL lower than any qualified batter in 2019. There’s a lot more room to go up when the initial projection is the worst hitter since Chris Davis / Alcides Escobar in 2018.
Yes a lot of it is just that the projections couldn’t get worse. My big thing with Baddoo is that 62% contact rate and 17% SwStr. That really is a dead ringer for Chris Davis! I want him to be a feel good story so badly but I expect he gets exploited and by July gets the high minors seasoning he probably needs.
He’s a rule V guy, so no minors for him unless Detroit dumps him
This is a big reason why he’s getting so much attention. “The one who got away” is a good story.
The key is that you have to hit lasers when you do connect, which is how Miguel Sano has made his money. But it’s hard to hit lasers like that consistently, which is why his other close comps are guys who have hot streaks and then fall back to earth, like Keston Hiura, Austin Riley, Keon Broxton, Jorge Alfaro. and Michael Chavis. I think there’s a chance he alternates enough hot streaks with not-hot streaks to function this year while he runs out his Rule 5 year.
One other thing to note about Baddoo is that he’s pretty heavily protected, only batting against right-handed pitchers. This means (1) his role is fairly limited right now and (2) in comparison with someone who is playing every day, he has an even more miniscule sample to work with.
& that is the main reason why I’d take the over on even the adjusted OPS projection. If & when he starts struggling, they’ll sit him down & Jacoby Jone will be the CF.
BTW, since this article, he’s hit 2 more HR’s. He has enough power where the .630 OPS seems really low. He really would have to turn into Chris Davis for that to occur..except he also is fast & not a pull hitter, so should have better BABIP’s.
“bitter Y realizes…”
I imagine Yandy Diaz going through a mid-life crisis realizing that his biceps will go for naught if he continues his groundballing ways.
Something similar crossed my mind, too. “How come they don’t love me as much as they love Ji-Man?”
Mike Trout is the only guy to appear on both leaderboards for beating his preseason ZiPs.
Mike Trout is also the only guy to improve upon an OPS greater than 1.000 in either leaderboard.
My projection system predicts that Mike Trout will discover the cure for COVID in his spare time
I enjoy the annual tradition of Mike Trout looking at his pre-season OPS projection of 1.000+ and spending the first few weeks of the season convincing ZiPS that it is still too low on him.
He’s on pace for a 16.2 WAR season. Normally I’d say that as a joke….
I know he just had a kid, but it’s hard to believe Trout is human sometimes.
The CW has Superman having a kid, so you never know.
How exactly do Kryptonians mate with Humans? Do Kryptons have super sperm? These are questions that need answers.
Have I got some good news for you!
http://www.rawbw.com/~svw/superman.html
You have made my week. Thank you
Excellent write-up. Like how the projections factor in production early on.
Hey Dan, love the hurricane tracking comparison. Pretty cool stuff. Although he’s known to be terrible with the glove, given his bat, is there any chance Mercedes sees the field occasionally (like later on when Eloy returns but may want to DH Eloy as precaution)? Or is he so bad that he should be permanently considered a DH-only player? Thanks
The position he came up through the majors at was actually catcher and he’s cartoonishly bad there. So probably not.
The thing about this sort of thing that I always think about is this:
Dan’s clearly 100% correct that projection systems are improved with even small amounts of this year’s data.
That doesn’t mean that individual human judgments are, because humans are chronically and probably incurably bad with recency bias and cannot accurately identify the both which improvements matter and how much they matter.
Zips is improved with this information. If you’re using zips you should use the updated projections.
There’s also a high probability that if you try and personally account for this in your own opinion you’ll end up more inaccurate than you were in the preseason when you were making your evaluation on a larger dataset from prior years.
It reminds me of how FIP remains one of the most accurate predictive statistics. We all know it misses a bit of information and is necessarily imperfect, but individual attempts to adjust for this by accounting for ground ball rates or the like are almost universally worse. The few things more accurate in predicting future performance than FIP are elaborate, closed box projection systems (like zips) and metrics that follow the principles of FIP while using a more granular measurements (I saw, for example, that a variant of FIP using barrel rates instead of HR was marginally better, though the range of time you can test for accuracy with it is necessarily narrow).
Just because something is flawed or marginally better in a controlled model doesn’t mean an individual human should attempt to account for it. The number of people whose personal opinion prediction can beat raw FIP is incredibly small: things that are more accurate are also systems.
Its virtually impossible to improve on systemic projections and models with raw human opinion, even opinion informed by those models. Even if we pay attention to them, it’s impossible to remove basic human cognitive biases from our perspectives.
Is that the Byron Buxton: MVP klaxon I hear?!
More a believer in Mercedes than Baddoo.
Individual stats don’t mean much this early in the year but the cumulative numbers add up to more than enough plate appearances to give an idea where the season is headed. The number of player below the Mendoza line is frightening and more than offset the hot starts which garner most of the attention. HR’s are slightly down. 1.20/ game against 1.28 in 2020 and 1.39 in 2019 but hits have not increased and batting average is down from .252 in 2019 to .236 after the first 264 games. I don’t know who thought taking some juice out of the ball with the concurrent lessening of the number of HR’s would result in an increase in singles and doubles but it was a flawed concept. It hasn’t happened and it is very doubtful if it will as the season progresses.
Well, yeah, but you can’t expect such a change to happen in one season. Hitters need time to adjust to the new situation before we’d actually see meaningful changes in results.
“But if you’re not a little more optimistic about him than you were two weeks ago, even after just seven games, you’re doing it wrong.”
I didn’t even know who he *was* two weeks ago never mind have more optimism for him. Going from zero to something slightly positive on the optimism scale wasn’t too hard!
That paragraph about Tim Anderson is the best way to put those projection numbers- you put it into words extremely well
I will never forget the great Chris Shelton April 2006.