The Meaning of Small Sample Data

We’re a week into the Major League season, which means most regular hitters have roughly 20 to 25 at-bats, and each team’s best pitcher has maybe thrown 13 or 14 innings. These are the smallest of small samples, and almost anything is possible over the course of five or six games. Right now, we have things like Jose Iglesias leading the American League in Batting Average and Kevin Kiermaier leading the AL in Slugging (.941). Among the many dominant pitching performance from the first week, you’ll find names like Aaron Harang, Tommy Milone, and Jason Marquis.

For years, the analytical community has strongly advised against reading anything into early-season results, making the phrase “Small Sample Size” into a term you’ll even hear on broadcasts. We have an entire entry in the FanGraphs Library devoted to sample size, and another on regression to the mean, which is a related concept. If you’re reading FanGraphs, odds are you’re probably aware of the fact that you shouldn’t jump to conclusions based on a week’s worth of data. The Braves are not the best team in the National League. The Tigers aren’t the ’27 Yankees. Over any given week, weird stuff is going to happen, and we just notice it more at the start of the season because it’s the only thing that has happened yet; if you look at any seven day stretch throughout the year, you’ll find similarly odd results.

But there’s a problem with just saying “Small Sample Size” all the time: it forces you to draw a line in the sand. At some point, you have enough data for it to not be considered a small sample anymore, but the terminology suggests that it magically transforms out of being a small sample at some point, worthless beforehand but useful afterwards. This assumption has been somewhat reinforced by the very-useful-but-often-misinterpreted research by Russell Carlton (and others) on what are generally called stabilization points; the number of trials — usually notated in plate appearances, balls in play, or batters faced for a pitcher — at which a given statistic only needs to be regressed about halfway back to the mean for it to contain useful information.

Unfortunately, the “Small Sample Size” creed and the availability of published numbers for these stabilization points have led to the idea that data is useless up to that point, and then useful after that point, which is wrong on both ends. In reality, every small piece of data contains a little bit of signal and a lot of noise, and as you begin to stack small pieces of data on top of each other, the noise begins to cancel out, making the signal more visible. If you have one piece of data, you probably have so much noise that you can’t even hope to find the signal in the result, but your hope marginally increases at data point number two, then again at number three, and so on, until you’ve collected enough pieces of information that the noise begins to fade as randomness cancels itself out.

In other words, the climb to a a reasonable sample size is more like stairs than an elevator. The fact that strikeout rate’s stabilization point is 60 plate appearances doesn’t mean that you have no information at 59 PAs and plenty of information at 60 PAs. That one extra trip to the plate doesn’t magically transform the previous 59 into being useful. You had almost as much information at 59 as you did at 60, and should treat the 59 PA sample not much differently than you do the 60 PA sample, even though one is north of the stabilization point and the other is not.

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

So while making conclusions based on the first week’s worth of results is a bad idea, so too is discarding the small bits of data we do now have for 2015 as if they tell us absolutely nothing. They certainly don’t tell us enough to where we should have radically different opinions than we do a week ago, but in extreme cases, the week one performances were so extreme that they should move the needle enough for us to notice. And because we now have rest-of-season projections on the site for both ZIPS and Steamer, we can identify a few of these instances where an extreme performance over five or six games does tell us enough to slightly adjust our expectations for the rest of the season.

No one had a more extreme week-one performance than Adrian Gonzalez, for instance. He’s currently hitting .609/.667/1.391, good for an .838 wOBA and a 478 wRC+. 13 teams still have fewer home runs this season than Gonzalez does by himself. Obviously, he’s not going to keep this up, and we shouldn’t take his opening week performance to mean that Gonzalez is going to launch 40 homers again. He’s still 33 years old, after all, and his power has been waning the last few years. But the fact that he was even physically capable of hitting five home runs in six days tells us that Gonzalez is unlikely to be playing hurt or to have suffered from an off-season decline that hadn’t yet been captured in his performance record.

His power spike in week one makes the worst-case-scenario outcomes — which drag down the mean forecast — less likely, and thus, his projection can already start climbing. And that’s exactly what we see when we look at the difference between his pre-season and his rest-of-season ZIPS numbers: Gonzalez’s wOBA is now projected at .358 for the rest of the season, up from a .341 mark at the beginning of the year. That 17 point increase is by far the largest upwards move of any player to start the year, which makes sense, given that no one has even come close to hitting as well as Gonzalez has in the season’s first six games.

Now, keep in mind, a .358 wOBA forecast isn’t so different from a .341 wOBA forecast that we should start dramatically altering our perceptions of him as a player. But 17 points of wOBA over 600 plate appearances is worth about eight runs, or almost an entire win worth of value. On a WAR/600 basis, ZIPS saw Gonzalez as a +3.1 WAR player before the year started, but after a week’s worth of games, now projects him as a +3.9 WAR/600 guy over the rest of the year. Gonzalez going bananas in the first week doesn’t mean he’s going to keep this up, but it does mean enough to add essentially one win to the Dodgers projected total for the season.

Of course, I’m highlighting Gonzalez because he’s the most notable change, and most players have seen their projections move far less; the second largest positive wOBA jump belongs to Miguel Cabrera, whose ZIPS forecast has only increased by eight points. Miggy has shown that he’s probably healthy enough at the moment so that the parts of his forecast that had him breaking down are less likely to occur in the early part of the year, and we can be somewhat more confident now that he’s still one of the game’s elite hitters than we were a week ago. It’s not a huge move, of course, but if we don’t think we’ve learned anything from Cabrera torching the ball the first week of the season, then we’re reacting too slowly to the data.

Reacting too slowly is certainly a better alternative to reading far too much into these kinds of samples, so if you had to choose between not changing your opinion at all or solely basing your opinion on a tiny number of data points, you’d be better off with the lack of change. But because guys like Dan Szymborski and Jared Cross have created algorithms that do the work for us, we don’t have to settle for either of those two incorrect options. By incorporating the most recent data into our projections, but still heavily leaning on the past track record of a player’s performance to guide his future forecast, we can climb the stairs to a reasonable sample size rather than ignoring everything that happens up to a certain point and then drawing conclusions after a magical platform is reached.

The 2015 data we have now shouldn’t be the basis of any kind of strong conclusions, and that will hold true for quite a while. But it’s also not entirely worthless, and shouldn’t be treated as if it is. When we see things like Jim Johnson or Joe Kelly throwing 96 mph, or Anthony Gose elevating and driving the ball for the first time in his life, we should be aware of the fact that these results matter a little bit. It doesn’t mean that these guys are going to be great this year, but that outcome is slightly more likely now than it was before. Incorporate the new data at a reasonable pace, and you’ll end up better off than if you ignore it entirely.





Dave is the Managing Editor of FanGraphs.

26 Comments
Oldest
Newest Most Voted
hscer
11 years ago

Regardless, I have decided that Bryce Harper will hit .261/.346/.522 all year. (Just kidding–although he could.)

I feel like there should immediately be a link to this piece on the Sample Size page of the Library.

Bookbook
11 years ago

Good idea. I’d like to see a piece examining the players who most exactly are meeting full year expectations through the first ten games or so. (seems like a Sullivan piece)

Eric R
11 years ago
Reply to  Bookbook

David Wright and Jeff Francoeur have the same wRC+ in 2015 as their career averages through 2014 🙂

Jason BMember since 2017
11 years ago
Reply to  Eric R

Which is alright if you’re David Wright, a little more problematic if you’re Jeff Francoeur.

everdiso
11 years ago

its important to be able to match up the eye test with this early season data to look for trends and confirmation

for example russell martin isn’t hitting right now, but he has a .111 babip and has looked good in a ton of his abs. david ortiz and mike napoli on the other hand have looked completely lost at the plate and anyone who watches baseball can see they’ve permanently fallen off a cliff

along the same lines but pitching wise, aaron sanchez’s early season stats may look average, but his stuff was phenomenal and he was mixing his pitches well but getting squeezed a crazy amount by the home plate ump. if we take these observations and data together, we can better determine noise vs signal

bdhudson
11 years ago
Reply to  everdiso

“anyone who watches baseball can see they’ve permanently fallen off a cliff.”

compelling argument, that.

everdiso
11 years ago
Reply to  everdiso

Ah the obsession continues. Hardest witking troll in the business.

Actually Sanchez was awful, and i have little hope for him ever being a good starting pitcher, let alone now at age 22.

His poor performance was no surprise, unfortunately. I’d rather see him in the bullpen where he has a chance of being an asset – though even that is questionable – and try to get by with a fungible redmond or estrada or hendriks in the 5 spot. heck i’d rather see osuna in the rotation instead of sanchez.

and don’t worry your pretty little head – i’m sure naps and ortiz will be just fine. then again, ort is 39, and that cliff always looms ominously.

everdiso
11 years ago
Reply to  everdiso

what the jays should be doing, ironically, is what the red sox refuse to do – package off some overrated prospects (leading with Sanchez) for one Cole Hamels.

K
11 years ago
Reply to  everdiso

That isn’t ironic.

everdiso
11 years ago
Reply to  everdiso

i believe there is some irony buried in a red sox fan using my name implying that i overrate a jays prospect when in fact i think he’s overrated and am willing to include him on a trade that red sox fans would refuse to consider giving up a similar ranked sox prospect (who i may think they in turn overrate) in.

but i’m not sure. maybe no irony.

TKDCMember since 2016
11 years ago
Reply to  everdiso

I’m not sure if that is ironic, but it is definitely confusing.

Dave
11 years ago
Reply to  everdiso

IT’S LIKE RAAAAIIIIIINNNNNNN ON YOUR WEDDING DAY!

BPhippsMember since 2017
11 years ago

I am curious. Why does a veteran like AGonz, who we have a large amount of data on, get a 1 war upward adjustment based on 1 week of results ? While a relatively new player like Kiermaier doesn’t see a much larger boost to his projection?
The Kiermaier data is 6% of his career at bats versus less than 1% of the available Agonz data.

Damaso
11 years ago
Reply to  BPhipps

Huh. Something felt wrong in this article and i wasn’t sure what… but you nailed it. It does seem more intuitive that guys with smaller mlb samples would have their projections change more at this point than guys with long track records.

AMartin
11 years ago
Reply to  Damaso

Yes and no; projections are based off of similar career paths; as players get older, they could have more divergent outcomes, meaning evidence one way or another has a larger impact on the mean outcome. (i.e. older players have more bimodal distributions of outcomes). There could be other reasons too, would be an interesting thing to explore.

Also, with established players (e.g. Cabrera), the author noted he was coming off an injury; the SS indicates that he is fully recovered, moving the projections much more strongly than say a prospect showing he deserved the call-up..

I would also speculate that prospects may typically perform better in their first weeks of the big-leagues, when scouting reports on them are less well formed, so a big performance is less indicative of future performance (sophomore slump in the micro).

Patrick
11 years ago
Reply to  Damaso

The A-Gonz example is the debated will he “make up” and hit less HRs the rest of the year or will he hit HRs as he previously projected rate.

Last year he hit HR 4 straight and 3 straight days. Could that streak just happen to be at the beginning of this year? That is very possible but his total expected HR output has to be increased because has a head start.

Drew7
11 years ago
Reply to  Patrick

That’s not really “debated” though. If I’m understanding you correctly, one of those is true and the other is Gambler’s Fallacy.

Vince Clortho
11 years ago
Reply to  BPhipps

I believe it’s because he’s ‘banked’ a win already. The projection doesn’t adjust for sequencing, basically (not that I think it should)

The FoilsMember since 2017
11 years ago
Reply to  Vince Clortho

That’s not it. The rest of the way projections have updated, not the overall projections.

Unique Commenter 2
11 years ago
Reply to  The Foils

Yet in essence, isn’t that what’s happening? Isn’t it *less* likely that A-Gon suddenly hits .000 next week than that he perhaps hits a little worse each day than during his first week? Isn’t that sort of the argument that Cameron is making: that week-over-week samples, while they obviously fluctuate wildly, aren’t *entirely* random, therefore hedging between throwing out the data completely and using the data as the gospel is the smartest option? Thus, he has in essence ‘banked’ a win?

Unique Commenter 2
11 years ago
Reply to  The Foils

(‘Banked’ a win in that his forecast is accounting for the mild possibility that his true talent level is ever-so-slightly nudged up toward his titanic first week. Projections are based on past performance after all.)

Matt
11 years ago
Reply to  BPhipps

Part of it is likely as simply as Gonzalez’ preseason projection included a drop due to aging – his .341 wOBA was lower than any mark he’s put up in the last 10 years. But with his strong start, it looks less likely that’s about to fall off the proverbial cliff, and so his mark improves by a lot.

For Kiermeier, more likely he simply does not have the history that would expect him to put up a big offensive season, so in many ways the projection systems treat him as someone who simply got very lucky for a week. If he continues to hit well for another couple weeks, then I assume the projections will catch up faster for him.

DNA+
11 years ago

“His power spike in week one makes the worst-case-scenario outcomes — which drag down the mean forecast — less likely, and thus, his projection can already start climbing.”

To be clear what we are talking about is sampling variation here, right? The algorithms do not consider his skill level to have changed, or his projected health to have changed? I assume he is projected to perform at the same level as before the season began, however, we know he has already had a lucky streak (i.e. a small sample drawn from the far right tail of the distribution of expected outcomes for a six game stretch). If not, I would really love to see how the algorithm is written. Is this information available?

dj_mosfett
11 years ago
Reply to  DNA+

No, absolutely not. Those algorithms are proprietary! Why would you ever want to analyze it on your own? Someone else has done all the work already!

Matthew Cornwell
11 years ago

Joe Kelly throwing 96 is not rare. He through up there all the time. Getting k’s with his 96 MPH fastball is what is rare.

Matthew Cornwell
11 years ago

Threw, of course. 🙂