A New Way to Look at Sample Size

Jonah Pemstein and Sean Dolinar co-authored this article.

Due to the math-intensive nature of this research, we have included a supplemental post focused entirely on the math. It will be referenced throughout this post; detailed information and discussion about the research can be found there.

INTRODUCTION

“Small sample size” is a phrase often used throughout the baseball season when analysts and fans alike discuss player’s statistics. Every fan, to some extent, has an idea of what a small sample size is, even if they don’t know it by name: a player who goes 2-for-4 in a game is not a .500 hitter; a reliever who hasn’t allowed a run by April 10 is not a zero-ERA pitcher. Knowing what small sample size means is easy. The question is, though, when do samples stop becoming small and start becoming useful and meaningful?

This question has been researched before — notably by Russell Carleton, Derek Carty and Harry Pavlidis. Each of them used similar methods of finding reliability to achieve a point of stability.

Our aim in this project is to extend the understanding of reliability and show a more complete picture of how additional plate appearances affect the reliability value of everyday stats for both batters and pitchers. We want to reinforce the idea that reliability is a spectrum, not a single point. There is no single point at which you can say a stat has stabilized. We also want to use the concept of reliability to regress toward the mean and to make confidence bands that give a better idea of a player’s true talent. (Throughout this project we define true talent as the actual talent level — not the value they provide adjusted to park, competition, etc.)

We used a similar approach as Carleton did in his latest reliability study: Cronbach’s alpha. There are, however, differences in our sampling structure, which we explain in more detail in the math post.

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

METHOD

Sampling

We used a data set that is also similar to Carleton’s most recent studies — Retrosheet data — but we used more recent data from a shorter time frame (2009 to 2014). We also removed intentional walks, bunts and non-batting events such as stolen bases.

From there, we broke the data into different player-seasons instead of just players: 2012 Mike Trout, for example, is different from 2013 Mike Trout. We then used these player-seasons to define samples for a given number of plate appearances (PA), at-bats (AB) or balls in play (BIP). So for 10 PA, we took 10 random plate appearances from each player-season with at least 10 PA; for 600 PA, we took 600 random plate appearances from each player-season with at least 600 PA. As you can imagine, the 10-PA sample has many more player-seasons than the 600-PA sample. Since we used player seasons, we maxed out our sampling at 600 PA, 500 AB and 400 BIP. For anything beyond those limits, the sample size became too small, and results become erratic.

We chose this sampling structure because we think this best represented the general question of the stats’ reliability. The most recent, smaller data set mitigated a bias we found associated with Major League Baseball’s changing run environment. We make no assumptions about a player’s talent levels being the same across years, so we separated each season for each player. This also allows for the comparison of players across years. We will detail the implications and effects of sampling in a future article.

Cronbach’s alpha

There are many different methods to measure reliability, which is mathematically related to correlation but is a different construct with different assumptions. We chose Cronbach’s alpha because it provides a good framework to measure the reliability of a full sample of plate appearances. Given the nature of the data — different parks, pitchers, time of year, etc. — there was no obvious single way to split the data. We used a method that split it in as many ways as possible. Once again, you can read more about Cronbach’s alpha and reliability in the math post.

Cronbach’s alpha’s calculation gives a value — alpha — that is a measurement of the reliability. The value represents the proportion of true-talent variance to the observed variance.

Alpha_Equation

This is not the same as r, r-squared or linear regression.

RESULTS

Alpha

Below is a data visualization of various batting stats’ reliability as it relates to the number of PA/AB/BIP. The lines represent the measured reliability at each 10-PA/AB/BIP increment for each stat. To calculate regression toward the mean and the associated confidence band, enter the stat’s value into the red box and select the appropriate confidence level. Then scroll across the line for the results of the calculation for each PA/AB/BIP increment.

The reliability of each stat increases when the number of PA/AB/BIP increases, and the curve increases at a slower rate as the value gets closer to 1.0. One of the goals of this project is to demonstrate how the reliability of stats changes with number of PA. Most importantly, there is no single point at which a stat becomes stable — every additional PA/AB/BIP simply increases the reliability. Even with a low reliability, there is information within the stat; it just has more noise than a stat with a high reliability.

Regression to the Mean & Confidence Bands

Reliability values are useful for comparison between different stats, but they don’t address the uncertainty of that stat in tangible terms. In other words, it doesn’t give you a likely point or range for the player’s actual skill. Regression to the mean and confidence bands allow us to estimate a floor and ceiling for this uncertainty.

CI_Diagram

This diagram demonstrates how to regress to the mean and create confidence bands from that regressed stat. Since we are estimating true talent from an observed stat, the first step will be to regress the stat to the mean. If a stat has a low reliability, the sample’s average is a better estimation of the true talent. A high reliability means the stat contains more true-talent information, and that it’s regressed much less toward the mean. Reliability provides an empirical method to regress toward the means in a manner similar to the mathematical approach outline in the appendix of Tango’s The Book (more on that in the math part).

The second part uses the sample’s total standard deviation to estimate the uncertainty and the upper and lower bounds. The higher the standard deviation, the wider the confidence band. (These confidence bands are not the same as the binomial standard error.)

DISCUSSION

All of the previous reliability studies and this one are based in math typically used for test evaluation — where researchers are trying to gauge how well the test is constructed. The basic idea is there is a true score (or in our case talent level), error (or noise) and an observed score (or observed stat).

TT_Noise_Observed_Diagram

What reliability attempts to measure is the ratio of the true-talent information to observed information. If there isn’t a lot of true information the reliability will be lower; if there is a lot of information the reliability will be higher. The noise term contains almost every factor that could be associated with affecting a plate appearance: pitcher, park factors, weather, injury and so on. The intent of this analysis is to create reliability measurements and confidence bands for everyday stats, which do not contain these adjustments, so we left all our data unadjusted.

Reliability is partly determined by the distribution of skills within the sample. As a result, sampling becomes an important factor in determining reliability. We tried several variations of sampling structure, including the one Carleton used in his most recent study. The results followed similar patterns, but there were some discrepancies due to different pools of players being used. Using a sample restricted to a high minimum number of PA will decrease the standard deviation because players with better statistics get more PA. This weeds out the lower echelon of players. The remaining players are all bunched tighter together. The larger the spread in talent, the higher the reliability; the smaller the spread, the lower the reliability. We discuss this more in the math post.

Conclusion

The most important conclusion to be drawn is there is no single point at which a stat becomes reliable or stable. The alpha reliability data visualization demonstrates that idea, using the reliability measurement to regress the stat to the mean and create confidence bands. The regressed stat and confidence bands are descriptive, rather than predictive, and are not adjusted for park factors, league adjustments, etc. This can provide an estimation of a player’s true-talent level based on how the player has performed. They aren’t intended to be projections.

NOTES: If you are comparing our results with results from Carleton’s analysis, we are reporting the alpha for the entire sample of PA. His previous analysis found a particular number of PA/AB/BIP associated with a certain value for alpha ( .70) and then halved the PA/AB/BIP value. The Cronbach’s alpha calculation finds the reliability coefficient associated with the entire sample, and it does not need to be halved. Our reporting method is critical to regressing toward the mean and calculating confidence bands.

The code we used — plus a .csv file of the results — are available on GitHub.





Jonah is a baseball analyst and Red Sox fan. He would like it if you followed him on Twitter @japemstein, but can't really do anything about it if you don't.

22 Comments
Oldest
Newest Most Voted
MustBunique
11 years ago

Wow. The interactive graph is fantastic. Well done you two. Extremely useful.

Candy LaChance
11 years ago

What is the practical relevance of within-season reliability tests like this, given the easy availability of projections?

Sean DolinarFanGraphs Staff
11 years ago
Reply to  Candy LaChance

The “this stat stabilizes at X PA” has been around for a while. This extends that. It also helps to understand the more uncertain stats like BABIP. This is backwards looking for people analyzing how seasons have gone.

But yes, if you want to know how a player will do, use projections.

Candy LaChance
11 years ago
Reply to  Sean Dolinar

Thanks, guys for the replies. Really interesting stuff. I should have made clear in my question that is was not rhetorical– I’ve long wondered how to think about these “stabilization” questions in comparison to the projection systems.

Doug
11 years ago

Excellent stuff! I’ve heard way too many analysts dismiss possibly useful data, because it didn’t hit some magical number.

To me the results seem intuitive, and intuitive is quite often the right answer.

MK
11 years ago

Any theories as to why doubles alpha is so low?

Sean DolinarFanGraphs Staff
11 years ago
Reply to  MK

There’s not much variation between players. Most are grouped closely together. Triples have more variation.

Eli Ben-PoratMember since 2016
11 years ago
Reply to  MK

I’d say it’s because you can get a lot of random doubles on ground balls down the line. I’d guess that Fly Ball 2B% is a lot more reliable then just 2B% overall and probably tracks close to 3B%.

Sean
11 years ago

Should it be HR%? If you put in any integer, the number gets heavily regressed and don’t seem to make much sense.

Youppi!
11 years ago

Peer review?

jake
11 years ago

Could this lead to a new row on player cards, in addition to seasons and career? I’d love to see each stat in a “most recent stabilized value” that can span seasons depending on sample size necessary

Russell Carleton
11 years ago

In defense of halving the PA estimates, it’s actually a key portion of the method. When I settled in Cronbach/KR-21, I wasn’t using it in the classical manner. It was actually a hack of the formula. Cronbach takes 300 observations and chops it up into two halves of 150 each, and then compares them to each other, and does this every which way that 300 observations can be split in half. For my purposes, it always meant that I was comparing a sample of 150 PA to another of 150 PA, so if the two results were pretty consistently correlated (which is what Cronbach tells you) and one PA is as good as another (a conceit of the research, but I’d argue a justifiable one), then I’ve got two samples of 150 PA that are acting the same way. With Cronbach, I can side-step some of the randomization “yeah but’s” that can be had. So, at that point, I feel pretty good about my sample of 150 PA.

Mechanically, I can see that you’re going to use alpha as a factor for regressing toward the mean. I would caution you that the general form of that formula calls only for a reliability estimate. (Cronbach is the most commonly used.) I don’t recall if I’ve ever tested whether the halves or the non-halved estimate performs better as a predictor, but I would at least encourage you to pay attention to that issue.

The Hat of the Three-Toed Man-Baby
11 years ago

Too bad for you any evidence of non-iid errors will invalidate all of these standard errors you constructed. Since “true” talent is almost certainly highly autocorrelated your exercise is pretty much worthless.

Corey
11 years ago

How do you estimate variance of their true talent? We don’t know anybody’s true talent, its an abstract idea, so what are you using as a proxy for it?

Sean DolinarFanGraphs Staff
11 years ago
Reply to  Corey

The math section addresses this more. But in short the regressed stat with the associated confidence bands are the estimated true talent (in the descriptive sense of true talent.) Once again these are all estimates.

Shaemoose
11 years ago

Personally don’t like the fact that you’re splitting each player up into certain seasons, I think a more helpful approach would be to look at where a skill suddenly changes (think K% and BB%), then measure approximately half a season out or until the player has a visible change of skillset. An example of this would be Salvador Perez being overworked at the end of last season. His averages would change if we eliminate the last month of the season from him. This, however, would be much harder to do.

Psingman
11 years ago

Just finished reading both pieces, great work guys. As someone who’s recently become code-literate, I appreciate the effort put in to both explain the theory and make the dataset and R code available to see/use. I imagine there’s close to a whole semester’s worth of knowledge to be learned here.

Thanks.

dacureMember since 2025
11 years ago

I have a bit of a problem intuitively with regressing to the mean of the sample. It implicitly seems to assume a null hypothesis of “player X is average at stat Y” (where average will change depending on who you arbitrarily pool into your sample). For instance, Tulowitzki is off to a slow start but still remains above average if you include enough players. Regressing him to the mean of that pool would actually probably shift his projection further then where we intuitively might expect something like Steamer to push him towards. Or take some bad player off to a hot start which leaves him below average, the regressing would push him even closer to MLB average. So it seems like regressing to the mean is only helpful for players that actually have true talent levels close to the mean, which if we imagine a normal distribution on talent might be a good chunk of them, but its not exactly the interesting pool of players that people are probably more interested in projecting. Maybe my interpretation is way off but it doesn’t seem as magical as people make it out to be considering “the mean” for each player is just his true talent level which is unobservable (obviously).

Sean DolinarFanGraphs Staff
11 years ago
Reply to  dacure

When you introduce prior knowledge of the performance of the player, you use a different framework. It would be interesting to pursue this under a Bayesian framework, but I’d have to learn more about it first.

Our initial interest was actually try to produce confidence intervals which weren’t your simple proportion standard error of sqrt(pq/N), which is the same value no matter how well that particular stat measures talent. We opened a can of worms and had to address reliability and regression towards the mean also fit into that. But there are definitely other approaches.

Chris WalkerMember since 2025
11 years ago

I’ve got your story here, Sean. Wow! I noticed you used a lot of big words. Nice. Good for you. It was quite long, so I didn’t read it all, but who cares ’cause I gave you an A.

(Just kidding, this is a quote from Orange County that seemed appropriate. Despite this, I am code literate and can appreciate your hard work and insight. THANKYOUSOMUCHKBYE.)