Gary Sanchez and the Persistent Belief in Small Samples
Perhaps no player in Florida is the object of greater expectations this spring than Gary Sanchez.
Consider: in the fantasy baseball world, only Buster Posey is being drafted earlier at catcher. Generally conservative projection systems forecast that Sanchez will be a star this season. ZiPS pegs Sanchez for 27 homers a 112 wRC+ and a 3.4 WAR season. PECOTA’s 70th percentile outlook has Sanchez recording 33 homers, a .504 slugging mark, and 4.8 wins. And the Fans’ average crowdsourced projection for Sanchez is a .274/.344/.488 slash line and 5.4 WAR season. The Fans believe, in other words, that Sanchez and Bryce Harper are going to produce similar value this season.
Last week, ESPN ran a poll asking respondents to guess how many home runs Sanchez will hit in 2017. Fewer than 10 home runs? That option received 1% of votes. How about 10 to 20 homers — i.e. the range within which he’s resided over each of his first five professional seasons in the minors? That seems like a reasonable wager, right? Only 4% agreed.
The most popular range was 21-30 homers, receiving 44% of votes. Forty-one percent predicted he will slug between 31-40 homers, and 10% think he will hit more than 40.
There’s much to like about Sanchez. This is a player with pedigree, who was regraded as the top catcher in the 2009 international class, and perhaps the second-best bat in that class after Miguel Sano.
The raw power, the arm strength, those are loud tools. They are very real.
The receiving appears to be about league average, but that’s perhaps the biggest question mark based upon his 2016 work. For RotoGraphs, Andrew Perpetua employed Statcast data to get a sense of what to expect from Sanchez in 2017.
Wrote Perpetua:
“These are just a few guys who seem to be most closely related to Sanchez in terms of these batted ball stats. Cespedes and Cruz appear to be the closest matches, with similar vertical and horizontal angles along with similar exit velocities to Gary Sanchez.”
Anyone for a Cespedes- or Cruz-type bat at catcher? You can lower your hands.
Everyone, it seems, loves Sanchez in 2017. But let’s pause and take a deep breath. Have we learned nothing?
Despite years of warnings about the dangers of small sample sizes, it’s possible we’re still, and always will be as humans, susceptible to making judgment errors from too little data. Michael Lewis’ most recent book, The Undoing Project, has some potentially important implications for the industry. It details the friendship and work of Israeli psychologists Amos Tversky and Daniel Kahneman.
The first paper Tversky and Kahneman wrote together was titled “Belief in the Law of Small Numbers”. The paper refuted the idea that people are, by nature, Bayesians, or that we have some innate understanding of probability and statistical principles and act in such a manner.
From the book:
“The power of the belief could be seen in the way people thought of totally random patterns — like, say, those created by a flipped coin… People seemed to believe that if a flipped coin landed on heads a few times in a row it was more likely, on the next flip, to land on tails — as if the coin itself could even things out. ‘Even the fairest coin, however, given the limitations of its memory and moral sense, cannot be as fair as the gambler expects it to be,’ they wrote.
“They went on to show that trained scientists — experimental psychologists — were prone to the same mental error… Even people trained in statistics and probability theory failed to intuit how much more variable a small sample could be than the general population… This failure of human intuition had all sorts of implications for how people moved through the world, and render judgments and made decisions.”
Remember, Sanchez hit .225 in September after a torrid August, a month he might never repeat in any 31-day period of his career. Sanchez’s 2016 Triple-A slash line was good — .282/.339/.468 (and 10 homers in 313 plate appearances) — but it also mirrored the slash line of his seven-year minor league career: .275/.339/.460.
He’s not going to sustain a 40% home-run rate on fly balls, of course, and I don’t think anyone is expecting him to.
While projection systems have greatly aided in our understanding of performance, while they are designed to be unbiased, they can still struggle with small samples and the translation of minor-league statistics to major-league equivalencies. There’s no greater leap for professional players than from Triple-A to the majors.
When I think about reasons to pump the brakes, I think about Brett Lawrie’s 43-game close to the 2011 season, when he hit nine homers and slashed .293/.373/.580. Many were predicting stardom based in part upon that sample of performance. While Lawrie’s had a major-league career, he’s never developed into the dynamic player that many, including evaluators, thought he would become.
Even former Baseball Prospectus prospect analyst Kevin Goldstein, when suggesting that expectations be tempered regarding Lawrie, still predicted Lawrie would become “a star.” In ESPN’s franchise player draft in 2012, Lawrie was selected 15th overall, two spots behind Mike Trout and only three spots behind Harper. (Matt Kemp and Troy Tulowitzki went 1-2, reminding us how quickly fortunes can change in what can be a cruel game.) Jim Bowden predicted Lawrie he would win a batting title early in his career, and ZiPS projected a .275, 27-HR, 79-RBI, 24-steal season in 2012.
Lawrie and Sanchez are different people and players and they will follow different career arcs, mostly likely. But Lawrie is an example of a recent player to really enhance his perceived value through a small sample to close a season.
Sanchez seems to stand a good chance to become a quality major-league player, perhaps even a star. But expecting stardom right away might be too much. And perhaps the lesson from the “Belief in the Law of Small Numbers” is, as humans, we cannot help ourselves from getting caught up in small samples. As you prepare and conduct your fantasy drafts, this is just a reminder Yasmani Grandal will be available about 100 picks later according to these consensus rankings.
Sanchez might be great. But will he be in 2017? One scout with whom I spoke very much believes in Sanchez’s power and throwing arm and says his work behind the plate in receiving and handling a staff has come a long way, but suspects expectations have grown too large for 2017. The league will make adjustments to Sanchez. He will have to counterpunch.
Yankees general manager Brian Cashman is one of the few curbing expectations this spring.
“I don’t know if you can repeat that type of year,” said Cashman recently. “That kind of is impossible… I’ve been around the game long enough to know that you can never count on anything.”
Including a small sample of performance to be extrapolated out over the course of a full season or career.
A Cleveland native, FanGraphs writer Travis Sawchik is the author of the New York Times bestselling book, Big Data Baseball. He also contributes to The Athletic Cleveland, and has written for the Pittsburgh Tribune-Review, among other outlets. Follow him on Twitter @Travis_Sawchik.
This is really important for people to remember. Thank you, Travis!
Ah Brett Lawrie. In my home league we even had to pass a “Brett Lawrie Rule” explicitly governing add/drops and keeper eligibility for guys picked up during the fantasy playoffs because he just destroyed it those last few weeks.
[/NoOneWantsToHearMyFantasyBaseballStories]
I mean, he hit .225 in September…with a .314 OBP and a .520 SLG and a 119 wRC+.
Which is, of course, right in line with the current Steamer and FAN projections.
That was with a 40% hr/fb.
Ironically, the “look at his September numbers” argument runs into the smallest sample size of all.
The drop to “only” a 119 wRC+ is entirely driven by his 20 wRC+ over 23 PA in the first week of September.
From 9/9 through the end of the season, he hit .244/.326/.610 with 9 HRs and a 143 wRC+.
And a .233 BABIP, despite some good batted ball data. Of course the 30% K rate isn’t pretty.
I’m calling a noisy .255/.330/.490 line.
Yup, all of the projections are expecting offensive regression…while accounting for the fact that Sanchez’s minor league stats came in some of the most pitcher-friendly parks in the most pitcher-friendly leagues in professional baseball, rather than Yankee Stadium.
The issue with the FAN projections is the +15 defense, considering that he’s not a finished product as a receiver and will see plenty of time at DH. The offensive expectation is optimistic, but entirely reasonable.
Obviously he’ll never have a full season like he had last year, however he does look pretty good so far this spring. His throwing has been phenomenal.
I wonder how many people, especially those who aren’t NYY fans, would bet the over on 5.4 WAR. Not me.
Betting the over on a 5.4 WAR projection seems difficult for anyone not named Mike Trout. Donaldson and Kershaw seem to be the consensus #2 and #3 projections in baseball, and both of them are just barely in the “comfortably over” range for me. Maybe I’m too risk averse to be making bets 🙂
Betting the over on 5.4 WAR is a losing bet for all but the very best players in the history of the game. Right now Trout is the only guy I would take the over on 5.4 WAR.
Seems like fan projections suffer from an inherent selection bias, since fans who are excited about Gary Sanchez are much more likely to go to the trouble of submitting a Gary Sanchez projection than fans who aren’t excited about Gary Sanchez.
Very solid point.
Especially for younger players like Sanchez.
This came up in the chat today, but the other player for which this reminder is very relevant is Trea Turner. Just as impressive a half-season, if not even better.
I was going to post this very same thing.
YES! I couldn’t agree more with this.
This article could of been written for him. I feel as if there is MUCH more risk at Turner’s current draft position (end of first round) than Gary Sanchez’s position of around 50. It would be shocking to see him not hit 20 home runs. The bar is very low for production at catcher, and barring injury I do not see how he does not end up near the top of catcher production.
And David Dahl, especially now that he is dealing with a back injury.
Definitely Dahl. He had a terrible K/BB spread and only managed a 111 wRC+ with a .404 BABIP.
The projections are not human though, right? Any reason to believe they do not properly evaluate SSS? I don’t see why that would be. Just because they got it “wrong” on Brett Lawrie doesn’t mean much to me.
Because Sanchez’s 2016 major league performance was so far beyond anything he had ever done in his life, including his 2016 minor league performance. It is the most recent performance and therefore it is possible that he made a Daniel Murphy or J.D. Martinez type adjustment that took his swing to new heights, but it’s still so far out of established character that it merits skepticism. He hit more home runs in 53 major league games than he did during his previous calendar year in AAA.
His 2016 was actually very similar to his 2015 AFL performance, and was actually quite similar to his AAA performance when you adjust for the fact that Scranton and Trenton are an insanely pitcher-friendly park relative to two of the most pitcher-friendly leagues in professional baseball.
http://www.statcorner.com/bat/596142/Gary-Sanchez
Note: Statcorner’s wOBA+ is a more compressed scale than wRC+.
Here’s Miguel Cabrera’s page for reference:
http://www.statcorner.com/bat/408234/Miguel-Cabrera
Even if that were true, what does it have to do with the projection systems (non-fan projections) and SSS?
Thank you for this desperately needed corrective! ESPN.com just ran a post inquiring “Is Gary Sanchez Already the Best Catcher in the AL?”, and I wanted to scream out at the computer something along the lines of what Travis just wrote much more eloquently. Part of the problem with Sanchez in particular has to do with a noxious syndrome I call “What Would Cincinnati Say?”: in this particular case, what would the narrative be if Sanchez played, say, for the Reds? Nobody overrates their athletes the way New Yorkers do, and, alas, that includes New Yorkers who work for national outlets. Bill James writes in his historical abstract that he had to learn to “let some air out of” descriptions of NY players in the NY press as early as the 1910s.
“Nobody overrates their athletes the way New Yorkers do.”
Is there any evidence for this? It’s a popular narrative, but Jim Rice is a Hall of Famer while Nettles, Randolph, Munson, Mattingly, Bernie Williams, Posada, and Mussina are not.
My (non-scientific) take on this is that oft-repeated hyperbole definitely favors NY athletes, but awards and HOF voting don’t, or at least haven’t very much for the last decade or two. I’d be very curious to see numbers on this, though.
My (also non-scientific) take on this is that there are simply more of every sort of New York fan (smart ones and dumb ones) because it’s a bigger city than everywhere else. Also the stereotype of the loud, obnoxious, New Yorker comes into play. It isn’t that New York fans overrate their athletes more than other fan bases. There are just more loud obnoxious homers because there are more fans in general.
Nailed it! No one would be asking whether Sanchez was the best C in the American League if he played for the Reds!
However, asking whether a player who had more WAR in 2016 than any other AL catcher is the best catcher in the AL doesn’t seem that hyperbolic to me (although the answer is “No, Jonathan Lucroy is the best catcher in the AL”).
It really speaks more to the sorry state of AL catchers than anything else.
I think you are misweighting the small sample sizes the other way if you think the under-20 HR’s bet is anywhere close to odds-on.
There’s a big difference between, “Gary Sanchez probably isn’t going to slug .700 in a full season,” and “Gary Sanchez probably isn’t the number one or two catcher for 2017.”
BP wrote a thing some years back that said FG’s live projections were overweighting recent performance. Later, they changed their own weightings, and declined to say which direction they changed them in. Sure, misweighting toward the SSS is more common, but it’s just a much a crime against accuracy to hand-wave away SSS’s as nothing.
While SSS is absolutely a thing, this was a signature performance – no one not very good could do it. Bill James noted this with *one* Roger Clemens game early – you strike out 15 guys and walk zero, and the needle seriously moves.
Arguing about it doesn’t mean much, though. I’ve made Gary Sanchez bets elsewhere and if anyone wants a piece, I’m easy to find. (Loser pays charity of winner’s choice.)
“this was a signature performance – no one not very good could do it.”
I mean, Sandy Leon looked like a stud out of nowhere from June through August, and I think most people would probably err on the side of “he’s not very good”. I obviously think Gary Sanchez is way better than Sandy Leon, but the point is, players who aren’t very good get hot as heck all the time over small time periods. Brandon Morrow K’d 17 guys, etc.
Of course, Leon put up a .113 higher BABIP during that stretch than Sanchez did on the season…and Gary’s offensive numbers were still a full step above, haha.
Beware small sample sizes.
Single player comp.
comment deleted
People are over looking his near 50% ground ball rate he had during this SSS, that can certainly improve, which could provide even more dingers. I think .260 20 80 is a super safe baseline. Joe Girardi already admitted that he’s the best hitter in the lineup and will get a ton of ABs. I have no clue what the AVG will look like but I can totally see a player who hits 30HR with 100rbi a year.
Gary Sanchez will not be on 1 of my teams, especially at the idiotic price people are paying for him. His minor league stats show me who he is. Also, at the end of the season he went 4-35 with 1 HR. You can have him
I mean…his minor league stats (which came in some of the most pitcher-friendly parks in the MiLB) show you that he’s a .275/.339/.462 hitter with 24 HR/650 PA.
His 129 MiLB wRC+ (which, again, isn’t park-adjusted) would have led MLB catchers last season. His 24 HR would have tied him for 3rd. His AVG and OBP would have been 6th.
If you look at only the last two years (again, in two of the worst offensive parks in the MiLB), he’s been a .278/.337/.490 hitter with a 135 wRC+ (not park-adjusted), and 27 HR/650PA.
Nice numbers for a catcher. But again, those are minor league numbers. Most players fair worse when they get to the majors. Add in pitchers adjust, which they might have already. Let’s see what he does to re-adjust. If you like him, pay the price in a league, myself, I won’t
So…you’re saying his minor league numbers don’t tell you who he is?
And his major league numbers don’t tell you who he is.
I might be temped to pay way less, and assume the Astros will come to their senses/assume Beltran or McCann gets hurt/assume Gattis gets 500 PAs.
That way he can build off his equally impressive second half. Just as Sanchez was tearing it up, Gattis was hitting .288/.358/.594, .396 wOBA and 153 wRC+ over the same second half of last year.
He doesn’t have the allure of that jaw-dropping small sample, but I’d give even money that Willson Contreras turns out to be more valuable by WAR in 2017.
Is it the higher BABIP with a worse Hard%, the lesser defense, the worse BB/K, or the lower scouting pedigree that puts Contreras over the top for you?
It’s not that I’m expecting Contrearas to light the world on fire. His projections (2.4-ish WAR) look about right. Sanchez is very likely more talented than that. But I’m wary of seeing his fluky luck balance out. I guess I think Contreras has an underrated floor and Sanchez’s coming down to earth might get rough.
Fair enough.
“But I’m wary of seeing his fluky luck balance out.”
But is that how these things work? If I get 10 heads in a row with a fair coin, aren’t my future tails odds still just 50%? We don’t expect extra tails going forward to balance it out, right?
I don’t think these factors are as simple as you are making them out to be.
– The BABIP was higher, but Contreras had a higher GB%, slightly higher LD%, lower FB%, much lower IFFB%, and higher center and oppo%. Contreras wasn’t consistent with BABIP in the minors, but at the very least it’s worth mentioning that he put up .370 in AA and .382 in AAA, which were his two post-breakout samples. Sanchez has no significantly abnormal BABIP history. Contreras projects for a .310 by depth charts, Sanchez .285.
– As for better defense, that’s very much up in the air. Both have excellent arms, we know. We don’t have much in the way of stats; Sanchez beat Contreras in DRS 4 to 1, but Contreras beat Sanchez by four runs in framing per StatCorner. These stats might mean nothing at this point, but it’s definitely not a sure thing that Sanchez is a better defender.
– The K/BB spread is also not that dramatic. Contreras had a 14.5 K-BB%; Sanchez, 14.4. Minor league track record gives Contreras an advantage. In particular, Contreras ran a nearly 1:1 K:BB in his two post-breakout MiLB samples, while Sanchez consistently hovers around 10-12% spread in the upper minors. Contreras projects for a 19.6 K% and 8.2 BB% by Depth Charts, while Sanchez comes in at 21.4% and 7.4%.
– Contreras has less historical pedigree but was, at the very least by MLB Pipeline and I believe by a few other sources, rated above Sanchez coming into 2016 as the game’s best catching prospect.
– Sanchez’s overall projection beats Contreras’ on the back of a much higher ISO – an ISO he hasn’t come particularly close to since A ball. Contreras’ projected ISO, on the other hand, rests between his A and AA numbers, well below his AAA and MLB – the former of which tops anything Sanchez did in the minors outside rookie ball, albeit in a very hitter friendly park. I’m not trying to argue Contreras has more power than Sanchez, but I don’t think the difference is nearly as decisive as the projections suggest.
So basically these two are a lot closer than your comment suggests.
Re: power, you have to account for ballpark. Very hitter friendly park vs. very pitcher friendly park is a big difference. I don’t think there is a reason to doubt the projections there.
Yeah, I’ve seen stronger arguments than “Player A had slightly better power numbers in the PCL than Player B did in the least hitter friendly park in the IL, so that should take precedence over Player A’s much worse power numbers when they were both in the FSL, AFL and MLB.”
Would those stronger arguments include the other 95% of my post? All the power section is meant to do is cast some doubt on the degree to which the projections favor Sanchez’s power, not on the fact that Sanchez has more power.
My greater point was that your dismissive list of Sanchez’s supposed advantages over Contreras was not a valid argument.
Sanchez had a .040 higher MiLB ISO, while spending his MiLB career in the least hitter-friendly parks in the least hitter friendly leagues in professional baseball (the SAL isn’t bad for offense…but Charleston suppresses HRs by over 50%), than Contreras did while spending time in the Midwest League, the Southern League and the PCL.
Even if Sanchez’s MLB ISO hadn’t been over .150 points higher than Contreras’s, the projected power gap would be appropriate.
I’m operating on the assumption he’s not going to be Piazza immediately (if ever). But, if he puts together a “disappointing” season of 25 HR, 90 RBI, and a .250/.350/.500 line with league average catching and 4.0 WAR…I’d take it.
Is this really about SSS, or hope, or not knowing what most other hitters really do to compare him to, or the halo effect? Or any number of things? Fans over rate players all the time, we know this from the Fans projections vs the computer projections. That’s not just about SSS at all.
speaking of small sample size, remember in 2015 when Charlie Morton had 1.62 ERA, 5 Wins and 33.1 innings in his first five starts (roughly 1 month of baseball)
…funny how that “extrapolated” to a 4.81 ERA, 9 Wins and 129.0 innings at the end of the year.
I get the excitement he created last year, but you have to step back an analyze was it a fluke, a sign of things to come or something in between. Ultimately to believe a career minor league player who never hit 20 HR’s in his first 6 seasons is going to be a 30 – 40 HR guy is crazy. Mike Trout is an exception and he did it after making it to the majors in only his 3rd season. Trout’s miLB slash was much more impressive than Sanchez’s (.334/.414/.499 to .275/.339/.460). I am not saying Sanchez isn’t going to be real good, just step back and take 20 HR’s for an entire season with a .270 BA and .450 slg from the catcher position.
And for a better comparison than Lawrie… How about Brennan Boesch? His first 63 games in the majors he had 12 HR’s in 235 AB’s for a slash of .345/.402/.600. Unfortunately he had to play the next 70 games of the season and it didn’t go as well.
This would be a much stronger argument if he hadn’t hit 25 HRs in 2015…or 17 HRs in 83 games back in 2011.
…and if he wasn’t moving from Trenton, Tampa and Scranton to YSIII.
That said: I’d be pumped with .270/.340/.490 and 25 HRs for a ~120 wRC+.
To hype the Gattis train again as a potential Sanchez alternative for those that dig big second halves. Top wOBA in the second half of last year, C eligible.
1. Gary Sanchez – 225 PAs, .305/.382/.670, .432 wOBA, 176 wRC+, 10.7 BB%, 24.9 K%, .325 BABIP
2. Evan Gattis – 241 PAs, .293/.361/.605, .401 wOBA, 157 wRC+, 9.1 BB%, 26.6 K%, .328 BABIP
I almost always stay away from second year players in fear of the dreaded sophomore slump. After reading your reflection on Lawrie and thinking about it, I wish I could go back and take a look at the projections for second year players, especially the ones hyped like this (top 100 ADP), to see how they have historically compared to their projections. How often are any projections for second year players close to the mark? I would be curious to see if that is different when compared to more established players. Do those projections get better when the player in question has a larger rookie sample in the bigs (200 PA v 400 PA)?
I totally have a sophomore bias and I am not sure if I should. I can tell you right now, I am staying away from Sanchez, Story, etc. like I did with Correa and Seager. I guess it is all relative to draft position and the perceived floor (demotion to minors!).
Anyone remember Joe Charboneau?
CAREER:
AB: 647
H: 172
R: 97
HR: 29
RBI: 114
SB: 3
BA: .266
OBP: .329
SLG: .453
OPS: .782
OPS+: 115
Looks similar to what we are projecting Cespedes or Cargo this year. Except he had one outstanding year where he played 80% of the games for Cleveland followed by complete ineptitude the following two seasons.
Injuries.
And a manager who, in spring training, told him he had to compete with Miguel Dilone for playing time.
MLB history is full of “joes”, not so much of “garys”. Regression is a given… just how far is the question. He is much more valuable to the Yankees than any owner in fantasy because of what he does with the glove and arm.
He had a great month in August. It looked like pitchers adjusted well in September although the power was still impressive.
Last year in ST he went something like 1-22 or 1-27 and folks said he was a bust. It seems as if he may be a pretty streaky player.
The MLB ball is certainly juiced compared to the minor league balls. I see him as at least a 25 HR guy. Maybe 40 HR if he can increase his FB rate. That presumes that MLB does not deaden the balls too much after last years HR barrage (especially in the AL where HR rates were at an all time high)
I don’t think drafting Sanchez as the second catcher off the board is overrating him. It simply doesn’t take a great offensive season to be the second most valuable catcher for fantasy baseball. Buster Posey ranked in that spot last year by hitting .288 with 14 home runs, 6 steals, 82 runs scored, and 80 RBI.
Normally I’m the guy screaming “beware of small samples” but Gary Sanchez can regress *a lot* and still be a Top-5 catcher in MLB. He put up 170 wRC+ at catcher. The top guys at catcher (excepting Posey who is amazing) the last few years have run about 120 wRC+. So he could lose almost 1/3rd of his offensive value and still wind up as a 4-win player. That, and he’s only 24 (so he might still get a bit better), he didn’t have some crazy BABIP-fueled run, and it wasn’t that small a sample (229 PAs).
It’s pretty obvious he’s not going to hit like Mike Trout next year but he doesn’t need to. If you were running a baseball team and wanted someone only for next year, who else would you take? Posey, Grandal, Lucroy, and then who? Martin? Cervelli? Leon’s likely to regress even more, Ramos tore his ACL, Gattis isn’t a catcher, and Contreras is kind of like Sanchez except he put up 2/3rd the WAR in more PAs.
So let’s not give Sanchez an MVP just yet but he could lose an incredible amount of offense and still be a deserving all-star.
Was Jesus Montero too obvious of a comparison?
Why would that be an obvious comparison?
I have no idea.
They’re not similar players.
You say projections “can still struggle with small samples”, do you mean they systematically tend to overreact like humans do? I’d be pretty surprised if they do that, since that’s just a matter of tuning the regression knob, and that’s what the designers do this for.
If they just “struggle” in the sense of have high error when running on little information, well, yeah, all of us on this earth.
“pump the breaks”
One of those rare typos or misspellings that actually conveys the opposite of what was intended.
While I’m being technical, also remove “he” following “Jim Bowden predicted Lawrie…”
I don’t think your argument is logically consistent.
To establish that fans may be overreacting to a small sample size, you refer to an ESPN survey, projection systems, a fantasy ranking and scouting. Two of those may indicate that fans are overreacting, but the other two suggest the exact opposite. While I don’t know exactly how ZiPS or Steamer arrive at their projections, neither should be influenced by “human nature.” They’re programmed by humans. But they’re not programmed to think like humans, and if sample size is not accounted for, that’s staggering. And you’d really need to prove that’s so.
Which renders this “twist” utterly perplexing.
“Everyone, it seems, loves Sanchez in 2017. But let’s pause and take a deep breath. Have we learned nothing?”
You’re assuming your own conclusion without supporting it. If indicators of fan enthusiasm, like the ESPN survey and Sanchez’s draft position, were opposed by systematic indicators, like statistical projections, perhaps fans should take a deep breath and think over their assumptions. But quite the contrary is true. Suggesting that you may be the one who is clinging to a preconceived notion, inflexible to the influence of contrary information.
I enjoy Michael Lewis and love the work of Tversky and Kahneman (Everyone should read “Thinking, Fast and Slow”. It’s a humanist Sermon on the Mount, kinda) but the quoted text doesn’t at all match your argument. Your point is, I guess, that in 1971 when “Belief in the Law of Small Numbers” was published, Tversky and Kahneman discovered that “Even people trained in statistics and probability theory failed to intuit how much more variable a small sample could be than the general population.” But you can’t just generally apply that, I don’t know, heuristic? to all projection systems. Which I think is your intention. Nearly 50 years ago, some people “trained in statistics” “failed to intuit” &c. Did that failure of intuition mar their work? Do the creators of modern projection systems suffer the same bias? The gap between that quote and who and what it’s being applied to is vast. You’re generalizing and the utility of Tversky and Kahneman’s findings are wholly lost in that generalization.
Speaking of which: Immediately after citing evidence of the human inclination to overvalue small sample size, you wrote “Remember, Sanchez hit .225 in September after a torrid August.” The rest of the post is marked with this strange, self-contradictory combination of assertion and evidence.
“[Projection systems] can still struggle with small samples and the translation of minor-league statistics to major-league equivalencies.” But where evidence to support this argument should be, we instead get a collection of anecdotes about Brett Lawrie. Goldstein and Bowden were wrong, as was the collective intelligence of the fantasy baseball community, but which projection system was wrong? And if it were wrong, why? and does Lawrie challenge some of the assumptions of those projection systems? or is he, looking at the big picture, one of a predicted number of improbable “misses” that must exist for the projection to be accurate? I.e., if there’s an 80% chance of rain, but it’s sunny, the prediction wasn’t wrong, only the less likely outcome occurred, as it must, like, a fifth of the time.
After this there is some padding, some hedging, and some interpretation of an out-of-context quote. I’m not trying to be rude. Maybe I am coming off that way. I am not full of hate nor bitterness nor resentment nor ill will, and wish you success. Only I’m terrified when people entrusted to enrich, educate and inform normalize faulty reasoning–especially when they’re assuming the tone of wise and impartial observer, employing reason to cut through bullshit.
(Of course, maybe I’m guilty of this very thing now! And the joke’s on me! This flabbergasting, inescapable human condition!)
People are praising this post. The top comment describes it as “very important.” I think what it is is a heuristic technique, an old one, misapplied and poorly supported. If you had written, as if a tweet, “Gary Sanchez’s sample of major league at bats is still small and we should not project too much from it,” we would be in agreement, but my optimism re: his future would not be greatly reduced.
What constitutes a good sample size varies a great deal. A computer could analyze one chess match and have a good idea how skilled both players are. One race is sufficient to determine if someone is fast. From what I’ve read, thousands of games are needed before quarterback interception rate stabilizes, and even then it’s a weak indicator. Gary Sanchez may fail. Mike Trout may fail. Best evidence suggests Sanchez should be damn good, so sterling was his brief time in the majors, so good was his performance in the minors, people are both rationally and (thank God) irrationally excited. That’s great. This time of year anticipation of the potential feats of a young superstar is about the best feeling possible for a baseball fan. Few doubt that he could fail.
Unfortunately, this post will play on the irrational nature of humans, so that if Sanchez struggles, it will seem prophetic. And if he doesn’t, you’ve written in enough safeguards. But, in truth, it’s already wrong, because it’s poorly argued, and the argument is the process, and the process is the post, and the post has both discouraged and misled.
But I’m an ass. So …
I regret writing this.
… but see no way of deleting it. All the links on the menu which appears when I hover my mouse over “Welcome John Morgan!” lead me to an Amazon listing of a Hewlett Packard computer.
Misplaced anger.
As Jay Peterman said…
“Wellll…that certainly is a lot of words.”
FWIW, I enjoyed your lengthy comment, and found it well reasoned and relevant. Many worthy ideas can’t be expressed in 144 characters.
http://www.fangraphs.com/statss.aspx?playerid=1007895&position=1B/DH
Imagine Sanchez’s April and May = Joc Pederson’s rookie year 2nd half. In NY. He’s as likely to fall off a cliff as maintain his happy yum-yum numbers, seeing pitchers a 3rd and 4th time.