Corey Kluber, the Cubs, and Small Sample Size
The World Series Notes column that ran earlier today included quotes from Dexter Fowler and Ben Zobrist on the subject of Corey Kluber. More specifically, their lack of success against the Indians right-hander. The sample sizes are small, but nevertheless real. The two Cubs came into tonight’s game a combined 1 for 20 against Kluber.
They weren’t alone in their woe. The nine players in Chicago’s starting lineup were 4 for 35, with 15 strikeouts, in their cumulative career against the Cleveland ace.
Do small sample-size results mean anything in a given game? Conventional wisdom says no. It is, after all, small sample size. That doesn’t mean it can’t hint at future performance. Some players simply don’t see the ball well against certain pitchers, which is something you can’t quantify. And if a player isn’t careful, the conundrum can go from his eyes to between his ears.
“It’s a mentality,” said Cubs outfielder Chris Coghlan. “Some pitchers, the more you face them, it domes you up. You’re like, ‘Man, I’m getting out all the time. I don’t feel good.’”
“Guys know if they’re comfortable against a pitcher or not,” confirmed Cubs hitting coach John Mallee. “They know how they felt in the box, and that’s something you can’t see in the numbers.”
Tonight’s numbers looked all too familiar to most members of the Chicago lineup. They went 4 for 22 against Kluber, with nine strikeouts. Eight of those punch outs came in the first three innings.
Did a multitude of Cubs lack confidence in the box in Game One of the World Series? The answer to that question is an unequivocal no. This is one of the best hitting teams in baseball. They aren’t about to be cowed, no matter how good the pitcher.
That doesn’t mean subconscious doubt didn’t begin to creep into a few heads. As for how well the NL champs were tracking the ball, the number of swings and misses, and called strikes, tell a story.
Which brings us back to sample size, which now stands at 8 for 47, with 24 strikeouts. Still too small to be meaningful in a certain sense. As much as anything, what it says is that Corey Kluber is very good.
The Cubs will face Kluber at least one more time this October, and while they’ll do so with stiff upper lips, it’s hard to imagine them being fully confident.
David Laurila grew up in Michigan's Upper Peninsula and now writes about baseball from his home in Cambridge, Mass. He authored the Prospectus Q&A series at Baseball Prospectus from December 2006-May 2011 before being claimed off waivers by FanGraphs. He can be followed on Twitter @DavidLaurilaQA.
Wow, kluber pitches yet another awesome postseason game but fangraphs still manages to make this article about the cubs. How about just officially declaring fangraphs to be a cubs’ fan site?
Probably the only time I will agree with a Giants fan.
Hey, that 9th inning of Game 4 of the NLDS sure was fun, wasn’t it? We need more articles on that, I think. You’d find a much more receptive audience if you went and whined about the Cubs at McCovey Chronicles or some other site where Giants fans congregate online. Here, where there are a diversity of fans, it’s just annoying.
There are two teams for which there is anything of present value to speak. Both are getting tremendous amounts of attention. If the Giants hadn’t lost, there would still be content on them. But they lost.
Don’t look at me, I didn’t mention them.
I sometimes like to use statistical significance as a quick check on whether tiny-sample numbers are or aren’t meaningful. The odds of going 8 for 47 or worse, if your true talent batting average is .250, are 13.5% (assuming independent, identical trials). That’s… not very significant. Obviously this is an extremely imprecise and non-rigorous test, but for what it’s worth it says there’s not a ton of reason to think there’s anything real here, especially since Kluber is good anyway.
(One thing this technique suggests, by the way, is that having really good numbers against a pitcher is probably more meaningful than having really bad ones. If you see someone who’s 4 for 5 against a pitcher, well, there’s only a 1.5% chance of a .250 hitter getting four hits in five at-bats–far more significant, in this sense at least, than hitting .170 over 47 ABs. I’d be curious whether this is actually true, that extreme success against a given pitcher in tiny sample sizes is more predictive than extreme failure.)
What about if you chose batting averages that are equidistant from their true talent? It seems intuitive that a .250 hitter hitting .800 in a small sample has a lower chance of happening than a .250 hitter hitting .170. What about a .250 hitter hitting .050 over 40 AB (2/40, I think) versus .450 over 40 AB (18/40). Are the odds of those things equal? Because then I’d think they’re equally predictive.
Yeah, the point is basically just that since batting averages are well below .500, there’s far more room to perform way above them than way below them. (If we were to shift to on-base percentage rather than batting average then this would not have been true of Barry Bonds for a while, because, y’know, Barry Bonds.)
As for your examples, the odds of a .250 hitter going 2 for 40 or worse are 0.10%; the odds of a .250 hitter going 18 for 40 or better are 0.47%. I’m not sure exactly why the former is higher than the latter; possibly there is some effect where getting close to zero (or, presumably, to one) makes things even more unlikely.
So yes, hitting .400 against someone in a small sample is fairly meaningless (though it can become meaningful in larger samples). Hitting .800 against someone, though: maybe not meaningless!
8 for 47 isn’t that extreme, it’s true. But 24 strikeouts in 47 abs is a bit much.
Anything COULD be. That’s the ultimate weasel word in an argument. In The Book, we looked at multitudes of players, collectively, who absolutely destroyed pitchers or were destroyed by pitchers in 15 to 30 PA per batter. What did they do in subsequent PA against the same pitchers? Almost identical results. (see pages 74-75 in The Book.)
So the “could” is really “extremely unlikely” and we have to assume that there is virtually ZERO predictive value in batter/pitcher matchups even for 30 PA which is huge of course as batter/pitcher matchups go. (As these things go, if you want to use them as a tie breaker no one will argue with you.)
It’s nice to think that destroying a pitcher in 20-30 PA would or could create confidence that would carry over, or that getting destroyed by a pitcher in that many PA could create doubt which would also carry over to future performance, but what we “might” think or what seems to be intuitive is quite often wrong when we actually do the requisite analysis in professional sports. In this case, there is overwhelming evidence that if there is any carryover effect, it is microscopically small, and if there are some batters for whom there IS a predictive effect, the percentage of those batters is also microscopically small and we have NO way of identifying them from the numbers, regardless of the sample sizes.
We got this “sabermetric” site where you find an article about pitcher vs team stats in 40ish PA. Just a waste of everybody’s time. Makes it harder to find good analysis. Probably by the name of the article you should know what you are in for, but I don’t know what’s the point of publishing this, when probably even the author knows that it isn’t really saying anything.
In a nutshell, what the article says is that Corey Kluber is a great pitcher, and that a small subset of players may not match up well against him. By no means was it meant as an argument against what we know about sample size. FWIW: It is also a short piece written in short order, timed to coincide with the end of a World Series game. That’s why it’s an InstraGraphs article, as opposed to a main- blog article.
I’m sorry if I came off too rude, my comments were probably out of line. And you are right this is a short piece written in a short time as a reaction of the game. I should see this with different eyes as HarryLives says, just as what players think of these kind of stuff. I apologize.
If you’d read any of David’s stuff before you’d know that he focuses on getting player reactions and trying to get inside their heads a bit, while often using some sabermetric concept as a peg. His stuff is different than much of the more analytical articles on Fangraphs, and as a fan, I find the player’s perspective on various sabermetric concepts to be interesting and informative. You’re criticizing this for not being what you expected it to be, which is a dumb reason to criticize anything. David’s too polite to say any of this to you, but I’m not.
Why is anybody discussing the substance of this article when they could be lauding Chris Coghlan for using/inventing the phrase, ‘domes you up’ for, I don’t know, thinking?