Clutch Baseball Teams aren’t Clutch Baseball Teams
You’re familiar with win/loss standings. I’m pretty comfortable with this assumption. You’re likely also growing familiar with BaseRuns win/loss standings. It’s something we’ve cited pretty often since we began to offer the data, and the idea behind BaseRuns is that it strips away sequencing. Actual standings show you performance plus variation. BaseRuns standings show you performance. It’s not quite that simple, but that’s the outline, so it’s interesting to compare how teams have done to how BaseRuns thinks teams have done.
Take the American League Central, right now. The Royals are leading! The Royals are leading the Tigers, by half of a game! Yet, BaseRuns preserves the Tigers’ winning percentage, but drops the Royals’ winning percentage from .542 to .480. BaseRuns doesn’t think the Royals are as good as the Tigers at all. So why are the Royals presently where they are? They’ve been clutch. Sequencing has been a strength of the Royals, and of course, sequencing can make an enormous difference if you under- or over-perform in high-leverage situations. The Royals have earned their current record by doing well when it’s mattered the most.
Clutch makes for a really good explanation for differences between winning percentage and BaseRuns winning percentage. Check out this graph of 2014 data, comparing that winning-percentage difference to overall team Clutch score.
This makes total sense — if a team isn’t performing to its BaseRuns numbers, presumably it’s because of performance bunching, or some better term. The five teams with the biggest positive winning-percentage differences have a combined Clutch score of +19. The five teams with the biggest negative winning-percentage differences have a combined Clutch score of -17. Clutch is basically why the Orioles have a massive division lead on the Rays; BaseRuns considers them the same, but the Rays have been lousy when it’s mattered, and the Orioles have been really good.
Fans can feel this stuff, too. Fans know when their teams have been clutch or unclutch, even if they’re unfamiliar with the statistics. And everybody understands the importance of doing well in stressful moments, because, say, a blown save can totally negate eight innings of excellent work. Fans of unclutch teams are going to develop trust issues, because they remember past instances of failing to come through. Fans of clutch teams are going to search for explanations of why their teams are so great. Usually, people will point to bullpens. Bullpens, of course, throw important innings, so they can make or break a lot of baseball.
Here’s the thing, though, and you probably already knew. You know all those studies about clutch performance in the majors? You know how they haven’t found much of anything? There aren’t clutch baseball teams. There are baseball teams that perform well in the clutch, but it’s not a skill; it’s just a thing that happens sometimes.
Following, you’re going to see a comparison between first-half team clutch score and second-half team clutch score, within the same seasons. I only looked at five years of data, between 2009 – 2013, but I didn’t see any need to go longer. The information here speaks for itself. How’s that relationship look, over the 150 team seasons?
There’s nothing. It’s not a completely flat line, and it’s not an R value of literally zero, but for all intents and purposes, this is randomness. You don’t even need the numbers to know it’s random — you can just eyeball the distributions. Clutch team for the first half? Great! Don’t count on that in the second half.
Of course, some clutch first-half teams have remained clutch second-half teams. Some unclutch first-half teams have remained unclutch second-half teams. But that’s what you’d expect from the sample. That’s how numbers work. The top ten most clutch first halves had a combined Clutch score of +52. Those same teams had a combined second-half Clutch score of +1. The top ten least clutch first halves had a combined Clutch score of -53. Those same teams had a combined second-half Clutch score of +6. It’s meaningless, but you’ll notice that’s even better than the most clutch teams in the first half. There’s absolutely zero predictive power at all.
Yeah, the 2012 Orioles were amazing. In the first half, they were at +5.6, and in the second half, they were at +5.2. But consider the 2012 Pirates. In the first half, they were at +5.3, and in the second half, they were at -4.1. The 2013 Orioles, in the first half, checked in at +5.3; the 2013 Orioles, in the second half, checked in at -2.4. On the other side, the least-clutch first half belonged to the 2009 Nationals, at -9.5. In the second half, they came in at +3.2. Last year’s Brewers turned it around. The year before, the Phillies turned it around. And so on. When you’re dealing with randomness, there’s some continuation and there’s some discontinuation, but that’s to be expected, because it’s randomness. Half the time, a team with a positive first-half rating will have a positive second-half rating, just because of basic mathematics. It’s not because the team has a special skill. It’s because a baseball season doesn’t achieve an infinite sample size.
This doesn’t mean that, say, the Royals don’t deserve to be in first. This doesn’t mean that the Orioles don’t deserve to have their big lead on the Rays. This doesn’t mean that clutch teams are about to collapse, or that unclutch teams are about to catch fire. That’s not the way regression works, and the season’s already three-quarters complete. All this means is that clutch teams aren’t clutch teams. It’s great to get a quality high-leverage performance, but it’s not something you can count on over and over, not when the performance exceeds the performance you’d already expect just from the players’ talent. Looking back, the Orioles have been a lot better than the Rays. Looking forward, one shouldn’t expect the Orioles to be a lot better than the Rays. You’re already well aware of the role and importance of randomness, but from time to time we all need to be reminded.
Jeff made Lookout Landing a thing, but he does not still write there about the Mariners. He does write here, sometimes about the Mariners, but usually not.


What it means is, everything averages out. Right now the Royals are starting to do a positive move toward what their overall average for the year should be and the Tigers are doing the opposite. Which makes sense because most would agree the Royals have underperformed on offense most of the year. The Tigers are in a full scale meltdown with injuries etc., but many didn’t think they were anything special at the start of the year and they seem to be playing the part now.
That’s… that’s not what it means at all.
Reading comprehension score: F
Math score: F
Royals fandom score: F— YEAH!
And honestly, as a baseball fan on the positive side of this, number 3 is all that matters.
“Regression to the mean” does not mean that a time that experienced positive luck for awhile will experience an equal amount of negative luck in the future. Rather, “regression to the mean” means that you should project average luck going forward for all teams.
Imagine you flipped a fair coin and came up with heads ten straight time. Projecting forward, you should not expect to flip tails ten more times than heads in order for your overall count to even out. Rather, you should project an equal number of heads and tails going forward.
team* in the first sentence
I wouldn’t bet on that.
And that is why casinos make a lot of money.
You’re dead
Tell that to the Giants.
Or you might conclude that the coin is loaded
Even with a 50/50 distribution, there is a non-zero chance that a coin flip will land the same way 10 times in a row. It’s not true to say that it has any greater or less chance to land that same way on the NEXT flip. It’s still 50/50, since they are independent events.
http://en.wikipedia.org/wiki/Gambler%27s_fallacy
If a coin lands on the same side 10 times in a row, be a good Bayesian and ask if it’s really a 50-50 chance.
Nope. Gambler’s fallacy.
Jeff, you are exactly right. You are welcome to come to Atlantic City whenever you want. Drinks are on the house.
“Wait til those moron baseball players and analysts see this!! Then, they will know how dumb they were to think clutch existed. One look at these baseball article charts and graphs and they will admit that we know more about baseball than they do! Plus, this gives us another chance to get a sabermetric boner over the Rays.”
Oh get a life
“Yeah, and we’ll make sure they see it by posting it on a website frequented by statistically-minded baseball fans and players!”
If there was a start for it, I’m pretty sure that this comment would max out the SPA — stupidity probability added.
This doesn’t explain why the Cubs have a negative clutch score year after year after year.
they were basically average between 2008 – 2011
Can we re-run this investigation only while including “curse” as a variable?
I blame Wrigley Field! The place is so beautiful and relaxing it is hard to focus on much other than the Ivy.
An example that past results do not attribute to future results in regard to regressing to the mean:
I once was at a casino playing roulette. The marble landed on red about 8/10 times. I thought to myself, it’s bound to land on black for a long streak, so I started betting every spin on black. The marble did not land on black but instead repeated nearly the same sequence of 8/10 spins on red. The odds did not change to favor black just because of a recent history of frequently landing on red.
That’s how casinos make their money. And how crooked ones make even more. Rig the wheel.
Why on earth would you rig the wheel and risk being shut down when you can run a fair game and keep getting a 5.26% win percentage (give or a take a bit) month after month, year after year, in perpetuity?
Yeah, if the wheel is rigged there is actually a chance for the customer’s to win. A fair wheel favors the house.
And it’s why casinos have boards to show the last ten spins. Trends attract the (un)smart money
Of course in roulette, you can bet that those all those red numbers will suddenly turn into blacks to “catch up”, or that the “hot run” of reds will continue. Doesn’t matter much, the house edge is still the same.
Smart money isn’t betting on roulette.
Gambler’s fallacy, dice have no memory, marbles aren’t sentient, roulette pockets don’t acquire repellent forces, etc…
Really great article. Keeps me going at work ha ha ha.
It is so easy to fall into the it of “This guy is very clutch, he seems to always get a timely hit”. I think the biggest reason this thought process is perpetuated is baseball announcers/broadcasters and the stats they show. They are in love with hitting with RISP. Hitting with bases loaded, hitting with 2 outs. I have to constantly remind myself that baseball players play baseball, and they do their best every chance they get.
My only qualm is momentum. There is something to players performing better during a “hot streak”. But again this is not predictive.
Ya got that right.
Momentum’s probably attributable to health. A guy’s hot because he’s healthy. Then he gets tired, sick, or microinjured and he gets cold. He gets hot again when his health improves.
You could probably make it predictive if you could measure health so precisely or reliably identify things that negatively affect or improve health (e.g. expand on the things we already know, like playing 4+ games in a row sees a performance penalty, etc)
Fielding and relief pitching. That’s clutchety clutch clutch.
Except when it’s not.
This reminds me of how the Cards were doing unbelievably well batting with runners in scoring position, for a good long time too.
And then, not so much. It’s possible to have a long streak of “clutch”, but it doesn’t mean you should expect better or worse luck going forward.
I guess its safe to say the mediocre O’s are awesome at being good even though they aren’t, and the obviously better-in-the-future Rays just suck at being good. Maybe theres been a disturbance in the spacetime continuum that created this unexpected success of the Orioles. Or, like Lefty Gomez used to say, “I’d rather be lucky than good”.
I’d rather be good. Luck runs out.
In the hypothetical luck is as persistent as good.
Not so sure ….on that O comment!
What I think can’t be emphasized enough is that the lack of real clutch is the central organizing principle of sabermetrics. The reason we can evaluate a hitter by weighting his various offensive events is because we assume that the actual value of those offensive events—e.g., whether a single occurs with the bases loaded and drives in two runs, or with two outs and no one on base and doesn’t lead to any runs–can be ignored in favor of an average value. At any one time (i.e., particular season, or maybe even as updated game by game throughout the season), a single has the same sabermetric value for any hitter, even though it may have very different results in different contexts.
Once this is accepted, anything that a hitter, fielder or pitcher does can be given a value that is context free, and independent of who the player is. That insight is really what has allowed almost everything the field has provided to follow.
“Clutch” doesn’t have to only be synonymous with “Luck”. Teams shift heavily against certain players for a reason. If there are runners on base, they can no longer shift. Why would be surprised then when a player who is normally heavily shifted against hits better with runners on base? It isn’t completely luck.
We already can control for this, wOBA is higher when there are runners on. It’s when a player exceeds those raised wOBA expectations that we call it “Clutch.”
The shift argument might have something to it, but it’s not super likely – shifts are in response to a hitter’s ground ball tendencies, and it turns out that for the most part ground balls are really hard to hit anywhere besides the pull and center thirds of the field (50% of all ground balls are pulled, 34% up the middle, and 16% oppo, data is from 2013 MLB gameday: http://gd2.mlb.com/components/game/mlb/). There are obviously differences for particular hitters who might get a few extra balls through, but I don’t know that it’s going to show a big “clutch” difference.
In a team context, it’s almost certainly some form of random variation.
Couldn’t a manager play a role in the ‘clutch’ score by playing a horrible relief pitcher for every save situation? Or is this accounted for?
It is to some extent, in that no MLB manager is that stupid.
Hi, I’m new here.
Incorrect. See Joe Nathan
Or a guy battling injuries, or a guy at the end of a road trip compared to a guy coming off a days rest in the first game of a home stand? Too many things to account for to come up with a reliable clutch rating or statistic, and yet only a fool would say there is no such thing as a pressure or clutch situation.
Christ’s testicles, no one is saying that there aren’t clutch situations. What is being said, and supported by mountains of data, is that performance in those situations is not a skill that some players have and others don’t. Given a large enough sample size, any given player’s “clutch” stats (BA w/RISP or whatever) will end up very similar to their overall career stats.
Performance in clutch situations has zero predictive value.
Does “clutch” depend on team composition at all? For example, if a team is made up of all league average hitters, are they more of less likely to be clutch relative to a team with 4 stars at the top of the order and 4 scrubs, when both teams’ totals are the same? The stars and scrubs would be more likely to leave runners on base at the bottom of the order (assuming Dusty Baker isn’t setting your lineup) but more likely to score at the top of the order. Does this make a difference?
Great study. I tried to look at something like this about 10 years ago but I think what you did more clear. My study was
Does Team Clutch Matter in Baseball?
http://cyrilmorong.com/teamclutch.htm
2 things I concluded were
Splitting a team’s performance into clutch and non-clutch helps very little, if at all, in explaining winning percentage.
Non-clutch performance has a greater statistical impact on winning than does clutch performance.
First half clutch? Is that sorta like Spring Training clutch?
Spring Training Clutch sounds like an excellent nickname to be distributed as need be.
The key here is that there is no predictive power.
One interesting aspect is that a clutch event is viewed in the isolation of one player, for an AB usually the hitter but not the pitcher. A “clutch” base hit might not be clutch at all if the pitcher delivered an “unclutch” pitch. Add in all the other factors such as defensive capability of the team, the pitches before the clutch event, where framing and umpiring could effect pitch sequence and location, and it gets harder to fully identify what clutch really is.
And over the course of 162 games teams are going to “concede” some games and situations for the greater good of the season. If a team is on a the tail end of a ten game road trip, in which they’ve gone 7-2, a 3-1 deficit versus a top bullpen in baseball is a questionable clutch situation. Maybe that team has to play at home the very next day and just want to focus on that game, especially say if they have a 7 game lead in the division. The pressure there isn’t very high to win or perform.
And maybe during that 10 game span they have spent much of their bullpen, and so have run an “innings eater” out there to just bridge the gap until they get home the following day.
Again, too many factors!
If Maybes and buts, were wins and Championships, then the Rays would be number 1.
If you get nine hits off me in a game, one in each inning, you will almost certainly get 0 or 1 runs. If you get nine hits off me all in one inning, you will almost certainly get 7 or so runs. That is the takeaway.
… and which teams get those run distributions does not appear to be repeatable. I.e., sequencing is random.
And my cry of “Too Many Factors!” in this thread was your cue nerds, to come up with a better formula. A season bottlenecks just like a game does. Context matters just as much in a season as it does a game. A team or a player is going to feel pressure relative to the context they are in. A single game is going to feel like “a big game” relative to the context it is in. Now get to it!
Dumb question… where are Team Clutch scores? Is the only way to find them to add the batting/pitching Clutch scores found in team stats/win probability together?
I think so, it’s annoying, but also not that bad
Sure but that shit is banked and I don’t see my beloved Rays making that ground up. As always, you would rather preach tempering your expectations rather than say just wait and you’ll see. In a league playing a finite number of games this stuff matters and it gives little solace to know that, but for a month of bad luck this team would be a no-brainer for the playoffs.
Well sure. I think it’s just a reminder not to get too down in the dumps over a month of bad luck (if you’re a fan of a losing team) because it will all reset back to 0 next year.
And for a fan of a team with unexpectedly good luck this year, enjoy it! It’s not likely to be repeated next year, so enjoy the postseason trip and hope that the coin keeps falling your way into October. Who cares if it’s not predictive, or won’t last forever. I would gladly take my team being lucky as all holy hell and still playing meaningful baseball deep into October, rather than sitting at home.
This has very little to do with anything, and I doubt anyone will read it, but I was playing around with the BaseRuns standings, and I noticed that the four AL West teams who aren’t Texas are all underachieving relative to their BaseRuns record. Of course, they’re also overachieving relative to preseason projections, but that’s somewhat explained by Texas being much much worse than expected. Is there something going on here that would allow so much “anti-clutch” to concentrate in one division? Or is it just a weird coincidence?
(I guess I’m specifically wondering if the BaseRuns model might break down in an extreme case like where baseball’s two best teams and its worst team all play lots of games against each other.)
So if say the O’s and A’s played the Red Sox? If you accept that the bullpens are equal, I wouldn’t expect it to break down. If the good teams have good bullpens and the bad team doesn’t, then I think the model would fall apart.
How do you deal with the residuals in that first fit? Might it be that a linear fit is an act of wishful thinking.
I propose that viewing this over a long set interval (half-season) is too crude an approach to rule out a predictive effect in the shorter-term. If one examines this over shorter intervals (say two weeks) I would expect to see a predictive effect for certain teams for certain years, whether one uses correlational analysis or time series analysis. I base this on my sense that “clutch” performance is related to confidence in one’s own performance and in one’s teammates’ performance. No one would deny the importance of self confidence in golf or the importance of positive feedback in many situations. Thus, the performance of baseball players and their teams can not be modeled simply as chance events with set probabilities as they actually do depend somewhat on recent prior events. A gambler can not really be “on a roll” at the craps table (in that the prior rolls do not effect the subsequent), however a baseball player and team can truly “be on a roll”.
I would posit the hypothesis that all professional athletes are either a) confident or b) successful in spite of a lack of confidence, and therefore confidence tells us nothing at all about the performance of professional athletes.
Two points on BaseRuns, both of which happen to work in favor of the Royals:
1) Percentage of base runners that score is not purely a question of sequencing, at least on the defensive side. We should expect the Royals to allow a lower percentage due to excellent control of base runners, viz. Salvy Perez and outfield arms. I wonder if it would make sense to subtract rSB, rDP, and rARM (or ARM and DPR) from the BaseRuns-calculated runs allowed.
2) Sequencing is partly determined by choice of relief pitchers entering the game. The Royals have a particularly high-variance bullpen, with Holland/Davis/Herrera yielding almost nothing and Chen/Coleman/Brooks giving up many runs in “garbage time.” I don’t see a simple way to incorporate this into the analysis, but I suspect it explains some of why the Royals are outperforming their projections.
Building on (1) above, it occurs to me that BsR could also be added to BaseRuns as well. This would account for superior base running, which leads to a higher conversion rate.
I hypothesize that after making these two adjustments you would see less W/L record deviation, as well as higher correlation in graph 1 above.
I am responding to the comment of Corey that “all professional athletes are either a) confident or b) successful in spite of a lack of confidence, and therefore confidence tells us nothing at all about the performance of professional athletes.”
This misses the point of my argument (which I should have made, and now will make, clearer). The point is that self-confidence varies depending upon recent events, and that self-confidence IS very important for performance. Thus, if one has had a series of poor at-bats one can begin to “press” or try too hard, to question one’s ability, and even, perhaps, to change one’s approach and style. Similarly, if one has had a series of good outcomes, one is more relaxed, less anxious and more able to concentrate on simply hitting the ball. This also extends to the team level, in that a team that has had a series of good outcomes will be made up of players who feel that if they fail a teammate will “pick them up”.
I used golf as an example in my original comment because I thought the phenomenon is most evident there and that it would be the arena where readers would be most familiar with the phenomenon on a personal basis. But more generally, the idea is that positive feedback is a real phenomenon that does make future events somewhat more dependent on recent prior outcomes than on prior outcomes averaged over longer time intervals.
Failure to take this into consideration is a crucial weakness to most sabermetric approaches.
Could someone share whether there is a good place online to do analysis similar to this one? I’m interested in looking at team batting splits (e.g. Platoon split) broken out into 1st half vs. 2nd half, and looking at correlation between the two halves.
Jeff (or someone), please do an analogous in-season analysis for W-L record vs. Base Runs projection. If teams like the Orioles are doing something measurable not accounted for in the model, it ought to show up in the same way that clutchiness doesn’t.