Yes, the Playoffs Are Still a Crapshoot

We live in an era where every team, to some degree or another, embraces modern analytics when assessing itself and the rest of the league. Wanting to know about the numbers that drive baseball has filtered into fandom as well, which is why you’re here on this very website! But despite all the progress the stathead crowd has made over the last quarter-century, when it comes to the playoffs and playoff results, many fans seem more inclined to defenestrate the numbers and attribute the losses to all sorts of causative elements beyond a surplus or dearth of players just happening to have particularly good games that week.
In the worst case, failing to win two of three games or three of five is attributed to some kind of character flaw. At best, the loss is because of some fundamental flaw in a team’s construction, typically something that sabermetrics is to blame for, no matter whether the team is sabermetrically inclined or not. Here’s one example of very common thinking along these lines just from the last couple of days. It’s from a fan on Reddit, so I’m not specifically attributing them, mainly because I don’t want to risk social media pile-ons:
The postseason is the real season from now on, so the focus shouldn’t be on sabermetrics as much as it has been. With the playoffs being expanded, 111 wins doesn’t mean squat. Shift some of the focus to bringing in guys who play with fire and aggressiveness and will situationally hit rather than live and die by the long ball. If it costs us some wins during the regular season, so what?? It’s the little things that win the most important games.
With three of MLB’s four 100-win teams already out of the playoffs and the 99-win team pushed to the brink but ultimately surviving, these types of incriminations will be common this season. The Dodgers didn’t lose a 60/40 matchup (the ZiPS projection for the series) because they were simply outplayed over four games, but because something was broken in how the team was built. Depending on who you listen to, you can hear the same type of grumbling about the Mets and Braves.
Since questions should be explored rather than dismissed, let’s look at playoff overperformers and underperformers over the Wild Card era (starting in 1995) and examine if there really are consistent patterns behind which teams are overperforming or underperforming in the postseason. And since this is illustrative more than anything, I’m trying to keep it as simple as possible, within reason.
The best place to start is the 950 playoff games played since the 1995 postseason [originally said 1,900 erroneously -DS]. (This doesn’t include Game 5 of the Yankees-Guardians ALDS since I was crunching the numbers while that game was being played.) Using Pythagorean wins plus home-field advantage and projecting games on that basis, here is how each team has over or underperformed in the playoffs:
| Team | Expected Wins | Actual Wins | Difference |
|---|---|---|---|
| New York Yankees | 105.1 | 119 | 13.9 |
| San Francisco Giants | 42.3 | 50 | 7.7 |
| Kansas City Royals | 14.4 | 22 | 7.6 |
| Boston Red Sox | 61.5 | 68 | 6.5 |
| Florida Marlins | 17.6 | 24 | 6.4 |
| Philadelphia Phillies | 26.2 | 32 | 5.8 |
| Cleveland Guardians | 47.3 | 51 | 3.7 |
| New York Mets | 26.2 | 28 | 1.8 |
| Chicago White Sox | 12.8 | 14 | 1.2 |
| Washington Nationals | 17.9 | 19 | 1.1 |
| Houston Astros | 62.0 | 63 | 1.0 |
| St. Louis Cardinals | 74.1 | 75 | 0.9 |
| Detroit Tigers | 24.7 | 25 | 0.3 |
| Pittsburgh Pirates | 3.7 | 3 | -0.7 |
| Atlanta Braves | 72.8 | 72 | -0.8 |
| San Diego Padres | 16.0 | 15 | -1.0 |
| Colorado Rockies | 11.3 | 10 | -1.3 |
| Baltimore Orioles | 16.4 | 15 | -1.4 |
| Milwaukee Brewers | 15.1 | 13 | -2.1 |
| Arizona Diamondbacks | 20.2 | 18 | -2.2 |
| Tampa Bay Rays | 30.8 | 28 | -2.8 |
| Seattle Mariners | 19.9 | 17 | -2.9 |
| Toronto Blue Jays | 13.1 | 10 | -3.1 |
| Los Angeles Angels | 24.6 | 21 | -3.6 |
| Cincinnati Reds | 8.9 | 5 | -3.9 |
| Texas Rangers | 25.2 | 21 | -4.2 |
| Chicago Cubs | 31.2 | 25 | -6.2 |
| Oakland A’s | 24.6 | 18 | -6.6 |
| Los Angeles Dodgers | 68.8 | 62 | -6.8 |
| Minnesota Twins | 15.4 | 6 | -9.4 |
The Dodgers finish near the bottom, though part of that is they’ve played in so many games; they’ve won about seven fewer games in the playoffs since 1995 based on what you’d expect from their Pythagorean record, the Pythagorean record of their opponents, and whether they had home field advantage in any particular game. But one thing worth noting is that Los Angeles’ deficit is almost entirely in games played before Andrew Friedman took over after the 2014 season and moved the team in a more sabermetric-oriented direction. From 2015 to ’22, the Dodgers won 47 playoff games against 47.9 expected wins.
Since eyeballing lists is a poor method of dimensionality reduction, I ran a series of logistic regressions (winning or losing a game is binary data) using a few dozen different variables. If these variables have relevance, then including them along with the Pythagorean projections ought to make them better at sorting wins and losses. I started off with an obvious one: bullpen strength. I want to avoid getting too far into the weeds talking about chi-squares and Wald tests, so here’s a method with a visual element, the ROC curve. Simply put, it evaluates how well we identify wins and losses ahead of time. The larger the area under the blue curves below, the better the model.

On the top is the model using only the Pythagorean winning percentage, and on the bottom, the model uses Pythagorean winning percentage and the relative FIP- of the team’s bullpens. The dashed line represents the results for a random number generator in picking wins and losses, with the blue line indicating the output of the model. In other words, we get a very slight improvement by knowing the bullpen strength in addition to team quality.
To put it into an easy number, if you picked the 1,900 wins-losses using just Pythag and home-field advantage, you’d have correctly picked the winner 1,018 times (53.6%). If you used the model that also included the relative quality of the team bullpens, you would have instead picked the winner 1,026 times. That translates to being correct one more time about every three years.
Another common “explanation” of why good teams underperform is that they’re too reliant on home runs. So I calculated a home run propensity for each team, starting with park-neutral seasonal stats and then calculating runs created both with and without every at-bat that resulted in a home run. The least reliant on home runs was the 2007 Angels (22.1% of runs), and the most was the 2020 Dodgers at 53.4%. I then included the relative dependence on home runs for both teams in each game. With this knowledge, you’d improve from picking 1,018 winners to… 1,085.

While the model improves by knowing which team relies more on home runs to generate runs, it didn’t confirm the hypothesis for one crucial direction: the effect was in the opposite direction. Teams with a greater reliance on home runs have overperformed in the postseason, not underperformed. For relative three true outcomes (walks, home runs, strikeouts) performance, you only improve from 1,017 correct to 1,056, but again in the direction of sabermetric-friendly teams, with high TTO teams overperforming in the playoffs slightly.
Contact rate is another post-hoc explanation for why certain teams win, but that was rejected from the model at the 95% level of significance. The same goes for attempts to include momentum (second-half performance, September record). Team speed doesn’t make the threshold. Neither does clutch performance during the regular season or postseason performance history of the players involved. In layman’s terms, knowing which teams have these characteristics has little to no value in telling us why certain teams won playoff games over the last three decades, apart from their effect on their Pythagorean record. Surprisingly, what was (barely) useful enough to make it as a variable was team age, but this was an advantage on the side of young players; young teams, and consequently, teams with less playoff experience very slightly overperformed over the last 28 postseasons.
In the end, I looked at about 60 team-related variables, and they all did almost nothing to explain better which teams won playoff games. ZiPS does a somewhat better job, but only because it “cheats” since it knows who is starting, who is healthy, and who is in the lineup, something which isn’t necessarily a team construction issue.
The Dodgers, Mets, and Braves lost because… well, they lost. That’s it. Enjoy the games, but sometimes a cigar is just a cigar, and that’s perfectly fine.
Dan Szymborski is a senior writer for FanGraphs and the developer of the ZiPS projection system. He was a writer for ESPN.com from 2010-2018, a regular guest on a number of radio shows and podcasts, and a voting BBWAA member. He also maintains a terrible Twitter account at @DSzymborski.
Great article, Dan. I understand that people are annoyed when their team loses and have trouble grasping randomness in these things. But it’s frustrating that we have to Keep. Having. This. Discussion.
Baseball is a situation where there’s both lots of randomness and a whole ton of data that allows us to say with some confidence that, yes, something unusual did just happen and we shouldn’t completely change our baseline belief. But you still can’t ever rule out the possibility that our analytical tools can be improved and that we missed something that could have improve our prediction. The discussion can’t be avoided. “How much should this unlikely event cause us to update our priors?” is a key question of life.
A few years ago, people were really tired of having the “hot hand” discussion, and yet that still goes on, with occasional interesting contributions to the literature.
Can’t be avoided even more in other areas of life, of course, which lack the type of regulated repeated games of baseball. But I understand that people are even more tired of the discussions when it comes to public policy or politics.
I don’t think its the randomness per se that has people frustrated. It’s that the counterbalances added to the format to compensate for adding more (ie less successful) teams into the mix seem to have done nothing to alter the randomness. Teams don’t even seem to have played or composed their roster differently for it. The BYE just ends up putting top seeds where they would’ve been under previous formats. It’s just one year but it certainly looks like the new format will undermine regular season success.
THis all presupposes that maintaining the integrity of the regular season is inconsistent with the goal of the new postseason format. If you think that the postseason is meant to reward the best team based on their regular season performance, that is probably inaccurate and the data doesn’t bear it out. If you’re cynical like me and believe that the postseason is meant to be watched on TV, then there is no inconsistency in injecting randomness into it, because that means upsets, cinderella series, and all sorts of other very watchable things. My horse was eliminated from this race a while back and I have to admit that teams that I hate (hate) have played some of the best, most tense games of baseball I’ve ever seen.
It’s too soon to tell, but if the postseason format retains viewers it will stick around. My point though is that it doesn’t have anything to do with rewarding the teams that play the best or most consistently over the course of the regular season. If that were the case we wouldn’t need a postseason, we would just end the regular season and crown the Dodgers or something.
It would seem to me that if it didn’t reduce randomness then those additional teams potentially deserved to be there.
I’m not saying that means further expansion should be on the table, because it shouldn’t. But it does mean that in this case 12 teams doesn’t undermine the competitiveness of the playoffs.
That would be because playoff format changes aren’t made to “address randomness” but to generate more revenue. In fact, the “randomness” is more of a feature, not a bug.
The Playoffs would be meaningless if you could predict the winners regularly just by the numbers.
I think the questions are good. The Reddit thing posited a theory based upon limited observations. He lacked the tools to explore it further. That it was disproven does not make the question stupid. Personally, I read this expecting the conclusion Dan reached, but I read the post because I was open to the possibility that their could be a correlation; that the Reddit thing could be right. Why else would I waste my time reading it? I note that you also read this.
Eh, I’m at work. What else am I supposed to do?
Is there any way to look at the effect of pitching depth? Both for starters and relievers. Because I think that still matters less in the postseason and teams that squeak into the playoffs with two great starters and three bad ones or three great relievers and 5 or 6 bad ones might overperform their regular season records since those bad pitchers won’t be called upon to throw as big a percentage of innings in a short series, especially if there are extra days off.
But yeah, most of it is just randomness. Nobody got worked up if the Dodgers lost 2 of 3 to the Diamondbacks in April or the Rockies in June, but lose 3 of 5 or 4 of 7 to a 90 win team in October and suddenly there must be some fatal flaw in the organization.
“We live in an era where every team, to some degree or another, embraces modern analytics when assessing itself and the rest of the league.”…well 29 teams and the Rockies.
It’s weird to think about, but I maintain that the current Rockies front office would be an above-average 1990 front office.
1990 was 32 years ago. unreal.
Barry Bonds was skinny and had a normal sized head back then!
Didn’t we all??
And well on his way to being an inner circle hall of famer until jealousy of McGwire and Sosa took over…
Greinke averaged 89MPH on his fastballs this season which puts his fastball velocity at bottom 4%.
Pretty sure that would be above-average in 1990 as well.
Not just an above-average 1990 front office. It would be a Top 5 front office.
I don’t think people have really wrapped their minds around how much front offices have improved. The best GMs of the 1990s were:
-Sandy Alderson (now in an administrative role because the game passed him by)
-Walt Jocketty (in retirement because the game had passed him by, probably around 2010)
-Andy Macpahil (in retirement because the game had passed him by, at some point in the 2010s)
-Gene Michael (who laid the foundation for the Yankees’ big recent successes in the late 90s and early 2000s)
-Jim Bowden (hopelessly out of his depth with the Nationals in the late 00’s)
-Kevin Towers (hopelessly out of his depth with the D-Backs in the early 10’s)
-John Schuerholz (smartly left before the game totally passed him by, although I think Dayton Moore gives a pretty good impression of where he would stand today–a bottom 5 GM)
-Dan Duquette (held his own with the Orioles for a while but he wasn’t great near the end; he would be a bottom-5 GM today)
-Brian Sabean (smartly left before the game totally passed him by; he could still hack it as a GM today, but he would be in the bottom third of GMs)
-John Hart (one of the only guys who could still do it who isn’t the game today, although after the Braves scandal he’s probably done)
-Gerry Hunsicker (who is still sort of involved in baseball, he advises Friedman in LA. If he wanted to I bet he could still do it)
-Billy Beane (who took over in the late 90s and is still fantastic)
Keep in mind, the guys who could potentially hack it as a GM today like Hart, Hunsicker, Duquette, Alderson, and Sabean–these guys learned a lot in between the 1990s and today. The Rockies would probably run circles around their 1990s selves. Maybe not Hart and Beane and Hunsicker, but probably the rest.
I’m not sure Sandy is being moved to an advisor position because the game passed him by he never wanted to return as a top decision maker, he just wanted to be on the business side and hire a pobo. It is difficult to judge his Mets tenure because he was handcuffed by the Wilpons financially and by Fred not allowing him to make analytics hirers. Besides for Granderson, Colon and Cabrera he didn’t do to well free agent signing wise.
Sandy drafted very well and obviously the Dickey and Beltran trades were grand slams. The Jay Bruce love fest was very odd and he even had Bruce block Nimmo and Conforto’s playing time at times.
I would rate him as the best Mets gm since Cashen, which isn’t saying a ton.
I think Alderson was great with the Mets for most of his tenure but at some point the game overtakes everyone. And I think he recognized it too.
John Hart is more of a transitional figure in the evolution of GMs from the “baseball lifer” to today’s “MBA” GMs/FO Executives, many of whom never played the game.
Hart’s bio lists him as a Minor League Catcher for three years with Montreal. Then he served as ML coach for Baltimore under Hank Peters who took him with him when he moved to Cleveland and groomed him to take over. He even managed the team for a few weeks as part of his training.
So he was more of a lifer who saw the value in the new ways early enough to be a pioneer in some aspects.
Duquette was an extreme upgrade over Lou Gorman when he was with Boston. Gorman himself had done ok with the Mets and other teams in the 70s and 80s, so the moving bar of having a “state of the art” front office has been going on since probably before Branch Rickey.
Duquette doesn’t get enough credit for the work he did in Boston. My understanding is that the Red Sox were a truly dysfunctional organization before he arrived.
Yup the ‘04 world champions were as much his team as Theo maybe more so.
Another way to think of it: If the Rockies front office more or less does the same types of evaluation that they did in the early 2000s in the Dan O’Dowd era (which seems about right), then it’s the same level of competence as a front office that in the 2000s (so not even the 1990s) was smart enough to build one of the best teams of the late 00’s.
If the Rockies have gotten better–a reasonable assumption, even if it is incremental–that means they are even better than the O’Dowd team that built a serious winner in the late 00’s.
So can you imagine what sorts of things this group could do in the 1990s? They. would. clean. up.
Well, Cashman started in the 90s and is still pretty OK.
Oh, come on. Even most teams in the 90s wouldn’t keep giving out extensions to their own mediocre players or sign a supposedly good but not great bat to play him badly out of position, just to keep having losing seasons and not even trade away expiring contracts at the deadline. Even if such teams back then still decided not to rebuild like they should, they’d at least try to sign some good free agents.
I am constantly amazed at how much time and energy people put into trying to come up with a formula for winning in the postseason. Meanwhile MLB keeps putting more and more teams into the postseason, which keeps increasing the randomness. There is just one formula and always will be one formula – be the team that gets hottest in October.
People hate random. Our brains are wired to find and manipulate patterns. It is what we do.
We really hate, “It’s random luck”. We hate it so much that otherwise sane people will figure it matters to the outcome of the game whether or not they wear their lucky socks, and they believe this when they are at home watching the game on TV.
Emotionally, there has to be a reason, and if logic says, “no reason we can understand or do anything about” then most people figure logic can go hang, because there must be a reason, and they are going to figure out that reason, even if the reason they figure out is something you can statistically state is almost certainly wrong.
I had this moment watching Game 5 of the ALDS the other day. I think it was when they showed the stat about the Yankees being <large number>-0 in history when having a certain lead at home.
At the same time that my logical brain realizes that no graphic they put up on the screen will affect the result of the game at all, my emotional brain is like “yo, don’t jinx it”.
The pattern-seeking behavior is how superstitions start — our human brains innately understand correlation, but kinda suck at causality.
If we can’t predict things then how the hell can we even begin to control them? And if we can’t control things how we know we’re not all going to die tomorrow?
In short, when betting favorites lose we are reminded that the screaming void awaits us all.
I think the formula is pretty solid:
-Have a 5 win starter, a 4 win starter, and a 3-win starter
-Have a crazy bullpen. This includes a lights-out closer, a guy who would be the best reliever on 20 other teams to handle the 8th inning, a guy who would be the best reliever on 12 other teams to handle the 7th, and then 2 or 3 other guys who you would definitely trust in high-leverage situations.
-Have a deep lineup of regulars. A lot of times teams can get by with a rotating cast of characters who step up during the regular season. The Brendan Donovan, Oscar Gonzalez types–the heroes that show up out of nowhere and are hot for a few games–that is good during the regular season. But you need consistent, grinding hitters for the postseason.
Is this hard, bordering on impossible to get? Yes. Is it foolproof? No, not really. But it’s the formula you need, and a little different than the regular season where both pitching and hitting depth is essential.
Add a big dollop of luck to account for injuries, bad umpire calls, and random brain cramps by the players and/or managers.
It’s a human game and humans are random.
Probably the best strategy would be to try to figure out a way to cheat and apply it only in the playoffs, so that the chance of it being discovered in the regular season is minimized. Like if the Astros had shown restraint with the trash can lid thing and only used it in the playoffs, they might have gotten away with it and be using it today.
Fried – 5 WAR
Strider – 5 WAR
Wright – 3 WAR
Bullpen- Iglesias, Minter, Jansen, McHugh
Lineup wRC+
Acuna 134
Swanson 116
Olson 120
Darnaud 120
Riley 142
Harris 136
Arcia 104
Contreras 138
Rosario 61
About average defense efficiency.
Almost the formula.
What you outlined will get you TO the playoffs and fairly easily. As Antonio laid out, you are damn close to describing the Braves who got bounced by the Phils fairly easily. There never will be a magic bullet for the playoffs.
This – you can’t “magic bullet” your way to winning a 3-game or 5-game series.
No magic bullet. Having better players increases odds. Being able to get better players to play a higher percentage of the time helps (whether through deep bullpens or deep starting rotations that can also have guys used in relief).
Well sure, I don’t think anyone would disagree that having better players will give you a better chance to win. This we can all agree on!
Are the results really that random recently? The teams reaching the LCS since the beginning of what I call the Luhnow era are remarkably consistent. The astros every year. The dodgers and Yankees. Most of the slots in the LCS go to these three teams.
The earlier rounds have all kinds of variability but by the time you get to the LCS, it’s the same teams. We are surprised by the Phillies and San Diego but with their commitment to marquee players and high payrolls, no one should be surprised if we start seeing their names pop up in the LCS more – especially the Padres who have just added Soto.
You can always find some consistency/apparent consistency in the randomness. Yah, the Astros have made 6 straight LCS, doesnt change the fact that higher seeds in the division series have a .527 winning percentage overall, and are .500 against teams that played in the WC game/series. Good teams have good runs, I suppose it’s no different than flipping a coin 100 times, you’ll have times where you get heads 6-7 times in a row, but total it will still be around 50 times.
So the Dodgers, Astros and Yanks consistently getting to the final four is still random ? That appears to be the opposite of random.
(Forgive me if my statistics terms are wrong; I get the concepts but forget the details sometimes.)
As Dan points out above, Pythag record (and home-field advantage) does explain some of the results of the playoffs. Those three specifically have had really good teams consistently over a (relatively) long time — which means they often have home field advantage and better Pythag records.
The coin flip analogy isn’t great here — it’s more like flipping a slightly weighted coin that comes up heads somewhere like 53-54% of the time.
No doubt my coin flip analogy is very simplistic example.
I think in the end, the best way to explain those three teams playoff performance comes down to they are consistently good teams, and good teams make the playoffs more times than not, so they’ll have more chances at advancing. Despite that these teams have combined for only 2 WS victories in the past 10 years. If these three teams have figured out a way to playoff success, I would think they’d have won a few more world series as well.
Provided the Astros don’t win, we are looking at 9 champions in 9 years with regular seasons wins ranging from 108 wins to 88 wins (or 87 if the Phillies win). Its a crapshoot, and will only get worse once MLB gets its way and has 16 teams in.
*regular seasons wins ranging from 108 wins to **43** wins
If HFA helps explain it, that doesn’t really make a difference. The same teams are going through repeatedly. When you get to LCS I can see how fate takes over. Put those three teams and others like the Braves in a grudge match and you’ll get random results. But the reality is that the same teams are consistently in the final four. 13 of the last 24 LCS teams have been those three and there are a couple others that have appeared twice.
If you go back to two teams or four teams in the playoffs, it didn’t really concern anyone whether “the best team” won. It’s understood when you get to the LCS that everyone is basically even.
What has now changed is that an outlier team making the playoffs has a miniscule chance of winning four straight series. In the past when you had to win two or three series, a weak team could go through. That won’t happen much now.
This is why as a Dodger fan Octobers are filled with skeletons. It doesn’t matter how carefully and consciously you build a team that can win most days during a 162-game season, or how clearly above the league standard your organization is in planning, instructing and recruiting, year after year; you’re going to head into the postseason and simply hope that everybody doesn’t go cold at the same time, or that you don’t run up against a team that’s playing spirited, unified ball exactly at the right time. Because if you do, you’re going to spend the next six months listening to people ask what you did wrong even though everybody knows the answer is: nothing.
MLB isn’t the NFL, where the best team wins nine out of ten times. There’s a massive variance between season-long success and your chances of winning any single game. And the postseason is the best place to see that difference underlined.
Lifelong Dodgers fan (so I’ve seen this movie a time or 10 in 70 years), but this encapsulates my sentiments perfectly. You can’t win ’em all, even when (especially when?) you’re “supposed to.”
Yeah, the only thing you can do is try to get to the post season as often as you can so you can have more chances and the Dodgers have done that well lately.
Probability is hard.
Nice article! Just for future reference for what IMHO is a nice visualization of the effect on AUC scores when subtracting independent variables from models, see pp 369 and 371 of this article https://journals.sagepub.com/doi/pdf/10.1177/0022343309356491
Thanks! Yes, that’s a good way to do it. This is always the hardest part for me: explaining some of these things in a way in which there will be utility for the typical reader.
Randomness, schmandomness….
Chuck Norris didn’t rely on randomness when he was kicking Commie ass!
Chuck Norris doesn’t have random factors happen to him. Chuck Norris happens to random factors.
This illustrates that you perform at the level you are expected too isn’t there is a difference in the required team depth that would improve your expected outcomes post-season compared the regular season?
In a 3 game series you only have 3 starting pitchers have a good 4-5 starter is useless, having a better backup catcher is useless, etc. That is have a team weighted towards star players is going to be better than a team with a lot of depth. Maybe there is very little to change and maybe this was already included in the analysis. However bullpen quality and the top 3 relievers’ quality seem different.
This is a complete overstatement; your 4-5 starter can provide relief depth if your 1-2 starter blows up or gets injured, your backup catcher can pinch hit or may be forced into action due to injury. Depth is less relevant in a short series than the full season, but it’s never irrelevant or useless, because rarely does even a short series go completely according to plan.
Sorry to quibble, Dan, but I think you might’ve doubled the number of postseason games since ’95– there are 949 wins shown in the table, not 1900
I see I did that, but I’m quite concerned about the missing win!
It shouldn’t change results, but I’m worried there’s a game there in which both teams have a loss.
I live in MN, and the 18-game losing streak maybe gives me more enjoyment than the 1987 and 1991 World Series wins. Each loss was frustrating, and after a few years it became maddening. But the accumulation of 18 consecutive playoff losses is so sublimely absurd; relish is the only proper reaction.
I wish I had your outlook on life
I definitely do not relish. Still maddening.
If each game were a true 50/50 coinflip it would be a 1-in-262,144 probability.
Even if they were only a 40% favorite to win each individual game during that 18-game losing streak, managing to lose all 18 in a row is still an almost 1-in-10,000 occurrence (1-in-9846, to be exact).
Could you please list the team related variables that were not significant?
I’m sure there are others, this is not intended to be an exhaustive list:
In response to a reddit post about the extra days off when sweeping a 7-game series might be a detriment, I took a look at all LCSs from 2000 onward and looked at how the series went for the team based on how many extra days off they had.
The teams that had exactly one extra day (n=11) won 45.5% of the games in the WS. The teams that had multiple extra days off (n=11) won 44.1% of the games in the WS.
Which means that the team that had fewer days off between the LCS and WS won 55.2% of the games.
People generally tend to think that extra days off is a good thing for athletes (and teams, when they have the opportunity to re-align their starting pitchers), but at least over the last 22 years, that doesn’t seem to be the case.
Well, it looks like those previous five years (back to the start of divisional play) make a pretty big difference on the sample set.
All five of those years had the team with extra off days winning the wrold series and going a combined 20-7.
This pushes the .448/.552 overall split to .500.
And now the teams that had a single extra day off are at .519, while those with multiple days off were at .479.
—
I suppose we’ll need a bunch more data to get this to stabilize a bit…?
Going back to 1985
1 extra day off: .549
>1 extra day off: .457
Beyond this point we’re in 5-game LCSs
LOL Twins.
May the Astros eliminate you again.
“111 wins don’t mean squat”…
I completely disagree here. Ask the Cardinals, Mets, Blue Jays, and Rays if they’d prefer the bye. The change doesn’t change the odds for the best 2 teams in each league compared to the previous formats since the division series was added. For the 4 other playoff teams it drastically changes it.
Could it also be that regular season and playoffs are different because of depth issues?
By that I mean that in the regular season your 4 and 5 starters, 5-to-8 relievers and the bench are important, and teams with larger payrolls can afford the best for these spots. So they end up winning more games. But in the playoffs it is usually a battle of the best 7to8 position players, 3 starters and maybe 3to4 relievers and maybe 1 bench player. The regular season record usually does not matter, as most playoff teams are comparable at this level of depth. Add in the odd unexpected playoff heroes (a la David Eckstein), maybe the playoffs are not really a crapshoot as much as what you would expect when comparable teams face each other.
It might be true that the bottom half of the roster plays more in the regular season, but it seems like it’s very few teams where that part of the roster adds wins. Basically the teams with the best 5-8 players win in the regular season and postseason. The Rays are probably and exception. Maybe the 2021 Giants.
If I’m understanding correctly, what they mean is that the Dodgers, Mets, or Braves back end starting pitchers. 4-5 (even 6-7 because you have to use those guys over the course of 162) are way better than the back end of the Phillies or Padres. On a random Thursday in July both teams may not be running their horses, which gives an edge to the team with more depth. Teams don’t always get to run their best players out there during a regular season series. But in a playoff series you’re not going to see those guys, evening the playing field a bit.
That is what they’re saying, but I’m questioning it. The WAR difference between the Mets (55.4) and Padres (40.8) is almost entirely on position players (35.1 vs 21.4). The back end starters aren’t the difference.
Braves (51.5 Total, 28.7 Position) and Phillies (44.3, 21.6) is a similar story. In fact the 5 best position players on the Braves had a 23.6 WAR (Swanson, Riley, Harris, d’Arnaud, Olson), vs Phillies 15.5 (Realmuto, Schwarber, Harper, Hoskins, Segura)
There’s a much bigger difference between the stars on the better team than the scrubs. Scrubs by definition are close to replacement level.
WAR is a cumulative stat. I prefer to use Rate stats to compare ability (wRC+ and the like). My thesis is that ability wise; the players being used in playoffs are comparable, as depth from the 16-26 doesn’t matter as much in the playoffs.
Yeah I see it in this way too.
I think there’s something to this. The problem is if you have injuries and no depth, you might not get to the playoffs at all. If you do, then you’re more likely to have a scrub play in place of a star in the biggest moments of the year.
To remedy this, you can invest more in depth, which takes away your ability to get a star to lead your rotation or play in the middle of the lineup. There’s just no easy answer. The best thing you can do is just acquire as much talent as possible and even that isn’t a guarantee as the Dodgers have demonstrated.
I’m saying this as a fan of the team that slipped into the playoffs with 88 wins last season then won the crapshoot… Is that good for the game? Baseball is a marathon sport. Is it good that a team can show consistent dominance for 162 then have it all be over with some cold bats and a couple bad breaks over the course of a weekend? Or you can stumble and bumble your way through for 5-6 months then got hot for a month and you get the trophy? I don’t know. It really goes against the nature of our sport to crown a champion this way.
I don’t disagree but I think you’re overstating it a little. I don’t think, for example, that you can pound sand for 6 months and get into the wild card with one good month. I agree to the extent that expanding the WC to include two additional teams in each division (combined with the small amount of daylight between, say, the 2nd and 3rd best teams in each division) was unnecessary and icnreased the amount of games to be played and therefore opportunities for bizarre upsets to occur, injuries to accrue, etc. That’s bad/consequential randomness.
Hey the Braves did last season haha. Maybe not “pound sand” but they treaded water at .500 for most of the season. I watched them every night. They weren’t very good. 4 months of mediocrity, then 3 months of good ball. Champions.
Yes to your point for sure. Adds to the drama, etc which is what MLB wants I guess. But not the best way to determine the champion of your league.
Edit to add that yes I did overstate it a bit in my previous post. You can’t do it for 5-6 months but can definitely do it for 4.
I think if you just crown a champion at the end of the season to whoever wins the most games? We could contract about 1/3 of the teams then. Maybe more. And the game would get a helluva lot more boring. Would we even have playoffs then? Just a World Series? Personally, that sounds like the NBA where, unless there’s a massive injury somewhere along the way, you know what four teams are gonna be there at the end.
While many of the rational folks here are spinning themselves like a top over the chaos in the NL, lots of other fans see that chaos as fun. And it’s honestly kinda fun to see favorites go down. Because it’s unexpected and a surprise!
It doesn’t have to make sense and it’s foolish trying to force it to make sense because it won’t happen. Embrace the unknown and the chaos of playoff baseball.
NBA is way less random than the MLB. Only 5 players on the floor, stars have the ball in their hands every possession, etc.
What is the “nature of the sport” though? And does it matter that we maintain it? If the playoffs are randomized, it makes every game of every series worth watching. The NBA and NFL both have much less randomness to their games, but the early rounds of those playoffs can be absolute snoozers.
I’m not coming at this from a “stats ruin the game” standpoint — not at all. But I do think that what’s “best for the game” and what most accurately rewards talent are not necessarily the same thing.
Yeah, “nature of the sport” is doing a lot of presumptive heavy lifting. The
long slog” through summer is a facet of the sport, sure.
But so is the team that seems like a dog but gets hot. Are people writing the Miracle Mets out of “natural baseball” for going 24-8 in September and having the wild luck of the Cubs going 8-17.
The Braves pythag showed a much better team. It wasn’t as much of a miracle as it seemed.
Contact rate seemed like it should matter in the postseason since run prevention is so dependent on strikeouts. But BABIP has fallen since 1995 also despite many more hard hit balls. FB% seems pretty consistent but HR/FB has jumped, so the FBIP% looks lower. I’ve no explanation. More GBIPIS%* maybe?
Ground Balls In Play Into Shift
“They were obviously missing the most important stat: TWTW.”
-Old Announcers Everywhere
When talking, listening or reading about the effect analytics have on the performance of players and the outcome of baseball games I all too often come away with the impression that the true believers just cannot bring themselves around to the idea that their analysis isn’t foolproof and guaranteed to predict the result. An accumulation of little things have led me to question much of what is placed in the public domain. Other entries in this discussion talk of randomness. That is baseball and by its very nature will always defy any attempt to quantify it. My number one area of doubt pertains to defensive metrics. The various systems continue to be all over place with one saying one thing and another just the opposite. Case in point, in 2021 Jacob Stallings and Yasmani Grandal were acclaimed as a great framers and defensive catchers and in 2022 both were rated as poor or even worse. This just defies reason and don’t get me started on attempts to put defensive metric values on players from a century ago. Analytics would be of more importance if they weren’t universally in use by every team but as it is now the edge that is gained is very close to insignificant. In addition, I still do not understand how player value as measured by WAR is ascertained. Just yesterday the Guardians sent out Aaron Civale, for reasons I will never understand, and he was lit up, as could have been expected. I didn’t need numbers to tell me that he probably wasn’t a very good choice but I dutifully looked to see his record and upon checking his bWAR found a -0.8 but his fWAR was a far less onerous 1.3. Which is it? It is fair to question how a pitcher who wasn’t even replacement level started the most important game of the year. Dan is right as usual. The playoffs are still a crapshoot. This is still baseball where the best teams lose 30-40% of their games and endure losing 3 out of 4 a few times every year and that will not be altered by all the analysis done in all the front offices.
You’re on Fangraphs and you don’t know the difference between fWAR and bWAR or that they’re asking the same question with different methodologies for answering? Or are you just trying to engage a rhetorical device to make… what point, exactly? That they aren’t the same. Everyone who knows anything about them knows that.
You think people who actually get paid to do analysis think their conclusions are bulletproof? What are you arguing?
The postseason being a crapshoot doesn’t change that you have to build a team to make it there, which in 2022 is going to require good analytics to do consistently, and that continuing to pursue new and deeper information through analytics may provide insight into variables that correlate with postseason success.
Please stop with this strawman about “true believers” in analytics. No one with any credibility on analytics asserts that analytics can or does answer all questions. At most a strictly analytically-inclined individual will reject or refute any assertion or narrative that has no supporting data–even if there is no contradictory data.
“This is still baseball where the best teams lose 30-40% of their games and endure losing 3 out of 4 a few times every year and that will not be altered by all the analysis done in all the front offices.”
It feels like if it were 15 years ago you’d have simply posted, “WAR! What is it good for?! Absolutely nothing!”
If fWAR and bWAR had entirely different names, the lazy “there are multiple ways to calculate WAR, so I hate them all” argument could finally go away. But people who say that would probably find another cloud to yell at anyway.
WAR is a framework, not a stat. It is like arguing that batting on base percent and slugging percent are flawed (namewise) because they are both labeled “<something> percent”.
The differences in those WAR numbers basically mean that Civale gives up a ton of hard contact, while his FIP (on which the much inferior fWAR for pitchers is based) is better but still mediocre. As such, you could predict that Civale will have a little bit of a rebound next year, but he really sucked this year and had no business being on a playoff roster in the first place.
“Case in point, in 2021 Jacob Stallings and Yasmani Grandal were acclaimed as a great framers and defensive catchers and in 2022 both were rated as poor or even worse. This just defies reason”
So I take it that you don’t like batting average either. In 2021, min 400 PA, Jeff McNeil was 35th percentile for batting average and then this year wins the batting title. Clearly that shouldn’t be possible.
You mean to tell me that every point the anonymous Reddit guy made, every one, is dumb? Really?
Hey Dan! Were all of the variables you considered from the same year? Did you try any multi-year data (e.g., previous year’s pythag)?
I did multi-year for performance history, not the other things. As bizarre as this may sound, I was trying to keep it as straightforward as possible.
what if you factored the team’s regular season but taking out things like the 5th starter, garbage time innings, games played after seeding was clinched, etc.? not all data are equally valid.
I think there are 3 parts to the perceived disconnect between regular season and playoff performance:
1) the games are different, depth isn’t as important during the regular season, top heavy teams benefit from the postseason
2) we take the regular season results as some sort of truth when they are not. Looking at actual vs pythag records is a start, but the pythag records are based on run scoring and that is impacted by luck too. Maybe, the Dodgers weren’t a 111 win team maybe they were a 104 win team that was very lucky when it came to BABIP?
3) short series increase the importance of good luck
#3 is the most obvious point though and gets most of the attention.
Why don’t we have an awesome trophy w/ other benefits awarded to team in each league with best regular season record?
Akin to the President’s trophy in NHL. We have division pennants but it would be easy to award a best record in each league.
That MLB President’s Trophy would quickly be considered every bit as cursed/undesirable as the NHL version…
I’m glad Dan is throwing cold water on the idea of postseason mojo or that the Dodgers or Braves are chokers or didn’t want ‘it’ etc. Randomness in the postseason is great but that’s why it needs the added context of meaningful qualifying rules. We all know that the impetus for adding more teams was motivated by $$.
Can’t really argue with the numbers (try as I might) but my surmise is that the frustration is because the randomness/crapshootery seems to go hand in glove with the postseason format changes. Now, again, the numbers don’t necessarily (or even remotely) back this up – I can complain all I want that the mets were better than the padres, which they were, at least in the regular season, but there’s no getting around the fact that all else being equal the format changes would have still had us facing San Diego, just in a single game instead of a best-of-three. And all else being equal we would have still been eliminated. Just not by Local Boy and the Greasy Ears Gang. And I’d be happy with that, but that’s beside the point.
My rather long-winded point is that people who don’t routinely do statistical analysis can only try and fish labels out of the gaps between more obvious data points, and slap them atop their feelings (this phenomenon is familiar to anybody who plays competitive video games). If your team overperformed in the regular season and got trounced by some grubby, .530 cinderella, you either struggle (and generally fail) to explain why, or you accept the randomness and hope for the best next season. The latter, despite being the healthier and more correct choice, is actually far more difficult to sit with as a fan.
Now, finally, the obvious implication of all of this is that we should either move to a Premier League-style relegation system and get rid of the playoffs, or get rid of the regular season and only have the playoffs. Neither are ever going to happen for fairly obvious reasons. The new postseason format (like the previous expansion) are explicitly designed to prolong the postseason and inject randomness, not reward the best team, because drama and narrative means better TV for casual viewers and that in conjunction with more teams playing more games on a longer timeline means more revenue for franchise owners (even if it means we have to stay up until 2 am to see west coast games, watch 18-inning assaults on our patience, or watch four games in a single afternoon because my fiancee is always happy to watch me growing moss).
Ultimately while I hate hearing “learn to live with the crapshoot,” I agree that that’s really the only thing you can do.
You gotta take it one game at a time, give it your best shot, and the good Lord willing, things will work out.
Make the division series a best of 5 with all games at the top seeds’ ballparks and no off days. The extra off days likely help the lower seeded teams more.
Tom Tango had an interesting post about the cumulative effect of home field advantage and how something like a 5 game series played entirely at the better regular season team’s stadium was equivalent to a 9 game series split 5 home/4 away for the better regular season team, if the gap between the 2 teams was large enough. The bottom line was that the split in site was far more important than the series length.
http://tangotiger.com/index.php/site/article/when-does-a-one-game-and-a-15-game-playoff-series-become-equivalent
Perhaps the answer is to adjust the home/away split based on the win differential between the two teams. for a 7 game series with 2 teams separated by 5 games or less a 4/3 home/away split, for 5 to 10 wins it’s 5/2 split, for 10-15 wins it’s a 6/1 split and > 15 games all 7 games are played at the better regular season team’s site.
This is essentially the logic of the current first round, though right?
3 games at home > 9 game split series for a team with a 50-point lead or >15 game series for a team with a 30-point lead.
Still the Phillies took out a +37 Cardinals, the Padres knocked out +74 Mets, Mariners, a +12 Blue Jays team. At some point you just have to let the teams play.
I really like the methodology used here.
One minor issue with it is that the games played are given, rather than a result of the previous matches.
To give an extreme example,
Team A: 2-0 in WC, 2-3 in DS (4-3 total)
Team B: 2-1 in WC, 3-2 in DS, 3-2 in CS, 0-4 in WS. (8-9 total)
Given similar talent levels, team A did a better job at exceeding their expected wins.
But of course, that sounds silly.
A proper simulation is probably needed to account for this issue and I doubt that the numbers would change that much, but just some food for thought.
p.s. As Dan pointed out, the Dodgers are 47-35 since 2015 during the postseason.
That is equivalent to a 93 win record during regular season facing other postseason teams. Hardly a shameful record, especially taking 2017 world series into account.
As a general philosophy, when the ultimate precision isn’t really important, I try to do things in the simplest way possible that makes the point accurately. There are certainly ways to do this in this vein!
It would be interesting to see how an article like this would go over with the subscribers of “The Athletic”.
Or worse sports talk radio hosts or callers, I’m not sure whether the hosts or callers are dumber.
Hey, when I was working for ESPN, I got away alive after an article insisting that the players caught for PED use so far had shown no pattern of underperforming over overperforming projections at any point of their career.
Wonder how many folks at ESPN viewed sudden “spikes” in performance as proof positive that someone was using PEDs.
If he were alive, I’m sure Michel Foucault would be nodding his head.
Of course if he were still alive, he would have just celebrated his 96th birthday and would be completely exhausted by the deluge of “alternative facts”.
Could you link or just give the article name here? I was just discussing this the other day with someone, and they were adamant I was wrong about this.
https://www.espn.com/mlb/insider/story/_/id/10922627/melky-cabrera-strong-start-helps-demonstrate-why-peds-really-enhance-performance-mlb
Unfortunately, they took some liberties with the title!
(I’ve continued to TRY and find something for the decade since and I still can’t find the slightest thing)
Over 27 completed seasons since the advent of 3-division play with 3 or more levels of playoffs (1995-present), the BEST record in MLB has won the World Series 7 times (25.9%). The best record has appeared in the World Series 13 times (48.1%)
In the 24 seasons before that (1970-1993), where there were only two rounds of playoffs, the BEST record in MLB won the World Series 7 times (29.1%). The best record appeared in the World Series 16 times in that span (66.6%)
Prior to that, the best record in the MLB appeared in the World Series every single year! 🙂
With MLB moving to a more balanced schedule anyway, does anyone think MLB should just get rid of the divisions and just let the six teams in each league with the best records into the playoffs? I can see arguments for either way.
Team rivalries = $$$
Division title races = $$$$$$$
What if we reframe the Reddit poster’s thoughts in a way that assumes modern analytics is fundamental to building a successful + sustainable franchise in 2022 (that may be too generous based on the actual comment, but I try to be if I can)?
I get it and largely agree with the premise that the playoffs are a crapshoot. But say we’ve finished the marathon of the regular season successfully, having built a winning team through a well-informed analytical method. Outside of continuing to trust the numbers that tell us what matchups to favor in our pressurized sprint of a playoff format, what else have we learned that can help us find an edge? Are there yet unexplored dimensions of the game that we don’t fully understand in this era of baseball (kind of like the before-times of modern analytics)?
This is where I can relate to some of what our Reddit friend is saying. Is there room for qualitative answers/edges in the post-season when we run into the level of error the playoffs present from a quantitative perspective? Or do we stop at the crapshoot and sing Que Sera Sera until the final out of the season?
Smarter people know, I’m sure. It’s just fun to imagine what the next revolution might look like.
Was absolute or relative payroll one of the variables tested ?
Starting with the Pythagorean winning percentage presumes that the regular season performance is equivalent to true talent level, which it is not. Luck impacts the regular season as well. I would look at factors like BABIP that might reveal teams that seem to have benefitted from luck during the regular season(like this year’s Dodgers).
That’s not really the assumption.
The point of this article isn’t *to* create the best postseason projector. If it was, Pythag wouldn’t be the base, nor is it in the ZiPS postseason projections.
Pythag is used here is because it’s straightforward and simple. The interest here isn’t how good Pythag is, but the *relative* improvement in a model by different variables being added *to* Pythag. If, for example, postseason experience was a useful predictor, then Pythag + experience should be a better model than Pythag alone.
Ok, but did you try variables like a team’s pitching BABIP which might serve as a proxy for “luck”?
The thing that jumps out the most to me in that list? The Pirates. Just. Ouch.
Had the same thought. Three wins in almost three decades – not series wins, just regular old wins. FFS
Math is everything. Luck is also everything.
As the old saying goes, sometimes you’d rather be lucky than good.
(Both lucky and good? Now THAT may be the key to postseason success!)
It’s funny how the Guardians have outperformed expectations but literally nobody would believe that.
Or, as it seems to be…the Yankees playing Houston.
I appreciate the attempts via data modeling. Such a shame we can’t get the AUC above 0.6!