Checking in on Pythagoras

Kiyoshi Mio-Imagn Images

This June 25, the Dodgers and Tigers both played their 81st game of the season. Both teams finished the day 50-31, sharing the best winning percentage in baseball at .617. The Tigers got there with a slightly better run differential, though; their Pythagorean winning percentage was a cool .608, while the Dodgers checked in at .595. Pythagorean record is implied by runs scored and allowed, and broadly regarded as a more stable measure of talent than simple wins and losses. Since that day, though, the Tigers have gone 35-40 (.467 with a .483 Pythag), while the Dodgers have gone 38-37 (.507 with a .556 Pythag).

I’m bringing this up – last data project for a while, incidentally, I just had a bunch of things in my queue and couldn’t resist tackling them all – because “how good is that team, anyway?” has been a hot topic this year given the various surprising teams who have, at times, taken up the mantel of “hottest in baseball.” Versions of this question – “This team is doing well/poorly now, what does that mean for next month?” – have been both interesting and top of mind in 2025. The Tigers and Brewers played so well for so long that they each crashed the best-team-in-baseball debate. The Mets did their hot-and-cold thing. The Dodgers have endured multiple fallow stretches. Sometimes, teams felt like they were getting very lucky or unlucky relative to their run differential. But what does any of that even mean?

I highlighted the midpoint of the season because it fits into my experimental method. I was interested in answering this specific question: If we stop at the halfway point of each season and consider a team’s actual record and Pythagorean expectation, which does a better job of predicting its record in the second half? I took every game from 2010 through 2024 and used that to construct each team’s record and Pythagorean record at the halfway point. I considered each of those as estimates of second-half record. Then I measured three things, all related: 1) the correlation between first-half record of the selected type (actual or Pythagorean) and actual record in the second half, 2) the root mean squared error of each option, and 3) the Brier score of using first-half metrics to predict second-half record.

If you’re well-versed statistically or just read my last investigation, you know that Brier scores are the accepted best metric for measuring questions like this, where you make a projection and compare it to the actual outcome. If we’re trying to come up with how good a team is and wondering whether actual record or Pythagorean record is a better representation of future play, measuring which makes a better prediction of future records feels like the exact way to go. The lower Brier score of the two will be the one that has less error in its estimates.

To that end, I first calculated every team-season, split into two 81-game halves, from 2010 through 2024, excluding 2020. To test a given metric, I noted each team’s performance in that metric (actual record, Pythagorean record, several contenders you’ll meet later) in the first half and its actual record in the second half. I threw in a coin flip version that predicts each team will have a .500 record in the second half, just for fun. Here are the takeaways of that investigation. In each table that follows, I’ve highlighted the best performance in each metric in yellow:

First-Half Prediction of Second-Half Record, 2010-2024 (Excluding 2020)
Predictor Correlation RMSE Brier Score
Actual Record 0.5466 0.0832 0.2483
Pythagorean 0.5595 0.0823 0.2482
Coin Flip n/a 0.0928 0.2500

That’s a nice first result, and one that matches with the existing literature. In The Book, Tom Tango and his co-authors found effects of similar magnitude; they found RMSEs of roughly similar size and reported that Pythagorean expectation did slightly better than actual record when it came to predicting future record. They used seasons rather than half-seasons as a measure, and used different years, but the similarity of results is still gratifying to me. At various points, Tango has also measured this effect via correlation coefficient and found similarly sized numbers, with coefficients in the mid-50s.

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

But overall this is not a very satisfying answer. Yes, Pythagorean expectation is a better predictor of future record than actual record. No, it’s not that much better. To use Brier skill score language, actual record is hardly better than just picking .500 for every team’s record, decreasing mean squared error by only 0.67%. That’s tiny! Using a team’s Pythagorean record to predict the future isn’t much better, though; it’s only a 0.73% improvement in mean squared error relative to pure randomness.

The drawbacks of using actual record as a predictor are fairly obvious. A team that has gone 19-1 in one-run games and been outscored overall probably isn’t as good as a team that has the same record but with a 10-10 record in one-run games. What are the drawbacks of using Pythagorean expectation, then, that make it only slightly better than simply caring about wins and losses? The easiest one to pinpoint is that formula’s insistence that every run is worth the same.

The Royals won a game 20-1 last week. They wouldn’t have been more or less likely to win that game if they’d stopped at 10-1. Those last 10 runs still feed their run differential, though, and thus their Pythagorean record. And, oh yeah, they scored all of their final 10 runs against a position player pitching, after both teams had stopped contesting the game. If we really want to see how good teams are at winning games, we probably need to come up with some adjustment for games like that.

I went with what I’d consider the simplest option. Since I already had every game’s score, I just took the Pythagorean expectation for each game independently. Win a game 10-1? Pythag assigns that game a .985 winning percentage. Win it 20-1? Pythag assigns it a .996 winning percentage. That’s exactly what we want – those last 10 runs do very little to change our estimation of a team’s talent. Then I summed up every single game “expected winning percentage” to get a team’s first-half game-by-game Pythagorean expectation. I calculated two versions of this metric, one that uses this exact formula with no modifications and one that adds a single run to each team’s score in each game to avoid counting shutouts as all the same. (The Pythagorean formula gives you a 100% winning percentage if you don’t allow any runs).

Whether you modify game-by-game Pythagorean record for shutouts or not, it beats our other estimators of future talent:

First-Half Prediction of Second-Half Record, 2010-2024 (Excluding 2020)
Predictor Correlation RMSE Brier Score
Actual Record 0.5466 0.0832 0.2483
Pythagorean 0.5595 0.0823 0.2482
Game-by-Game Pythag 0.5499 0.0778 0.2474
Game-by-Game Pythag (Adjusted) 0.5560 0.0771 0.2473

But again, it beats it by so little! Skill score says that my two methods each reduce mean squared error by about one percent compared to just guessing every team’s record will be .500. Did you expect more? I expected more. Shouldn’t looking at a team’s Pythagorean record do a much better job of predicting its future than looking at its actual record? Shouldn’t my fancy, hand-calculated version that has special accounting for blowouts do even better? Those skill scores are so tiny. The error terms are still so high. I decided to try one more method: Splitting the difference. I took the average of a team’s actual record and Pythagorean record at the midway point and used that as my estimate of future record. It did better than either alone, but still worse than my modified game-by-game method:

First-Half Prediction of Second-Half Record, 2010-2024 (Excluding 2020)
Predictor Correlation RMSE Brier Score
Actual Record 0.5466 0.0832 0.2483
Pythagorean 0.5595 0.0823 0.2482
Game-by-Game Pythag 0.5499 0.0778 0.2474
Game-by-Game Pythag (Adjusted) 0.5560 0.0771 0.2473
50/50 Blend 0.5650 0.0810 0.2480

I’m satisfied that there’s no way to use either actual record or the runs scored in each game to beat randomness by all that much. I’m also satisfied that you shouldn’t use either actual record or Pythagorean record alone; blending them improves their performance. All of the options I tried beat a naive expectation of every team being equally skilled, but none beat it by all that much. I wasn’t quite stumped, though. I have one other source of high-quality team data: Projections. Instead of stopping 81 games into each season and using those games to come up with some estimate of future winning percentage, I stopped 81 games into each season and simply looked up each team’s rest-of-season projected winning percentage. Yet another improvement:

First-Half Prediction of Second-Half Record, 2014-2024 (Excluding 2020)
Predictor Correlation RMSE Brier Score
Actual Record 0.5633 0.0831 0.2483
Pythagorean 0.5761 0.0823 0.2482
Game-by-Game Pythag (Adjusted) 0.5785 0.0756 0.2471
50/50 Blend 0.5820 0.0809 0.2480
Projections 0.6098 0.0737 0.2469

(Note that the numbers are slightly different because we only have projections starting in 2014.)

For the record, this is using FanGraphs mode projections, which take ZiPS and Steamer for player talent and Depth Charts for playing time. Use season-to-date mode instead, and our projections perform worse than game-by-game Pythagorean.

The takeaway from all of this, at least for me? It’s really hard to make projections of future record, what with so much randomness baked into baseball. Using actual records or Pythagorean records to come up with estimations beats guessing randomly. Averaging those two does even better. Going game-by-game and handling blowouts differently is better still. Even that method can’t beat computer-driven projections. And yet that projection-based method, the best in our study, still only reduces mean squared error by 1.3% relative to pure chance.

None of this means that you shouldn’t watch a current season and try to guess the future, of course. That’s why we all like baseball so much. But next time you hear that a team is unsustainably over its head because its record and Pythagorean expectation don’t match, or that a team “can’t keep getting this unlucky,” remember that none of these methods are all that much better than random chance. Can a team keep playing over its head? Sure, and we’re not even that great at measuring where its head is, to continue the analogy. Can a team keep getting this lucky or this unlucky? Obviously! Baseball is a sport governed by randomness at the game level.





Ben is a writer at FanGraphs. He can be found on Bluesky @benclemens.

12 Comments
Oldest
Newest Most Voted
jcutigerMember since 2024
10 months ago

Looking at fWAR, the Mets have way underachieved win total. Is there a formula that uses a team’s fWAR (positional, pitching, etc) to predict win total?

David Klein
10 months ago
Reply to  jcutiger

Struggling with runners in scoring position most of the year and going 0-67 after being down after the seventh or 8th I forget doesn’t help. I believe base runs has them like 7-8 wins better.

3cardmontyMember since 2025
10 months ago
Reply to  jcutiger

Replacement level is around .300, so add WAR to 30% of games played to get expected wins

Paul BoisvertMember since 2021
10 months ago

Hi, Ben,
Your results didn’t surprise me at all. A half season isn’t nearly “consistent” enough to get a particularly good estimate of future performance from either wins or runs. And the future team is actually a different team from the current one! Teams change (“in effects”) every day in random ways, depending on: the schedule, home field, who is playing for them or their opponents, particularly who is pitching, who’s been called up from or sent down to the minors, who is injured (esp. when “playing through injuries”), who has adjusted their batting stance or pitching strategy, what the weather is, random strategy decisions (“guesses”) by the managers, and of course all the other innumerable chance outcomes in the “game of inches”. The second half “team” (and its opponents’ second-half overall “opposition team”) simply aren’t the “same (unitary) team” as the first-half team. “Teams” are huge supersets composed of “temporal slices” (on any given day) of versions of the players that make them up, and no player is exactly the same from one day to the next…

Just out of curiosity, which “Pythagorean” formula are you using (exponent, run-context related, etc?) Not that it matters, since with all your correlation coefficients r being around .58, we get that r-squared, the coeff. of determination, is around 1/3. So around 2/3 of the variation in 2nd half performance is due to factors OTHER than the first-half data, for all the predictors. Yes, baseball is full of chance… 🙂

However, if you are interested in getting very slightly better results for your game-by- game analysis, you might consider a few additional things: a) adjusting extra inning games as you did shutouts. Since winning an extra inning game on a grand slam is really no different than on a suicide squeeze, list extra inning games with the winning team having only one more run than the loser actually scored. And/or b) perhaps only add 0.5 runs to each team’s score for shutouts, or maybe even a smaller decimal. The formulas don’t require whole numbers of runs, and presumably the less distortion of the actual run totals the better.

And/or c) Yes, definitely cap blowouts, as you mentioned. Maybe score games as no more than n+5 to n, where n is the losing number of runs. Or n + k to n, where you could try various k values to see where the best “cap” lies… My guess is this idea is the only one that would make “much” difference, but much still wouldn’t be very much, I suspect… And/or d) for the “actual record”, whether predicted or given, regress one-run games a bit, as they clearly have a substantially higher likelihood to have had the win or loss be due to “luck”. Perhaps regress a 19-1 record in one-run games halfway back to the mean, making it 14.5 – 5.5, and treat that as the “actual record” in the relevant half-season, to see if now the correlations to the predictors (other half’s “adjusted” record, or Pyth.) improve.

I have no idea if any of the above would help, and no inclination to investigate myself, as I actually don’t bother speculating about future performance–being a Tigers fan, I just expect them to self-immolate automatically, as is happening this season. Then I am very pleasantly surprised, about twice per century, when they prove me wrong… 🙂

Jeremy FoxMember since 2026
10 months ago

It’s striking that even the best projection systems are only a bit better than actual record at the season’s halfway point. Dan S. did all that work to make ZiPS, and it’s only a slight improvement on saying “Eh, the future’s going to be just like the recent past.” I feel like you see this in various other sports, and in other domains of life too. For instance, Nate Silver’s election forecasts go to great lengths to slightly improve on naive averaging of polls. As another example, you need a massive amount of data on all sorts of atmospheric variables, and a supercomputer-sized simulation model of the atmosphere, to make weather forecasts that are only a bit better than going “eh, tomorrow’s weather will probably be the same as today.” My hat goes off to people like who work so hard to squeeze just a little more juice from the orange.

brentdaily
10 months ago

What I hear you saying is, “don’t bet on baseball.”

Welp, there goes that sweet, sweet avalanche of DraftK*ngs revenue. (*No free ads.)

thecoracleMember since 2020
10 months ago

I’d be curious to see how something like BaseRuns record stacks up. Probably better than Pythagorean record but also not super-great? Perhaps a topic for a follow-up article during the long offseason.

indianbadger1Member since 2018
10 months ago

Instead of just run differential, maybe add compression factor to blowouts. Maybe after 100 games or so. Like any result 5+ runs is treated as a 5 run victory.

3cardmontyMember since 2025
10 months ago
Reply to  indianbadger1

This is essentially what his game-by-game method is doing.

sadtromboneMember since 2020
10 months ago

I don’t think I have ever thought that Pythagorean is predictive the same way a projection is. It tells you *how well a team has been playing* by considering whether they’ve been just squeezing by opposing teams or whether they are beating the tar out of them.

You can imagine that is predictive, but mostly because underlying talent predicts both run differential and future performance. And the projections are modeling underlying talent directly, which is why it works better for that purpose. Run differential or BaseRuns is more for us saying “man, that team was great!” (or terrible).

Travis LMember since 2016
10 months ago

Joe Sheehan is a fan of looking at extra innings and 1 run games to see how much variance has swung records for a team. I’d love if you’d release your code so others could quickly test these types of hypotheses out.

With tools like Google collab, it would be fun to crowdsource a model from the readers to see if someone can come up with a bigger edge.

With AI, it’s even more accessible to everyone, translating natural language into code is relatively straightforward! Maybe we could do something like this for a fun offseason reader competition?

bndj88Member since 2017
10 months ago

Win a game 10-1? Pythag assigns that game a .985 winning percentage. Win it 20-1? Pythag assigns it a .996 winning percentage.

I’m confused about this part. Pythag says you are only 99% likely to win a game that you definitely won, and won by 9 runs? Pythag is run differential based, so how could it have anything other than 100% win probability for a concluded game with a positive (and, in this case, large) run differential? This sounds more like a BaseRuns expected winning percentage than Pythag.