The Meaning of the Standings So Far: Adding On
Monday, I wrote a post, building off of another post. Now this is a post, building off of Monday’s post. To review, this is said post, wherein I examined the relationship between early-season team performance and rest-of-season team performance. How much might the current standings tell you? Going back 10 years, and choosing an appropriate date:
Right, this has already been published. As has the following plot:
As far as predicting the rest of the season was concerned, I observed a far stronger correlation with preseason team projections than with early-season team records. Which isn’t to say that the projections did great, but they did a hell of a lot better than actual wins and losses through roughly two months. Right there, you’re given an important reminder never to dismiss what the projections are saying. Or, were saying, as it were, in this case.
After I published the article, though, I realized I could’ve taken another step, and it was pointed out by a few people in response. While I looked at early-season wins and losses, I could’ve easily also looked at early-season Pythagorean wins and losses, based on run differential. That could help to eliminate some of the noise. After all, many of us recognize that run differential is a bit more meaningful than record, and you never want to penalize too much for luck.
So this following plot is the same as the first plot, except instead of actual winning percentage on the x-axis, we’ve got Pythagorean winning percentage. How well might this predict the remaining four months of the year? Does this give us a real step forward?
No, it doesn’t. Not really. Relative to the first graph, you do see a slightly stronger relationship. And, relative to the first graph, you do see a slightly steeper slope. But this still falls well short of the predictive power of preseason team projections. Erasing some noise didn’t make enough of a difference. You still shouldn’t put all that much stock in early run differential. Not if it’s in disagreement with what was expected, statistically, earlier on.
To keep consistent with Monday’s article, here’s a glance at the extremes. Again, this is through June 7. I identified the 25 biggest under- and over-achievers, comparing early Pythagorean record to projected record.
25 biggest early-season over-achievers
- Pythagorean Win% through 6/7: .597
- Projected Win% through 6/7: .486
- Actual Win% after 6/7: .499
25 biggest early-season under-achievers
- Pythagorean Win% through 6/7: .396
- Projected Win% through 6/7: .511
- Actual Win% after 6/7: .502
You see the same thing as before. The over-achieving teams have tended to strongly regress. The under-achieving teams have tended to strongly regress, in the other direction. Doesn’t make that much difference whether you’re looking at actual wins and losses or theoretical wins and losses, based on run differential. Ultimately, you’re still better off more heavily weighing the projections. Two months of team performance isn’t irrelevant, but it’s far from the whole picture of everything you need to know. Each team is made up of players, and all those players have longer performance histories.
It’s somewhere around here that I run into the limitations of my own mathematical ability. While my job is to be a baseball analyst, I’m mostly untrained in the art of advanced statistics, and when it comes to that sort of analysis I’m a total fraud. Other people could dig into these numbers even more deeply, but I’m content to stop at “trust the projections more, and the early performance less.” I think even these simple statistics can get the right point across. To put it all together, here’s a table, on forecasting the last four months:
| Approach | R^2 | Slope |
|---|---|---|
| Early Actual Record | 0.13 | 0.37 |
| Early Pythagorean Record | 0.15 | 0.43 |
| Preseason Projection | 0.27 | 0.93 |
At some point, it stands to reason you do become better off trusting the season-to-date performance, but that point is nowhere in June. For every individual team, there might be a player or two where you think the projections are just wrong, because of some change in true talent, but those aren’t actually easy to spot, nor do we know when they’re sustainable, and they also don’t make huge differences. Projections might seem too simplistic, but actual performance creates this illusion of meaning. Over smaller samples, we can’t help but underestimate the influence of noise.
There are two steps I can’t take. Here, I went from looking at actual early record to Pythagorean early record. It would be even better to look at BaseRuns early record, but that doesn’t exist. Because BaseRuns strips out more noise, I assume it would be more predictive than the run-differential stuff. Maybe even almost as predictive as preseason team projections.
But then, that’s the other step I can’t take: I can’t access historical in-season projections, updated for more recent events, because those also don’t exist. This whole time, I’ve been comparing between updated statistics and projections missing two months of data. The projections we have now are better, because we have updated numbers and updated depth charts, taking into consideration injuries and transactions. At this point, Pythagorean record is less predictive than preseason projections. My assumption is that BaseRuns record is less predictive than updated projections. That part, I can’t prove, but even the older projections clearly have worth.
It’s not real fun to keep writing “projections” over and over. I know it makes a good number of you tune out. And that’s fine — I might tune me out, too. But there’s a reason we rely on the things. They’re pretty good, and baseball is always lying to us. Believe me, I’m not invested. I don’t have my own projection system, and I think it would be equally interesting if the projections didn’t hold up. But here we are. You’ve seen the information. Make of it what you will.
Jeff made Lookout Landing a thing, but he does not still write there about the Mariners. He does write here, sometimes about the Mariners, but usually not.



Interesting article,
One of the advanced statistics I think it would be interesting to see would be to see if there is better predictive power if a linear model was performed when win%after june7 is correlated to “projection record” and confounded by “pythagorean win%before june 7.”
Do the projections not effectively contain some regression in them already? That the overachievers and the underachievers end up at .499 and .502 suggest that you might do better where the projections and the actual record (actual or Pythag) differ significantly on either side of .500, a simple .500 projection would do about as well as using either of the above.
If you used both the preseason projection AND the early-season Pythagorean W/L record (a 2-variable regression), would that bring the R-squared much higher than .27?
There’s some information to be had in the first two months of results, but it would be good to know how much weight that would carry vs. the pre-season projections.
I’m not sure why Fangraphs continues to dismiss data points away from the mean when they are just as important and valid as any other data point. They happened. It’s one thing to say that most teams will regress to the mean, quite another to say that a specific individual team will do the same.
The basic premise is true, that most teams will regress to their inherent talent level as assessed by the projections, and fans of the Twins and Astros should expect their teams to come back to earth. But some teams don’t, as the charts show, and there is no way of knowing which team you have. Each team is not the mean; the combination of all teams provides the mean.
Perhaps I’m mis-reading your comment, but you seem to be answering your own question.
The Astros might be a .700 team the rest of the way, or they might be a .300 team through the end of the year. You’re entirely correct in saying that for a single team, there are too many unknown variables to make any concrete statements. The projections are a statement of probability, rather than prediction.
How can FG “dismiss” the outliers? They aren’t. They just don’t know which teams will be the true outliers at the end of the season and which are short-term mirages. Probably, most teams significantly over or under their projections will play more like those expectations going forward. The wins (and losses) in the books will change the playoff odds dramatically, but they will change the Rest-of-Season projections much less so.
Certainly every year has its big surprises, and that, as they say, is why they play the games. But anybody who thinks they can predict *right now* which teams have made legitimate changes (for better or for worse) and which will regress back to the expectations is just playing roulette.
(Again, sorry if I’ve misinterpreted your comment.)
I think you’re addressing a bit of a strawman. It’s obviously not impossible that the Twins could keep overachieving, but that’s not the point. The point is that they’re not *likely* to keep overachieving.
There’s a difference between absolute knowledge and a reasonably strong belief built on past evidence (which is what the mean values in the regression represent). Saying that Jeff’s conclusions are invalid because it’s possible that the Twins could continue to play well is sort of like saying that you shouldn’t dismiss hitting on a 20 in blackjack because you might flip an ace.
Also, an OLS regression (what Jeff’s using) doesn’t ‘dismiss’ points far away from the mean–it minimizes squared residuals over the entire dataset. That’s basic statistics, not some editorial scheme cooked up by FanGraphs.
My comment was directed toward PackBob–apologies for any confusion.
thanks for taking this extra step, jeff.
good stuff.
Yup, agreed. Thanks.
Tune you out, Jeff? Never!
“At some point, it stands to reason you do become better off trusting the season-to-date performance”
This string of articles has got me thinking about just this concept, and I’m skeptical of whether or not this is true. Here’s another way of phrasing that statement: “It’s reasonable to expect that once X% of the season is complete, we believe that the remainder of the season, teams will tend to win games as frequently as they have to this point. We just don’t know what value X is.”
Well that’s hard to believe. Taken to an extreme, that would suggest winning percentage for the month of September alone could be predicted by the five months preceding it. Or that predicting the last week could be done by looking at the first 95% of the season. Or that teams with winning records are really likely to win the last game of the season.
Plainly speaking, the shorter the sample of baseball games, the more erratic and noisy the data looks. Of course the pre-season projections correlate more closely with the last two thirds of the season (after June 7th) than it does the first third of the season: there are twice as many games played after June 7th than before. I would be surprised if there is any point in the season where the portion of the season that precedes it can reasonably predict what will happen after it. Maybe the season’s midpoint….?
The valid statement is that ‘at some point, it stands to reason you do become better off trusting the season-to-date performance’…for predicting the full-season winning percentage (rather than the rest of season winning percentage, or pythagorean, or base runs, etc).
No, the valid statement is exactly what the author said: at some point of the season, using YTD pyth record alone will be superior predictor of rest of season record when compared to using pre-season projections alone. You’ve said nothing to disprove this theory.
P.S. If you use BaseRuns (esp. after adjusting for before/after SoS), the point in the season this occurs is much earlier, as the author correctly speculates. Source: 10 years of profitable baseball betting.
I wonder if it would make more sense to look at teams individually – and especially teams at the extremes – and see if we can find any signal in the early season which might outweigh the projections.
So, say there’s a team whose record, run differential, base run differential, and war (and perhaps other stats as well) after 2 months all line up to tell a similar story….and a story much different than the preseason projections told….is there possibly enough confirmation there that the early season sample does in fact become as or more predictive than the projections for this specific scenario?
aka is there any scenario in which we have a good reason to change our opinion of a team based off a couple of months performance?
This would be one way of testing the models. It feels to me like the models have been tested (and found imperfect) at predicting playoff series results. It turns out teams with certain attributes (dominant bullpen, etc.) did outperform in the postseason. So the models have been adjusted.
If there is any systematic reason we might know beforehand that a team is likely to outperform its projection substantially, the projections can be improved to capture that change. Fortunately, much of baseball, as life, can’t be predicted.
Could you look at observed winning percentage versus projected winning percentage from season start to June 7th? Could you also simulate/calculate a projected winning percentage based on each team’s schedule up to this point? Surely strength of schedule has had some impact on what we’ve seen thus far?
I’m not really familiar with the magnitude of the R2 usually seen in baseball stats, but what kind of stands out to me is how low they are for all the methods. I mean, at best only 27% of the variation can be explained by the model? Maybe I’m spoiled by the stats work I do in biomedical research, but an R2 that low makes it seem like the model is only just barely informative. None of these methods seem particularly good. Which I think is interesting! Baseball records are hard to predict!
I posted a link to June 7, 2013 on the article yesterday, but just wanted to say that you can use the Internet Archive/Wayback Machine to find updated ROS projections for many dates throughout the 2013 and 2014 seasons. You may even be able to go back further if the projections page used to have a different URL.
As a couple comments may have touched on, I’d expect the base runs pythagorean win % to be more predictive than the actual pythagorean record. In fact, it’s unsurprising that the pythagorean win % is less predictive than preseason predicted win %. Many of the over- and under-performers in the first two months are probably driven by unlucky/lucky sequencing in addition to unlucky/lucky run allocation. I suppose it’s too difficult to get historical base runs splits?
This is a poorly constructed article. Sullivan writes: ‘Projections might seem too simplistic, but actual performance creates this illusion of meaning.’
The stats are the stats and the projections are what they are. But anyone who has watched baseball enough can know that the two months of awful Seattle and Boston baseball is not ‘noise’. These are very badly constructed teams. I would gladly bet against the projections for their ROS record.
Disagreeing with the conclusion of the article doesn’t mean it’s necessarily poorly constructed. It just means you disagree.
If you think the article is poorly constructed from a writerly point of view, you’ve done a poor job of expressing why.