The Year in Projecting
Hello. There is still regular-season baseball taking place, even literally right now, but I am an impatient person. So, here’s a post about the year in projections, even though the year isn’t finished. The year is basically finished, and that’s good enough for me. We could re-visit in a month, but I don’t know if I’ll see the point.
You know, at least anecdotally, that it hasn’t been the best year for team projections. We had the Rangers projected as one of the worst teams in baseball, and they’re currently leading their division. We had the Twins projected as one of the worst teams in baseball, and they’re still alive in the wild-card race. We had the Royals projected as just about a .500 team, and they’ve got the best record in the American League. Then there are the Mariners, and the Red Sox, and so on and so forth. It seems like it hasn’t been a banner season for the numbers.
We can do something with that. We can compare, firstly, actual winning percentage against projected winning percentage from just before the season. Do that, and you get a simple R^2 value of 0.27. On its own, that seems like a fairly low relationship, but it’s better to try to put this in context. Thankfully, I have some team-projection data stretching back to 2005, and while many of those projections came from other sources, ultimately all projection systems work similarly so we can use this. Here’s how the projections have done over time, comparing actual performance to expected performance:
On that graph, you see a couple peaks, representing particularly unremarkable regular seasons. This season is nowhere near the peaks. As a reminder, this year’s value stands at 0.27. The average over the whole window is about 0.36. So, in that sense, this has been a bad year for the projections, indeed. They’ve under-performed their usual baseline.
That much, you all probably could’ve guessed. But now let’s look at this again, only in place of actual winning percentage, let’s substitute BaseRuns winning percentage. To be brief, we’re under no delusion that we perfectly understand everything about baseball, but to this point no one’s shown a consistent ability to over-perform its BaseRuns record. (Or the opposite.) It seems like mostly randomness, and randomness can’t be predicted, by definition. We shouldn’t expect projections to capture all the noise. Here’s a different version of the above graph:
It’s almost a perfect mirror of the first graph, but for 2015. As noted, between actual and projected winning percentage, there’s an R^2 of 0.27. But between BaseRuns and projected winning percentage, there’s an R^2 of 0.40, yielding the biggest difference between the two observed. For the above graph, the overall average since 2005 is 0.39, so in that sense this has been a totally normal year for the team projections. They’ve done as well as usual on the BaseRuns. They’ve done worse than usual on actual, meaningful, real-world record.
So that gets to this recent post by Dave. BaseRuns has proven to be a pretty good estimate of performance. Teams, historically, haven’t strayed very far from their BaseRuns records. This year, relatively speaking, has been nuts. Consider the following graph, tracking the average differences between actual records and BaseRuns records. This is expressed below in terms of winning percentage.
Used to be, the average difference was a hair above three wins. Topped out at 3.8, in 2009. This year, the average difference is on track to be 5.5 wins. Not only is that the biggest; it’s the biggest by a huge, huge amount. As Dave said, this really is the year that BaseRuns hasn’t been up to task. It hasn’t been nearly as accurate as before.
That doesn’t mean the equation is no longer applicable — this isn’t proof that BaseRuns is broken. Weird things do happen, and sometimes they clump together. Before this year, BaseRuns did the worst in 2009. Then it did the best in 2010. This doesn’t have to be the beginning of something, and even if you wanted to believe in a trend, there’s not much evidence of anything going on before this year. At this point, it’s just a blip. An interesting blip, a blip to monitor moving forward, but you shouldn’t get ahead of yourself. Baseball usually doesn’t make the math obsolete overnight.
But maybe teams are learning about clustering. Maybe they’re way ahead of the rest of us. That’s why this is a thing to pay attention to. Next year’s numbers will be revealing. It’ll be a long wait.
Switching gears real quick, I figured I might as well put the following in this same post. We knew in the first few months this year’s projections weren’t looking great, especially in the American League. But, have things gotten more normal over time? I decided to split at July 31, right after the trade deadline. Things we knew on July 31: actual team record, projected team record, and projected team record based on season-to-date statistics. Thing we know now: team performance since the beginning of August. This is mostly just for fun, but let’s look at some plots.
First, team performance since August 1 against team record at the deadline:
All right, there’s some relationship, but it’s obviously noisy. Now, team performance since August 1 against projected team record at the deadline:
What do you know — that’s much stronger. Still very obviously far from perfect, but it beats the hell out of the first plot. Anecdotally, some wins for the late-season projections: the Mariners, the Red Sox, the Dodgers, and the Indians. The projections never gave up on these teams, and lately they’ve played more like they were supposed to.
Not that that excuses the first few months. That’s why I said this is mostly for fun. Lastly, team performance since August 1 against projected team record at the deadline, based on season-to-date statistics. If I’m not mistaken, this would fold in various trade acquisitions.
It sucks! It’s a sample of one season, but it sucks. If I run a multi-variable regression, this tells us basically nothing, and actual team record tells us only a fifth as much as projected team record. Because we’re looking at only one year, there’s not a whole lot we can conclude based on this, but consider it a little more evidence that you can never afford to just abandon what the projections are saying. That’s an over-statement. This is sports. You can totally afford to abandon the projections. Just, expect to be wrong, if being wrong matters to you.
Overall, the projections contain good information. They do to some extent tell the future, and while this year the projections haven’t been great at projecting record, they’ve been much better with BaseRuns, and deviations between BaseRuns and record seem at this point to be almost entirely random. That’s going to be a thing to watch, because it’s possible that baseball teams are figuring out how to sequence, or something along those lines. The industry is always ahead of the non-industry, and this would be a big area of study. Lots of teams would be interested in figuring out how to beat the expectations. But you don’t want to defer too much to the experts, because they don’t know that much more. This year could absolutely be a blip, and we’ll know more in another 12 months.
Myself, I’m mostly happy with where the projections are. There’s no sign of their getting better, on the team level. Maybe that comes as a surprise, I don’t know. No real progress, over the years. I like how much they’re able to say, and I also like how much uncertainty remains. There must exist some sort of sweet spot, where we can know just enough without knowing too much. I have to think that we’re in it. I imagine we’ll stay in it for the foreseeable future.
Jeff made Lookout Landing a thing, but he does not still write there about the Mariners. He does write here, sometimes about the Mariners, but usually not.






The Twins would be closer to the top if they had brought up Berrios a couple of weeks ago…
Looks like we have the weekly explanation of why the projections are not broken. I’m not even going to argue for either side, but I feel like I’ve been reading variations of this article here all summer with nothing new ever added to attempt a further understanding. That’s the frustrating part.
you’re asking for brainstorming? Analyzing what has actually happened is the first part of figuring out what to do better next time.
Good thing the first part has happened in many other articles here before, especially recently. Now where’s the second part?
+1
I don’t think that “clutch” explains any of the variance this year; I don’t think that the theory or method of projecting is poor; I don’t think that anyone on FG has an agenda (beyond the one we all share: embracing statistical analysis as a means to understand and forecast baseball).
I do, however, think this year has made some commentators look a bit foolish in their arrogantly repeated appeal to projections, and constant chorus that what we have is correct and the rest is noise. “What we have is good enough” is what this movement was born against. I’d rather have some creative hypothesis explaining divergence — even if the hypothesis was ultimately proven wrong by the piece itself — than I would yet another article showing how at best we have .36r2 and that that’s good enough thank you very much.
Creative thought, criticism, and rigorous analysis to explain things our numbers have yet to accomplish is what brings me here. Every time a commentator arrogantly pronounces how right they are (blaming everything that didn’t come true on variation) I fear we’re veering ever-closer to the old timers, just with a new set of idols to cling to no matter what we see.
Nobody is saying that what we have is good enough though. I’m sure many of the authors on this site have ideas or guesses as to what causes or explains some of the divergence. However, rather than publish wild speculation that ultimately gets proven to be incorrect, they exhibit patience and wait to prove that their findings meet the rigors of science and basic statistical analysis. The null hypothesis is always no effect and until we can extract signal from the noise then it is irresponsible to pretend that we can. Advances are still being made but they are much slower than they were many years ago because we have already taken all the low hanging fruit.
My hope is the above comment is understood and appreciated. As long as the results of these studies are stating facts and providing analysis on facts, knowledge advances. Also, the “Low hanging fruit” comment is spot on. We are just not going to get a groundbreaking “That’s IT!” on one of these things in all likelihood, at least without more advances in technology. The rest of this stuff will be grinding.
First line of the last paragraph: “Myself, I’m mostly happy with where the projections are.” So, the author of the article appears to be disagreeing with your opening statement.
We already have a pretty good theory of what causes the variance, though you inexplicably dismiss it in your first sentence.
“Clutch” (i.e., sequencing) is highly correlated with winning, as Jeff as shown in previous posts. “Clutch” also appears to be almost entirely random, as Jeff has also shown in previous posts, so it seems likely that even a hypothetically perfect model would be unable to incorporate much or any of it in projections.
I agree that writing based on this stuff appears overly confident, if not arrogant, but heavily qualified opinions tend to be less fun to read and those qualifications are already implied to knowledgable readers given the sources the writers transparently refer back to.
“constant chorus that what we have is correct and the rest is noise.”
You made this position up to make your position look better.
Through the name changes, you’re as tiresome and pedantic as ever.
I agree. To move this discussion forward I recommend a quantification of the uncertainty.
I’d be much more interested to see how accurate individual projections are. Because just based on my own cursory glance the answer appears to be “not very”.
It would seem logical that individual player projections are far less subject to noise and volatility than projections that need to account for 25 – 40 players and their respective playing times.
Not very accurate compared to what?
Shirtless George Brett’s predictions, of course.
Its not to hard to check. I pulled all the preseason steamer projections for wOBA then ran a correlation between that and all the qualified hitters so far this year and the R^2 was .36 for wOBA.The standard deviation of the projected wOBA and the actual wOBA is .0295.
I don’t know if that’s good or bad, but it seems better than nothing.
I hate to be that guy but if you can come up with a better way to forecast HUMAN performance over the course of something like a baseball season, please publish your work and then find a job with a MLB team that will pay you a boatload of money. If you can’t, I don’t see why you are complaining.
Couple things…
1) Who is complaining? I asked a legitimate question about the accuracy of individual projections. I would be very curious to learn the answer regardless if its “good” or “bad”
2) You are essentially saying that any system is invunerable to criticism unless you first come up with a better system. Which is just absurd. Thats like saying someone cant rate a movie until they make their own, better movie.
1) Hopefully my comment above gives you some idea how well they do. Though I guess I should have told you that the R^2 between last year’s wOBA and this year’s wOBA is .28. So .36 is better, but maybe not be a huge amount. Though interestingly, the slope is a lot flatter in the 2014 vs 2015 regression (.6 vs .75), which suggests to me that’s an important .08 increase in the R^2.
2) I wouldn’t say you have to make your own projection system to have criticisms of a projection system, but I will say its nothing like your movie analogy either. Your question of how good are they is certainly worth knowing, however your criticism based on browsing things isn’t a meaningful or helpful criticism.
“You are essentially saying that any system is invunerable to criticism unless you first come up with a better system.”
The word “essentially” here is being used to denote the word “not”.
The word “essentially” here is being used to denote the word “not”.
Not sure why it would be. Thats is pretty much exactly what he said. He said if you cant come up with a better system then he doesnt see why you are complaining.
How else should that be interpreted?
Too many commas, I say. I say, you use too many commas.
No. Most people use too few commas.
It’s as painful to watch as Bartolo : running a 100 yard –
For all the “knocks” on projection systems, it’s really the AL that’s been totally wacky this year.
In the NL, projection systems said the Dodgers, Cardinals, and Nationals would win their divisions, the Pirates would get a wild card, and the Cubs, Giants, and Mets would battle for the second wild card. The Phillies, Braves, and Rockies would be terrible and everyone else in the middle. The only miss there is that the Nationals well under-performed, “elevating” the Mets from WC contender to division winner.
Jeff, how do those graphs look for “NL only”?
I think a lot of it has to do with the parity in the AL. There aren’t any really bad or really good teams, so sequencing, clutch, luck, and having a bullpen that’s performing really well, is playing a larger role than normal in w-l outcomes.
I agree that it is mostly the AL that has “broke” the system, but at the same time, almost any objective look at the A’s and Red Sox rosters this spring could see that the projections were over ambitious. For some reason the system went heavily towards best case scenario for those teams and ignored the reality that many of their players just aren’t as good as they “should” be. If you take those 2 flop predictions out the rest is largely just variation around .500 by average teams.
But the A’s still have a run differential that’s nearly identical to the Rangers, and it was one of the best in the AL before they traded away some good players at the deadline.
Then you have the Twins, whose combined pitching and position player WAR is 22.5, who have a slightly better record than Cleveland, whose combined WAR is 39. Which isn’t really a case of the projections breaking, it’s a case of WAR breaking
I’m no statistician, but 0.36 doesn’t seem like an incredibly strong correlation, especially when the rest of the graph makes the two ‘best’ years look a lot like outliers.
2015 looks more like a return to the 2008-2012 normality than it does an outstandingly bad year – the interesting question seems not to be ‘why were 2015’s projections unusually poor?’ as much as it is ‘why were 2013’s projections unusually good?’
Do note that the 0.36 is a R-squared, not just a correlation. R-squared is just the percentage of the dependent variable explained for by a linear model. Basically how well does the data fit the line.
Does the data fit the line well or not, then?
This season is pretty unique. It has the highest rate of K+HR+BB+HBP ever. I think that means there should be an increase in the amount of error we could expect between baseruns and actual runs. That’s because this season has the lowest proportion of balls in play ever. With fewer BIP events, sequencing plays a greater role in the actual outcomes. Thus, there is a higher chance of error between baseruns and actual runs.
It might not be so much that certain teams are learning about clustering – just that its impact on team record is greater this year than before because there’s so much parity, especially in the AL
Do you want to customize your sport banners or team banners? Why wasting time and money to have someone else design soccer banner for you. At Team Sport Banners you can easily create your own soccer banner designs.
Do you want to customize the banner or sports team banner-why spend time and money to someone to design a banner for your football. In the banner, you can easily create your own banner Soccer Design.
Seriously, flip a coin. You would likely get the same results.
Jeff, have any of the models changed over the last ten years? Because the projections don’t appear to be getting any closer to being accurate than they were then, I’d be interested to know if, say, 2005’s Steamer algorithm is any better than 2015’s.
What I hope this shows is we can’t just go “ZiPs predicts this guy for 5 WAR” or “Steamer predicts this team will get 90 wins” because we’ve seen the models fail pretty badly this year (no matter how insistent you are that they’re fine, really). If they can’t do any better than they’ve done this year, I’m really not sure why we’re all wasting so much time poring over them.
I don’t understand this. Why is “ZiPs predicts this guy for 5 WAR” worse than “I think this guy will have a 5 WAR season?” It’s information, so we use it. Baseball is entertainment, it’s supposed to be a waste of time.