Why Are the Orioles’ Playoff Odds So Low?

At this point, it’s becoming a meme. The Orioles chug along, at or around .500, and our playoff odds continue to say that they’ll almost certainly miss postseason play. Across the internet, sites like Baseball Reference and FiveThirtyEight give them a higher chance. The headlines write themselves: “Why doesn’t FanGraphs believe in the Orioles?”
Just to give you an example, after the games of July 29, the Orioles were 51–49. Baseball Reference gave them a 34% chance of reaching the playoffs; we gave them a 4.6% chance. Ten days later, on August 8, Baseball Prospectus pegged them at 22.2% while we had them at 5.4%. On August 11, FiveThirtyEight estimated their playoff odds at 16%; we had those odds at 5.7%. Another week later, on August 19, Baseball Reference pegged them at 35.5% to reach the playoffs; we gave them a 4% chance. You can snapshot whatever day you’d like and you’d reach the same conclusion: we don’t think the Orioles are very likely to make the playoffs, while other outlets do.
Now, we’re getting down to brass tacks. The Orioles are 68–61 after Wednesday’s games. Baseball Reference thinks they are 43.6% to reach the postseason. FiveThirtyEight isn’t quite so optimistic, but still gives them 23% odds, while Baseball Prospectus has them at 29.9%. Here at FanGraphs, we’re down at 6.6%, even after they called up top prospect Gunnar Henderson. Why don’t we believe?
The answer comes down to the way we construct our odds. Our methodology is time-consuming but fairly straightforward. At its core, it relies on the two projection systems we use on the site: ZiPS and Steamer. We take updated rest-of-season projections for every player in baseball, as well as rest-of-season playing time projections from RosterResource. We stitch those together into estimates of a team’s overall line on both sides of the ball, then use BaseRuns to convert those lines into runs scored and allowed per game. Finally, we use a Pythagorean estimation to estimate a true talent level for that team. Repeat this process for every team in baseball, and it gives us the matchups for every remaining game in the regular season.
That’s the core of our disagreement with other outlets about the Orioles’ playoff odds: we don’t think they’re as good as their record. Mostly, that comes down to pitching. So far this year, the O’s have scored 4.21 runs per game while allowing 4.13. We think their offense will actually get better down the stretch, scoring 4.33 runs per game. The pitching fares far worse, though: we think they’ll allow 4.91 runs per game, which places them squarely below .500 in terms of expected record — .437, to be precise.
Those forward-looking odds take what teams have done this year into account, but they also take previous performance into account. Take Nolan Arenado, for example. So far this year, he’s hitting .307/.370/.571, good for a 163 wRC+. You could project that to continue for the rest of the year, but our projections don’t because they look at his prior career line as well. Instead, we project him to hit .273/.337/.494 the rest of the way, a still-excellent 132 wRC+. That offensive gap means we project the Cardinals to get less from him for the rest of the year than they have so far.
Baseball Reference uses a simple methodology — literally. They use Simple Rating System with some mean reversion thrown in, detailed here. In essence, they look at how the Orioles have performed in their past 100 games, add in 50 games of .500 baseball, and use that as their projected winning percentage. That’s quite similar to our Season-to-Date odds method, which sees the Orioles as slightly better than .500 the rest of the way; those odds have the Orioles with a 40.2% chance of reaching the playoffs this year. Their method is the easiest to replicate, so I’ll focus on that today.
That’s the central question to be answered in any consideration of the Orioles’ playoff odds. Should we focus exclusively on what they’ve accomplished this year, or take each player’s prior baseball history into account? There’s one further complication: both our Season-to-Date mode and the SRS method ignore personnel changes. They don’t know that the Orioles got better in June when Adley Rutschman was called up (though he’s been up long enough that his contributions are fairly well baked into the pie at this point), but they also don’t know that the team traded Jorge López and Trey Mancini at the deadline. That’s on purpose: the beauty of those two methods is that you can replicate them easily, no opaque projections necessary.
That raises an interesting question: why do our projections think that the Orioles pitching staff will perform worse than it has so far this season for the remainder of the year? If they keep pitching like they have, our odds are too low. To try to explain what the projections are doing, I’ve prepared a simple chart. Here are the top 10 Orioles pitchers’ ERAs and FIPs this year (using RosterResource to pick the top 10), as well as our ERA projection the rest of the way:
| Pitcher | 2022 ERA | 2022 FIP | Proj ERA |
|---|---|---|---|
| Jordan Lyles | 4.25 | 4.28 | 5.10 |
| Kyle Bradish | 5.63 | 4.90 | 4.62 |
| Dean Kremer | 3.24 | 3.52 | 4.77 |
| Austin Voth* | 2.72 | 3.68 | 4.83 |
| Spenser Watkins | 4.26 | 4.28 | 5.44 |
| Félix Bautista | 1.58 | 2.95 | 3.92 |
| Dillon Tate | 2.70 | 3.09 | 3.87 |
| Cionel Pérez | 1.64 | 3.10 | 3.98 |
| Joey Krehbiel | 3.02 | 4.59 | 4.88 |
| Keegan Akin | 2.78 | 3.76 | 4.35 |
Other than Kyle Bradish, everyone is projected to get worse. Mostly, that’s because they’d put up poor performances before this year in the majors. For rookies, their minor league numbers didn’t herald a breakout. If you believe in the projection systems, Orioles pitching has been playing over its head. That’s not to say it can’t continue to do so, but I think systems are right to be skeptical that an entire unit can continue to show a heretofore unseen level of performance.
Another way of thinking about it is that there have been plenty of times over the past eight years (that’s how long we’ve run our projections in their current form) where our season-to-date odds and projection-based odds have regarded teams differently. Which one has done better? We can test them using a “gambling” method that I used earlier this year.
My method: take each team’s playoff odds on every day of August for the 2014–21 seasons, excluding 2020. I used August to highlight the time of the year we’re most interested in here, when plenty of games have been played, but we’re not quite to an endgame scenario yet. No one’s interested in what each system says when there are three games left in the season, or when there are 130; it’s right now, when we’ve seen a lot of the season but there are still plenty of games on the docket.
From there, I had each system wager against the other system for every team, every day. Let’s use a hypothetical example: say that on a given day, the projections mode gave a team a 5% chance of making the playoffs, and the season-to-date mode gave them a 25% chance of making the playoffs. The projections mode would “bet” against the team making the playoffs, with that 25% chance from the season-to-date model forming the odds. If the team subsequently missed the playoffs, I’d credit the projections model with a positive 0.25 score. If the team made the playoffs, I’d debit them 0.75. Likewise, the season-to-date model would “bet” on the team making the playoffs. They’d lose 0.05 if the team missed the playoffs or make 0.95 if the team made the playoffs. I did that for every game in August, every year other than 2020, starting in the first year of our current methodology, 2014.
If our model is systematically messing things up in August, you’d expect the season-to-date mode to rack up a hugely positive score by consistently fading the excesses of the projections. Instead, the results were lopsided in favor of the fancier, projection-aided model. Over 6,510 projections (one per team per day of August), the projection model posted a score of +282. The season-to-date model posted a score of -25. In other words, if you see a disagreement between the two modes, you should lean in the direction of the projections, even if you don’t use them as gospel.
Maybe we’re not interested in August as a whole, though. Does anyone actually care that the Cardinals are 97.3% likely to make the playoffs based on projections and 99.6% likely to make it based on season-to-date statistics? Not really. I ran a separate cut — every day where the two methods differed by at least 20 percentage points.
I expected both modes to put up positive numbers here, because when two things are that far apart, the likeliest true outcome seems to be in the middle. I wasn’t disappointed; across 4,920 observations, the projections system accumulated a score of +940, and the season-to-date system accumulated a score of +502. Another way of thinking about that: our projection-based odds are generally picking up on something useful when they differ from simple season-to-date odds, which explains their high score. They’re too confident, though. They move too far from the baseline, which explains the smaller positive score achieved by betting against them.
For completeness’s sake, I also looked at only the times where projection-based odds were 20 percentage points lower than season-to-date odds, rather than when they were lower or higher by that much. Is there something specific going on when we think a team is unlikely to make the playoffs? Nope — the projections racked up a two-to-one score advantage, +387 to +190.
This has all gotten rather long-winded, but let me give you the high points. Our odds give the Orioles a lower chance of reaching the playoffs than some other popular sites. That’s because we bake in pre-2022 performance in making projections, rather than just their run differential this year. Over time, our odds have done pretty well, both late in the year and when they differ markedly from season-to-date-style odds. If you wanted to be cute, you could get an even better estimator by using two-thirds projection-weighted odds and one-third season-to-date odds, but that sounds a lot like overfitting to me. That would give the O’s a 17.8% chance of making the playoffs, which probably wouldn’t do much to placate fans.
Truthfully, I’m not surprised by this outcome. When I looked into our odds’ accuracy last year, I found that our odds were better than season-to-date odds even late in the season. I do think we’re too harsh on the Orioles; the true answer is likely somewhere in between the two. We tend to underestimate slightly teams with a low chance of making the playoffs; of the 2,201 observations with between 5–10% chances of making the playoffs, we predicted 7.4% would make the playoffs when in reality 11.5% did. Combine that with the fact that our method’s edge over season-to-date models fades late in the season, and it’s reasonable to think we’re low on the Birds.
But only a little low! For the most part, our odds do a pretty good job of estimating how hard the road ahead is for teams. That doesn’t mean, by any stretch of the imagination, that the Orioles won’t make the playoffs. All it means is that they have their work cut out for them. Plenty of players on the Orioles have exceeded expectations so far this year. If they manage to keep that up and make the playoffs, they will have defied the odds — whichever set you choose to use.
Ben is a writer at FanGraphs. He can be found on Bluesky @benclemens.
Off topic, but has anyone else noticed how favored the NL is favored about 60-40 to win the World Series this year in the projections this year? Every NL team that has a greater than 0.1% chance of making WS has a greater than 50% chance of winning it if they make it that far. Assuming this is based on relative strengths of the league, but that still seems like a large difference in perceived talent.
This is because the NL team is not likely to meet the Orioles in the WS according to the model. If they do, they’d have no chance
Damn straight.
You’re looking at it the wrong way. The FG odds are making the LADs and NYMs significant favorites, and I suspect ATL would be favored over HOU as well, thought that’s close.
Unless there are a LOT of upsets on the NL side of things, any non-HOU team from the AL is going to be a dog. And even HOU is going to be a dog most of the time.
The interesting part would be if the Dodgers are considered a favorite over the Astros or Yankees despite the iffy pedigree of their starting pitchers and possible crappiness of Kimbrel. Yankees and Astros have warts but frontline starters and relievers matter more in a short series
LADs have a 63% projected chance of winning the WS if they make it there. And give that 55% of scenarios involve them facing either HOU or NYY, I expect that means they are favorites over both of them.
What does this mean? They’ve got the third-best rotation by FIP. Gonsolin’s not this good, but he’s perfectly startable in a playoff series.
It would be hard for Gonsolin to be this good, but after a while you sort of have to look at what he has done imo
It means that that third-best rotation by FIP includes Buehler (who is out for the year), Kershaw (who is old and has battled injuries for a while and is about to come back from an injury, and we shouldn’t take for granted he’ll be the same again), and Gonsolin (who just went on the IL with a forearm strain, which is always an ominous injury). They’ve still got Urias and May and knowing them they’ll just transform Heaney into a postseason ace, but I think it’s pretty fair to say that the pedigree of the starters is currently iffy.
>It means that that third-best rotation by FIP includes Buehler (who is out for the year)
Dodgers rotation ERA/FIP
ERA 2.82 / FIP 3.40
Buehler ERA/FIP
ERA 4.02 / FIP 3.81
As weird as it sounds, Dodgers are third-best by FIP despite Buehler, not thanks to him.
List of Yankees/Astros starters with a lower ERA- than Dodgers rotation:
Justin Verlander.
That is it.
Yes, ERA has its issues, but we are talking 1100+ innings from the Dodgers rotation.
It means they’ve not been great until 2022, and their “stuff” according to observers isn’t great
While the “stuff” might not be great, there’s obviously more to pitching than stuff – there was a fantastic piece of analysis on pitch tunneling put out earlier this year that highlights why some pitchers outperform middling stuff, and why others underperform outstanding stuff: https://www.prospectslive.com/prospects-live/2022/8/23/the-mystic-art-of-pitch-tunneling
There are a lot of Dodgers pitchers mentioned in this article, but in particular it highlights why Gonsolin has basically perfect tunneling and how the Dodgers helped Heaney learn a gyro slider to improve his tunneling. The latter is at least a pretty meaningful change for performance in 2022 versus previous years.
Just backing out the odds the model “thinks” the braves are the best team in the NL by implied WS probability.
Yeah, I am likely interpreting them incorrectly but it also looks like the Yankees win the WS less often than the Jays, of the times they are projected to make it?
That’s how I read it too. The controversial take here is not that the top of the NL is so heavily favored, it’s that FG’s odds seem to like TOR more than NYY.
Hey, whatever happens – they’re a lot of fun with more reinforcements on the way. Jays and Twins keep sweating and M’s and Rays gotta keep an eye out. Projections are fun, but the game is more fun!
As an Orioles fan who knew more or less how the BR and FG models worked, I’ve never really gotten upset about the low odds here for the O’s. It made sense given that players like Dean Kreamer and Keegan Akin etc were still on the team following awful seasons. But the flipside of that is, that projections might not be catching those could have been the outlier years and maybe they aren’t actually that bad. I’m not convinced this pitching staff is going to stay as good as they are, but I don’t think they’ll regress nearly as much as zips/steamer thinks. And let’s be honest… those FIP numbers are in much disagreement with ERA.. for the starters two of them are almost spot on, 2 come close to canceling out and the other isn’t far off either.
I do think BR has been way over optimistic, and think projection models make a lot more sense; the human element and challenge for a GM is in evaluating if your players have really figured it out this year after an abysmal one or if its a fluke. I’m most curious about the 538 model in comparison though.
I kinda wish they had Mancini instead of Odor in the lineup, both because he’s a lot better and because he’s a heartwarming story. Lopez also is better than their worst reliever. Their statistical odds of a WS Win May be minuscule but it’s a great story to root for.
I continue to be amazed at how Roughned Odor not only keeps getting MLB jobs, but also significant playing time
At times it seems like the Orioles keep him on the roster just so they can play him whenever they’re in Toronto.
Jesus Aguilar is gonna get a shot. I would think he would be an offensive upgrade over Odor.
I think Odor will see some reduction in playing time between the expanded rosters and also the addition of a lot of lefties (Stowers, Vavra, Henderson) in the last month, but Brandon Hyde really loves him, and he’s going to keep playing more than you’d think. It didn’t help that when Vavra finally got a start a second he f’d up a double play, a skill Hyde repeatedly praises in Odor (I know defensive metrics think Odor is no good, but fwiw Hyde clearly thinks otherwise, he’s said it a bunch).
The only thing I can think of is if you look back at prior years, the O”s were playing in a different ballpark. The Mets had different ownership. I think if everything is stable in the franchise like Cards or Braves it works better. But if there’s a major franchise change like new ballpark or philosophy then that might lean towards this year’s performance.
It is looking at the players’ track record, not the teams’. The change in Mets ownership may affect how the team as a whole performs because they have acquired more talent, but Pete Alonso didn’t change as a player when Cohen bought the team. The projections are based on the talent on the roster not the history of “lol Mets” or “the Cardinal Way”
I think star players are more predictable but it’s the other guys, the one-year deal guys, where a team can make a diff and have them doing something they may have never done their career. if it were up to stocking up on the perceived ‘more talented player” the Rays would be in last and no powerhouse team would get burned on a free agent or trade.
I’ve thought about the success of the O’s rotation a bit because it really is odd that they’ve done so well. They’re a bottom 3rd rotation by K%, but a top 3rd rotation by BB%. They’re one of the most fly ball prone staffs in the game, but have one of the lowest HR rates. This is despite having below average contact suppression. I wonder if what we’re seeing (if it’s anything besides luck) is some strange intersection between a weak offensive environment, a newly pitcher friendly ballpark, and a team that allows a lot of balls in play without putting a lot of free runners on base. Whatever it is, I don’t blame the projections at all for not buying it, but it is fun to see them pitching so well since I’m not sure there is a single team in the MLB that would swap rotations with the O’s.
I think what you describe with their pitching results is likely part of the issue and may be tied , plus some “good luck” (BaseRuns suggests they should have a -22 RDiff and a 62-67 record). The lower walk rate may be because of catcher setup as mentioned here: https://blogs.fangraphs.com/introducing-the-new-orioles-bullpen/ and detailed here: https://www.baltimoresun.com/sports/orioles/bs-sp-orioles-reset-catchers-change-set-up-more-strikes-20220418-t2dllbzrrzcibe6mwbgftyr7ii-story.html
Anyway, much in accord with Ben’s suggestion that the truth is likely somewhere between the forward looking and season-to-date odds, I think that the core FG playoff odds likely catch some of the truth in the results by using BaseRuns while also naturally being unable to know when a strategy change, physical ballpark change, or player development gain has become an enduring gain versus just a fluke of baseball randomness until it persists long enough into the future for the projections to believe it. Since, as I think Dan Szymborski would say, a good projection system is designed to doubt what it sees at first blush in a “prove that wasn’t a fluke” type way.
Maybe I don’t understand enough about statistics and statistical modeling, but how can any projection system claim to be “accurate” when there’s no way of actually testing the data? What does it really mean to say that the Orioles have a 6% chance (or whatever) of making the playoffs?
Oh yeah, we can only test it retrospectively by looking at teams that have had similar odds to the O’s at similar times in the past. If the projections just stopped working suddenly, we’d never know until later. By “accurate”, I’m just trying to say that in the past our forecasts appear to have been well calibrated.
Trust me, as a PhD probabilist, no matter how much you knew about statistics and statistical modeling, that would be a question for debate. There are large philosophical camps with opposing views about what that really means, and if it’s even a well-formed question that has an answer. The philosophy of probability actually ends up impacting what sorts of statistics you might use and consider valid.
It’s absolutely bonkers how deep this gets. It’s partly why statistics is way more fun than math.
i always took things like this as an example of absurd accuracy or ludicrous precision. i know the difference between the two terms, these are just the two names i usually see when people refer to processes that produce unreasonably ‘firm’ estimates for things.
basically “Don’t get hung up on the number 6%, just accept it means they think their odds are low” is how i always view it.
or maybe Amazon really does know Gleyber Torres is exactly 28% likely to reach base with a certain count :p
Had a reviewer pull the “we can’t even talk about cause in any sort of model” card at a general interest journal last week. This was in an experiment! I’m assuming it’s some old Rudin student who’s gone off the deep end.
Real pleasant review /s
I don’t understand why so many fans get upset about this stuff, and take it so personally, as if these rankings and ratings and odds are somehow a reflection on them. It’s weird. Why do you care whether projection models think your team is likely to make the playoffs, or whether your team’s best player or top prospect get ranked #1, 3, or #47? Your team doesn’t get a prize for having the best playoff odds, and your favorite player doesn’t immediately get enshrined in Cooperstown if they’re ranked #1. If Fangraphs only gives your team a 4% chance of making the playoffs, that doesn’t mean they’re judging you to have a 4% chance of success in life. Besides, doesn’t it suck when everyone expects your team to coast to the playoffs, and they implode instead (I’m sorry, White Sox fans)? Isn’t it more fun when your team defies the odds and shocks the world, or when that 9th-rounder Dutchman-by-way-of-Florida who wasn’t on anyone’s radar turns out to be the best pitcher on the planet? (NO, VERLANDER ISN’T BETTER THAN DEGOAT YOU SHUT YOUR FILTHY MOUTH!!!)
Sports fans take things too personally…NEWS AT ELEVEN!
Well, I certainly take these thumbs-downs personally.
It turns out that “fan” is short for “fanatic”. Who would have guessed?
Which is why I like the Italian word, tifosi, i.e., infected with the (typhoid) virus and experiencing a fever from it.
Oh, I always thought it was short for “fan-damn-tastic.”
As long as I’m poking this particular bear, it’s objectively funny that fans refer to their favorite teams as “we”/“us,” despite not being affiliated with the team in any way, other than as a customer and well-wisher, really. Like, if your favorite band is, I dunno, White Lion, and you go to see them in concert, you don’t later say, “Dude, when we played ‘When the Children Cry,’ I was so into it I almost set the guy in front of me on fire with my lighter during the guitar solo! [NOTE: I assume there’s a guitar solo.] I can’t believe VH1 only ranked us #46 on their 50 Greatest Hair Bands of 1991 list. THERE SO BASED!!1!”
I’ve always thought this criticism was silly. Teams are tied to a city and bear that city’s name. Those ties matter too–attendance, crowd volume, and revenue all affect team performance. And of course in many instances residents are literally subsidizing the team. People identify themselves with their cities (New Yorkers, Chicagoans, etc.), and the team is part of the city. Bands are nothing like that.
Feel free to think the criticism is silly.
But I’m not talking sports with anyone who says “we” when referring to a team they passively watch 10-12 games a year of. These are often the types of people who say “Joey Gallo for Mike Trout seems fair, right”
And then they scream about how the team’s GM is a total idiot for not swinging that trade. “We coulda thrown in I Say A Kinda Falafel, too!* The Angels need a shortstop, and a good shawarma place! Fire Cashman!”
*Yes, of course I just wrote that just so I could shoehorn my stupid little pun in there.
Except for Boston, Chicago, Kansas, Nazareth, Berlin, Toronto, Calexico, Beirut, Chilliwack, L.A. Guns, New York Dolls, Nashville Pussy, Miami Sound Machine, Manhattan Transfer, Tokyo Police Club, Dresden Dolls, of Montréal, Alabama, Alabama Shakes, Blind Boys of Alabama, Texas, Ohio Players, Kentucky Headhunters, Georgia Satellites, Buffalo Springfield, Florida Georgia Line, America, Asia, and Europe.
Those are only the ones I can think of off the top of my head. Several of the aforementioned bands do not hail from the city, state or continent for which they are named. Regardless, even those bands are tied to the city, state or other geographic location that bears their names.
I actually just heard a guy do this – refer to the band as “we”
It was very confusing and weird.
OMG that is amazing. Was it a well-known band? (Was it White Lion? 🦁😆)
The following post is long because I have spent a significant amount of time thinking and working on a projection system.
For baseball analysis, I tend to limit myself to Fangraphs. I find it easy to see the information that I value such as game logs, and team statistics more than other sites. This is why I’m a member and have gifted memberships.
That said, I have found the Fangraphs projection system not helpful in asking the question I value: how does a team in a division compare with a team in another division. This comes from the fact that it is pretty clear AL Central and NL Central teams tend to be overrated by writers and fans, while teams in the AL East tend to be underrated. Hence I built my own Bayesian-probability-based projection system using data from scoring performance of teams during the season. (It is comparable, but more teams based than the 538 model.) Employing the metric that Ben Clemens used to analyze the success of projection systems, last year my team performance based system was more accurate than the Fangraphs system post All-Star-Break. For example, it calculated a division win by the Rays long before the Fangraphs model did. (Other seasons have also shown it to be more accurate, but not always by a significant amount.)
I believe the success of this approach is because it values team performance over performance of individuals. It produces evidence that some teams perform better than the sum of their parts, e.g., Rays, while other teams do not, e.g., Padres. Other examples of calculations from this model include it has consistently calculated a view the Yankees and Padres (and now the Mets) are, or have been, overrated, while producing a view that the O’s and Mariners are, or have been, underrated. (M’s not at this point.) Currently, this model calculates a probability of 1 out of 3 for the O’s making the playoffs.
Notice, even though the simulation is run a million times, I do not report a 3 digit value. Error bars, while shrinking during the season, are significant., usually overlapping any information in the 2nd significant figure. The reality of random error is rarely included in graphs, though I have noticed Dan Szmborski writes about it on occasions. A tip of the hat to Dan.
Is this system perfect? Nope. Does the Fangraphs approach do a better job earlier in the season? Yep, but not by much. Could it be improved? Sure, but given the vast amount of randomness in any baseball game, I suspect the improvements will be within the error bars that are inherent in a projection system. A task for when I retire.
Write it up, show your work, and make it a Community Post!
Fair enough. From it, I’ve made some predictions here. However, a full write up requires time, which lack at the moment. see comment about retire.
There are several points to my post. One being the the model of having a team being the sum of its assumed parts has problems. The other being the lack of including random error in data presented.
Looking at the rest of year projected ERAs for a lot of the pitchers- not sure if this is factor- but could there be something there with park factors and how frequently they are updated? The dramatic shift in run scoring environment at Camden could be resulting in lower pitcher ERAs for returning players than they had in the past. I thought I recall park factors were based on like a three year rolling average or something which would be distortive for this year. Of course, conversely that might mean the increase in runs scored per game projection for rest of year is also overstated. Not sure how/if that would impact these projections.
This gets to some big gripes I have about FG’s predictions in general. It just places way too much faith in the better team winning. Virtually every team loses a third of its games and wins a third of its games. But the game by game predictions are very confident before the games even start and the team projections for the season are also very confident. Banked wins are certain and FG takes the projected wins as certain too without any sort of regression back to .500.
I agree that season to date is an inferior method to using projections but the projections just don’t account for enough error. FG thinks the O’s are the fifth worst team in MLB going forward but it’s crazy to treat that as banked wins. The Orioles should have higher playoff odds not because they’re a great team but because they’re so close it wouldn’t take much of a run to put them in because they have actual banked wins.
I guess I don’t follow what you are saying. Just looking at today’s game probabilities, the only one over 70% is Braves over Rockies. The 72 win Mariners are 49-51 underdogs against the 50 win Tigers.
The O’s currently have 6.4% playoff odds which reflects the fact that they are close and it wouldn’t take much of a run to put them in. Your issue appears to be with the .412 projected ROS winning % for the O’s, not the low playoff odds, because if they are actually that terrible team, it is very unlikely that they leapfrog the Blue Jays or Mariners who have the 2nd and 3rd best ROS projection in the AL and a 2-3.5 game lead on the Orioles. If they were being treated as certainties, then the O’s would have 0% playoff odds. There is about a 6% chance that a bad team goes on a run and one of the best teams falls off. That doesn’t sound unreasonable.
To put it another way, how likely do you think it is that the D’backs outperform the Blue Jays by 2 games the rest of the way?
There is an issue here which may be what sadtrombone was getting at: if the projection systems have the Orioles as a true-talent .412 team, the playoff odds assume that that’s correct and use that in all the runs of the simulation. Now, a .412 talent team might win 30% of its games, or 55%, or something else in any give simulation run, but the system doesn’t vary its estimate of the talent level at all. That tends to cause it to underestimate the odds for teams like the Orioles: if their talent level is worse than .412 then they can’t lose more than 6% playoff odds, but if they’re better then there’s a lot more scope to increase the odds.
I know Dan’s ZiPS playoff odds simlator actually builds a distribution of possible talent levels for a team, and I believe the Baseball Prospectus system does something similar. That might explain why BP, at least, has the Orioles with a better chance of making the playoffs than FG does.
Thanks for the explanation. Looking at Pecota, they don’t publish their ROS winning %, but they appear to basically project every team extremely close to .500 the rest of the way (Mets 16-14, Dodgers 17-15, Orioles 16-16, Pirates 15-17, Tigers 15-16) so their playoff odds are very similar to the FanGraphs coin flip odds.
I guess I just don’t see the value in varying the estimate of talent level. There is already variance in the performance around the true talent level estimates as you describe, but if you start moving the starting talent level baseline that the results deviate around, then it is little more than guesswork. ZiPS/Steamer thinks the Orioles are a .412 talent team the rest of the way, moving off that value as the most likely outcome is just disregarding those projection systems’ analysis. If someone just wanted to use a different projection system as a baseline or believes that the standard deviation of the results around the baseline should be greater, that I could understand, but building in a “what if our numbers are just wrong” with no basis for it just seems like a way to get less accurate results.
The ROS winning % in the ZiPS/Steamer projections are between the Rockies at .382 and the Mets at .618, essentially saying that every team will be between 100 wins and 100 losses in true talent, which is generally what we have seen historically. That seems like a reasonable place to start and work from there. Regressing everything towards .500 would just be saying that the actual distribution of results should be in a narrower band, which doesn’t align with anything we’ve seen historically.
The idea of varying the estimate of the true talent levels is to produce more realistic probabilities. On average we’d still use .412, so we’re not ignoring the projection systems, but each individual run of the simulation would use a slightly different value. That accounts for the fact that we know that the projection systems aren’t perfectly accurate, that playing time estimates can be off due to form or injuries, and that (at least pre-deadline) teams might acquire or lose players in trades.
To answer your question I’d say the chances are much higher than 6%. Probably 3 times that amount. Bad teams win games against good teams (and average teams, and other bad teams) all the time and good teams lose them. There are strong forces pulling teams towards 500 ball.
The forces pulling teams towards .500 are already baked in, that is why no team is projected to be over .700 or under .300 the rest of the way. Good teams tend to win more games than bad teams, that is why there is a very small, but non-zero chance of a bad team overtaking better teams.
Of the 15 teams with negative BaseRuns run differentials (an extremely rudimentary proxy for team talent) two have winning records over the last 30 games. Of the 15 teams with positive BaseRuns run differentials, two have losing records over the last 30 games. That would indicate that there is roughly a 13% chance that a good team can have a losing month or a bad team can have a winning month. So yeah, bad teams can win and good teams can lose, but over larger samples it’s unlikely.
If it comes down to that final head to head series between the Orioles and Blue Jays with the winner of the series going to the playoffs, then yeah it will be close to a 50/50 proposition because team quality is less important over small sample sizes, but over the larger sample size of a month it is significantly more likely that the Blue Jays have a better record than the Orioles if the Blue Jays are a true talent good team and the Orioles are a true talent bad team.
@ben, thanks for this analysis, seems pretty reasonable. One question I have is whether the models tend to miss more on teams with a lot of young players. Is there a good way to measure this or does it slice the data too finely?
I’m wondering if that’s why the season-to-date model does better in some cases. Some earlier-than-expected breakout teams would fall in this category (eg recent Braves) as opposed to journeymen type teams. Although the O’s do have a couple journeymen types in their rotation (Lyles, Watkins, Voth)
The ZIPS estimate is, from what I recall, somewhat higher on the Orioles than the FG estimate, despite broadly speaking similar methodologies. Only somewhat, though.
Fellow O’s fan chiming in to say that I appreciate the explanation (most, but not all of which, I knew) and the ultimate conclusion seems right too–the truth likely lies somewhere in between, but probably closer to the FG range than the others’.
One thought it raises is that the projection systems will probably get better as they learn to incorporate more of the statcast data. I’m thinking about the O’s pitchers especially. I can absolutely buy that they’re overperforming, but I don’t buy that they’ll regress to their career norms. There are a bunch of factors at play there like the new park dimensions, Adley’s framing, and also Adley’s gamecalling (which is particularly hard to capture, but I became more open minded to its importance after Kevin Goldstein constantly emphasized it).
But it also just seems odd to me that we’d treat pitchers as similar to their career norms when they’re…just obviously not? Take Dean Kremer for instance. He’s a very different pitcher. His four-seam usage has gone from 55% to 39%, with most of that difference going to his cutter and a sinker (albeit just 4%) that is a new pitch. His curveball usage is about the same, but now it’s got an extra 300 rpms(!). Similar story for guys like Watkins and Voth too–different usage of old pitchers, altogether new pitches, and some big spin rate differences too. Again, I don’t think these guys are suddenly legit #3 starters who won’t see any regression, but I also view with some skepticism the idea that we should lean heavily on their past performance when that performance came from very different pitchers.
I realize this reflects limits of the current models, so I’m just kind of thinking out loud here that it’ll be interesting to see those projection systems continue to improve over the year as they take into account updates in players’ (especially pitchers) underlying data.
I agree with this.
I think the granular FG approach is the right one, but the projection system might be over-regressing to past seasons, particularly in light of the fact that we should be able to use statcast that to better gauge the durability of changes in performance. Are the in-season ZIPS and Steamer as robust as the offseason models?
It’s also worth pointing out that the system “hates” the Orioles but it also “loves” the Blue Jays, which is also suppressing the Orioles playoff odds because they play 10 out of 33 games against each other and that’s probably the Orioles best chance to catch someone.
How much volatility is added because the Orioles changed park dimensions this year? Are you using single season park effects? This is going to generally favor pitching stats given equivalent skill.
So the projected regression is largely based of off pre-2022 historical pitching performance . Problem is that there is very little historical data for most of the Os pitchers. Of the 5 current starters, Lyles is the only one with more than 22 career starts coming into 2022. Career starts for the other 4 starters: Voth 22, Kremer 17, Watkins 10, Bradish 0. As for the bullpen, Bautista, Krehbiel, and Baker are all rookies. Cionel Perez had 51 career innings over 4 seasons before 2022 and Keegan Akin is in a completely new role. Dillon Tate’s ERA lines up very solidly with his FIP, xERA, and SIERA on the year. Also Bryan Baker, Keegan Akin, and Dillon Tate are solidly underperforming there FIP, xFIP and SIERA over the last 2 months. Their ERA is just as likely to go down in the next month. Finally Bautista has improved massively since April/May. FB velo is up 3 mph, xwOBA is way down, Ks are up, BBs are down. He’s got a 1.57 xFIP, and 1.21 SIERA since mid June.
Well clearly Fangraphs just has it in for the Orioles! (Sarcasm)
As a Jays fan, I hope FanGraphs is correct but with Baltimore soon to be just 1 1/2 games back of the Jays for the 3rd Wildcard I’m not so confident.
Baltimore is winning the season series and there’s lots of games between the 2 remaining. Sure, the Jays should be favoured but it’s closer to a toss-up than the FanGraphs probabilities show.
The difficulty in these models, like many models, is always the step from correlation to causation. When an entire staff performs well above expectations, it could be a random event with a nearly certain regression or there could be an underlying cause that prior year’s performance is not able to pick up. More often than not the error is assuming causality when the real answer is random variation, but I think this Orioles staff may be an exception.
While my modeling experience is not primarily in the baseball arena, I don’t for example think that anyone could actually have watched Felix Bautista pitch this year and expect him to have an ERA of nearly 4 down the stretch. I personally think there is just something different about pitchers like Voth, Perez and Bautista than what their past record would indicate, but I also acknowledge the near impossibility of a past performance-based model to pick that up.
Interesting stuff, Ben! One question: does the rest-of-season projection take into account quality of opponent? (Don’t know if you follow football, but DVOA from Football Outsiders sparked this thought)
As the ink was drying on this article Brandish went out and pitched another seven scoreless. Bet the model wasn’t expecting that. I’m not a Luddite and I’m aware of the value of statistics, but I would agree with those who are saying that the O’s overall performance this year is showing that the models missed something. I think they are still missing.
Models will always miss some things, but the quandary is that in most cases the attempt to capture and adjust for that “something” missing will generally make the model worse on an overall basis by overfitting and modeling for noise and random variation. The pursuit to understand what might be missing is a valuable exercise but identifying those inflection points when something really does change is really hard.
A good modeler should always have the humility to recognize that there will always be a gap between a model and reality but closing that gap when the model is good overall is a treacherous path.
I think Fangraphs projection undervalues a lights out bullpen. It hated the mariners last year who had a mediocre offense and a lights out bullpen. It mistakes smart bullpen usage for luck and doesn’t understand that a mediocre offense doesn’t blow out teams as much as it can get blown out hence the run differential.
I think so, too.
And I think because they use the Pythagorean Record estimation method which measures total runs scored by total runs allowed. A lights out bullpen allows you to win close games – rack up a lot of wins with a small difference in runs scored vs runs allowed.
Yes, another point to note is that good bullpen management will often result in losses by a higher differential, for example letting the starter “wear it” after giving up 5-6 runs in the first, only to give up four more in the second; or letting the “white flag” reliever pitch multiple innings even though he is letting things get worse.
Look at the Rays bullpen means this year and you’ll see this in effect.
However, watch the Rays handoff a one run lead to the ‘pen in the fifth and they typically surrender only one or two hits, no walks and no runs.
I beleive that the Fan Graphs proejction method underrestimates the intangibles represented by (1) Adley calmingthe young pitching staff and (2) improved coaching. It also does not consider remaining strength-of schedule and home-vs-away distribution, and does not count health/rest vs. overwork. These refinements probably improve the Orioles’ odds to something closer to the other sites.
how could a model account for “calming the pitching staff” and “improved coaching”? A model that accounts for coaching effects would be interesting (I’d be especially interested to see it in hockey or soccer moreso than baseball, personally), but I’m not sure how you could possibly differentiate a player’s expected/actual production from what they were “coached” to do.
It seems far more reasonable to leave the model as-is and do a qualitative/narrative analysis of why the model could be underrating this particular team, which might include those factors.
The yankees were like 99+ percent to win the division a few months ago. Lets see how far that falls
I think the most bizarre Fangraphs playoff prediction recently was not the low probability for the Os, but the ridiculously high probability that FG kept giving to the Red Sox.
By early August the Orioles had shown that their 10-game winning streak in July was not a fluke, and the Red Sox were done, finished.
Yet to pick a random date, after the games of August 9 Fangraphs estimated that the Red Sox had a 14.3% chance of making the playoffs, and the O’s 6.6%, despite the O’s having a 4 1/2 game lead over the Sox!
That’s a low probability for both teams of course, but who has been thinking that the Red Sox have had a better playoff probability than the Os at any point since July? No one AFAICT.
Even 538 had the Os over the Red Sox, 17% to 12%.
Pecota had Os at 31.9%, Sox at 13.9%.
Baseball Ref was, as the article points out, probably over-optimistic on the Os, 50.0% to 2.9%.
PlayoffStatus.com like BRef seems to rely very heavily on current season records, and had the Os 40% to 10%.
It’s fine that Fangraphs’ playoff probabilities are based on projections of players’ performances, but those projections should be updated as the season progresses. From the article it sounds like Fangraphs just keeps using the same pre-season projections rather than using Bayesian updating or some other updating technique. Or if they are updating them, they’re putting too low a weight on the current seasons’ stats.
Most of these systems of calculating playoff probabilities should do more regressing to the mean, but that’s a whole other discussion.