Opposition Quality Seems Hardly a Factor
Remember when you thought baseball statistics were easy? Now, analytically, a statistic is hardly worth anything if it’s left unadjusted. You know all the adjustments that go into numbers. Adjustments for ballpark environment. Adjustments for era. Sometimes adjustments for league. As far as WAR is concerned, there are adjustments for position. There’s one adjustment we still don’t make, though: that’s adjustment for quality of opposition. In theory, if there were a pitcher who only ever faced the best teams, and a pitcher who only ever faced the worst teams, that wouldn’t be accounted for. That’s something you’d have to figure out yourself.
Related to this, James Shields has obvious selling points: first and foremost, he’s been good. He’s been durable, and he’s experienced, and he’s pitched in the playoffs, and everything. Then there’s one other thing I don’t think has gotten much attention: Shields has, relative to the average, faced a fairly tough slate of opponents. In the past, I’ve manually calculated average opponent wRC+. Baseball Prospectus has its own version, oppRPA+, and the big advantage of oppRPA+ is it’s already been calculated for me. The average, as usual, is 100. Last year, Shields’ opponents came in at 105. The year before that, 105. The year before that, 105.
Seems like this should be a good thing for Shields’ market; seems like, if you adjust for this, Shields’ numbers would get a boost. But how much does this matter, really? Coming up soon, a first attempt at an answer.
The research itself wasn’t difficult. I set a season minimum of 100 innings. I looked for pitchers who threw 100 innings in consecutive years, and I tracked their season-to-season oppRPA+. Simple math yielded the change in oppRPA+, and I sorted by the change, in descending order. Then I linked the pitchers to various FanGraphs stats: ERA-, FIP-, and xFIP-. In the end, I had the full slate of numbers for Year 1, and the full slate of numbers for Year 2. I had a sample exceeding 1,100, so I split that into five groups. This table shows you the performances, on average. Group 1 includes the pitchers who, in Year 2, faced considerably tougher opponents. Group 5 includes the pitchers who, in Year 2, faced considerably weaker opponents. The oppRPA+ for Group 1, on average, increased six points. The opposite occurred for Group 5.
| Group | Y1 RPA+ | Y2 RPA+ | Change | Y1 ERA- | Y1 FIP- | Y1 xFIP- | Y2 ERA- | Y2 FIP- | Y2 xFIP- |
|---|---|---|---|---|---|---|---|---|---|
| Group 1 | 97 | 103 | 6 | 95 | 96 | 96 | 100 | 99 | 98 |
| Group 2 | 99 | 101 | 2 | 95 | 95 | 96 | 98 | 98 | 98 |
| Group 3 | 100 | 100 | 0 | 95 | 94 | 95 | 96 | 97 | 97 |
| Group 4 | 101 | 99 | -2 | 96 | 98 | 98 | 99 | 99 | 99 |
| Group 5 | 103 | 97 | -6 | 95 | 97 | 97 | 97 | 97 | 97 |
To make that easier to understand:
| Group | Y1 RPA+ | Y2 RPA+ | Change | Change, ERA- | Change, FIP- | Change, xFIP- |
|---|---|---|---|---|---|---|
| Group 1 | 97 | 103 | 6 | 5 | 3 | 3 |
| Group 2 | 99 | 101 | 2 | 3 | 3 | 2 |
| Group 3 | 100 | 100 | 0 | 2 | 3 | 2 |
| Group 4 | 101 | 99 | -2 | 3 | 2 | 1 |
| Group 5 | 103 | 97 | -6 | 2 | 0 | 0 |
Absolutely, you see different numbers in the columns. They’re just not extremely different numbers. All the groups saw a little Year 2 ERA-related regression to the mean. Those pitchers who faced tougher opponents saw their average ERA- go up almost five points. Those pitchers who faced weaker opponents saw their average ERA- go up just over two points. You observe similar changes with FIP- and xFIP-. Differences do exist. Differences ought to exist — it’s not like opposition quality is completely irrelevant. But at least based on this, it’s a decidedly small factor. Maybe a run. Maybe two or three. And that’s toward the extremes. Most of the time, this isn’t worth thinking about for even a couple of minutes.
In case you’re curious, the most difficult season-to-season adjustment: 2012-13 Joe Saunders. In 2012, Saunders faced maybe the weakest group of opposing hitters in baseball. The next year saw an increase in oppRPA+ of 15 points. Sure enough, Saunders was far less effective the second year. On the other hand, Bud Norris saw his opponents also get a lot better between 2012-13, and he lowered his ERA and FIP.
And the least difficult season-to-season adjustment: 2010-11 Madison Bumgarner. Somehow, in 2010, Bumgarner led all of baseball in oppRPA+. The next year, he finished second from the bottom, a massive drop of 16 points. What happened? His ERA- actually got worse, but his FIP- and xFIP- got significantly better. Jordan Lyles also saw an improvement between 2013-14, as his oppRPA+ fell by 12 points.
Bigger differences might emerge as you isolate the very biggest season-to-season changes in oppRPA+. We can set a threshold of a change of at least 10 points. There are 14 examples of pitchers who saw oppRPA+ hikes of at least 10 points. In Year 2, their average ERA- got worse by 14 points. Their average FIP-, meanwhile, got worse by 10 points, and xFIP- by 4 points.
And there are just nine examples of pitchers who saw oppRPA+ drops of at least 10 points. In Year 2, their average ERA- got better by 12 points. Their average FIP-, meanwhile, got better by one point, while xFIP- got worse by 3 points.
So you wonder how much that means. That’s a huge ERA implication, but it’s more mild by FIP, and it’s virtually non-existent by xFIP. It’s also a small sample of just 23 total pitcher season-pairs, and such changes are uncommon, with an average of fewer than two per year. Of course, it only makes sense that the biggest changes in opposition quality would lead to the biggest changes in pitcher performance. But we don’t have enough data yet to really be able to say how significant this is. Last season’s lowest oppRPA+ figures belonged to Brad Hand, Clayton Kershaw, and Gerrit Cole. Among available pitchers, Cole Hamels was close to the bottom. So there’s that small amount of risk in an AL team dealing for Hamels. But a team like San Diego would have relatively little to worry about.
I suppose all this probably just confirms what you would’ve already thought: opponent quality does make a difference, but almost all of the time it makes a small difference, so leaving it unadjusted for doesn’t throw anything way out of whack. On a case-by-case basis, it’s worth checking if a given player is in for an unusually extreme adjustment, but those cases are rare. I should also acknowledge there are shortcomings with the method. This was already selective for pitchers who threw 100 innings in consecutive years, so guys who bombed in Year 2 are excluded. Also, this gives no consideration to platoon issues, and it doesn’t have the hitter numbers adjusted for pitcher quality. It turns out adjusting for competition is surprisingly hard. Thankfully, it’s also usually unnecessary. That’s the convenient bit.
Jeff made Lookout Landing a thing, but he does not still write there about the Mariners. He does write here, sometimes about the Mariners, but usually not.
interestingly, the burgeoning advanced stats movement in hockey has found a similarly negligible effect of quality of competition.
Interesting about hockey – can you site any examples?
on my phone, but here are some references:
http://nhlnumbers.com/2012/7/23/the-importance-of-quality-of-competition
http://www.arcticicehockey.com/2011/7/7/2264529/does-qualcomp-matter
and more recently:
http://hockey-graphs.com/2014/01/06/gauging-the-relevance-of-quality-of-competition-on-a-players-stats-toronto-maple-leafs-d-edition/
in general the consensus in hovkey is that ehile qualcomp is significant in theory, in reality the actual differences in qualcomp are too small to matter all that much.
On your phone you say? 🙂
(I kid, appreciate the links. Good info.)
I find it interesting that the whole sample got worse overall between the two years. You called it regression, but I don’t see why the first sample would have had a bias towards unexpectedly high ERAs, which is the kind of sample from which you would expect regression the following year. If anything, it’s probably age-related decline, as at any given time, the majority of qualified pitchers are probably at an age where they should be declining.
I think that is exactly right. It is decline. One additional point on this is that if there is a really bad year, it is likely to be the second because the really bad year is likely to result in the next year not happening. This could be because of injury or just declining performance.
I’d be curious how the results would change if you only looked at pitchers that also had 100 innings pitched in year 3.
Not only potential expected age-related decline among the qualified.
There may also be some element of ‘absolute decline in MLB batter quality’ vs ‘relative batter quality in MLB forced to stay constant cuz of the calculation’ – depending on the years included in the sample.
Yeah, aging curves for pitchers seem to basically be a constant downward slope, so it makes sense to expect a slight decline from Y1 to Y2 for basically any sample.
This raises an interesting question to me. We are pretty sure that in even a long, 162-game season, there is plenty of random variation affecting team record. I wonder what any given baseball season would look like if the season lasted 1,000,000 games. Usually, the best team in baseball in a year has 97-103 wins, but I bet their true-talent winning percentage is lower than that. I bet the best teams in a given season are usually true-talent 92-wins teams.
My point here is that quality of competition might be a lower factor than it seems like it should be because quality of competition doesn’t actually vary as much as team records vary at the end of a season.
Team talent is roughly one SD = .060 wins per game.
Observed win percentage is roughly one SD = .072 wins per game.
How is team talent being measured? Projected WAR prior to season? Combined WAR after the season?
If a season was 1,000,000 games, it would take over 2739 years to complete, if they played without any days off. Even daily doubleheaders would take nearly as long to complete as years have passed since Muhammad did his writing. Over the course of such time, players would come and go, injuries would arise and age would take its hold. Managers, GMs, even owners would turn over. In a way, we’re in the middle of a 1,000,000-game sample; indeed some 5% of the way done. I’d imagine the true talent of many teams over this sample would land around .500.
Or just have a time machine and observe a season 3000 times. It would be like simulating a season in a video game until you get the result you wanted.
It is perhaps worth noting that a lot of the big adjustments in opposition strength you found were when guys switched leagues (i.e., Blanton and Norris in ’12-’13 & Lyles in ’13-’14).
I’m not sure, but i think he said Saunders, not Blanton.
Wait here a second, let me check.
Yep. Saunders.
good catch.
Saunders kind of switched leagues – he had 21 starts with the DBacks in 2012, then was traded to the Orioles, then in 2013 was with the Mariners all year. I guess he still fits…
Seems like this would open a can of worms. Would you then have to adjust all the hitters based on the pitchers they faced. Then how could you do that when you haven’t adjusted the pitchers for the hitters faced.
Then how about adjusting for handedness of opponents faced. RHP will do better facing more RH hitting lineups and vice versa.
Just seems like a lot of adjustments to make on the margins.
That’s a typical analysis problem. There are equations for it. Basically, you need to put all of their data in a matrix and adjust everyone at the same time using matrix algebra. The downside of such analyses is that you basically need to update them for the whole league as soon as anything changes.
But I agree: it’s a factor that is minor enough that it only makes sense if you’re trying to do a big predictive model (e.g., something with a LOT of features, where their little contributions will add up to new insights). On its own, it’s a big “Meh” compared to the standard linear models.
“I looked for pitchers who threw 100 innings in consecutive years”
So you intentionally selected a sample that would be exclusively starting pitchers. This sample is exactly the type of group that would minimize the difference in opponent quality because the best strategy to counter SP is stacking the lineup with hitters with favorable platoon splits. Then you measured differences without considering platoon splits. Your results are interesting, but I don’t know that they have much value due to the sample and evaluation methods that were used.
Any robust research of opponent quality should include platoon splits. The groups most affected by opponent quality might be the hardest to study due to smaller samples (RP and PH).
“It turns out adjusting for competition is surprisingly hard. Thankfully, it’s also usually unnecessary.”
This topic interests me greatly. Thank you for placing some focus on it. I agree that opponent quality is a very difficult and complex task, but I don’t think you have shown a reason that it is unnecessary (or a trustworthy range of its impact).
I think I have an idea how platoon splits could be accounted for, but it would be a bitch to do manually, so we’ll see what I can do in terms of getting Appelman’s attention. Longer-term project but I feel like we should have something available on this site.
Good luck with Appelman.
Also, it would be amazing to have opponent quality stat(s) on the site.
Worse opponents?? The Rockies were not good last year, but they were much better with Tulo in the lineup, it’s all kind of relative.
Its not who you play but when you play them. Facing a good hitting team in a slump is easier than facing a bad hitting team that’s hot. Unless you can adjust for hotness, the quality of opp stat may have too much noise
hotness is noise.
FWIW, I think Baseball-Reference pitcher WAR does adjust for quality of competition.
http://www.baseball-reference.com/about/war_explained_pitch.shtml
This is a much bigger factor prior to the expansion of 1969 and subsequent expansions. Whitey Ford never faced the Yankees who were the best or second best offense in the AL for nearly his entire career. But consider someone like Alex Kellner who pitched 291 games for the A’s from 1948 to 1958. He never got to face his own poor offenses. There were so few teams that Kellner faced NYY 58 times, 12 times more than any other opponent. He gave up a .795 OPS to them, second worst among his opponents, and a 4.70 ERA. Ford faced the A’s 59 times with a .626 OPS against, third lowest among his opponents and a 2.54 ERA. If the tables were turned, Ford would look a fair bit worse and Kellner somewhat better.
We could play this out with guys like Gomez, Miner Brown, and Matty as well. Perhaps Feller and Lemon or Pennock?
With more teams, interleague play doubling the number of teams faced, but (roughly) the same number of games, pitchers face a more diverse slate of opponents today since they have 29 possible opponents not just seven or nine. Not facing their own team isn’t as influential on their stats. But over a career it’s possible this could add up. Someone like Pettitte, Glavine, or Smoltz might show a career long trend that might not be present for, say, King Felix with his anemic M’s offenses.
It may also be true that, like Casey with Ford, before strict rotations emerged in the 1960s and 1970s a manager might manipulate a pitcher’s usage to start his ace against the better teams and his worst against the lesser ones. That might amplify the affect of smaller leagues as well.
And Whitey Ford was pitching ion the old Yankee stadium, the one with the 463ft to CF and the LCF power alley at 430. I remember once a proposed trade between the Bosox and Yankees, DiMaggio for Ted Williams. Bosox backed out in a hurry when they realized how short the RF porch was in NY, and Ted was a dead pull hitter.
Second order effect: the best pitchers should have a lower oppRPA+ (or whatever) than the worst pitchers just because good pitchers, by definition, make hitters worse and vice versa. So, if a pitcher lowers his oppRPA+ compared to the previous season, he may have gotten better rather than his average opponent gotten worse. If the pitcher gets better (or luckier) in year 2, but faces exactly the same batters as in year one (or batters of exactly the same skill), his oppRPA+ will go down! This leads me to wonder:
When Baseball Prospectus calculates oppRPA+ for pitcher A, do they calculate each batters’ stats ignoring ABs against pitcher A? This word alleviate this admittedly minor problem.
I would sure hope they adjusted for it. Otherwise they’re just lazy.
I’ve looked into the issue regarding adjusting batting statistics based on the quality of the opposing pitchers, and the magnitude of the adjustments is even smaller.
One adjustment that is not made is adjusting bequeathed runners to the league average % that are scored. This made a material difference for example in comparing 2014 Felix Hernandez to 2014 Corey Kluber because the Seattle bullpen let a much higher % of Hernandez’ runners score than was true for the Cleveland bullpen with Kluber’s runners. This doesn’t affect fWAR but will affect the runs allowed version of fWAR and affects b-refWAR.
I find it very telling that you continually point out Cole Hamel’s quality of opposition but never point out that he’s surrounded by NL East studs such as Strasburg, Fernandez, etc.
Telling of what, the fact that Hamels’ value is being discussed since he’s an obvious trade target?
Hamels has been a very good MLB pitcher, pitching in that bandbox stadium in Philly that made Ryan Howard look like Lou Gehrig. Id like to see Hamels in Petco or Safeco. Hed likely be another King Felix if he pitched half his games in Safeco.
Ryan Howard has hit the same amount of Homeruns at home as he has away. That’s like revisionist history. Howard sucks now but offensively he was second only to Pujols in his prime.
Seems like its a useful descriptor…kind of like park factors only less influential. Probably not much predictive value as it will get drowned out by randomness, change in true talent, etc.
Interesting to have these numbers. This issue is often phrased as: “But how would he do in the AL East?” It’s an aspect of Shields’s record has in fact been mentioned a good deal, at least in AL East cities; and it’s a doubt about Hamels.
I enjoyed this article and agree that quality of opponent often washes out, particularly over the course of a season. But the easiest way to look at this is to compare cFIP (which adjusts for park and stadium) to FIP- (which somewhat adjusts for park and not at all for competition).
cFIP is maintained at Baseball Prospectus and FIP- is at Fangraphs.
cFIP is more reliable overall, and notable differences in cFIP and FIP- tell you that the pitcher is facing challenges beyond his control, quite possibly including the quality of his competition.