I Voted for Justin Verlander
I submitted my American League Cy Young Award ballot at the very beginning of October. The results were just released yesterday, almost smack in the middle of November. A funny thing happens between the beginning of October and the middle of November: A lot of time passes, time that includes the entirety of the MLB playoffs. As I focused on other events, I mostly forgot about my selections. I was reminded yesterday that my own ballot read:
I was one of 13 voters to put Verlander in first. The other 17 voters, though, put Snell in first, and as such, Snell won, and Verlander was, once again, the runner-up. Clearly, it was a close race, and I think it should have been a close race. I don’t think that Verlander got robbed, and I don’t think that Snell is an undeserving winner or anything. But one of the perks of being an award voter is that voting grants you automatic editorial content. So on the off chance you care about my own thought process, allow me to quickly explain why Verlander was my first-place pick.
For so many observers, this race was going to revolve around two statistics. Verlander threw 214 innings, or 33.1 more innings than Snell. Yet Snell finished with an impressive 1.89 ERA, more than a half-run better than Verlander. Both Snell and Verlander were charged with just three unearned runs apiece, so if you want to explain why Snell wound up the winner, you could say this: In those 33.1 extra innings, Verlander allowed 22 extra runs. That works out to an average of almost six runs per nine innings. Should that really be a point in Verlander’s favor? That’s replacement-level pitching. That might even be worse than replacement-level pitching.
There’s no question that Snell finished with excellent results. He timed them well, too — his ERA in the second half was 1.17. Verlander’s was 2.95. As the award spotlight got brighter, Snell turned in quality start after quality start. But the first half matters, too. The first half is quite a bit longer than the second half. And I do give Verlander credit for simply pitching more often. Compared to Snell, Verlander pitched to 133 more batters. That’s good for the bullpen, and that’s good for roster management. Pitchers who eat innings can have a positive cascading effect. But I don’t want to go into every single detail. I doubt you want me to go into every single detail, either. I prefer to just present the core of my argument. To what extent was Snell responsible for his results?
This table provides a simple overview:
| Pitcher | GS | IP | ERA- | FIP- | xFIP- |
|---|---|---|---|---|---|
| Justin Verlander | 34 | 214.0 | 62 | 67 | 72 |
| Blake Snell | 31 | 180.2 | 46 | 72 | 75 |
Verlander started more games than Snell. He threw more innings than Snell. He finished with a better park-adjusted FIP than Snell, and he finished with a better park-adjusted xFIP than Snell. If the park-adjusted ERAs matched up with the peripherals, Verlander presumably would’ve won the Cy Young running away. But Snell’s lead in the ERA- column is enormous. Too enormous to ignore, it turns out.
How did Snell end up with such a sparkling ERA? He wound up with a BABIP of .241. But it goes even beyond that. For both pitchers in question, check out their wOBA splits:
| Pitcher | Overall | Bases Empty | Runner(s) On | RISP |
|---|---|---|---|---|
| Justin Verlander | 0.259 | 0.273 | 0.230 | 0.246 |
| Blake Snell | 0.246 | 0.270 | 0.203 | 0.175 |
Essentially the same with the bases empty. But with the bases not empty, Snell got more outs. And with runners in scoring position, Snell blew Verlander away. That’s a margin, in the last column, of 71 points. Those are higher-risk plate appearances, so any pitcher would look better if he’s pitching his best with runners on second and/or third.
You know there’s a “but,” though. Snell would deserve credit, indeed, if he pitched better with runners on base. But with runners on, Snell’s strikeout rate dropped, and his walk rate got higher. Verlander struck out more batters when the bases weren’t empty. And for Snell, with the bases empty, he allowed a BABIP of .292. With runners on, that dropped to .161, and with runners in scoring position, it dropped to .118. If we’re dealing with anything here, we’re dealing with a question of contact quality. And that question always sends me to Baseball Savant, so I can check out what Statcast has to say.
The previous table showed you their wOBA splits. In this table, I want to show you their xwOBA splits. I know that xwOBA is still considered kind of experimental, and I wouldn’t want to base awards voting on xwOBA alone, but why shouldn’t we consider the actual batted balls these pitchers allowed? We’re always trying to strip away the effects of defense and luck. Let’s do some stripping:
| Pitcher | Overall | Bases Empty | Runner(s) On | RISP |
|---|---|---|---|---|
| Justin Verlander | 0.237 | 0.239 | 0.232 | 0.227 |
| Blake Snell | 0.272 | 0.273 | 0.269 | 0.248 |
Overall, Verlander had the lower xwOBA. With the bases empty, Verlander had the lower xwOBA. With the bases not empty, Verlander had the lower xwOBA. With runners in scoring position, Verlander had the lower xwOBA. And these margins aren’t even all that small. Again, I know xwOBA isn’t perfect, and I know pitchers sometimes try to pitch to their ballparks or defenses or whatnot, but I went looking for a reason to believe Snell really was blowing Verlander away under more pressure-packed situations. By looking at xwOBA, I could find zero evidence. xwOBA indicates that Verlander was, if anything, the better pitcher than Snell, and he threw a good deal more innings than Snell. I don’t worship at the altar of 34 starts, not in this era of baseball, but Verlander was durable and terrific. With Snell, I couldn’t get myself over the top.
There were 121 pitchers who faced at least 500 batters. Snell had the fourth-most positive difference between actual wOBA and expected wOBA. There were 83 pitchers who faced at least 250 batters with runners on. Snell had by far the most positive difference between actual wOBA and expected wOBA. And there were 137 pitchers who faced at least 100 batters with runners in scoring position. Snell had the second-most positive difference between actual wOBA and expected wOBA. It’s very possible, if not probable, that’s not giving Snell enough credit. Maybe he was in some way able to pitch to his defenders. The name of the game, after all, is run prevention. It’s also very possible xwOBA underrates the extent to which Snell suppressed the best contact. I’m not upset that Snell wound up the winner. I just think you need convincing evidence if you’re going to vote for the guy who pitched less often. The evidence, as I see it, doesn’t convince me. Snell was outstanding, but his ERA had some help. Help that Verlander didn’t get.
Realistically, this was always going to be a two-pitcher race, so you’ll be less interested in how I filled out the rest of my ballot after the top two spots. I don’t want to go too deep on those choices, either, because ultimately they don’t matter very much. Gerrit Cole was incredible. Had I weighted xwOBA a little stronger, relative to actual wOBA, I could’ve made an argument for moving Cole in front of Snell, too. But I didn’t want to ignore Snell’s actual results entirely. As good as Corey Kluber was, I docked him for pitching a bit worse with runners on, and with runners in scoring position. I included a closer in Blake Treinen, and the league’s two outstanding relievers were Treinen and Edwin Diaz, but Treinen faced more batters. Also, Diaz never went more than 1.1 innings, while Treinen recorded at least five outs on 14 different occasions. On a per-batter basis, Diaz was probably better by a hair, but Treinen had more on his shoulders, especially over the first few months. Diaz probably would’ve wound up sixth or seventh on a deeper ballot.
At last, while it would’ve been nice to include Chris Sale — he was amazing — he threw 158 innings over 27 starts. He threw all of 17 innings in August and September combined. It’s one thing if a pitcher is expected to throw something like 158 innings, because of a preexisting plan. But the Red Sox weren’t prepared for Sale’s absence. He put the team in a bind (even though it was running away with the division). Treinen might’ve been a closer, but at least the A’s always had him available. Sale was basically missing for almost a third of the season. I couldn’t bring myself to write him in anywhere. Five spots isn’t that many.
That’s what I’ve got to tell you. If you were at all curious about my explanation, that sums it up. I voted for Justin Verlander over Blake Snell. The rest of the voters, collectively, voted for Blake Snell over Justin Verlander. It was pretty close, all things considered, and no other pitchers really factored into the race. I’m open to the suggestion that I might’ve made too much of xwOBA. It does still kind of remain in its infancy. But all I needed was a sign that Snell in some way deserved those minuscule BABIPs. The more I looked, the more I believed in Verlander’s case. It was a good case, but congratulations to Blake Snell, anyway, on a magnificent season.
Jeff made Lookout Landing a thing, but he does not still write there about the Mariners. He does write here, sometimes about the Mariners, but usually not.
Snell’s numbers are backed up by peripherals that are almost certainly not sustainable going forward. But the Cy Young award is for success in 2018, not for how sustainable that success is going to be for 2019 and beyond. Therefore, I think Snell was the right choice
I don’t think I buy the argument that peripherals are purely predictive. To me, they’re a clearer description of what a pitcher actually did, as opposed to what happened after he let go of the ball.
But the two are connected: The reason why it is predictive is *because* it is a much better descriptor of what the pitcher actually did.
If you have two measurements:
1) One of which is narrow but filters out all unwanted noise,
2) And a second one that is broad but lets in a ton of unwanted noise
The first one often tends to perform better in terms of, well, everything–and it’s because the first one usually captures the construct of interest better. This case is no exception, and it’s demonstrated by its predictive validity.
All you are doing is giving Snell all of the credit for run prevention, when that credit should go to him, his defense, and in this case, a huge amount of luck.
I think that’s a bit dismissive.
Would it be fair to say that “all you are doing” is giving Verlander all the credit for things that could have, might have, and perhaps even should have – but didn’t – happen?
No, you’re giving credit to Verlander for things you can actually measure.
Snell had a good year, but Jeff presents a really compelling case why a lot of Snell’s credit should go to his defense while Verlander was just awesome and struck people out. paperlions’ comment is justified by Jeff’s work here.
I’d love to know the luck vs. defense split. I’m actually happy to give the pitcher credit for the year’s luck, in an award like this. But defense belongs to other people.
My wild guess is this is more luck than defense. But we could calculate it for real, with some assumptions…
Can we please retire this straw man? When people use stats like xwOBA or even FIP for Cy Young considerations, rarely are they looking ahead to the future or thinking about “sustainability.” The objective is simply to credit the pitcher for only his own contributions to run prevention, and not those of his teammates or any other outside influences.
Some day, people will be using nothing but spin rate and pitch location to vote for Cy Young. Shortly thereafter, the award will be retired.
The PitchFX award
Awards are so much better when we vote on them because of storylines and other nonsense. AMIRIGHT??
F
This really misses the whole point of sabermetrics.
Peripherals are not supposed to just be predictive. The reason they happen to be more predictive is because they do a better job describing what already happened.
What about Bauer’s tweet today? Not even in the top 5?
Bauer was very good, and he would’ve been in my top ten, but he just didn’t face enough quality opponents. He was pretty good against quality opponents! But almost no starter in the AL faced weaker bats on average. According to Baseball Prospectus, anyway.
I had to include the latest Bauer controversy. Regardless, does that model utilize current streaks of players? What if Bauer was facing a hot KC team vs Snell facing a slumping BOS team? I agree SOS matters, but every great team has bad weeks.
Downvote or come at me. I love to argue!
Further: the average position player Bauer faced in 2018 had a wRC+ of 94.6. The average position player Snell faced in 2018 had a wRC+ of 102.0. (League average for position players is 100.)
Thanks for your replies. Is that the end of the year wRC+ or the average wRC+ of the batters when they pitched against them?
I’m sure it’s end of year wRC+. Why would he use a mid-season wRC+?
I feel cumulative wRC+ for opponents is imprecise. Lineups change, splits change, wRC+ can vary between weeks depending on the streakiness of hitters, etc. For example, Snell could be facing a NYY team with an annual wRC+ of 110, but was facing a cold lineup with backups. Regardless, I was being facetious, but to keep it simple culmulative wRC+ works and Snell was superior in that regard.
You can use the Splits Tool to figure this out (or a script in your preferred language) if you’d like. Annual wRC+, however, shows a hitter’s overall talent level in a season so should be an acceptable answer.
Is there any evidence that “cold lineups” is a real thing?
When they face Verlander, sure.
As an Indians fan, there should be. Look at their first two months. Even then, say the NYY are missing Sanchez and Judge for 2 months. Is that current lineup as wRC+ heavy as their peak wRC+? Or say you have a Kipnis that is like 70 wRC+ for the first two months, thus deflating team wRC+ significantly, but having to be played for contract purposes. Or say you have Alonso playing 1B for the last two months because he adequately performed earlier in the year and is also under contract, and thus must play despite also hovering well under 100 wRC+? Even then, why do some teams produce far better wRC+ some months than others? That you can’t deny. That happens. Would that not be a hot or cold lineup? Why do some players have better averages during certain months? The cumulative wRC+ is very flawed in this regard. Ideally, you would rate SOS for a pitcher based on the current wRC+ of a team, but that would take quite a few calculations.
Or how does an Indians offense that was rather potent during the regular season cumulatively, produce nothing during the postseason? Is that not a cold lineup? Cumulative wRC+ would say that the pitcher did well against a good-great lineup, but that is not always the case. There are variables missing and not calculated here. What if? What if! Snell faced the worst lineup each of the best teams had to offer? Or what if he faced them when they were having a down month? Just questions that escape cumulative wRC+.
>Is there any evidence that “cold lineups” is a real thing?
No. A player’s recent career line (say previous three years) is more predictive of the result of any given plate appearance then any “slump” or “streak” they might be on within the season.
Right. Because if there’s one thing that the past three seasons has taught us, it’s that players like Jackie Bradley don’t have “slumps” or “streaks.”
In any given month, we should expect him to have a wRC+ of exactly 100!
Culmulative wRC+ as a comparison between two pitchers is is deeply flawed, especially when the “lesser” pitcher actually has more “effective” batters utilized by the author as outlined below.
Jeff specifically mentioned that it was the average position player wRC+, not average lineup wRC+, which I don’t think is a thing that anyone uses
Sure, but that is still simply an average and still too loose. Look at the tale of two Jose Ramirez’s. Even regular players face fluctuations, thus making average position player wRC+ misleading. There’s still ground to uncover.
For many of the reasons you articulated (although I do not have the same command of advanced stats), I would have voted similarly- Verlander, Snell, Kluber, Trienen.
(another note, Snell was almost always protected from the 3rd time through the order penalty, and Verlander more often pitched the third time around)
I know WAR isn’t everything but a difference of 2.2 wins is pretty stark. I think Verlander was the clear choice.
You’re assuming fWAR, Snell wins in bWAR by 1.3
IIRC fWar is based mostly on FIP and bWar is based more on RA9. As Verlander has more Ks and fewer BBs, his FIP is way better.
IMO, FIP is a better predictor going forward, but a step further removed from actual results.
yes but see Jeremy’s comment above. It comes down to how much you think pitchers should be held responsible for outcomes on balls in play.
How long until we’re filtering out catcher framing & sequencing and umpire trends/preferences from individual pitcher strikeouts and walks as it pertains to FIP?
I’d like to think that I’m being somewhat sarcastic in regards to that situation, but it’ll become a reality sometime in the future.
dammit!
according to bWAR, Aaron Nola had the best season of any pitcher in the last 16 years.
That’s because DRS often gives you really extreme numbers for defensive values, thought the Phillies’ defense was the worst one ever, and consequently, thinks that Aaron Nola gets all the credit for everything.
I don’t know exactly how much they weight DRS in pitcher performance or for position players’ value, but for whatever reason (overweighting, DRS being a little fidgety) you get some crazy things in there like Nick Ahmed being a 3-win player (yep, he really did…look it up) or Aaron Nola having the greatest season of the last 16 years. It’s hard to take too seriously for that reason.
Does DRS have more variance than UZR?
If I’m not mistaken they were extremely similar until last offseason when UZR was “dampened” and extreme seasons were regressed. Now the ranges of best fielder to worst fielder (before position adjustments) look like this:
DRS
Chapman +29
Blackmon -28
UZR
Simmons +20
Andujar -16
The dampening was announced in a short post with essentially no supporting data so it’s hard to say which system is better. From the above comment it would seem we’re supposed to think that the idea of Ahmed as a 3 win player or Nola as a 10 win player is facially absurd, but I’m not quite sure why.
DRS is jumpier, and always has been. It tries to break down fielding into very small parts of the field, and is very aggressive about assigning credit and demerits as a result.
But 3cardmonty is right, UZR was “dampened”, which does make some of the results a little more reliable.
As far as whether it’s facially absurd–the problem is that bWAR has a *lot* of those cases. According to bWAR, Bryce Harper was apparently backup-quality in 2018 (bWAR: 1.3), or Rhys Hoskins being essentially replacement-level (bWAR: 0.5), and Matt Kemp going from impossibly bad to merely below-average in 2018 (bWAR: 1.1), or Miguel Rojas being a solid starter (bWAR: 2.4). I don’t really have a big problem with the majority of bWAR numbers, but those extreme defensive stats really do something weird to some players.
I got interested in this a few years ago when Jay Bruce was crushing the league at the all-star break, and I clicked over to see what his WAR was and according to BR he was actually providing negative value despite being a Top 10 hitter at that point. Something about DRS and/or how it is included in bWAR can create some very extreme values.
Maybe the problem is less that bWAR uses DRS but that it weights its defensive component – whatever metric they choose to use – too heavily, especially in this era where there a fewer balls in play than ever before.
“but that it weights its defensive component ”
Not sure I follow — isn’t there not really a weighting? The components are boiled down to runs and then added up. If you want the defensive component to have less weight, you have to revise the defensive component.
In 2014, when voting for the Player of the Year Award, you said ” I wanted to prioritize performance over playing time. That is, I didn’t want to dock people too much for an injury”. You then voted for Clayton Kershaw, who made only 27 starts, for first. What made this case so different that Sale not only wasn’t first on your ballot but didn’t even make the top 5?
Not only do I think about baseball differently now from how I did four years ago, but also, importantly, Kershaw threw 40 more innings than Sale just did. He faced 21% more batters than Sale. That’s enormous.
A characteristically compelling Sullivan piece: thoughtful, well-reasoned, well written. Not 100% certain I’m convinced. But I’m lot closer to being a Verlander supporter now than I was before I read. (And bonus points for the responses to the feedback).
BUT I found the dismissal of Sale for Trenien – despite the enormous workload gap (that’s clearly a meaningful part of the pro-Verlander narrative) because of “availability” to be totally peculiar. Would love an article-length explanation of this.
Is being always “available” to work 3 innings per week 100% of the time really worth that much more than being available to work 7 innings a week 80% of the year?
I should say I’m not totally settled on where I am with a case like Sale’s, and a year or two from now I might reevaluate yet again. But the way I looked at these things: Sale faced 617 batters, while Treinen faced 315. But Sale’s average leverage index was 0.91 while Treinen’s was 2.21, and with some multiplication Treinen faced more “effective” batters by a margin of 696 to 561. Accounting for stress, Treinen carried a heavy load.
I understand the pro-Sale argument, absolutely. I know what his WAR and peripherals were. But in my current frame of mind, I feel like half the job is simply being available to pitch in the first place. Missing time like Sale did, then, feels like a major hurdle to overcome. But I’m open to having my mind changed.
Oooh effective batters. Now multiply by WRC and you can have effective run-production batter leverage environment or some insane acronym. To progress!
How would “effective” batters relate to strength of opponents? Bauer would have 709 effective batters compared to Snell’s 679.
I get this, but I’m more convinced by the fact that when Sale wasn’t able to pitch, he had to be replaced, essentially by a replacement pitcher. To me, you have to average that in. Treinen was (for a closer) always available, so you don’t need to average with a replacement reliever.
Jeff – Here’s my abbreviated/off the cuff attempt to change your mind:
(A) LI is a little too much like RBI for my liking (at least for this framework).
IE, we’ve generally moved past giving players credit for the context in which they are used. A batter who has more opportunities with runners on base doesn’t get extra credit (at least in SABR circles) for the same hit vs a player who bats with different timing. Why should the way in which Sale was deployed be held against him? Is there any evidence that he would have pitched worse if used in higher leverage spots?
[One might argue that LI is more analogous to WPA than RBI – But I didn’t see that argument being made with regard to Blake Snell’s 5.0 to 4.7 WPA lead over Verlander …]
(B) This “availability” concept strays a little too far from a basic principle for my liking: a win is a win is a win. Value added in April is no less important than value added in October. To me, the value of Sale’s performance is not diminished by the timing of it. Nor enhanced – ie, if he pitched two innings 70 times evenly spaced throughout the year, I wouldn’t give him (much?) extra credit for it. So how can I give Trenien bonus points for doing that for 70 x 1 inning?
[I appreciate there’s some nuance here that a part time pitcher requires an extra body/extra roster spot to make up for what he’s missed. But what Trenien missed – relative to Sale – is 60 innings of work].
So leverage index matters, but you make no effort to account for schedule differences (other than a cursory glance at BP’s massively flawed SoS numbers)?
Curious.
“I think about baseball differently now from how I did four years ago”
Care to elaborate? I’m always interested in how people’s opinions evolve over time.
What about quality of opponents for Snell vs Verlander, or is that already factored into some of the metrics cited above?
It’s not already factored in, but I did look at it. For most pitchers, it hardly made a meaningful difference.
Not sure how Sale isn’t in your top-5. 6.5 wins is 6.5 wins. And anyone – in theory – that replaces him is probably going to be above replacement, right? You can hide the actual player that replaces an injured player out in the bullpen as your seventh arm that only pitches in low leverage situations (if at all) if you have to anyway. So hypothetically the number one slot in the Red Sox rotation (Sale + hypothetical replacements) should have actually outproduced Verlander over the course of a season….?
The replacement for Sale probably isn’t replacement level, but the replacement for that guy, or the replacement for the guy that replaces that guy, is. And even if the AAA guy brought up to fill out the rotation or bullpen is slightly above replacement level, Sale shouldn’t get more credit just because the Red Sox are deep. Any time a player misses ought to be combined with a replacement player (which is why their WAR would remain the same as it was, since you would be adding 0.0).
But in this case, Jeff is penalizing him. He’s worth 6.5 WAR. It should be left at that in my opinion. He was as valuable as any pitcher in the AL and is no doubt top-5. Because he was much more dominant than all of them in fewer innings, he’s being penalized.
How do you feel about this, Jeff?
I know that time is undefeated but it sure seems like the surest bet in baseball is Justin Verlander taking the mound every 5 days and delivering 6+ innings of premium gas.
JV had seven more quality starts — 26 vs 19.
QS isn’t normally what I would consider, but seven is A LOT of starts.
Actually, that wasn’t true last year. He’s definitely had to make adjustments, because he was slipping. He just made them way better than anyone else.
Prediction: When all is said and done, and both players careers are finished, I think we’ll consider Verlander just as great as Kershaw.
IMO, he’s getting pretty close.
Their career WAR are pretty much the same at this point, and I think Verlander has a pretty good shot at beating Kershaw actually. Kershaw’s been dealing with injuries, and Verlander seems like an iron horse ATM.
Kershaw is like 8 years younger with identical career war. Nothing JV has ever done compares with 2013-2016 clayton Kershaw. One of the best peaks in history, especially on a per inning basis
https://www.fangraphs.com/graphsw.aspx?players=2036,8700&wg=2
I’m casting my vote against “pretty good shot at beating Kershaw”
Here are the SPs within 15% of Verlander’s IP through age 35 and also within 0.5 his per 200 IP fWAR rate with any active pitchers removed:
25.4 Lefty Grove
23.6 John Smoltz
15.0 Mike Mussina
14.3 Andy Pettitte
13.4 Kevin Brown
12.0 Bob Gibson
0.1 Bret Saberhagen
0.0 Dwight Gooden
-0.7 Roy Halladay
Now the same for Kershaw — with the same parameters, we only get Roger Clemens. We can add guys who were more than 0.5 fWAR/200IP away, though perhaps unfairly as it adds only one player up and all the rest down, but it is what it is.
70.2 Roger Clemens
58.8 Greg Maddux
31.7 Fergie Jenkins
28.9 Rick Reuschel
22.7 Pedro Martinez
7.2 Camilo Pascual
0.1 Bret Saberhagen
0.0 Sandy Koufax
0.0 Dwight Gooden
Taking the medians from each list, we’d get 77.0 for Verlander and 84.3 for Kershaw — though it’d probably be fair to take the two career comp lists and eliminate some names based on their last couple of years
IE, based on their career trajectories at those ages, Gooden and Saberhagen look like bad comps for *both* Verlander and Kershaw.
https://www.fangraphs.com/graphsw.aspx?players=2036,8700,1004852,1011355
Verlander has always been in the harder league, as well.
“I didn’t want to ignore Snell’s actual results entirely.”
Careful, looking at actual results is a dangerous line of thinking at Fangraphs. I hope you keep your job.
I would say the argument here is which results are you looking at. If you want to say the result of the pitcher ends when the hitter makes contact with the ball, gets walked, or strike out, then the results favor Verlander. If you say results of the pitcher ends only after the runner reaches base or is out on the field, then result favors Snell.
As an older fan, I personally don’t mind people essentially giving all credit or blame to the result in the field to the pitcher, since it’s pretty much how it’s always been done. It’s how I looked at baseball growing up.
But at the same time I can also understand that it’s not really the pitchers fault if he gives up a weak grounder that turned into an infield hit, while giving him credit for a hard line drive that goes directly into an outfielders glove.
Once you go down the rabbit hole of trying to reward what should/could have happened, it becomes a bit pointless even to have an award.
I don’t think the idea should be, “Let’s try to reconstruct the entire season based on what might have happened had all your pitches occurred in a vacuum.” Such logic introduces all kinds of unknown error from making spurious assumptions and using unproven metrics. As a result, it takes away from the historical value of the award.
Agree. This is a great thinking blog, but the enjoyment factor is reduced if you truly want to visit the rabbit hole. I think Snell should have won, to keep it simple, and he did. Good for him.
Lazy take – awards aren’t awarded for what *should* have happened. Rather, what *did*.
Verlander’s lower FIP and xwOBA suggest he should have been better than Snell in the season.
I know. But he wasn’t – that’s my point. Awards are given for what DID happen, not what SHOULD have happened.
I’m as pro-analytics as anyone, but common sense dictates the award should go to the production, not the hypothetical.
Keep in mind, above, I am calling my take a lazy take, not Jeff’s.
It’s not hypothetical…. FIP tells you what the pitcher actually did, not what he AND his defense did.
We cannot competently gauge how much of the plays outside the three outcomes a pitcher directly controls, are influenced by the pitcher and how much by the defense. We only know the play, the result of the play, and a very limited interpretation of the key drivers in creating that play.
To effectively handicap someone for something we don’t actually know how to quantify correctly, and place priority in that over the actual outcomes, is poor logic and far too reliant on variables.
We are advancing, but we are also getting ahead of ourselves in applying data we don’t really understand yet.
@CamH Well said.
It’s hard for me to believe than a Cy Young voter is using metrics that have not been shown to provide any reliable information (xwOBA).
Snell didn’t pitch as well as Verlander. His defense played better than Verlander’s. That’s what DID happen.
No credible analyst thinks FIP is the best way to strip out defense from pitching performance. It’s easy and lazy to use FIP.
What is the best way, then?
SIERA and/or DRA I believe may be marginally more predictive of future results. Of course, Verlander wins by both of those measures as well. Which makes sense as they largely track with FIP.
Did arm hair factor into your vote or have the tenets of Jeff been shattered by fame?
Good job, I agree with you!
Honestly, after last year, with Kluber and Sale totally dominating, this year was kind of a letdown the AL. I think that Verlander would have been a worthy choice, and there are just a whole lot less questions about what he did vs. the defense bailing you out of runners on base than there are with Snell. There’s kind of a big gap there in my mind, and I’m not at all convinced that Snell performed better this year than Chris Sale or Gerrit Cole.
Wait so you didn’t just ask some former player what they thought and call it a day like your distinguished colleague from San Diego? The fact that that guy has a vote is just irresponsible.
And the Oscar for “Most Deserved BABIP” goes to…
Do you have any evidence — and I’ll even take crappy evidence — that xwOBA is predictive or more stable than what normal stats give you?
That’s not the point of xwOBA (to be predictive or stable). The point of xwOBA is to describe a hitter or pitcher’s results when you give each batted ball outcome the average of all other batted ball outcomes in that bucket, which has the benefits of stripping the effects of defense and park out of the equation.
I hate hearing ERA referred to as actual results. Wins are actual results as well, but pitchers shouldn’t get credit solely for them as a lot of people are involved. ERA isn’t that far removed from wins, and it takes more than a year for batters and defense to even out. FIP is an actual result as well. XFIP is actual results. WAR is actual results.
xFIP is actual results? Sure it is but it’s assuming that a pitcher has no control of homers…. Sorry but that is making a pretty big guess and not what actually DID happen….
Also, the FIP on here counts IW as regular walks. So how in the hell can a pitcher control IW’s but they can’t control for instance a double into the gap…
While the Win’s have been diminished by voters(at least Negatively- think there is some voters who look at 20 wins differently)- they definitely haven’t diminished ERA in the least. I remember what Jayson Stark said I think 2-3 years ago- it’s the Cy Young Award. It’s not the Cy Whiff Award.
Was surprised to see Treinen and Diaz do so poorly. Seems the voters are moving away from Saves even as Wins and ERA remain highly valued.
You know you kinda convinced me that Verlander might actually deserve the Cy Young, but considering his wife, the fact he is a sure thing for the Hall of Fame, and that Snell plays for the Rays, I am glad he was just ever so slightly robbed.
I think the stat that is amazing to me for Snell. He won 21 games…..
In those 21 games, he never gave up more than 2 runs. There wasn’t a single cheap win in his bunch this year.
What is even funnier when you think about it- deGrom actually had 1 of his wins that he gave up 3 runs(2 earned)…..
I’d be curious to hear you elaborate on why Treinen or any other modern reliever has any business on the ballot when he’s not even good enough to start. Sure he had an all-time relief season but aren’t they all inefficiently deployed busted starters?
I hate how people think FIP is a super advanced stat. It isn’t. It doesn’t measure for quality of contact for example which is pretty essential if you’re trying to say who is lucky and who isn’t. So someone who is actually hit hard because they have bad stuff is viewed as being unlucky. Someone who can control quality of contact is viewed as being lucky. It’s a useful stat that tells you something but there is a lot of luck there too.
Snell had better defense but also had better numbers pitching in a tougher division. Snell had a much harder schedule than Verlander. The top 3 offenses in baseball were NYY, Bos, and Cleveland. Verlander had 3 starts against them. Snell had 9. Verlander is a completely fine choice as CY Young as is Snell, I just hate the reductionist view of stats of many people. “His WAR is higher, so he’s a better player.” “ERA doesn’t matter but XFIP does”. Bill James has drawn a lot of heat recently but his article (and Posnanski’s) on the Verlander/Porcello Cy Young is essential reading. For example, in that race, Porcello clearly had a much better defense so Verlander got a huge bump in WAR for having a vastly inferior defense but when you actually looked at the real numbers, it is pretty obvious that Verlander was not actually hurt by his poor OF defense so Verlander got credit for something that didn’t actually hurt him. Not saying it’s the case here, I have no clue. Maybe the deeper numbers show Verlander was more hurt by his defense and Snell was more helped by his bullpen. The point of this is I wish people would stop pretending flawed general stats are definitive truths.
Verlander has led his league in one of the WARs 6 times and has a 7.7 win season on top of that. Despite that he only has 1 Cy, talk about having some bad luck.
Great discussion. I like how you balanced the data with narrative considerations.
Blake Snell didn’t finish the inning in 14 of his 31 starts. Justin Verlander, on the other hand, didn’t finish his last frame in only 7 of his 34. Snell was simply yanked at the sight of trouble much sooner than Verlander. His short leash led to a sub-6 IP/Start average. If that doesn’t speak to who is perceived to be the more dominant starter, I’m not sure what else will, short of diving deep into the analytics (like above), which will also confirm Verlander as the more dominant starter.
they both only had 12 bequeathed runners though.
I’m all in favor of the term “bequeathed runners.” We so often hear only about “inherited runners” from the reliever’s perspective. Love it.
I am not arguing that the Cy Young should just be the WAR leader but there were 7 pitchers in the AL worth more than Snell. Clearly the ERA won him the award and it appears really no other numbers will support his candidacy given that he finished 8th in WAR and basically 2 full wins behind both Verlander and Sale.
Congratulations to Blake Snell. Whatever you might think, he had one heck of a season, and I think the award was given to the right player.
As a pitcher with a shifted defense, you are expected to pitch inside to get the hitter to pull the ball. The hitter knows this as well. So if Snell decides to go along with it and pitch inside, he’s going to get harder contact, but more outs. Verlander might decide to eschew the shift and pitch outside more to get the strikeouts, but in the process, allow more hits/runs. Assuming that’s what is happening (big assumption), I think I’d go with Snell over Verlander.
One thing interesting- the Tango Tracker for the Cy Young was perfect for the top 3 spots in each league…..