The Math on Letting Lance Lynn Hit
There were a bunch of turning points in Game 2 of the NLCS, including three late-game home runs that allowed the Cardinals to walk-off as winners. Yadier Molina‘s exit, due to a strained oblique, also looked like a big moment, especially when backup catcher Tony Cruz couldn’t handle Trevor Rosenthal’s game-tying wild pitch in the ninth inning. But, given the change in expected outcome, the biggest moment of the game might have actually occurred way back in the bottom of the fourth inning.
Already up 1-0, the Cardinals mounted a rally against Jake Peavy, with Matt Adams drawing a leadoff walk and Jhonny Peralta following with a single. Yadier Molina then laid down a bunt, which wouldn’t have made any sense if he was healthy, but it seems like he very well may not have been, which would help explain why he gave himself up to move the runners over. With first base open, the Royals easily decided to walk Kolten Wong, but then Randall Grichuk singled to drive in a run while also keeping the bases loaded.
At this point, the Cardinals had a 2-0 lead and three runners on with only one out. Their win probability had ballooned to 86%, in part because the run expectancy of a bases loaded/1 out situation is 1.5 runs, so while the Cardinals led only 2-0 at that point, the WPA graph was assuming that the inning would end with them either having a 3-0 or 4-0 lead, most likely. And that would make them overwhelming favorites to hang on and win.
However, Win Probability doesn’t know about which hitters are due up, or in this case, which pitcher is due up. Becasue Grichuk was hitting 8th, Lance Lynn was the next batter, and Lance Lynn is not a good hitter, even by the standards of a normal pitcher. 31 pitchers have hit at least 200 times over the last four years; of that group, Lynn’s .102 wOBA ranks 29th. His 49% strikeout rate is the highest of any of those 31 pitchers, and his .006 ISO is second worst. The only decent thing he can do at the plate is draw walks, as he’s somehow managed 11 of them despite a complete inability to do any damage with the bat. In the 195 plate appearances he’s had where he wasn’t walked, he reached just 11 times, an .056 OBP.
In other words, as long as Jake Peavy threw strikes, Lance Lynn was almost certainly going to make an out. And probably not even a run-scoring out. With the force play at home and the double play in order, any ground ball would probably lead to, at best, an out at home, and at worst, an inning-ending double play. Of the 59 non-bunt ball in plays he’s managed in his career, 39 of them — or 68% — have been hit on the ground. Of the 20 balls he’s managed to get in the air in his career, two of them didn’t even leave the infield. The odds of Lynn driving a fly ball deep enough to score a run were extremely thin, and the overwhelming likelihood was that either he was going to strike out or he was going to make an out that either forced all the runners to hold or hit into a double play.
As mentioned, the run expectancy in that situation was 1.5 runs; making one out cost the team 0.77 runs, while a double play would have erased the second half of that total as well. Of course, the out wasn’t actually guaranteed, so we have to calculate some probability that Lynn could have unexpectedly come through in order to find out the expected run value of letting him hit in that situation.
First, let’s deal with the walks. Lynn’s career 5.3% walk rate is actually not so bad, as mentioned, while Peavy has a career 7.3% walk rate, virtually tied with what he did this year. However, those are average walk rates in all situations, and with the bases loaded, Peavy’s career walk rate is just 4%. Not surprisingly, none of those six bases-loaded walks came with a pitcher at the plate. While it’s not impossible that Peavy would have lost the strike zone and issued four pitches out of the zone to Lynn, there isn’t much in the way of historical precedent for it. I think we can probably set the chance of a walk in that situation at somewhere around 1-2%, and given how rarely Peavy hits batters (just 74 out of 8,870), that doesn’t move the needle much either.
Now, for the chances of a hit. Lynn’s career average is .065, a little lower than the .096 mark pitchers have put up against Peavy over his career, which makes sense because Lynn is a terrible hitter even for a pitcher. But we probably shouldn’t even assume overall average production, since this is the situation where Peavy is least likely to just groove one to Lynn in order to keep his pitch count down.
In his career, Peavy had faced pitchers with the bases loaded 15 times; he gave up just one hit, recording 14 outs, seven of them by strikeout. Only one of the eight pitchers who put the bat on the ball managed to lift it into the outfield — Travis Wood, last year, when he hit a grand slam off Peavy — as Peavy was basically all strikeouts and groundballs in these situations. So, given Peavy’s extra incentive to get Lynn out, and Lynn’s own futility as a hitter, let’s assume a 5-6% chance of a hit.
We also have to account for the chances of an error by the Cardinals defense, since that would also get the run in. On the whole, the Giants defense had a .984 fielding percentage this year, exactly equal to the league average. In just over 4,000 bases loaded plate appearances in MLB this year, there were 43 reached on errors, so using a 1% chance of an error seems about right.
Finally, there’s the chance of a sacrifice fly. Lynn has hit one in his career, and Peavy has allowed one to a pitcher, so it’s not a zero probability outcome. It’s probably pretty close, though, given Lynn’s extreme K/GB tendencies. Again, Lynn has hit 18 air balls to the outfield in his 205 career plate appearances, and we’re already accounting for most of this in the probability of him getting a base hit, so there’s not much left here. But let’s give him another 2% chance of hitting a deep enough fly ball to score the runner from third.
Add it all up, and you’ve got about a 10% chance of scoring a run with Lynn at the plate in that situation, and of that 10%, some of the time the team would score multiple runs, either because Lynn shoots a gap or the Giants defense really screws up and boots the ball around. So, let’s estimate the run expectancy of Lynn hitting in that situation at around .15, and in situations where he manages to only make one out, the Cardinals still have Matt Carpenter hitting with a .73 run expectancy. That puts the net loss of having Lynn hit, versus having an average hitter at the plate, about .6 runs, and that’s without including the possibility of the double play. Add that in, and we’re closer to a .7 or .8 run difference in letting Lynn hit versus an average hitter.
Of course, one could argue that the Cardinals didn’t have any average hitters available to pinch-hit at that point. Oscar Taveras was the only pinch-hitter Mike Matheny used yesterday, and he was atrocious at the plate for the Cardinals this year, batting .239/.278/.312 in 248 plate appearances. The 2015 Steamer forecast — which includes his much better minor league numbers — for Taveras’ has him doing much better, but still only posting a 102 wRC+, and that’s without factoring in the pinch-hitting penalty. So, let’s assume that a pinch-hitting Taveras is a below average hitter, which pushes the net difference between he and Lynn back down to closer to .5 or .6 runs. Just to give Matheny the full benefit of the doubt, let’s call it .5 runs, since that makes the math easier too.
How much worse would the Cardinals relievers have to be relative to Lynn to justify punting a half run of offensive potential by letting him hit there? Well, it depends on much longer you’d expect Lynn to pitch, but let’s say that the goal was to get two more innings out of him, which would leave just three innings for the bullpen. Lynn had a great year this year, posting a 2.74 ERA, though his career mark is a less great 3.46. Let’s split the difference and say that Lynn is a true talent 3.10 ERA pitcher right now — dramatically better than Steamer’s 3.77 ERA projection, but that projection seems a little odd and is maybe a topic for another post — so he’d be expected to give up 0.7 runs over two innings pitched.
To balance out the half run loss of letting Lynn hit versus using Taveras, the Cardinals relievers would have to project to give up 1.2 runs over those two innings, which translates to a 5.40 ERA. That’s a far below replacement level performance, and is the kind of pitcher that no playoff team actually carries on their October roster. The Cardinals have plenty of good relief arms, and almost certainly could have increased their odds of winning yesterday’s game by pinch-hitting for Lynn and using their bullpen to cover the final five innings.
However, there’s a bit of a catch here. The two most likely pitchers to handle the 5th and 6th innings would have been Marco Gonzalez and Seth Maneess, but both were used fairly extensively the day before, with Gonzalez throwing 30 pitches and Maness throwing 19. With Michael Wacha being relegated to extra-inning duty, the Cardinals really only four relievers available for significant work last night, and one of them is an extreme platoon-split LOOGY. Is it fair to suggest that the Cardinals should have asked Carlos Martinez, Pat Neshek, and Trevor Rosenthal to get 12-15 outs between them in order to get Taveras at the plate in the bottom of the fourth inning?
Maybe. The math says it would have been a better option from the standpoint of winning Game 2. The potential for big hit there that could have essentially ended the game outweighs the marginal value of Lynn getting a half dozen more outs. And down 1-0 in the series, one could easily argue that Matheny’s sole focus should have been on winning yesterday’s game, and then he could figure out his bullpen usage for the rest of the series once that was accomplished.
But I don’t think we can shrug off the Cardinals relief situation entirely. Having just four relievers, one a specialist, to cover five innings isn’t a great situation, even with a significant lead, even with Wacha in reserve. The math strongly suggests hitting for your pitcher in that situation, but there is some value to be gained from not having to push Martinez, Neshek, and Rosenthal too hard in Game 2 of a seven game series. How much value we put on keeping their workload reduced is a matter of opinion, and it’s very difficult to think that the benefit is large enough to justify letting Lynn hit, but it’s at least another factor in the process.
Matheny, of course, would point to the fact that Lynn was pitching extremely well at the time, and very few -(if any) managers would remove a starter after four shutout innings. Of course, Lynn went on to show just how predictive being “on a roll” is, as he faced nine batters after hitting in the fourth inning, and gave up hits to four of them, recording just five outs and surrendering the lead in the process. But you’ve heard the times-through-the-order lecture from me enough lately, so I won’t say much more about this part of the decision.
Bottom line? Letting a pitcher hit with the bases loaded in a playoff game is almost never going to be a good idea. The best way to preserve a 2-0 lead is to make it a 3-0, 4-0, 5-0, or 6-0 lead, and managers put too much emphasis on trying to preserve small leads rather than being aggressive in situations that could make them much larger. Sending Taveras up to hit for Lynn would have likely improved the Cardinals chances of winning the game. But the team’s bullpen usage the day before was a legitimate extenuating factor, and Taveras was available to pinch-hit for the pitcher later in the game, so the gap is not as large as the simple run expectancy calculations might suggest.
Still, though, if that scenario comes up again, pinch-hitting is probably the better call. You’re better off protecting a bigger lead with even overworked relievers than you are trying to protect a small lead with a tiring starter.
Dave is the Managing Editor of FanGraphs.
Probably shouldn’t delete Choate from the bullpen equation – the next three up for the Giants in the fifth were Belt, Crawford, and Ishikawa. Lynn has a big split, and there’s a times through the order penalty. I can’t see him as a 3.1 ERA pitcher at that point.
A very fair point. I was trying to give Matheny as much of a benefit of the doubt as possible, but yes, three straight lefties due up in the 5th pushes this even further towards pinch-hitting.
Math is so overated. In this day in history a man became famous because his inability to do math caused him to believe he could sail from Spain to India. It worked out good for him.
I hope that you’re joking
Very good article. I think you gave too much benefit of the doubt to Matheny with the numbers, as you say, which makes it MORE of an argument towards lifting Lynn.
There is no reason to bias (give a benefit of the doubt) your numbers toward one side or another. If the numbers say .5 to .6, then use .55 not .5. If you bias your numbers and the final outcome/conclusion is close, it is going to wrong. If the final outcome is NOT close, then if you bias the numbers, the actual outcome might actually be close or REALLY not close, all of which makes a difference in your conclusion.
We can never be 100% sure of anything like this, but at the very least, try and keep your numbers unbiased. We can respect the uncertainty of the numbers, but once you start changing them one way or the other, you defeat the purpose of the analysis. You are not trying to prove one side or the other going in and then being “conservative.” You are presumably going in with no horse in the race or no axe to grind. Let the numbers take you where they take you.
If you want to show us the worst and best case scenarios, that is fine since there is uncertainty in some of the numbers. But don’t give us the best OR worst case scenario just because Matheny did one thing or another.
For example, when a manager thinks that he’d like to get 2 more innings out of a starter, the average number of innings is more likely going to be 1.6 or so. So don’t use 2! That changes the outcome a lot. The correct decision at that moment should be based on the average number of innings your starter is going to get, not the best case scenario.
Oh, and what about Wacha? He is on the roster, no? He could have simply pitches those extra 1-2 innings. He has not pitched in forever and the team has the day off today. I think you are overstating the bullpen issue. There are not often pen issues in the post-season with teams having days off every 2 or 3 games.
And the notion that Lynn’s is a true talent 3.10 pitcher is preposterous and not based on any evidence. Just because YOU decided to put a lot (too much) weight on this year’s performance. Steamer has 3.77. I have around the same. Why do you unilaterally decide to make him more than .5 run better than that? Do you do projections?
And then you mention the TTO penalty, but do you include it in the analysis? If Lynn is allowed to pitch another 2 innings, he will face the order for the 3rd time where he is now .3 runs worse than his overall true talent.
While I love the framework of your analysis, I believe that if you are intellectually honest in your assessments, you will find that the impact of Matheny’s decisions is much greater than you conclude.
Great article, greater comment.
I think the best takeaway from what MGL said here was this:
> If you bias your numbers and the final outcome/conclusion is close, it is going to wrong.
Nonetheless, I agree with Dave (and disagree with MGL), in that I think the decision to let Lynn hit was nominal when weighed against bullpen usage and other issues.
Remember, you use Taveras there, you don’t get to use him later, and if you’re piecemealing your bullpen together you’re going to need to make more moves, not less).
But if you’re going to do a granular analysis and split those hairs because you’re curious, rounding (and the other things MGL points out) make the analysis less useful, if not outright wrong.
Your method is good for reaching the truth about a particular situation, but it may not be so good at persuading. If you’re pitching your article at people who are inclined to believe that managers know a lot more than statheads, you can validly choose to adopt the argumentative posture that Dave chose here, which is to make “at least” statements by adopting those reasonable assumptions most supportive of the manager’s chosen course of action, and then demonstrating that the choice was suboptimal even under those optimistic assumptions.
If this were intended as a sabermetric research piece, it would of course be necessary to either use best estimates or to describe best and worst case assumptions for the choice. But since it’s more of an argumentative/persuasive piece, that’s not really necessary.
Well said Anon21. This exactly what I was thinking after I finished reading MGL’s comment. You also have to look at the audience Dave is trying to reach. My guess is that a good portion of the readership here is familiar with a lot sabr concepts without necessarily being mathematically inclined. Rounding and making assumptions for simplicity’s sake is probably a good idea in this setting.
I disagree and think MGL is right on point. Dave already made many run value and percentage estimations that were just as mathematically essential as the rounding issue. Why go all out for some calculations and not others? And MGL even points out that Dave could easily have included both the right and the best case, which would still give Matheny due mention. And some of the other decisions Dave made were flat out wrong. He invented his own Lynn projection seemingly just on a whim, and he completely ignored the TTO penalty. These two factors can dramatically influence the end result, and despite doing mostly good work up to that point, his argument developed serious flaws. I understand you’re thinking that the readership must be taken into consideration, but I don’t see how this can be, as you say, declared an argumentive piece rather than a research one. More relevant here is that It would be a disservice to every reader if the post provided the wrong information.
Your approach picks unnecessary fights. There’s a good reason to use conservative estimates: it sacrifices the magnitude of the finding in exchange for increased certainty of the conclusion.
Okay, but take a step back and ask yourself where those numbers come from. Numbers aren’t “truth” in the sense that they’re perfect. These are all just statistical models. Models exist to help one understand a situation. Dave tweaked the numbers at times because he figured the numbers themselves didn’t quite capture reality, as for example in this quote:
“But we probably shouldn’t even assume overall average production, since this is the situation where Peavy is least likely to just groove one to Lynn in order to keep his pitch count down”
Dave was doing what a good statistician does – use the numbers to examine a situation and help tell a story.
So basically, we have over 2,000 words and tons of tedious math to say pinch hitting might’ve slightly improved the chances of winning by a few fractions of a point, which then may or may not have been canceled out in future games by overworking the relievers that were needed to do it.
Not trying to be a dick, and I love your work Dave, but an article like this is an example of why so many of my friends/colleagues can’t embrace sabermetrics.
Imagine if he hadn’t rounded to .5 and instead went with .55 like mgl insisted. I’d probably being taking a nap right now, given the sheer tonnage of tedium.
I normally like Dave’s articles, and I do think it’s an interesting question on the surface, but yeah, I could barely make it through this piece. And even then, you have hedges with “maybe” and “probably”…. so I’m just left thinking, “what’s the point?” It almost seems contrarian for the sake of wanting to be contrarian.
Fair enough. This seems like pretty standard Cameron fare though – goes through the various possibilities, assigns run values, compares/contrasts. There’s math and uncertainty involved. This is unavoidable.
This is such a bizarre comment. All of the decisions in a single baseball game, or for that matter over an entire season, only change the likelihoods a small amount. We’re talking about half a run. What did you expect, that it would swing the likelihood three runs?
Randomness means that the right decision will often not end up producing the better result, but making the right decision every time will improve the probabilities that it’ll work out, and that’s the best we can do.
If we interpret it as an added 50% chance to score at least 1 run. That sounds pretty significant to me.
If you find this kind of analysis tedious, you’re very probably in the wrong place.
This isn’t some agenda-driven diatribe; it’s an exploration of the relevant math underlying decision-making in playoff baseball.
Don’t pretend any of us here to explore anything important. We like baseball; we like statistical analysis. This is good, relevant material.
Dave’s equivocation is the only way to be intellectually honest in the context of a counterfactual. If your friends/colleagues are looking for strong, reliable claims backed by overconfidence then I’m sorry they’re looking for love in all the wrong places.
“The best way to preserve a 2-0 lead is to make it a 3-0, 4-0, 5-0, or 6-0 lead, and managers put too much emphasis on trying to preserve small leads rather than being aggressive in situations that could make them much larger.” This x1000
I think you meant “the Giants decided to walk Kolten Wong” in the 2nd paragraph.
Crafty Ned Yost is already putting his stamp on the NL side of the postseason; instead of stealing signs, the Royals are inserting their own to influence who they’ll meet in the World Series. Suck it nerds: this is the new market inefficiency!
I’d pay to see Spy vs. Spy with Royals and Cardinals trying to perform sabotage.
I would be more likely to give Matheny the benefit of the doubt here if he hadn’t already seen this before. 2012 NLCS, game 5, Lynn comes up in a 0-0 game with the bases leaded in the bottom of the 2nd and hit into a DP. Good chance for the Cardinals to blow the game open, avoid going back to San Francisco, and close things out in game 5. Alas, managers not named Buck or Bruce love to treat October like June.
That was the bottom of the 2nd inning! There’s an argument for it in this situation after four innings, but you’d be crazy to remove your starter after just two innings in a scoreless game just to take advantage of an early scoring chance while your bullpen then has to cover seven innings of work!
Depends on the depth of your bullpen – if you have a rested starter in there then you’re golden.
The first two innings are a starter’s most effective. Pulling him at that point is fantastic if you have the depth – you could even pull a beard! Not to mention your pulled starter would probably be ready to start again on 1-2 days’ rest after only throwing ~30 pitches.
“To balance out the half run loss of letting Lynn hit versus using Taveras, the Cardinals relievers would have to project to give up 1.2 runs over those two innings, which translates to a 5.40 ERA. That’s a far below replacement level performance, and is the kind of pitcher that no playoff team actually carries on their October roster.”
It might be true that no team carries a pitcher with a 5.40 ERA on their postseason roster however on any given night any pitcher on any team is capable of giving up 2 runs in 2 innings. Just look at what Strickland, Machi and Romo did in that game. 1.2 innings combined, 3 runs. This article makes it sound like an extremely unlikely occurance that 2 innings of middle relief would result in >1 run being scored but in fact it happens all the time.
This entire analysis seemed skewed to put Matheny in the best (or I guess I should say least worst?) light.
– Tavares with the platoon advantage and a projected 2015 wRC_102 (overall) is assumed to be a below average hitter because of the PH penalty? Any chance having the platoon advantage possibly makes that a terrible assumption and he would be an above average hitter in that situation (or at the very least average)?
– the Lynn RE is .15 because you came to 10% and then just assumed 1/2 of those opportunities are multirun events? Is that based on anything or just a SWAG (the .15 part)?
– as MGL points out if you do get the RE is between .5 and .6, there is no need to round that unless you want to skew the analysis. Why not choose .6 and not give Matheny the benefit of the doubt?
– And with an off day tomorrow (now today), the “don’t want to overwork the relievers” reasoning rings a bit hollow and again seems like something just trying to skew this toward Matheny’s benefit. Folks, including yourself, rail on Ned Yost continuously for having a 7-8-9 formula and not moving off that, especially in playoff time, yet we should assume the Cards need to follow that plan and use lesser pitchers in the 5th and 6th? Or just deploy the ‘good’ relievers if the leverage presents itself in the 5th or 6th? Or ask each of them to get 4 outs such that you only have to cobble together the 5th inning?
If Ned Yost made a decision like this in the WS would similar “math”, rounding, language and assumptions be used?
Apologies – not meant as a reply.
Hank, Dave isn’t tryingt o absolve Matheny with his assumptions. He is trying to win an argument with the status quo. To do so, he is allowing as much reasonable doubt as he can with his analysis and still manages to show that Matheny is likely wrong. In one corner you have 100+ years of tradition and in the other you have sabr approaches. You aren’t going to win that battle by reaching for every fractional run and making every assumption in your favor.
“win an argument with the status quo”
And therein lies the problem with an increasingly large portion of the SABR community.
SABR should be about good, objective analysis. These days it has increasingly turned into “winning an argument” and telling the old school dinosaurs how dumb they are. And as a result the analysis is poor and in some cases biased.
I don’t want him to “reach” for every fractional run… how about just objectively quantifying it regardless of the name of the manager involved? Nor do I want him to make every assumption in favor of the opposite direction… what assumption was I looking to skew? I was just looking for Dave not to skew every assumption in one direction – that doesn’t mean you have to skew them in the other direction.
The issue is summed up well in your last sentence and the need to “win a battle”… what battle? At what point did the us vs them mentality become more important than the actual objectivity and accuracy of the analysis? It’s the Brian Kenney-ization of the SABR community.
I find this a rather odd response. It’s an article that attempts to understand the real conditions that a real game of baseball is played in. There is no exact or ideal set of numbers to work from, it’s merely attempting to analyse one particular decision and the effects it has on run scoring potential.
I’d say the bigger problem with the sabermetrically minded is thinking that mountains of tables and equations is better than an understanding of how human beings play the game.
“This article makes it sound like an extremely unlikely occurance that 2 innings of middle relief would result in >1 run being scored but in fact it happens all the time.”
Yes, it does. Obviously pitchers over the course of 1 or 2 innings will often not match their ERA over a career or season. Because runs are whole numbers, it’s actually pretty much impossible for a pitcher to match his expected ERA if he pitches 1 inning. He can end up either a 0.00 ERA, or a 9.00 ERA (or 18.00, or 27.00, etc.) But he can’t out of 1 inning end up with an ERA in the normal range of say, 2.50 to 4.50, that we see from major league pitchers.
As it happens, Lynn gave up 2 runs in the next 1 2/3 innings that he did pitch. But we don’t analyze him based on the ERA for those 5 outs – which is 10.80 – but based on what we can reasonably expect based on his career.
The exact same argument could be applied to Lynn, could it not?
Its a good analysis and while admittedly its the right thing to do. A MLB manager is unlikely to do this. Nonetheless this is the kind of thinking that needs to change. but in the defense of Dave I agree with his opinion of mitigating criticism on Matheny. The reason is its not only workload management but also the possibility of an extra inning game. While if they do increase the lead that decreases that chance. With a limited amount of pitchers available using them early could be troublesome for an extra inning game since we’ve had plenty of them in this series. Personally I think that Cards should have lifted him and put Randy Choate against the next 3 lefties and give him a shot against them. But I think the penchant for extra innings was a concern.
Interesting to me how the article and all the responses so far just seem to take it as a given that whichever Cardinals relievers were brought in from the bullpen to cover the next 2+ innings would have pitched effectively. Why is everyone assuming that the only consideration that needs to be balanced against going to the pen early is a potential negative effect in future games if the pen is overworked? When you’re relying on 4 or 5 different relievers to get you throught the last 5 innings of a game, what are the odds that any one of those relievers can come in without their best stuff or location and blow it even if the others are all lights out?
Seems to me there is a pretty good reason why no manager in the game today would make that bet and I’m not sure the math used in this article really gets at the all the right questions. To begin with, I would love to see splits of what the league wide bullpen ERA is according to how many relievers a team uses in a game. It’s easy to look at any particular reliever and say that the odds favor him getting out of 1 inning unscathed most of the time. But if you have to trust basically ALL of your available bullpen arms to be successful on the same night I suspect your odds actually go down more than just what a rough calculation of their combined ERAs might suggest. Because an ERA is just an average over time. It doesn’t tell you that on any given night any pitcher can come in and not have it, and that the more pitchers you set yourself up early in a game to HAVE to use, the more you are increasing your odds that one of them is going to come in and roll snake eyes.
I think everyone here agrees that there’s an element of risk involved (and still not appreciated enough by MLB managers and mainstream commentors) in leaving a starter in for the 3rd and 4th turn through the lineup. But I suspect there is also an element of risk involved with every new pitcher you have to bring into a game and I would like to see some research around that before I discard the notion that Matheny’s “old school” thinking in this case might still have been the best approach, all things considered.
It’s not the odds that one of them is bad – which is the same per inning no matter how many you use, and on some level applies to the pitcher already in the game as well – but the increasing cost as recovering from that cascades through the bullpen. If you’ve got five good relievers, then if they have to pitch three innings, and one of them has a bad day, you can just use another. But if they have to pitch five innings, then suddenly you’re scrambling to find someone to get those outs, and putting pressure on guys to get more outs than they’re used to, and increasing the likelihood that you have more than one bad day on your hands and eventually a disaster.
My bigger question is why a team that has essentially zero effective pinch-hitters has this problem at all. Seems like they should have punted the bunch to the point of minimal injury replacement and loaded up the bullpen further.
That’s an interesting question to ask when hindsight lets us say it was brilliant to add a third catcher for this series (and suspect Cards knew there was a problem with Yadi). If they’d not had three catcher’s they’d be facing the painful question of dropping Yadi for the rest of the postseason to have a better backup catcher than Descalso.
Also, their pinch hitters are below MLB average hitters, not of zero effectiveness. As a certain seventh inning AB proved, especially since OT would not be among the “minimal injury replacement” since he’s not much of a defender.
Also, as the Cardinals board discussions would show, there’s no available pitcher who inspires confidence. Siegrist was bad and hurt, Masterson and Motte were just bad, Lyons wasn’t allowed to pitch in September, Greenwood was pretty bad. Freeman was on the roster and lost all Matheny confidence with two nervous walks.
At least Cards have 10-11 pitchers they trust as opposed to Dodgers who seemed to have three starters and one reliever.
The important and worthwhile thing which Dave did, more or less correctly, was to show how the decision affected the win expectancy of the game. That is a very critical point. If you want to criticize or support the decision you must do so with facts and evidence.
While I quarrel with some of the estimates, the framework is correct.
I would have liked to see a final number expressed as win expectancy which represents the gain from PH. If I have a chance I’ll do that later on.
Once, and only then, you are armed with that number, e.g., “not PH cost 3% in WE,” then you can factor in the future cost by potentially taxing the pen. Without knowing the direct cost of not PH there is simply no way to have a credible discussion about which alternative might be best.
Basically that is where managers reside. They have no idea of the direct cost of their decisions and thus they have no way of making proper INFORMED decisions.
Let’s say that we recognize and appreciate that using a reliever for 1-2 innings has some cost? Well, what if the PH decision was only worth .5% in WE one way or the other? Would that balance the bullpen taxing cost (and I’m not convinced there is any – using a pen MORE does not necessarily automatically imply a cost)? What about if the WE cost for not PH were 5%? Would that outweigh a potential bullpen cost?
Surely there has to be a BE point? So again, we must know the numbers in order to have a credible discussion about what the correct play is. We even need to show or assert some evidence that there likely would be a bullpen cost and what that cost might be, also in terms of (future) WE. M
Just asserting that, “well it might be the right move on paper or in isolation, and it doesn’t take into consideration the cost of taxing the pen,” doesn’t make the decision wrong or right. It is one more piece of the puzzle, but the puzzle still has to solved with data and evidence!
This is a really long winded and in MGL’s case, obnoxious, discussion of something that I think might be dealt with quite a bit more easily.
Matheny elected to give up an out. To be honest, in that inning and with four shutout innings in the book, I’d have done the same thing with a two run lead even though I am a pretty firm believer in not letting pitchers bat much in the playoffs. My primary criticism of Matheny would be that he allowed Lynn to swing the bat, which I think creates more risk of DP outcomes than run scoring ones (as Dave’s math suggests).
But it is simple to assess the harm, isn’t it? -0.046 WPA if he just gets out on purpose by crossing into the opposing box as the first pitch is made (fastest way to be called out). If he’s not allowed to swing, maybe that goes less negative by a few ticks due to potential walks, HBP, WP, etc. If he’s allowed to swing maybe the overall effect is a few ticks (yes, 1/1000’s of a win) in either direction.
The lack of league average pinch-hitters means the total loss here was a little less, and one issue I think Dave omits is that using OT here (the obvious PH choice) probably leads to a pitching change that gives both OT and MCarp the platoon disadvantage unless Cards go through another position player in the fourth inning (and wind up even worse off in the 7th inning PH situation). So the impact on the game becomes even less calculable.
Add it up and the choice to leave Lynn in seems much more excusable than the selection of Cruz as the only PH to appear in Game 1. That was Matheneging at its inexplicable worst.
BTW, MGL told us in a previous comment that we were idiots for trying to quantify a decision on sac bunting because we can’t be precise enough and then here calls us idiots for not using all available precision to formulate an answer to another complex question with lots of logic tree branches if you think it all the way through. Okay MGL, I get it. We are all idiots and you are the only non-idiot. Why do you read this website? Did MENSA kick you off their board for being a pain?
I wrote the above before I saw MGL’s second post, which is at least a little gentler tone, I will acknowledge.
Let’s write an equation:
Dumbass Manager Impact= Sum of:
(1) WPA cost of wrong batter taking this AB
(2) WPA adjustment for right batter (difference from league average)
(3) WPA effect of all implications of a PH subsequent to this AB.
In the regular season (3) becomes much more important because of the carry-over effect of bullpen exhaustion on WPAs in future games. Here, especially with the off day, we pretend there are no items in (3) that affect the WPA in any other game (though this was a reason to pull Kershaw in NLDS1 because you should know you’re likely to want him on short rest in NLDS4).
I think we know (1) pretty well because Lynn is so close to an automatic out that swinging or not swinging will only change the WP when he came to the plate so few tenths of a percentage point from where it was after he made his out that we can debate whether he should even be allowed to swing.
We know (2) pretty well because none of the Cards bench players have been above league average, so the question is how much less than 0.047 the PH AB would have been worth.
(3) is about
(3a) choice between Lynn and more bullpen use, and the math is what Dave focused on
(3b) choice about using up STL position player (or two) early, and
(3c) impact on SFG choices, like bringing in a lefty there to fact PH, Carpenter and Jay if anyone else reaches, and effect of that on their bullpen.
The inscrutable stuff in here also includes STL disinclination to pitch Wacha, their understandable worry that they need a short reliever behind Rosenthal because he might do exactly what he did, and fact they are going with a short pen compared to the typical playoff roster.
Hmm, it’s an interesting analysis. It’s a decision that could’ve gone either way.
Although, I wish people would stop harping on Taveras’s lousy overall numbers. He was much better in September once he received semi-regular playing time and was able to adjust to Major League pitching, especially against right-handers like Peavy. Of course, Bochy could’ve brought in a left-hander to face Taveras (and Carpenter afterwards) and relied on his bullpen for longer than expected, a situation you’ve neglected from your analysis, so who knows?
Also, who says the pinch-hitter would’ve been Taveras? With Molina still in the game at that point, Matheny very well could’ve used Pierzynski instead.
I find the bullpen argument unconvincing in this case. I do, BTW, find the “save the bullpen” argument to be reasonably convincing for why the “by the book” move during the regular season is to leave in Lynn rather than pinch hit for him. The key differences in the postseason are the number of off-days and the heightened importance of each individual game.
Regarding the specifics of the Cardinals’ bullpen, I agree that Gonzales presumably wasn’t available for any significant action in Game 2 because he’d thrown 30 pitches in Game 1. As for Maness, he’d thrown 19 pitches in Game 1, but he threw 12 pitches in the entirety of the NLDS. His last outing prior to NLCS Game 1 was 6 pitches in NLDS Game 4, after which he’d had 3 full days of rest. Neshek didn’t pitch in Game 1 of the NLCS, so he’d had 4 days of rest after pitching an inning in Game 4 of the NLDS. Game 2 would be a good time to get more than 1 inning out of Neshek if needed. Martinez threw 13 pitches in Game 1 of the NLCS, Choate threw 1 pitch, and Rosenthal and Wacha weren’t used.
The Cardinals have is bullpen depth with 5 reasonably trustworthy pitchers (Rosenthal – who might be the scariest because of his walks, Neshek, Martinez, Maness, and Gonzales), 1 LOOGY (Choate), and 1 emergency long relief/extra innings option (Wacha). Everyone except Gonzales should have been available for at least 1 inning. It’s disappointing that Matheny didn’t want to deploy it to put the team in a position to score more runs.
I’ll note a couple of other factors, which aren’t completely logical but provide context, why this move bothered Cardinal fans.
First, Cardinals fans saw La Russa’s use of the bullpen early and often during the 2011 post-season, including the NLCS against the Brewers. In that series, the Cardinals’ bullpen pitched more innings than the starters, without any extra inning games. I don’t recall if aggressive pinch hitting decisions were part of that calculus, but it was a great example of paying attention to factors like the times through the order penalty and extra postseason off-days to make use of a deep bullpen.
Second, there are moves like Mathey’s awful and inexplicable decision to use Cruz as the first PH off the bench in a high leverage situation (7th inning, 2 on, 2 out) in Game 1. Cruz is statistically the worst hitting player on the Cardinals’ bench, and 2 better RH batters (Bourjos and Kozma) were available. Because Matheny makes decisions like that in relatively simple situations, Cardinals fans aren’t inclined to give him the benefit of the doubt in more complicated situations.
hhttp://www.vivaelbirdos.com/2014/10/12/6965639/tony-cruz-pinch-hit-seventh-inning-of-nlcs-game-1-why-gob-why
Forget to include that I find the argument of not tiring out the bullpen to be much more of a consideration in Game 3 – knowing that games are scheduled for the next 2 consecutive days – than in Game 2 – where there’s an off-day the next day.
Here are some interesting numbers I gleaned from my sim:
With an average hitting pitcher at the plate, the Cards score an average of 1.21 runs that inning (bases loaded and 1 out).
With a bad hitting pitcher, it is 1.09.
With Taveras PH verus Peavy, it is 1.80 runs, so an additional .71 runs assuming that Lynn is a bad hitting pitchers or .59 assuming he is average (for a pitcher).
If the Giants bring in Affeldt to face at least Taveras and Carpenter, Cards score 1.50 runs that inning, .41 more than if Lynn is a bad hitting pitcher and only .29 more than an average hitting RH batting pitcher.
If the PH Cruz or Kozma, they score 1.68, so the difference between them and Taveras is only .12 runs. Descalso is in almost the same as Taveras, at 1.78 runs.
So it looks to me that there is not much merit to the saving your pinch hitter theory as Descalso is just as good as Taveras and Kozma or Cruz are still almost .6 runs better than a bad hitting pitcher.
I don’t think that Bochy would have taken out Peavy if Taveras PH’s, but if he does, you can easily PH Cruz or Kozma who are going to score 1.56 runs.
Interestingly, because the Cards are already up 2-0, the marginal runs are not worth that much in WE. Normally an extra .5 runs is worth 5% in WE.
Here are the extra WE numbers for the various scenarios:
With Lynn batting if he is an average hitting pitcher they win 89.3% of the time. If he is a poor hitting pitcher, 88.9.
With Taveras or Descalso v. Peavy, it is 92.05% or a gain of 3.15% if Lynn is a poor hitting pitcher (we will assume he is from now on).
With Cruz or Kozma facing Peavy, it is 91.1% or a gain of 2.2%.
With Taveras v. a lefty like Affeldt, the Cards’ WE is 91.1%.
With Kozms or Cruz v. a lefty, it is 91.5%.
So we are looking at around 3.1% gain with Taveras and if the Giants counter with Affedlt it is a 2.2% gain if they leave in Taveras or 2.5% of they switch to Cruz or Kozma.
I think there is going to be an additional gain by bringing in a reliever for Lynn, no matter who it is since Lynn only has 4 more batters the second time through the order.
I think that is going to be worth around another .2% in WE. So total of 2.5 to 3.5 gain in WE by PH.
That is a huge gain for any one in-game decision. You won’t get much bigger than that. But if you sincerely, with evidence, believe that the loss from using an extra reliever for 1-2 innings is equal to or greater than that, then the PH is not a clearly correct move.
Personally, I have no reason to believe that using a reliever for 1-2 innings has any cost at all, let alone 2-3% in WE. But I am speculating with that.
I’m curious why your WE numbers deviate so dramatically from Fangraphs’ numbers. Inverting them to Giants WE, Fangraphs said the Giants had 13.7% before Lynn batted and 18.3 afterwards. You’re saying they had 8-9% with a PH (I think we agree below league average) batting. I’m guessing your sim includes all the actual players in the game, location in the batting order, etc.
BTW, I feel very confident that this was NOT the worst decision Matheny made in the game. What does your sim say about WE of Maness facing Sandoval instead of turning him around with Gonzales? That AB had over four times the LI of the Lynn AB.
I’d also expect Bochy’s decision to allow Peavy to face Carpenter a third time also was pretty bad, though with your Giants WE at only around 11% at that point, it won’t be that impressive an effect.
I’ll check the other numbers but yes my sim uses all the actual players with projections. FG just uses a generic WE I think.
First of all, Gonzalez threw a lot more pitches than Maness did in Game 1. Second, Maness’s primary bullpen role is that of a groundball specialist who excels at coming in mid-inning with runners on base, regardless of which way the batter swings.
It obviously has at least some cost – relievers are finite, and the earlier they’re used the earlier you run out. I have no immediate idea how to go about quantifying it, though. 3% does seem like a lot, but that’s purely intuitive.
Why not Pierzynski? He’s the second best pinch-hitter on the Cardinals’ bench, at least before Molina’s injury turned him into either the starting or primary backup catcher for the rest of the series.
And I think you meant the Giants defense in the tenth paragraph—unless you are implying some sort of grand conspiracy theory.
I’m relatively new to Fangraphs, so apologies if this is off base. These manager decision analyses and comments are fascinating, but obviously take some time to produce. On the other hand, managers typically have +/- 5 minutes to make their decisions. Is the desired solution: an assistant in the booth running probabilities, better-informed game managers, or something else? Or are these articles just for fun and I should be thankful for the entertainment?
Very fair point indeed. I think that’s one of the problems when MGL goes deeper into game specific scenario analysis. It’s fair to criticize a manager for sac bunting (I know, MGL’s least favorite example!) when we can say that our statistical analysis tells us a relatively simple approach that would generate better (albeit not perfect) decisions than what he’s currently doing.
It is unfair, as you suggest, to criticize a manager for not doing an extensive statistical analysis of each decision on the fly. Time doesn’t permit it. So when Lynn came up, I thought:
“If we (yes I give away my rooting interest here) were tied he should be crucified for not pinchhitting. And if Grichuk’s hit had scored two, I wouldn’t have a problem with leaving Lynn in. But up two, hmmm……not sure. But I sure wouldn’t let him swing away!”
And then I thought “Boy this will be fun to look at afterwards and wonder how the WEs work and whether the difference is enough to make PH worth the weight on the bullpen.”
I submit that if a Fangraphs savvy manager existed, (s)he would have a WE calculator in the dugout, but not MGL’s customized simulator. So (s)he would have been able to:
1. Know that Lynn striking out would cost -0.047 WE versus an average hitter.
2. Know that a (Cardinals, hence below MLB average) PH would avoid most but not all of that WE loss.
3. Ponder a guesstimate of the value tradeoff of needing more outs from the bullpen based on which pitcher(s) would be asked to get them, the value lost by using one or two of their PHs now, and the value of the possibility of SFG making a pitching change as a result of the PH decision.
This would lead to an informed “gut” judgment by the Fangraphs manager (and I think these comments should reasonable Fangraphs managers could differ), who would also then go back after the game to assess whether it was the right choice under a more careful statistical analysis than could be made on the fly. Then the next time a similar situation arises, the Fangraphs manager would make an even more informed decision.
We should celebrate that the 2011 Cardinals and Larussa showed the world that pulling your pitcher (especially for a PH) very early when tied or behind is the right move and you can win with your bullpen pitching the majority of innings in the playoffs, so now we get to say “he should know better” when a manager fails to do it. But those decisions, in WE terms, were probably much more obvious than this one.
I doubt we’ll ever convince managers (or even reach a consensus in our group) that it makes sense to PH for an effective starter who is still in the second time through the lineup when you’re ahead in a playoff game. The definition of the situation is just too specific (what inning? how many ahead? how many batters left in the order? what do you think about your bullpen?).
On the other hand, these playoffs seem to have given some additional powerful examples to support the simple rule “pull your starter after six no matter what, since TTO is so meaningful.” Though, like any decision about probabilities, you get counterexamples (Zimmermann). I do have some hope that organizations are going to reduce the frequency with which starters are asking to finish up close playoff games, though it means Carpenter/Halliday may wind up the last classic pitchers duel. Last year’s playoffs featured three 1-0 games and none of the starters pitched 8 complete, so Zimmermann would already have been an oddity.
I don’t know if we believe that it’s predictive, but Lynn actually has a reverse order split in his career. I saw Jonah Keri mention in a piece at 538 that Lynn did this year, and it’s also true for his career. It surprises me, but here are the slash lines against Lynn based on times through the order as a starter:
1st: .267 / .326 / .409
2nd: .252 / .332 / .395
3rd: .216 / .299 / .313
The sample sizes are between 686 PA’s (3rd time) and 871 PA’s (1st time).
In my study on “times through the order,” there was no evidence that a pitcher’s OWN pattern has any predictive value. Pitchers who got BETTER or especially WORSE through the order in any year regressed completely toward the average penalty in any other year, so we MUST assume that Lynn’s unusual TTO pattern was just a fluke. Jonah Keri’s article on 538 was a combination of weak sabermetrics and mainstream nonsense.
Remember that if something has little or no predictive value, it will still follow a normal curve when looking at samples of performance. All samples of performance follow a normal curve with plenty of random outliers.
“It is unfair, as you suggest, to criticize a manager for not doing an extensive statistical analysis of each decision on the fly. Time doesn’t permit it. So when Lynn came up, I thought….”
My problem with giving managers a break based on the “realtime” aspect of these decisions is that they have plenty of time before the game (in fact they have all damned year) in which to go through the more reasonable “what if” situations, think them through, then apply them in game as needed.
It should be pretty easy to come up with some general rules about pulling a starting pitcher for a PH based on inning, score differential and base/out situation. Just put a simple matrix up in the dugout for anything reasonable. This is their job. They have all year to figure this stuff out and with potentially millions of dollars at stake there should be plenty of motive to put some science into the process. Instead we rely on “gut” feelings because I guess thinking things through before hand is hard….? Sorry, that I don’t understand.
I was hoping someone would run the numbers on that decision – so thank you! A very fun, sober-minded take. What bothers me is that there’s probably not a big-league manager today who would pull Lynn in that situation, and I’m guessing we’re several years away from that kind of consideration becoming the norm. What’s odd to me, however, is that you don’t need to have an amazingly abstract, dispassionate mind to come to the conclusion that it may be a good idea to PH for Lynn there. The gut can get there too. I mean, imagine you’re Bruce Bochy and Tavares strides to the plate in that situation. Granted, Tavares isn’t Barry Bonds or anything, but still… You’d shit your pants, right?
Naw, you’d think, “well, this is the ballgame right here” and go to the lefty Lopez, which helps with both OT and Carpenter, and makes the whole rest of the game play out differently if you get through it. Then you probably don’t use Affeldt to start the fifth since he’s the last lefty and you want him to face PH/Carp/Jay later in the game. And then maybe you get two shutout innings from someone else instead and Affeldt prevents one of the three late inning HRs by LH v RHP. Or not.
Brian – you raise a good point, and I think that a lot of the problem in this situation is that managers follow a set of principles that are optimized for the regular season (and possibly for an era when bullpens were smaller). In the middle of June, sticking with your starter in situations like this one makes some sense. In the postseason – with more off-days for the bullpen and the 5th starter moved to the pen – strategies should change, making pinch-hitting a more attractive option.
In the regular season, no manager would pull the starter. In the post-season, I think some of the more progressive minded ones would.
As far as managers being able to do the “math” on the fly, no, of course they can’t. But certain rules of thumb can be taught to them. That is what we do in The Book. We go through the math in many common in-game situations and then provide a “rule of thumb” for a manager to use.
The interesting part of this particular decision is that of all the in-game decisions, by far and away, the strongest one is lifting your non-ace starter in the NL when it is his turn to bat in a high leverage situation or with runner on base when the RE difference is great. It is a no-brainer for 2 reasons: One, the run and usually win difference is huge. Two, it has the added benefit of preventing your non-ace starter from suffering the 3rd and 4th time through the order penalty. In this case it was a little early (Lynn had 4 more batters the second time through the order) and the run impact was not that large on the win impact because the game was already 2-0 and it was only 1 out (even with Lynn hitting, the Cards still had over a 1 run expectancy in that inning).
So, the “rule of thumb” for the manager is simple: If your non-ace starter is due to bat in a high leverage situation and he is facing the order for the 3rd time or later, pull him.
We also have a section in The Book where we did research to suggest that relievers can pitch more than they do with no adverse effect. I am also a believer that managers, using optimal decision-making can streamline their bullpen usage in order to save arms, especially for times like this where you call on your pen to pitch 1-2 extra innings. For example, managers do not use mop up relievers enough in low-leverage situations. I am an advocate of having at least one reliever in the pen who is your multiple inning “garbage man.” If he gets worn out, you can send him to the minors and get someone else. Managers sometimes make too many pitching changes when they are down 4 or 5 runs late in the game. Obviously you CAN and sometimes DO come back in those situations, but by definition when you are up or down a lot late in a game, and the leverage is low, the difference between a good and bad pitcher is small.
For the heck of it, here are Lance Lynn’s career OPS splits by times through the batting order:
1st .735
2nd .727
3rd .613
4th .484 (Just 51 plate appearances, but included for fun.)
Also, I love when MGL posts here; he makes me think of Wallace Shawn in Princess Bride and I smile.
See my comments above. Using those for predictive value would be like using a pitcher’s results on odd and even days for predictive/decision-making value.
Great article! Thanks Dave.