Simulating Batting Order: Just Games?
Certain sabermetric analyses, such as those of batting order, are based on simulations. Some will argue that things based on simulations are less likely to gain practical acceptance in baseball. This may be true, but does it really make sense to rule out using simulations for baseball strategy given the prevalence of simulation in the contemporary world?
I’ve addressed batting order and related issues recently, and some may think it is out of proportion with their importance. Of course, that significance is relative to one’s perspective. Indeed, baseball and professional sports in general aren’t “significant” from a larger perspective, so why blog about then at all? More to the point, I have additional thoughts that were brought back to mind after reading Scott McKinney’s stimulating manifesto on sabermetric managing earlier this week. One common response to a posts like McKinney’s is to point out that managers will be reluctant to try new strategies because even if they are the right ones, if they don’t obviously “work” (or even if they do!) they are likely to face a backlash from traditionalists. Even if the front office backs them up, so much of the manager’s job depends on perception that it might be “not worth it” to step off of the beaten path. I’ve discussed the problem of player attitudes and reactions to applied sabermetrics before and what might be done about it. That is a separate discussion; here I will briefly discuss how one might respond to some (not all) predictable objections to using simulations to figure out the best batting order.
I won’t review the batting order recommendations made in The Book yet again, as they are outlined in some of the links above. While those recommendations are given as general rules, as the authors acknowledge elsewhere, those general rules (e.g., the best hitters should hit first, second, and fourth) aren’t necessarily always the case, they are just the most common results of Markov analysis and simulations. The results will vary given the projected specific abilities of the players. Why are simulations (understood broadly to include simulations proper as well as Markov Chains — strictly speaking there is a difference) necessary? By way of contrast, we can test the results of a model-based (e.g., Markov chain, Base Runs) way of deriving linear weights against empirically-derived linear weights. However, in the case of batting order, there isn’t an direct “empirical test,” since we can’t re-run the whole season with the same players against the same opponents all at the same true talent level but with different batting orders even once, let alone enough times to get an adequate sample. So we need to model or simulate different options based on probabilities.
Objections to making baseball decisions based on things that look like simply “games” are understandable (if a bit ironic). However, one could also note a similar thing about using a traditional batting order against just a random one — it isn’t based on a comparable empirical test, either. Moreover, in fields other than baseball (and to be fair, there are likely some sort of simulations and modelling used in baseball front offices, if not necessarily for traditionally “managerial” decision-making) simulations and modelling are used as the basis for decisions all the time. “Wargaming” is one obvious example going back in history before computers, but now adapted to computers. Contemporary non-sports uses of simulations can be found in fields from healthcare to urban planning to engineering. Of course, there is some empirical background to those simulations, but that is also true of simulating baseball — not only using the historical performance to estimate player true talent, but to estimate the various outcomes of base-out and game states.
Naturally, the simulation has to be constructed properly (not to mention having good player projections!), and there is room to debate just how to do that (e.g., Markov versus a strict simulation). I’ll leave those discussions to the smart people. In the meantime, if a manager or fan wonders whether or not it is right to use a “silly game” to help make a decision about how to play a totally non-silly game involving balls, sticks, gloves, managers in uniform (wish they had this in basketball, am I right Stan Van Gundy?), and hot dog-tossing lions, I would simply respond that those silly games aren’t all that different from those used in industries just as, perhaps even more non-silly than baseball.
Matt Klaassen reads and writes obituaries in the Greater Toronto Area. If you can't get enough of him, follow him on Twitter.
The problem with using simulations is that the programmer’s assumptions about how baseball works will be “validated” by the program. A baseball sim will have to either include or not include a protection effect, for example. Using the sim to then test lineup orders to see whether or not protection exists will just reveal the rules of the sim, not of baseball.
You can’t use sims to test fundamental laws of baseball without being INCREDIBLY careful.
You could try to measure “protection” in real life. Simply compare hitter A hitting in front of a good hitter B versus hitting in front of a average hitter C. I’m sure there are enough natural cases of this happening to be able to at least get in the ballpark estimate of the effects of protection.
Yup. It’s been done, and they’ve found no substantial evidence it exists.
A question about those protection simulations – do they find a reduced incidence of intentional walks? Even if a star player hits no better with another good player behind him, doesn’t the fact that he gets more at-bats (and more at-bats in high-leverage situations) mean that protection provides some value?
Thanks for the comment, but does it boil down to anything other than “the simulation needs to be a good one?” Because I’m pretty sure I agree with that, and I’m not sure who wouldn’t.
Not necessarily. If you make a simulation environment and then do a train/test approach, at that point you’re mainly capturing the assumptions implicit in the data. Certain categories of naive algorithms in machine learning literally require zero assumptions by the programmer. This doesn’t mean the answers are right, but it means that they’re the result of the data being biased- not the programmer seeing what they wanted to see.
I personally don’t think that naive algorithms are the way to go with baseball, but one could employ some basic rule and physical constraints and then use data to train into them, without introducing much (if any) assumptions. A constrained Hidden Markov Model is a simple example of this sort of thing.
Another thing to consider in making more efficient lineups is someone like Carlos Ruiz hit .302 last year and had a .400 OBP. But he did that hitting in front of the pitcher last year. If you’d put his “.400 OBP” leadoff, he would cease to have that high walk rate.
I think this toolset in general works better in the AL, where you don’t have nearly as steep of drop-offs in batting talent, so the effects of protection/pitching around are minimized. The pitcher just screws everything up.
you are assuming the conclusion there.
Number 8 hitters in the NL don’t generally have .400 OBPs or 12.7% BB rates. And further more, you could assume that *some* of the OBP “lost” by Ruiz moving up in the order would be gained by whomever moved down; that is to say – the relative effect on the overall run scoring on any set of lineups is ~constant.
Unless you are suggesting that Carlos Ruiz has some ability to draw walks above and beyond the norm ONLY when batting in front of the pitcher!
The evidence that linear lineup simulations “work” is that that the correlation of simulated games using linear weights in run scored is very high. So high, that any “second order” lineup effects (i.e, protection, batting before the pitcher) are not going to influence the actual conclusion.
You would be better off arguing that psychologically batters do best in whatever batting order position(s) their manager derives from them.
Great point Ryan. The Reds toyed with moving Ryan Hanigan’s .436 OBP up in the order last year, but in the games where they did his OBP went down by over 100 pts. Of course its a SSS, but it is only one example of what should be fairly obvious. Guys who use being pitched around to inflate their OBP should not be moved to a part of the order where they won’t be pitched around.
Simulations are only a good indicator of what to do with a fixed level of production to maximize runs. That isn’t what you’re doing with a lineup. You’re trying to maximize production first and foremost. Utilizing that production is somewhat important, but certainly secondary in its effects on the number of runs scored.
Of course, a good projection will take account of that sort of issue. Now, if you find manager who magically knows how to move players around appropriately just when they are starting hot and cold streaks, um, yeah, that would be a good idea.
But to finish my response: an optimized batting order is about “maximizing production” and “scoring more runs.” And obviously, the better the simulation, the better it will do in doing this.
Does anyone more familiar with the math know if a simulation would just be asymptotically equivalent to a Markov approach, anyway?
Well, it depends what you mean. Almost any simulation is going to be a Markov process (you start it at one state and allow it to run through some path of states, subject to transitions). In that way, there’s definitely a strong connection between the two.
In practice, the approaches are different though. One typical Markov chain approach is to formulate the problem mathematically and then use proofs to determine things about it. Unfortunately, most problems don’t lend themselves to closed form solutions when you do this. You typically can’t do this with a simulation, since simulations usually don’t represent their transition probabilities in a way that you can evaluate them without actually running them (this a variant of the halting problem).
Then you have the more CS approach, where you’d estimate Markov states and transitions from data. This sort of approach IS useful and in theory one could represent any computer simulation using an equivalent Markov chain. With that said, given practical constraints this sort of approach won’t excel at the same things a simulation would. Using a train/test approach, you can get a decent idea of states and state transitions- given the parameterizations of the states you set up (i.e. what they consider). You can estimate these from data, then use the derived weights to simulate or apply math proofs.
A Markov model with the standard form of states/transitions should really be considered one form of representation of a Markov state generator. A typical computer simulation is, in fact, simply another way of representing the same sort of problem. With that said, Markov models tend to require a lot of explicit definition of states- which could quickly get out of hand for a rich model. You could do the same thing a lot more sparsely using a simulation approach.
With that said, I think that both simulation and Markov models are useful tools for looking at baseball. If I had all the time in the world, I’d definitely try to spend some time modeling this sort of stuff. Unfortunately, given that it doesn’t provide a lot of money or social utility (i.e. benefit to the world), I just try to comment on it and hope somebody else has more free time than me 🙂
Simulations that can’t be directly compared to empirical data are hard to prove to be valid enough to trust the results. This is true for all fields. You have to fantastically explicit that the underlying logic (or physics) is accurate and complete. The “and complete” is probably the harder step for any simulation that contains a useful amount of complexity and large numbers of simulators don’t know how to distinguish between “the best we can do right now” and “good enough to generate accurate conclusions”.
That said, baseball is much closer to a state-based system than a lot of other simulations of systems that are being used to come to conclusions even though the simulation cannot be tested for accuracy.
And of course, no team uses one lineup for the entire season anyway. Players move up and down depending on who is starting, who is getting a day off, how a manager feels about a particular batter-pitcher matchup. It is possible that by playing day-to-day matchups, a manger could hit on a more optimal lineup than the optimal lineup generator.
Unless that generator took in to account handedness and other such factors.
Simulations account for huge sample sizes (we’re talking hundreds of thousands of instances). Compared to a simulation, a manager has such a tiny amount of games to work with, so the variance will be much higher. If a player happens to homer from the #2 hole when he gets put there, isn’t the manager much more likely to keep him there even if it’s worse off in the long run?
In the NL, the batting orders can change considerably by the middle innings and later as pinch hitters and double switches are used. For this reason, I question whether the magnitude of the lineup optimizing improvement (in terms of projected runs scored) will be overstated, to the extent that the simulation is based on a fixed batting order throughout the game.
I find it amusing that we’re willing to use simulations to predict nuclear reactions, but balk at the concept of using them for baseball. Methinks one group is underestimating the value of simulation and one group is overestimating it…
^This is very funny, but somehow, still not convincing. That’s dogma for ya.
Hmmmm, a baseball simulation thread that I somehow missed out on. 🙂
The way I test mine for accuracy is to compare its results vs those of Vegas. Vegas (closing lines) is pretty smart (smart but still beatable) and is always what any simulator should be first tested against imo. Some in/out of sample testing can protect against back-fitting to a large degree.
Speaking from experience, there is a TON of stuff that goes into a good baseball simulator.
It is truly a great and helpful piece of info. I?m glad that you simply shared this useful info with us. Please keep us informed like this. Thanks for sharing.