Archive for run estimators

Team-Specific Hitter Values by Markov

In my first article, I wrote about the limitations of the linear weights system that wOBA is based on when it comes to the context of unusual team offenses. In my second, I explained how Tom Tango, wOBA’s creator, also came up with a way of addressing some of these limitations by deriving a new set of linear weights for different run environments, thanks to BaseRuns. Today, I will tell you about the next step in the evolution of run estimators — the Markov model. Tom Tango created such a model that can be accessed through his website, and I’ve turned that model into a spreadsheet that I’ll share with you here.

I’ve told you that the problem with the standard run estimator formulas is that they make assumptions about what a hit is going to be worth, run-wise, based on what it was worth to an average team. That means it’s not going to apply very well to an unusual team. What’s so great about the Markov is that it makes no such assumptions — it figures all of that out itself, specific to each team. And when I say it figures it out, I mean it basically calculates out a typical game for that team, given the proportion of singles, walks, home runs, etc. the team gets in its plate appearances. It therefore estimates the run-scoring of typical teams better than just about anything, but it also theoretically should apply much, much better to very unusual or even made-up teams.
Read the rest of this entry »


Linear Weights + BaseRuns = Good

In my last article, I explained how wOBA’s current implementation changes the value of walks, singles, home runs, etc., annually due to changing league characteristics.  Does this mean that the value of an event is the same for every team in the league each season?  Or in every park in the league?  No way.  If you’re talking about a weak offense in a high-offense era, then the overall constants for a weak offensive era are probably more applicable to that team.  However, it’s not really the point of standard wOBA to guess the run-producing contribution of a particular player to a particular team; I think it’s probably more accurate to say it’s about his probable productiveness in a typical team (although park effects aren’t taken into account, so not exactly… that would be more true of wRC+).

Anyway, Tom Tango realized this limitation, and produced a table that shows how the values change depending on a team’s runs scored.  He accomplished this system of “Custom Linear Weights” (“a necessary offshoot” of linear weights, he says) by making use of David Smyth’s BaseRuns formula, which is, in simplest terms, Runs Scored = base runners * (% of base runners that score) + home runs.  Home run hitters are not considered base runners, in this equation, by the way.  Makes perfect sense, right?

Tango realized that BaseRuns had a better handle on the team run-scoring process than his basic linear weights system (and all the other run estimators), so he translated the results of BaseRuns in various run environments into linear weights.  Specifically, the BaseRuns formula told him how many runs the team should score, and the linear weight value of each hit came from how many additional runs BaseRuns expected the team score if it had one more of that type of hit (the marginal value of each hit type).  Here are just the basics of his results, in graphical form:

Read the rest of this entry »


Adjusting Linear Weights for Extreme Environments

Well, it’s my first assignment as a real writer, having been promoted for my Community Research articles on pitcher BABIPs and ERA estimators, and I’ve been thrown into the deep end of the pool: linear weights.  It’s a tricky subject, but I’ll try to walk you through both the problems with linear weights and how they can be overcome.  This article series mainly draws from various works of Tom “Tango,” a.k.a. “tangotiger,” the creator of wOBA and FIP, as well as from David Smyth’s BaseRuns.  I’ll go deeper and deeper down the rabbit hole of stat geekishness as the series goes on, eventually emerging with a spreadsheet version of Tango’s Markov run modeler that I made for you all to play with.  Where the Markov mainly shines over wOBA is when it comes to extreme run environments, such as unusual offenses or extreme ball parks.

Who cares about extreme run environments?

Nerds like me, I guess?  Tom Tango cared enough to come up with ways to address the shortcomings his original wOBA formulation.  If you’ve ever wondered how valuable a certain player is to your favorite team, maybe you should care too; that low-OBP slugger might be more valuable than wOBA might suggest to your low-OBP team.  On the other end, a typical walk last year was worth considerably more to the high-OBP Cardinals than it was to the low-OBP Mariners (around 0.04-0.065 more runs each… which adds up over a season).

Read the rest of this entry »


Is Run Estimation Relevant to Free Agency?

Sometimes there seem to be two separate branches of saber-oriented blogging: one that uses sabermetric tools to analyze current events (player transactions, in-game strategic choices, etc.), and another which focuses on more theoretical issues (e.g., specific hitting and pitching metrics). Obviously, the latter is supposed to ground the former, but there still seems to be something of a disconnect between the two levels in popular perception. I say this because I was recently part of a discussion in which some were pointing out the superiority of linear weights run estimators for individual hitters to the approach of Bill James’ Runs Created. Someone then made a comment to the effect that this was simply a nit-picking preference for a “pet metric” that really did not make that much of a practical difference.

Sabermetrics is far from being a “complete” science in any area. Debates about how best to measure pitching and fielding are obvious examples of this. With respect to run estimators, there is a greater level of consensus. However, because of the progress (at least relative to pitching and hitting) that has been made with run estimators for offense, that also means there is less of a difference between the metrics. However, it does make a difference. Rather than arguing for one approach to run estimation over another, I want to simply look at a few different free agents from the current off-season to see what sort of difference using one simple run estimator rather than another would make on a practical level.

Read the rest of this entry »