Archive for tom tango

Team-Specific Hitter Values by Markov

In my first article, I wrote about the limitations of the linear weights system that wOBA is based on when it comes to the context of unusual team offenses. In my second, I explained how Tom Tango, wOBA’s creator, also came up with a way of addressing some of these limitations by deriving a new set of linear weights for different run environments, thanks to BaseRuns. Today, I will tell you about the next step in the evolution of run estimators — the Markov model. Tom Tango created such a model that can be accessed through his website, and I’ve turned that model into a spreadsheet that I’ll share with you here.

I’ve told you that the problem with the standard run estimator formulas is that they make assumptions about what a hit is going to be worth, run-wise, based on what it was worth to an average team. That means it’s not going to apply very well to an unusual team. What’s so great about the Markov is that it makes no such assumptions — it figures all of that out itself, specific to each team. And when I say it figures it out, I mean it basically calculates out a typical game for that team, given the proportion of singles, walks, home runs, etc. the team gets in its plate appearances. It therefore estimates the run-scoring of typical teams better than just about anything, but it also theoretically should apply much, much better to very unusual or even made-up teams.
Read the rest of this entry »


Linear Weights + BaseRuns = Good

In my last article, I explained how wOBA’s current implementation changes the value of walks, singles, home runs, etc., annually due to changing league characteristics.  Does this mean that the value of an event is the same for every team in the league each season?  Or in every park in the league?  No way.  If you’re talking about a weak offense in a high-offense era, then the overall constants for a weak offensive era are probably more applicable to that team.  However, it’s not really the point of standard wOBA to guess the run-producing contribution of a particular player to a particular team; I think it’s probably more accurate to say it’s about his probable productiveness in a typical team (although park effects aren’t taken into account, so not exactly… that would be more true of wRC+).

Anyway, Tom Tango realized this limitation, and produced a table that shows how the values change depending on a team’s runs scored.  He accomplished this system of “Custom Linear Weights” (“a necessary offshoot” of linear weights, he says) by making use of David Smyth’s BaseRuns formula, which is, in simplest terms, Runs Scored = base runners * (% of base runners that score) + home runs.  Home run hitters are not considered base runners, in this equation, by the way.  Makes perfect sense, right?

Tango realized that BaseRuns had a better handle on the team run-scoring process than his basic linear weights system (and all the other run estimators), so he translated the results of BaseRuns in various run environments into linear weights.  Specifically, the BaseRuns formula told him how many runs the team should score, and the linear weight value of each hit came from how many additional runs BaseRuns expected the team score if it had one more of that type of hit (the marginal value of each hit type).  Here are just the basics of his results, in graphical form:

Read the rest of this entry »


Adjusting Linear Weights for Extreme Environments

Well, it’s my first assignment as a real writer, having been promoted for my Community Research articles on pitcher BABIPs and ERA estimators, and I’ve been thrown into the deep end of the pool: linear weights.  It’s a tricky subject, but I’ll try to walk you through both the problems with linear weights and how they can be overcome.  This article series mainly draws from various works of Tom “Tango,” a.k.a. “tangotiger,” the creator of wOBA and FIP, as well as from David Smyth’s BaseRuns.  I’ll go deeper and deeper down the rabbit hole of stat geekishness as the series goes on, eventually emerging with a spreadsheet version of Tango’s Markov run modeler that I made for you all to play with.  Where the Markov mainly shines over wOBA is when it comes to extreme run environments, such as unusual offenses or extreme ball parks.

Who cares about extreme run environments?

Nerds like me, I guess?  Tom Tango cared enough to come up with ways to address the shortcomings his original wOBA formulation.  If you’ve ever wondered how valuable a certain player is to your favorite team, maybe you should care too; that low-OBP slugger might be more valuable than wOBA might suggest to your low-OBP team.  On the other end, a typical walk last year was worth considerably more to the high-OBP Cardinals than it was to the low-OBP Mariners (around 0.04-0.065 more runs each… which adds up over a season).

Read the rest of this entry »


Is Zimmerman a Better Fielder than Longoria?

Like many wannabe saberdorks, I love Joe Posnanski’s work. It’s not just because he’s so much better than, say, [horrible-and-inexplicably-award-winning columnist for major newspaper] or [rumor-mongering baseball reporter prone to bouts of self-righteousness]. This isn’t a Posnanski tribute, but in short: Posnanski is great because he tells an engaging story and incorporates good baseball analysis without confusing one for the other.

This doesn’t mean that I always agree with Poz.* I disagree with many things written by sportswriters. In Posnanski’s case, I think highly enough of him that it’s worth quibbling over minor points, unlike, say, with [arrogant breaker of stories for your dad’s favorite sports magazine that we would have found out about anyway], who is only worth refuting because of his [alleged] influence. I hold Posnanski to a higher standard (not that he knows I exist).

*Or “JoPo”; has a sports journalist ever had so many different nicknames?

Which brings us to today’s Poz post on likely future Hall-of-Famers currently under 30. It’s an entertaining (if unsurprising) read. One claim in particular caught my eye. Posnanski writes that Ryan Zimmerman is “probably better defensively” than Evan Longoria. Now, Longoria didn’t qualify for the list (hasn’t played 500 major-league games), so while I do think he is the better player, that isn’t the point here. The issue is whether Zimmerman is “probably better defensively” than Longoria, as Posnanski claims.

Although he doesn’t cite specific defensive numbers in this piece, Posnanski has used Dewan’s plus/minus system in the past (although he has increasingly cited UZR). Here are the Dewan numbers for Zimmerman and Longoria in seasons in which they’ve both played (2008 and 2009):

Plus/Minus 2008:
Zimmerman: +10 plays (+11 runs) in 910.2 innings
Longoria: +11 plays (+9 runs) in 1045.2 innings

Plus/Minus 2009:
Zimmerman: +28 (+22 runs) in 1337.2 innings
Longoria: +21 (+17 runs) in 1302.2 innings

Over the last two seasons, Zimmerman has been 7 runs better in about 100 fewer innings according to plus/minus. Seven runs is seven runs, but given everything that is rightly said about the large error bars on defensive metrics, the gap isn’t as significant as it looks.

Given the various issues with defensive metrics, looking at other systems will give us a more perspicuous overview. Here at FanGraphs, UZR is used to measure fielding. I’m not qualified to argue which metric is the best; I’m simply using them as separate data points. UZR has a helpful “rate stat” version, UZR/150 (runs above/below average per 150 games). I’ve included the “non-rate” runs in parentheses.

UZR/150 2008:
Zimmerman: +3.4 (+2.1)
Longoria: +20.1 (18.5)

UZR/150 2009:
Zimmerman: +20.1 (+18.1)
Longoria: +19.2 (+14.9)

Suddenly things are less obvious. While 2009 was practically even, in 2008 UZR has Longoria almost two wins better. Their career UZR/150s: +12 for Zimmerman, +19.6 for Longoria. It’s a smaller sample for Longoria, but if you check Jeff Zimmerman’s regressed and age-adjusted 2010 UZR/150 projections, Zimmerman is at +10, and Longoria +12.

Defensive stats are obviously important, but when estimating fielding skill, in particular, we need to weight visual evidence — scouting — heavily. I’m not a professional scout, and unlike Posnanski, I don’t have access to them. Perhaps legendary scout Art Stewart, who told Poz “You will remember this day for the rest of your life” after Royals great Chris Lubanski’s first batting session at Kauffman Stadium, thinks Zimmerman is way better than Longoria. Jokes aside, scouting is essential for estimating defensive ability.

While most of us don’t have access to professional scouts, we do have access to the
Fans Scouting Report. In both 2008 and 2009 Longoria was rated as (slightly) better than Zimmerman.

Given that plus/minus seems to “prefer” Zimmerman — and UZR, Longoria — does this make the Fans Scouting Report a tiebreaker in Longoria’s favor? No. Given the relative closeness of the rating, neither the numbers nor the testimony of observers has the degree of reliability for us to make that kind of call. However, contra Posnanski, I do not think we can say that either player is “probably better defensively” than the other.

Molehill converted to mountain? Check. Happy American Thanksgiving, everyone!