FanGraphs Audio: Dave Cameron and Jeff Passan, Amicably
Episode 482
Dave Cameron is the managing editor of FanGraphs. Jeff Passan is a baseball columnist for Yahoo! Sports. Together, they are the two-headed monster guest on this edition of FanGraphs Audio, which concerns very much the methodology of Wins Above Replacement.
Don’t hesitate to direct pod-related correspondence to @cistulli on Twitter.
You can subscribe to the podcast via iTunes or other feeder things.
Audio after the jump. (Approximately 44 min play time.)
Podcast: Play in new window | Download
You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.
Subscribe: RSS
Carson Cistulli has published a book of aphorisms called Spirited Ejaculations of a New Enthusiast.
needs more reggaeton horn
Was I the only one a little disappointed by this?
It mostly boiled down to, yes WAR is the really the only logical framework for analyzing value, but maybe FanGraph articles should on a more frequent basis point out known shortcomings of WAR. And oh, maybe regress defense 20 or 30%.
I was really hoping this would get into greater depth, things like the issue Jeff Zimmerman brought up about how the fielding component of WAR is surprisingly sensitive to outlier performances at a position, or get into the nitty-gritty on Passan’s article’s points on Lagares.
From Passan’s article: http://sports.yahoo.com/news/10-degrees–why-war-doesn-t-always-add-up-030133203.html
That’s all based on Inside Edge data FanGraphs has. So when the article was published, Lagares had made 12 plays judged to be between 0% and 90% likely to have been made, with only 5 of those being less than 50% likely. The article also says that DRS at that time was crediting Lagares with 23 runs saved.
Obviously DRS and Inside Edge aren’t going to match up 100%, but it’s hard to imagine how 23 runs could come from something in the neighboorhood of even 20 plays (let alone 12 plays). Doubles only have a linear weight value of 1.3 (singles are 0.9, triples are 1.6, HR are 2.1), and Lagares should only be getting credit for at most around 50% of those run values in this situation, if I understand things correctly.
So either this demonstrates just how far apart stringers can be in their defensive evaluations, or points to a systemic problem in how run values are doled out. (Or some of both, not to mention other things.)
To add a couple more points on Lagares, there is value in being perfect on all those routine plays, but according to Passan, there were 5 other players (presumably CF?) that were perfect at routine plays, and made more of them. (I guess one of those 5 flubbed something — now we’re down to 4 CF with more routine chances than Lagares and converting at 100%: Jennings, Ellsbury, Gomez, and Span.)
There’s also arm value, which I presume Inside Edge doesn’t account for? But after a spectacular year last year, Lagares isn’t getting as high of marks for his arm this year.
As I write this comment, Lagares is in first place in CF DRS, now up to 29. His ARM rating is just 4.3, so if DRS is rating his arm like so, that’s still around 25 runs purely from ball-catching.
Leonys Martin is in second for CF, with 15 DRS (with maybe 8 of those runs coming from Martin’s arm). Jackie Bradley, Jr. is next with 12 DRS, 4.7 ARM.
That’s close to a 20-run gap between Lagares and the best of the rest. Lagares is now up to 17 Inside Edge plays between 0-90% (still with none between 0-10%). Even if Lagares was the only CF making those 17 plays (which he’s not, by definition), I’m not sure that could be worth 20 runs.
Maybe it’s better to average together UZR, DRS, Inside Edge and/or other defensive systems for the defensive component of WAR? That way, if someone is an outlier in all of them, he won’t be punished/rewarded by regression.
If you take the average of each of the categories (95% for 90-100%, 75% for 60-90% ect) and use 1 runs per play (.7 RAA for double vs -.3 RAA for outs), Lagares is 17.5 RAA on the year. Add in 4 RAA for his arm and you are at 21 runs. How do you get to 29 runs? Switch the average play probability from
5% 25% 50% 75% 95%
to
5% 20% 45% 70% 93%
But we’re talking runs above average here. An average fielder gets to some portion of those plays, by definition, thus only some of those plays should accrue toward the RAA score.
It’s possible that Lagares only prevents doubles/triples/HR and fields only 10%, 40%, and 60% balls, but even then it’s hard to reconcile getting that far above average. The other thing I didn’t mention is that there are other CF with what look like similar or better overall Inside Edge profiles, but are rated much worse.
Could just be that Inside Edge isn’t a very good product.
The calculations I made includes plays he didn’t make as well.
We’ll use the 10-40 bucket as an example. Currently Lagares is 8 for 16.
If you use an average PlayProb of 25%, then you give him +6 runs for the plays he made (75%*8) and -2 runs for the plays he didn’t (25%*8), so 4 RAA total.
If you use an average PlayProb of 20%, then you give him +6.4 runs for the plays he made (80%*8) and -1.6 runs for the plays he didn’t (20%*8), so 4.8 RAA total.
What I was trying to show is that when you are looking at ranges as wide as 10-40%, the actual PlayProb assigned by UZR/DRS can make a significant difference.
Yep, definitely understand that point. But what I’m saying is that an average player should by definition be converting 5% of the 5% plays (Lagares doesn’t have any between 0-10% in his 16 chances for those plays, putting him in a hole to start with)… 20% of the 20% plays, etc.
He probably, compared to the average fielder, has 3 extra 10-40% plays, 1 (maybe 2) extra 40-60% plays, and extra 3-5 net 60%-90% plays. That seems like 10 extra plays over an average fielder in those ranges, at best. And none in the 0-10% range.
Ellsbury is one CF who actually looks somewhat better by Inside Edge data.
I would be disappointed too if I thought Passan had any real insight into the topic and wasn’t just rehashing the same arguments we’ve heard a million times and cherry picking examples to make WAR look bad.
Cistulli needs to do a regular episode in which he moderates a discussion between a saber-y dude and a regular baseball writer. Be the fucking shirpa!!
“My guest on this episode, this episode, of Fangraphs audio is indeed again two guests. Two guests. The first guest, guest number one is our own managing editor for Fangraphs, analyzing all sabermetrics, Dave Cameron. Our other guest, the guest in the other corner, the other corner of the proverbial ring, is none other than the highly respected, highly respected columnist for the Boston Globe, Baaaab Ryan. These gentlemen, these gentlemen of the game, will partake in what can only truly be described as, uh, a Battle Royale. A Battle Royale. There were moments of civility, and there were moments, moments of combativeness. You, as a loyal listener of Fangraphs audio, get to, to experience all of it. In it’s entirety. You get to hear Baaaab Ryan suggest the following: “How am I a Masshole?” to Dave Cameron plunging into deeper waters with the following “I’m pretty sure I bought a three pack of squeaky toys, but I can only find one”. This is Fangraphs audio, featuring Dave Cameron and Baaaab Ryan, and it starts…it starts right now.”
This was lovely!
The synopsis of Passan’s argument: 1)because WAR is not a perfect model and has limitations, lets be cautious when using it to derive conclusions and also make the limitations very clear in all works where it is referenced 2)because WAR is more accurate than older stats (which he agrees with, besides his stated skepticism of the defensive component),it needs to be disseminated to the larger public but in a way that is palatable, accessible, and understandable on a large scale. These are all fair points.
Where his argument loses ground is when Dave poses the question “If not X then what?” He has no answer for this other than retreating to point out known or potential flaws. A model seeks perfection but understandably does not achieve it. Assumptions are made because educated assumptions are better (for informative) then absent components (an concept that would be backed in any academic discipline). So in essence, an understanding of the limitations is part of the value because it creates a more complete understanding. Basically,from my perspective Jeff’s argument is that in his mind there is a margin of error that is acceptable to get his buy in and the current margin does not meet his smell test standards.
All good points. Another thing — I got the distinct impression that Jeff wants current-season WAR to show Mike Trout atop the leaderboard, period. Because that’s what his eyes tell him.
He added that the Mets front office doesn’t see Lagares as a 5-win player. So what? I have my reservations about how Lagares supposedly generated his wins (see above), but by itself, the Mets front office opinion is a meaningless criticism of the 2014 WAR leaderboards, since we’re not talking true talent here.
I wish Dave or Carson had pressed further on these related points, because despite Jeff paying lip service to the idea that single-season WAR is not true talent, Jeff really really seemed to be going back to the idea that it should be.
WAR is the summation of a component which Jeff considers trustworthy (Off) with a component which Jeff does not consider trustworthy (Def). I think Jeff is just making the reasonable point that when two players are roughly equal according to this metric, the one deriving much more value from the more trustworthy part of WAR is probably much more valuable in reality.
I would not accuse him of staking everything on what his eyes tell him, he’s not an old-school scout or something, just a writer who is making a pretty benign point about adding.
I posted this below as well, but I’ll put it here, too:
Kyle, please tell me how Passan’s comment starting at 8:28 of the podcast is anything but him saying Alex Gordon shouldn’t ever be ahead of Trout in WAR. Because of what Passan’s eyes tell him.
“As great as Alex Gordon has been this year, and I say that having watched him literally every night … I could not see how Alex Gordon was worth more than Mike Trout.”
Indeed, I’m as frustrated as you to have to point this out. But that’s my whole point. Since Passan, who’s supposedly stat-friendly, is saying this, that’s a big problem.
@ralph Are we listening to the same thing
Just because he mentioned that he watched a lot of Alex Gordon this year, suddenly that is his entire argument, and nothing else he said matters?
Thanks for checking back, Kyle.
I think the part of the podcast I quoted tracks what I think of as Passan’s WAR thesis, which I’d summarize as follows:
Because Passan doesn’t trust the defense part of WAR very much, he won’t be happy with what WAR tells him until he almost entirely agrees with how WAR leaderboards look.
That’s not an unreasonable position, since it’s what drives people to improve things like WAR, but “watching Gordon literally every game” is the only “evidence” of WAR’s shortcomings he brought up in the podcast. This greatly disappointed me, because it really limited the opportunity for dialogue.
I acknowledge Passan paid lip-service to the idea that there can be single-season defensive flukes, but that’s all it seemed to be — the least possible acknowledgement before he went to on to criticize defensive WAR for not conforming to his expectations.
What was your takeaway message from Passan in this podcast? If you could point me to a timestamp or two that made you feel pretty good about Passan here, I’d appreciate it.
There is perhaps a point to be made here about the error in the equation for WAR. There is indeed measurement error, and I think from what I’ve seen in these discussions is it’s highly likely that there is more error in something like defense (and maybe positional replacement value) than there is in hitting. If we were to apply error bars to Gordon and Trout, Gordon’s error bars would likely be much greater because so much of his value is based on his fielding metrics and his position.
However, you almost never see a sabermetric discussion where sources of error turn into something quantifiable. Maybe the topic of measurement error really is missing from the whole field of baseball research. Even in articles I read about projections, where one would think there would have to be some concept of plus/minus for each statistic, alas: we are always presented with a single number.
I wonder if, for example, Tango or others have put work into quantifying the potential error in measurement of the linear weights values? How much could we really be off by?
In fields which are more rigorous than baseball research, they seem to get this a little better.
ESPN is in a position to educate the public and their leaderboards are BA, HR, RBI, Wins, ERA, Saves, WAR in that order. Same thing for Yahoo, expect replace WAR with strikeouts. We’ll never reach the mountaintop lugging up these behemoths, they need to help us out.
don’t forget errors 🙂
Jeff quite clearly believes that it is a mistake to use WAR to compare players. He does not want WAR on a sortable leaderboard. Does Dave agree or was he unwilling to debate this point?
It sounds to me like Jeff Passan had no idea what he wanted, just that WAR shouldnt be a stat used anymore. I thought he was going to offer some actual insight besides “well no one thinks this guy is worth 5 wins”. Very disappointed in his side of the argument. I’m not too interested in a guy who has no possible solution to a problem he brings up
He does not have an alternative. That doesn’t mean his criticism isn’t valid or that it is not worth voicing.
WAR is the best we’ve got but it still isn’t reliable. Fine. Comparing players only via WAR values is not a good way to go about discussing value, especially when players are separated only by decimal points. Fine. But by presenting WAR as this singular value for a player, it has the potential to be treated as more than it is. Jeff didn’t have to go very far to find recent examples from sportswriters (Poz and pizza cutter) who were writing articles using WAR as that kind of end-all stat without apparently noting that that isn’t what it’s useful for. At the same time that we’re seeing Dave writing that WAR only is good for separating players into tiers, and that we shouldn’t look at detail more than a half a win or more, you have to admit that it’s still very tempting to make WAR-based arguments about players, and to do reasoning based on those numbers.
I have agreed with most of what you posted in this thread Kyle but…
I think the main response to this is that, up until about 10 years ago most sportswriters were using RBIs as the sole determinate of player’s value. Most sportswriters are frankly idiots who I wouldn’t let calculate a tip much less evaluate player value. It is not worth wasting energy worrying about what they are going to write.
Pasan’s criticism seems like complaining about his 2 year old spills a 1 %of a glass of water when drinking from a cup, two weeks after he was just spilling 90% of the water. It is just inappropriate and shows a total lack of appreciation for the situation.
WAR is such a colossal improvement over the type of analysis that was, and frankly still is considered standard in many circles. Even though it is very incomplete and flawed, it is still basically perfect in comparison.
I have an extremely bright math savvy friend who is just a couple years older than me who just refuses to see how something like WAR could be descriptive. In his mind the only potential descriptive things are stats at least 70 years old. He has been totally brainwashed.
I try to get him to understand that if you have say:
1) Two identical pitches
2) Hit with identical speed and placement by different batters
3) One center fielder makes a great catch and the other doesn’t
You can still grade those hitter performances equally and call your system descriptive. He just cannot make that leap. “But what happened on the field was different!”. Yes…but what happened on the field vis a vis the hitter was not different. The performances were the same. How he cannot see that is beyond me.
I think Jeff and Dave care more about what sportswriters write in their articles because they are themselves sportswriters and this is their craft. What their peers are doing is of interest. If we take it that way, Jeff Passan is making a point about his peer group and that’s why it’s important to him and to Dave.
I think you’re the first one to bring up RBIs in this discussion, which is yet another sign of progress :). It’s nice that we have an online community where people talk in depth about WAR and it’s strengths and weaknesses. Most folks on Fangraphs have moved well past the quirks and flaws of the older statistics, and that means it’s time to welcome discussion about the new ones. Sabermetricians have made a ton of progress on measuring value, it’s true. And that is why now is the wrong time to become dogmatic about what we have. There’s still work to be done.
I just don’t understand why people in this community are so sensitive to criticism of WAR. There was a commenter earlier in this thread who trotted out the old “what his eyes tell him” nonsense. Good god, are we seriously still fighting these battles with the old school – like this is a FJM article from 2004 or something? Jeff Passan isn’t old school. He’s new school – and his criticism is reasonable.
Is responding to Jeff Passan’s article by bringing up the same old pro-sabermetric slogans really helpful? No. Responding with new discussions about WAR is helpful. I love that Jeff Passan’s original position led to Jeff Zimmerman’s excellent article about Alex Gordon’s positional value. I really learned something there. I’ve also enjoyed Dave’s written rebuttals to Jeff. This isn’t an us vs. them, it’s an us vs. us. And it’s been some of the most interesting discussion I’ve read on this site all year.
Like I said I more or less agree with everything you have said. I am just advocating for the devil and trying to describe why “Wah WAR isn’t perfect” can come across as pretty crass in 2014. There is probably only 50% market share for advanced stats among well educated fans, and lets not even talk about the general public.
Kyle, please tell me how Passan’s comment starting at 8:28 of the podcast is anything but him saying Alex Gordon shouldn’t ever be ahead of Trout in WAR. Because of what Passan’s eyes tell him.
“As great as Alex Gordon has been this year, and I say that having watched him literally every night … I could not see how Alex Gordon was worth more than Mike Trout.”
Indeed, I’m as frustrated as you to have to point this out. But that’s my whole point. Since Passan, who’s supposedly stat-friendly, is saying this, that’s a big problem.
It felt like Dave was trying hard not to be argumentative, and kind of pandered to Jeff’s comments. I wasn’t hoping for a heated argument or anything, but let’s be honest here, if we made some of the comments Mr. Passan made, during one of Dave’s chats, he would’ve rightfully ripped us apart.
Agreed.
It seems to me many of the critics of WAR assume that those who utilize it view it as flawless and complete. Now part of the problem is the lack of an explanation from those of us who utilize WAR that it is an approximation and that some of the components (especially on the defensive side) are problematic. But on the flip side, these critics are holding WAR to a higher standard than other metrics, formulas, and statistical indicators. But maybe that’s the case with newer stats and formulations. Maybe non-traditional (for lack of a better term) stats and formulations just require more explanation, whether we like it or not.
I also think there are people out there who view sabermetrics as some religious cult, and anything that comes out of sabermetrics is viewed nefariously. For these people, they don’t judge things like WAR on the merits. So the more baseball media members who aren’t perceived as sabermetricians or who don’t fit into the sabermetric box who understand the legitimate usefulness of things like WAR, and espouse its virtues, the better.
I just got around to listening (I apologize to everyone I let down) but regardless; can I get an estimate of how excited I should be that my comment on Dave’s article was referenced in the discussion?
http://www.fangraphs.com/blogs/a-discussion-about-improving-war/#comment-4715954
6 EAR (excitement above replacement)