A Few Thoughts On Evaluations
Over the last few weeks, I’ve been kicking around a few thoughts in my head, and so today, I’m going to try and turn them into a cohesive post. I can’t guarantee I’m going to succeed, given that I’m writing these sentences before I write the rest of them, but after pondering these in my head for a while, it’s probably time to put them down on virtual paper and get some feedback on the ideas presented.
The primary genesis of these thoughts are spurred by the fact that, with three weeks left in the season, the battle for the second AL Wild Card is being contested by the Twins and Rangers. Entering the year, our Playoff Odds gave the Twins a 4.6% chance of making the postseason (based on a 74-88 projected record), and the Rangers a 3.5% chance (with a 73-89 projected finish), and now it seems quite likely that one of the two is going to end up playing the Yankees in the Wild Card play-in game. The American League as a whole has been pretty weird this year, or if you take a different perspective, our preseason projections have performed poorly in forecasting AL team records this season.
Rangers fans — or a segment of their fanbase who use Twitter, anyway — have been particularly loud in their objections to our evaluation of their team, and understandably, 140 games of their team winning games at a .525 clip has reinforced their belief that our methodology of team evaluation is incorrect. Or that I personally have a bias against their team. Or some combination of the two.
From my perspective, though, the disconnect is mostly just a philosophical choice, and is the same choice that drives a lot of the disagreement around many of our less popular evaluations. Primarily, we evaluate players and teams by their inputs, not their outputs.
One of the primary goals of analytical research has been to attempt to isolate individual performance, as many of the more traditionally accepted measures of player value were based on taking a metric that involved contributions from many players and ascribing them to just one person. This has always been the core problem with pitcher Win-Loss records or RBIs, and has been one of the main reasons we’ve advocated for moving away from ERA as the primary way to evaluate a pitcher as well. These numbers are the result of a lot of variables working together, and we’ve generally preferred to move away from those kinds of metrics and towards things that focus more on measuring just the contributions of one player.
It’s a generalization that isn’t true in every case, but I tend to think that one of the main differences between how we often see things around here — versus how the mainstream baseball fan sees things — is that we tend to try to construct the whole picture from individual pieces, while traditionally, evaluations have more often been done by taking the whole and attempting to deconstruct it into individual pieces from there. We start with inputs and try to build up; I think most mainstream fans start with the output and try to tear down from there.
I do think starting from the building blocks of what have historically shown to be sustainable skills is probably a better way to evaluate player and team performance, because metrics that take a number of inputs are difficult to unravel after the fact. This is why, despite all the focus on analytics in baseball over the last 20 years, separating out the credit that should go to a pitcher or fielder on any single defensive play remains difficult. When you take a bunch of inputs, put them in a blender, and then try to identify all of those individual components by simply looking at the final product, it’s not so easy to tell what happened to get to that final product.
But there’s also a cost to starting from the inputs; we don’t know all of the variables, and by building models that are based solely on the ones we think we do have a decent understanding of, we leave things out. Including things that have pretty large impacts on the output metrics, especially at the team level. The most notable thing that gets excluded from using input measures? The order in which events occur.
For example, let’s take Joe Kelly. By ERA, Kelly was absolutely miserable in the first half, but has really turned his season around since the All-Star break. But take a look at how different that picture looks if you strip out the effects of sequencing and just look at how opposing batters have done against Kelly in the first and second halves:
| Kelly | BA | OBP | SLG | wOBA | ERA |
| 1st Half | 0.268 | 0.338 | 0.418 | 0.329 | 5.67 |
| Second Half | 0.271 | 0.338 | 0.427 | 0.334 | 3.45 |
Based on simply the number and type of baserunners Kelly has allowed, his first half and second half performances don’t really look any different, but the order in which those events have occurred have had a dramatic impact on his results in differing directions, which drives a 2.22 run difference in ERA between his first half and second half results. If you’re using outputs to look at Joe Kelly’s season, you’ll see a huge change, but because sequencing is not an input that has been shown to be something that players or teams have sustainable control over, it doesn’t end up in the kinds of metrics we favor around here.
In reality, the FIP/ERA argument is the same as the BaseRuns/Pythag/Actual Wins arguments, and also drive the difference in our perception of the quality of the Twins and Rangers rosters. By the inputs that make up a team’s win-loss records, neither the Twins nor Rangers are actually having that great of seasons.
| Team | Winning% | Pythag% | BaseRuns% |
| Twins | 0.521 | 0.495 | 0.442 |
| Rangers | 0.528 | 0.480 | 0.479 |
When looking at the inputs without regard for the order in which they’ve occurred, the Rangers have played like a slightly below .500 team, while the Twins have played like one of the worst teams in baseball. Both are in the Wild Card race because they’ve maximized the value of their positive plays by stringing them together in optimal fashion, with the Twins scoring far more runs than we’d expect from a team with a .248/.304/.400 batting line, and the Rangers simply scoring their runs at the right time, turning a -27 run differential into roughly the same record as the Astros (+105) and Giants (+78).
At the input level, I think our ability to forecast player and team performance is pretty decent. It’s not perfect, certainly, but when we just look at individual performance, we do alright. When it comes to projecting the order in which those performances are going occur, however? We’re completely useless. Forecasts don’t even try to take sequencing into account, because we have almost no insight into predicting when events will occur. We can kind of get towards some decent predictions when we just say that, over some large period of trials that we should expect this type of thing to happen this number of times, but the order in which those events occur is a complete mystery.
But in baseball, the order of events is a big variable in output metrics, so there are going to be plenty of instances every year where simply taking the inputs we have a decent handle on and incorporating them together isn’t going to match the outcomes of what actually happens during a season. So, while outcome metrics have the drawback of being difficult to disentangle, input metrics have the drawback of being incomplete. It is, to some degree, a pick-your-poison situation, as both types of measures have strengths and weaknesses.
And the downside of mostly ignoring output metrics is that, when sequencing drives dramatically divergent results, we’re not going to see things the same way as those who are just looking at the outputs. And it’s basically a certainty that our models, based on transforming individual forecasts into a projected team output, are missing things. There’s no way we’ve solved for every important variable at this point, and in the future, the models will account for things that we aren’t accounting for now, and we’ll look back at our current tools and wonder why we missed things for so long.
But I think the preponderance of evidence suggests that evaluating players and teams by their inputs is generally more accurate than simply going with the outputs. As far as anyone can tell right now, sequencing really is mostly just randomness, and while it can have a big impact on the results, it isn’t really something that can be predicted ahead of time.
So, Rangers and Twins fans, I promise that we don’t hate your teams, and aren’t out to try and rain on your parades. We just don’t put nearly as much stock in output metrics as most people do, and your two teams happen to be this year’s examples of the power of the sequencing variable that we don’t really even attempt to take into account. Last year, it was the Orioles, and next year, it will be someone else. The order of events has a big impact for a few teams every year, and so every year, input-driven analysis is going to look silly.
But it’s the cost we’ve chosen to pay in order to attempt to isolate individual performances a bit more accurately. And I think it’s a price worth paying overall. This doesn’t mean our analysis is perfect, or that we never get anything wrong, but when the difference in results is simply driven by the order in which events occurred, I’m okay with that, and I don’t think an odd distribution of events means that the projection that assumed a normal distribution was incorrect.
At this point, all we can really do is gets as close as possible to projecting the quantity and quality of the events that will occur. We’re nowhere close to being able to predict the order of when those events will occur, and I’m not sure we ever will be. We’ll get closer, but sequencing is likely always going to be a variable with enough randomness to make projecting the order of events all but impossible.
Dave is the Managing Editor of FanGraphs.
Might I suggest that these evaluations provide a range of wins totals instead of a single value.
I’m guessing the value of wins will just be normally distributed around the expected win value – so it would show the Rangers or Twins getting have a 3% chance of getting 90+ wins but also a 3% chance of getting 60 wins
While this is true, we still never have discussions on kurtosis or heaviness of tails. The emphasis on mean is too simplistic.
There is something more to this. There are rosters that may project 80 wins but have a low ceiling even if lots of things work out while another team may project 76 wins but 90 wins if some things work out right. Perhaps floor and ceiling need to be better represented when doing future projections.
Moreover, I think there are more things we miss than just sequencing from projections. I think we don’t fully grasp the value of defense as a part of run prevention. Likewise, I think we also underrate relief pitching and other things like speed, youth and depth. While some of these are accounted for in some way at the player level there is probably some team effect that gets missed. And that’s before we even get to more esoterica like manager effect, ability of team to add salaries mid season and quality/quantity of prospect currency in the organization. All these things play a role in team outcome. And if we’re limiting analysis to only individual player data, we’re leaving a bunch of relevant things out or underrepresented.
I’m waiting for Dave to personally assure every single fanbase that he doesn’t, in fact, hate them personally. I think we’re up to 3 now, with the Orioles, Rangers, and Twins; if there are others that I’ve overlooked, please contribute.
Oh yeah definitely the mets, I wonder how a mets vs national similar analysis would look like
Probably a lot like this piece Jeff Sullivan did:
http://www.fangraphs.com/blogs/where-did-the-nationals-go-wrong/
I don’t think people feel like he’s biased or hates their team because of the projections, though. It’s the self-assurance he has in his projections when talking about these teams. At the end of May, it wasn’t just that he didn’t think the Rangers good run was sustainable, or a sign that they could be decent…
It was that they “had no chance” to be good. This year OR next, there was no way they were going to compete. So while a mea culpa that really boils down to him saying they got lucky is nice, I think the bigger take away is to always remember that it’s not just randomness that has an effect on these projections, there are truly things that can’t be measured in them.
So if he cares about people feeling like he hates their team, he should probably remember that what he knows also needs to be tempered by what he doesn’t.
This exactly.
He was 100% dismissive of the Rangers in May. There was no “well, if this, or if that”.
My feeling about that is it’s just too much work. Especially in the age of the second wild card, literally any team can win more games than their talent would lead us to believe they would and challenge for a playoff spot. There is an unspoken assumption in any discussion of any team that, yes, of course, they could be the lucky one this year, and demanding that someone states that contingency clearly at the beginning of any discussion is terrible and time-wasting.
To a degree, maybe. But there’s a difference between saying that he doesn’t think a team is riding more than luck, with the inherent caveat that they might remain “lucky”, and basically expressing the opinion that they’re trash with no chance of winning this year or next. Also, add in the little dig that the Rangers would be the only team “dumb enough” to give up their top 2 prospects for Hamels (who they got without touching their top 3), and it’s not too hard to understand where the perception of hate comes from.
Which, if he doesn’t care about, is fine. But then don’t whine about it when you’re wrong and catch flak.
What is the research on sequencing showing? Is it 100% luck or is there some attribute (contact rate, ground ball rate, speed etc) that increases yield on converting base runners into runs? Perhaps if Victor Martinez is on 1B he has less likelihood of scoring than Jarrod Dyson would.
on the other hand, Jarrod Dyson has a much worse chance of being on base in the first place, so maybe in the aggregate Martinez would still be expected to score more
Base running is an integral part of the inputs that go into the team projections.
But I assume it doesn’t consider things like “does this excellent baserunner bat in front of TTO-types who neutralize the effect of baserunning or in front of contact hitters who intensify it” or “does this batter’s GIDP tendency reduce the value of the preceding batter’s ability to reach 1B” or “are these linear weights adjusted by expected lineup context? (e.g. a 3-hitter will benefit comparatively less from BBs than a 1,2,4, or 5-hitter would)”. Maybe we need to start considering team environment as well as run environment.
right, and it doesnt even matter specifically what kind of batter bats behind a speed guy, just the fact that speed is a consistent skill that has wide variation in the population. A single behind dyson will have a higher average run value than one behind martinez. so, is that skill consistent enough and does it have a wide enough variation in the league to make a difference in team specific run values? I think its a good question. The Pirates, for instance, have 4 batters with career speed scores in the top 10 % of the population! Do league average linear weight values for singles for bucs hitters accurately predict the expected runs scored?
True and true, but lineups vary from day to day, and there are probably too many variables to accurately predict this and translate it into a projection. That, and you’re getting into predicting sequencing (that player A gets on base before player B makes contact or also reaches base). Projections are good at estimating a total rate of things, but not when they happen or how often they happen together in the right runners and outs condition to result in increased run scoring.
Nice try, Dave. We know you’re just trying to cover up your system’s failures!
Imo you’re 100% right about inputs being more relevant….but i’m not sure we necessarily have to ignore the rest.
We’ve had so many great analytical minds working so hard honing and improving in this one direction that i wonder if maybe we’ve simply neglected trying to discover new statistical insights in the other direction.
It’s been a while since i’ve read much about any attempts to find any kind of team effect that could possibly improve our input based projections. You guys did such a good job proving that most of that was just random variation that i wonder if we just gave up trying to do anything in that area anymore. Given how our understanding of the inputs has improved so much since even a few years ago i’d think we might even be able to try better experiments now.
because, on the surface at least, it does seem we’ve had a number of teams in recent years who have defied the projections with at least some measure of consistency….at least enough to maybe think there’s something worth looking into with a fresh, open eye.
i also think in general, especially with the vast increase in the amount of modern data we now have, that we might want to regularly put even our most basic assumptions back through the statistical gauntlet to see if all the new data is strengthening or weakening the basics….and to make sure that what we thought permanent truths weren’t just temporary phenomena. This might be an especially good time to doublecheck since the run scoring environment has changed so dramatically over the last decade.
All good and true. But it’s tricky to take “defying the projections” and ruling out crazy luck, instead of a missing piece of the puzzle.
very very difficult.
but i mean – what are we here for really?
If i were the GM of an mlb team and saw other teams besting the projections with any type of regularity i’d be spending as much money as i could to reassure myself it was just luck. I don’t think i’d be comfortable just throwing up my arms and saying “baseball!”.
especially since like i mentioned it seems that this might have become just a neglected area of study rather than necessarily a lesser important one.
The Orioles are a good example of exceeding expectations a few years ago when they made the playoffs and underperforming this year with a better run differential. I think we can rule out some teams doing this with regularity and attribute this to randomness of the schedule and sequencing. And sequencing not only from the team you are studying but also sequencing from the opponent.
I think it’s still very much up for debate whether we’ve seen any teams “consistently defy their projections.” At least, I don’t think there are enough instances of this to draw any meaningful conclusions.
Count me in this camp, reminds me of the narrative that some manager’s teams outperform their pythag/base runs consistently, but it just doesn’t hold up upon review.
I think Dave hit the nail on the head in saying that next year sequencing will matter a lot to a few teams and predicting those few is impossible at this time. While we beat up the projections for the AL, the NL is almost dead on, sans the WAS debacle.
absolutely very debatable – which is why i said “seems” – but my point is that we’re barely even trying to delve into this side of the issue, other than trying to dismiss it.
Damaso: as far as I can tell, no one has any idea how to approach doing so. It will take a Voros/DIPS-like insight to change that, I think.
Input- and output-based are both defendable. The approaches should be merged. Take the difference as the residual to be modeled. Sounds like a fun competition. I start by comparing the distribution of runs/inning to the expected distribution.
I am also curious if there is an offsetting benefit to underperforming early (as long as you stay in the race). Do Blue Jays still add Tulo/Price if they performed at first half baserun levels?
I think the more relevant question is “did Anthopoulos add Tulo/Price because he believed in the metrics that told him his team was better than the output up to that point?”
well, both inputs (baseruns) and outputs (runs) told him the team was one of the best in mlb.
I think there is an argument to be made that good managing (optimal batting order, pitching changes, bunting, stealing, etc) could somewhat explain sequencing, although I wouldn’t make it, and it’d be almost impossible to quantify.
Good managing and strong bullpens, among other things, have been suggested as ways to improve sequencing and outperform BaseRuns, but the fact that there’s so little period-to-period correlation between optimal sequencing seems pretty telling to me. Whatever skills or qualities there may be seem to be swamped by the noise.
This was a good piece by Dave, but almost too deferential. Just seems like a lot of fans refuse to believe that randomness can play a big role in game outcomes.
The Book did a study on batting order and found a small difference between optimal and sub-optimal batting orders. (I think it was, on average, less than a win per year, though it’s been a while since I looked at it.) This was, however, done with modeling based on outcomes (like a huge strat game), and so if there are psychological effects to various batting orders, it wouldn’t have been captured.
Nationals fans, take solace. For the remainder of the season you are projected to be the best. Inputs.
It’s stupid to even respond to this post, because it’s super trolly, but it’s worth noting that the Nationals are *not*, in fact, projected to have the best record the rest of the way, which in of itself is good evidence of projections adjusting to results. Probably not adjusting ENOUGH, but still adjusting.
No need to apologize, by any means. Comparing Texas and Minnesota with say, Oakland (-11 R differential) just goes to show that the standard deviation is high for every team on expected vs. actual wins. One run records play a part in this (Oakland at 16-31 seems particularly snakebitten), along with when you play a team (even bad teams have hot streaks) and when you score runs or sequence within games. This comes down to their are 32 teams, the top 3 or 4 have big separation and so do the bottom 3 to 4. The remaining 24 are usually quite closely clustered in terms of talent, which is also variable as is the sequencing. That’s why they play the games.
There are 30 MLB teams.
Sorry, was thinking football:) the comment still applies.
Question about run differential: Why does the teams record become a better predictor for winning% than run differential towards the end of the season?
Does that suggest that there is some sort of measurable way to sequence so that a team is able to reliable out perform their pythag? Or is it more likely that it is more of a function of how a team has changed throughout the course of the season; whether it is due to injury, trades, or other variables that would change a team’ make-up throughout the season?
One reason is that teams that have better records added good players at the trade deadline while teams with bad records subtracted good players. Or more generally, the rosters have changed enough that the historical performance (baseruns) is less predictive than the expected actions of GMs.
That makes logical sense. I think the way to analyze this is to see how run-differential ROS follows team’s current win% vs pythag.
That way you can see if that correlation is because a team has actual changed, which would mean the run-differential ROS would be different than season to date, vs if they actually have a specific way to sequence, which would show up if they continue to over/under-perform their more constant run-diff.
As a Twins fan, even the common eye can see the Twins aren’t very good even though they keep winning. This analysis on this site is beyond my level of mathematical intelligence, however, as I watch my lowly beloved Twins I wonder how they do it? What metric am I missing on fangraphs that tells me how they win? Now I know, there isn’t one its random luck.
I think Dave’s point is that it isn’t *just* luck – there are components that we haven’t identified yet. Maybe it is luck, or partially luck, but we’re not throwing up our hands and saying “oh well!!!”
I’ve been wondering about a different divergence – but perhaps it’s really the same thing. Listening to ex-ballplayers doing play-by-play, I usually get worn out with all of this discussion about the mental aspects of playing the game – ballplayers seem to be obsessed with “experience” and “mindset” as determining factors that underlie performance, rather than “skill” or “talent” or “ability”.
What if “experience” and “mindset” and other intangibles are inexact code words for whatever processes underlie a ballplayers (limited) ability to influence sequencing?
Isn’t that really what the idea of “clutch” is? A greater-than-average ability to influence the sequence of positive events in a game? I know that “clutch” is not something anyone has ever been able to reconstruct from data, but perhaps if or when someone figures out how to do so, it will lead to breakthroughs in analyzing sequencing…
Somebody else will likely be able to explain this better, but as far as I know there is very little predictive value in “clutch” statistics, which you would not expect if it was a repeatable skill. Maybe you have a point if you have a problem with how things like WPA and leverage index define clutch?
I understand that we have found little predictive ability in clutch statistics. My point is that if/when that changes; if/when we ever figure out how to separate the noise from the signal and discover some little tiny bit of predictive ability – then perhaps we’ll the root of that ability will have something to do with those factors we currently speak of as “intangible”.
Basically this boils down to “There are things in baseball that we can’t predict and that’s why are predictions are wrong”
You can’t predict baseball, and you know what? That’s awesome! Imagine how boring baseball would’ve been had the projections and predictions in the preseason were 100% right 100% of the time. That’s lame! Long live random variance!
But sequencing isn’t just randomness. Outcomes of one event can and do make subsequent events more likely.
Sticking with Joe Kelly as an example, in his May 9 game against the Jays, Kelly had a throwing error, a wild pitch and a walk in the first inning. And you could see watching the game that it rattled him. The second inning started with a walk and an interference call on Kelly. He was so flustered that the large number of walks he gave up after that seemed pretty predictable.
I’m sure someone has studied sequencing, no? Is a walk more likely than you’d expect to be followed by another walk? Is a perfect inning more likely to be followed by another? I’m not at all convinced that it is basically random.
We have found that streaky-ness is basically non-predictive, if that helps any. (Though I think they were looking at it on a game-by-game basis rather than event-by-event.) IIRC, good pitching performances (for a pitcher) are only slightly more likely to be followed by similarly good pitching performances (for that pitcher). And good hitting performances in one game had no predictive utility (beyond putting them into the projection machine and updating accordingly).
Yes, I’ve seen studies of streakiness on a team, game-by-game level. But I haven’t seen one at the individual level within games.
But managers base their decisions within games on an assumption of streakiness. A pitcher is pulled because “he just doesn’t have it today.”
MGL looked at this, using the framing of “when managers leave starters in (vs pull them), do they pitch better or worse than you would expect?”
It’s a really tough question because you can’t know how a pitcher who got pulled would have pitched, but he found that, overall, pitchers allowed to stay in the game don’t pitch any better than an estimate based on talent plus the times-through-the-order penalty — which is to say that managers’ assessments of who is “on” or “off” during a given start don’t lead to any decision-making insight.
I, for one, don’t want to live in a world where baseball predictions are always accurate.
You don’t welcome our new insect overlords?
2 things.
1. ERA – FIP differences aren’t about sequencing. Or at least not only about sequencing. So part of the differences have to do with different ways to evaluate pitching. Obvious differences are guys like Glavine or Buerhle, where over a career there’s a 10+ WAR difference. BR credits them differently than Fangraphs, though both focus on inputs rather than outputs.
2. Sequencing is a big part of it, clearly, and clearly there’s a lot of luck in that. As has been noted, bullpen and managing can perhaps impact sequencing a bit. Year to year correlation may be hard to establish in large part because personnel and inherit volatility of relievers varies so much, but it’s not hard to see that teams like the Yankees and Royals can more optimally sequence late inning high leverage situations than, say, the Tigers or the 1st half Blue Jays.
How to make the projections better to adjust for that I’m not sure, but I would say it’s more than just pure noise.
I’m a fan of the cardinals and orioles, two teams that have benefited from sequencing in the last few years. I don’t really understand why some fans get so offended when it’s pointed out. So the baseball gods are smiling on you right now? Great, enjoy it. You know they are eventually going to dump on you. I guess people interpret “your getting lucky” as an insult and indictment. I always have just thought of it as being fortunate.
The additional knowledge this site provides also helps my fandom – I knew the cardinals were not really a 105 win juggernaut this year, so the recent 2-8 streak wasn’t a huge shock.
What exactly goes into baseruns?
Pretty amazing that it suggests that both the A’s and Royals should have .517 winning percentages but one team is 84-58 and the other is 61-82.
Per the Book – “Many different versions of Base Runs have been introduced by several sabermetricians, to accommodate different datasets and philosophies. However, all Base Runs formulas take the form:
A*B/(B + C) + D
A represents baserunners; B represents advancement of baserunners; C represents outs; and D represents guaranteed runs (usually just home runs). Thus, Base Runs adheres to the true identity that Runs Scored = baserunners * (% of baserunners that score) + home runs. B/(B + C) is an empirical estimate of the percentage of baserunners that score, and is the main source of potential improvement for Base Runs formulas.
Base Runs was designed to adhere to real-world constraints on runs scored, and succeeds in doing so to a greater extent than other simple run estimators. Base Runs recognizes that a team must score at least one run for each home run it hits and that the number of runners scored cannot be greater than the number of baserunners.”
I can’t seem to find the specific method used by FG, but I think they use team wOBA and a table format to derive run expectancy. I could be wrong, anyone else want to help?
Isn’t it a little of, well, everything? No one foresaw the Twins owning Chris Sale this year and going 13-6 against the White Sox (the Twins are 61-62 against everyone else in the AL/whoever they’ve played in the NL). Even at 61-62, the Twins are out performing what were a lot of people thought they’d be, but it’s closer to where people thought they’d be.
I also think back to, as a Sox fan, the 2000 Sox, who went on a 17-2 streak in June. They were 78-68 (including playoffs) outside of those 19 games. Still good, but by no means great. So how do we factor in/figure out the team that plays .800 baseball for a few week stretch?
Same with the team that has three or four guys give career years all at the same time.
I think they key is for teams like the Rangers and Twins to not fool themselves into thinking that they’re going to “do it again!” in 2016. It’s one thing if you’re the Dodgers or Jays; it’s another if you’re those teams. The Twins and Rangers should be looking to upgrade from the Tommy Milone’s of the world in 2016.
I think you have already seen this. The Twins didn’t make a splash at the deadline and refused to risk Berrios. They acted like a team with more “growth” assets than a team who believes in 2015. Ditto for the Rangers, whose noteworthy deal will help through 2019 and complements a (hopefully) Darvish-led rotation in 2016.
I have a major, major problem with this, Dave, because what Fangraphs continues to do is misrepresent the value of its analyses in some areas. For example, you present single season UZR numbers despite the fact that we know the samples are too small to mean ANYTHING. Also, you rely on FIP for your WAR component when we cannot be sure that pitchers do not have control over babip. And while you and the staff do give qualifiers, you list these numbers on player pages and on the front page with NO qualifiers. And then people take these numbers and cite them as gospel. Here you are using BaseRuns to disqualify teams success when you dont know if there are other factors BaseRuns is missing in its analysis. Yet you and the staff cite it as gospel! You even say the Dodgers are far and away the best team in baseball based solely on the premise of BaseRuns. You need to give qualifiers and admit that there are some things we don’t know WAY more often than you do!
“single season UZR numbers despite the fact that we know the samples are too small to mean ANYTHING” – we “know” no such thing.
Actually, we do. He is alluding to the fact that three years’ worth of defensive statistics to form UZR equate to about one season of offensive data.
I believe that’s only a critique of the usefulness of the metrics to estimate a player’s “true talent” (i.e., their predictiveness). That’s a separate issue from how well the metrics estimate the defensive value a player has actually provided in a season or partial season. For sure, defensive metrics are fuzzy even for that purpose, but including them in models is better than omitting them at this point (i.e., they make WAR-based models better correlate with actual wins).
A year’s UZR does tell us what a player did that year. What it doesn’t tell us is the players actual defensive talent level, but it doesn’t make the yearly info worthless. It still tells us how the player did that one year.
Kind of like how one year’s offensive performance doesn’t tell us what a player’s offensive skill really is but does tell us how he performed offensively for that one year.
The worst thing fangraphs did when explaining defensive metrics was reminding us that one year worth of data doesn’t give us the whole picture of a players true talent level. Now everybody and their brother uses that to dismiss annual defensive metrics as meaningless, which they arent.
UA – Actually, small sample size UZR does NOT tell us what actually happened. The fact the you and many others think that is one of the greatest failings of Fangraphs.
Evidence please that single season does not tell us what did happen ….
I’ll wait.
Its an exceptionally straight forward statistic, was the ball in the zone, did the play get made, nothing to subjective there. I think that it is silly that people here can’t accept variation in Y2Y defensive performance, yet offensive variation is normal.
PH – you are wrong in your framing of this. It is your side who must prove that UZR provides reliable data in small samples, and your side have not done that! What are you waiting for?
For example, if I come out with new scientific theory, it is not up to my peers to prove to me that this theory is wrong. It is up to ME to prove to my peers that my theory is testable and observable.
You can’t come on here and say single season UZR is reliable unless you prove to me that it’s not. It is up to the creator of the formula to prove that it is reliable.
You do understand that the expressed intent of the article is giving qualifiers and admitting there are some things we don’t know?
Yes, but in practice these qualifiers are not often stated in chats, articles and player pages.
You seem like the person who sues McDonald’s because it didn’t sate on the coffee cup that the contents were hot. At some point, you have to assume that your audience has a decent knowledge of the data that you are using for analysis. I for one don’t want to see a chart or table with 15 foot notes on it stating all of the disclaimers.
Ha, funny you should say that. Because like most people you don’t understand that lawsuit. That McDonalds had a history of heating up their coffee SUPER hot. So hot, in fact, that the woman in that case was burnt so bad she had to have skin grafts. So let’s remember what the facts of that case actually were, not the water cooler story that’s not true.
I think the Rangers thing comes from the fact that Dave wrote that the Rangers had no hope of contending in 2015 OR 2016 and should sell Beltre, Darvish, and any player with value.
And that the Rangers would be the only team dumb enough to trade for Hamels.
Even the .480 winning percentage with Baseruns is better performance than Dave was expecting from the team. It was a very different evaluation of how much injuries cause the awful 2014 Rangers season and affected their performance this year.
Not that it justifies stupid twitter harassment, but the passive-agressive “I’m not wrong, your team is just lucky” post doesn’t exactly tell the whole story. The current Rangers team is a bit different than the early season projections and early season performance.
Deductive reasoning is virtually foolproof in all the stories! But maybe that’s the point…
In agreement, 2 words
Mr. October
I’m a Rangers fan and Fangraphs stat head and long time reader and admirer of Dave’s work. I don’t think Dave hates the Rangers, and I do begrudgingly admit that run differential paints a bleak picture for the Rangers, but the Rangers were never really the Rangers until these last two months. I felt the Giants and Cardinals were teams that barely made the playoffs by two strong months in the years they beat us in the WS. But, their teams were strong as presently constructed and not a reflection of their injury laden seasons their first halves. Same applies to Texas now.
Holland and Perez are our real rotation stalwarts, and Hamels being added just further skews the roster toward a postseason capable team. We are still getting schtick for run differentials and metrics that involved freaking Wandy Rodriguez and Ross Detwiler starts with Neftali Feliz and Kyuji Fujikawa planned for big roles in the bullpen. I refuse to believe there would even be a race right now had the injury bug not caught the rotation yet again (and we still miss Darvish). I am proud of this team, without our best player, Darvish, they may still take the AL West. Thats pretty amazing to me.
Firstly, I appreciate this more introspective slant some of your recent articles have taken. And I think you’re a fine writer.
But. You say at one point that you’re pretty good at predicting some of the inputs; however, as this season has shown, those inputs don’t appear to be terribly useful at predicting wins (one of the things we as baseball fans are interested in), as the random chance of sequencing has drowned out all that work. If random events mean predictions are way off, then perhaps we all ought to be much more honest about that.
I can predict at least one response – “we are honest, we don’t use these predictions as the be-all and end-all”. But the problem is, you do. We all do. How many arguments (especially online) have someone weighing in with “well, player X is predicted for 3.5 WAR this year” and expecting that to be the capper? How many pre-season prediction posts are there which just go “here’s ZiPS, and…er…”? It’s used in place of analysis by so many writers, here (both articles and comments) and elsewhere. I don’t read the SABR journals, so this might go on all the time there, but I would much rather read articles after the season which measure various predictions (team, player) and discuss what happened, and why the predictions were out (if they were). Because I keep seeing ZiPS and Steamer and I keep seeing players and teams wildly under and over-performing them. Why? If randomness causes too much noise, why are we still sticking with these predictors?
I love that there are stats which accurately measure what people actually did, and it allows players who don’t hit the flashy box score numbers to get their due. And that’s on sites like this and authors like you, so well done. But the other side of it is…trying to think of a good way to describe it…writers getting annoyed when reality didn’t match their predictions. BaseRuns seems to be a common one for this, and people who spend a lot of time tinkering with this particular formula occasionally forget that wins and losses are still kinda important. Let’s be honest after the fact about what the predictions missed, and why.
I really enjoy the FanGraphs site and I make it an early read almost every day. The insights, especially individual player breakdowns with GIFs, are remarkable and educational.
Still, I have to agree with Mortimer – the defense of partial models in light of reality is an issue. It’s not only irritating but it also dampens additional research. Continually defining the combination of intentional sequencing, random variability and other unidentified factors as “luck” is a rationalization and not science.
Sabermetrics is now solidly on the map. I would like to see the practitioners advance from simply defending their worth “on average” to looking at what makes outliers tick (or try to improve on BaseRuns +/- 8 win confidence level).
For example, just a simple linear regression between either reliever WPA or Clutch v. Base Runs differential shows a significant correlation over each of the past two years. Isn’t anyone curious enough to dig deeper? The Angels outlie BaseRuns over a ten-year span and it gets written off as expected? Several starting pitchers show high performance (by WAR and WPA) and high “clutchiness” over 2, 5 and 8 year periods and nobody thinks its interesting?
I think one of the problems is the idea that baseball players are no more than Ping-Pong balls in a Powerball lottery. In fact, they’re humans. As such, sabermetricians have to realize they are practicing a social science. Techniques like cluster studies and stratified sampling could prove both entertaining and insightful.
Again, thanks for all of the good work. And thanks for the frank article.
Dave consistently argues that individual player psychology is not an important variable. His argument is that MLB is already self selecting for very confident/capable people. I disagree with this completely and think there are tangible differences in approach, stress compensation, and anxiety that could have very real effects.
Great points. “Clutch” isn’t a repeatable skill and yet…David Ortiz.
amen, dbminn.
we all fell in love with fangraphs due to their aggressive curiosity…..but there seems to be an unwillingness 5o push boundaries anymore.
Team chemistry and culture are real, but will never be quantified or usable as inputs. This is a good thing. If luck were the only unknown variable, then we might as well give up baseball and pass our time by flipping coins.
My objection is never to the projections; it’s silly to argue with the results of a formula. My objection, rather, is to the conviction in which members of this site (Dave specifically in recent chats) often dismiss what may be improbable, but still possible, because the projections say so. It was not too long ago (maybe around the trade deadline) where I recall Dave completely dismissing the Rangers (and Twins) as possible playoff teams. Maybe I should take less stock in what is said in a chat environment, but I still found it odd that he would completely dismiss them, especially when the AL has been so lousy this year with nearly every team having a legitimate shot the entire year. That’s why you have “hindsight” posts like this one. This post wouldn’t have been necessary if he would have taken a different approach back in July and admitted that playoff odds in the American League were a bit of a joke. When you’re in a situation where any team that goes on a short 5-game run puts them back at the top of the standings, playoff odds are useless to me.
Yeah it’s a little bit disturbing when playoff teams are regularly dismissed as garbage or merely lucky when other teams that have been fighting to avoid last place all season are more revered and regularly considered as merely unlucky.
There is a point at which winning has to have more meaning that it’s being given credit for. It’s the reason that the baseball season is so long. Results do matter. And some teams demonstrate that there may be a skill to winning. Not saying teams can overcome marginal talent to win regularly…. but I think some teams can erase base runners better than others; and there are some teams that can cash in base runners better than others. And that maybe we are conflating randomness of “sequencing” with efficiency.
What you are saying is true of what did happen, but has been shown to not be a predictive repeatable skill, thus suggesting that it is indeed cluster luck/timing.
“Suggestion” being the key word. Dave, in his written words, takes projections and formulas as rigid certainties. A little humility would serve him well.
What ought to be noted on the ROS numbers is something like “Hey, we have no freakin’ idea what’s going to happen because…well, baseball. But if you punched a bunch of inputs into a computer, here’s what the computer thinks is going to happen”.
Funny though, I don’t see any such disclaimer.
LOL
How about the fact Cole Hamels and Derek Holland are now pitching instead of Nick Martines, Wandy Rodriguez, Ross Detwiler, Matt Harrison, Phil Klein and Anthony Ranuado…who have combined to start 50 games for the Rangers this season.
Nineten paragraphs to say “Our forecasts are great? It’s just pure luck! Sequencing!” while failing to mention the massive upgrade in the SP.
We’ll keep laughing at #6 org guy.
Just wanted to say that, as an Orioles fan who is analytically-inclined, there is a reason that some fan bases feel “dumped on” by FanGraphs when their good season is disregarded as randomness/luck. It’s not because the fans are dumb (well, not most of them).
The problem lies in the over-confidence the writers here express sometimes in their analysis. You and I know that there is an unwritten disclaimer, but because it’s not written casual fans/readers can read statements like “the Twins are playing like one of the worst teams in baseball” and assume the writer is more interested in trolling for clicks than serious analysis. A truer statement would be, “based on our proprietary models, which we frequently change and don’t include all factors that lead to winning, we find that the underlying performance of Twins players has been below average and therefore it is more likely than not that they might regress in the second half, leading toward a final win total within X range.”
But the writers don’t say that. They are much more declarative, which leads to assertions that over-interpret the level of confidence they should have in their analysis, which can lead to fan feeling like their teams (and their happiness and their teams’ success) is being targeted.
I mean, I don’t want to repeatedly read sentences like your example on this site. Declaratory is fine for the most part when it’s not presented as some universal truth.
Agreed Jake. And that’s why most/all regular FanGraphs readers, who know the implicit disclaimers, don’t take issue even when the writers are criticizing the performance of their teams.
I maintain that most of the outrage comes from casual readers and fans who take the words at face value.* They are not wrong – over-confident presentation of analysis is a problem – but neither is FanGraphs – it would be burdensome/impossible to properly caveat every single claim. So, it’s just one of those things.
*Yes, there are uneducated fans who genuinely don’t understand, and very old-school fans who simply want to enjoy the game the way they always have. Not my cup of team but who am I to judge?
Given my limited understanding of baseball analytics, I believe the original purpose these metrics that are based on inputs was to arrive at measures of individual talent: repeatable skills that would accurately project future performance for *that individual*. Trying to sum those individual projections into an accurate team projection – especially given the dynamic nature of MLB rosters – seems an exercise in futility.
I think Cameron has all but admitted that team projections built on individual projections is going to produce unreliable results. Sure, given a large enough sample size over multiple seasons, you can get results that look pretty good, but trying to project one season’s outcomes is not yet within their grasp. But I hope they keep trying. At the very least, it makes for interesting, often heated, chat rooms and comment sections.
September 3rd’s issue of Nature has a nice review article on the science of weather prediction. It’s also on their podcast. I think the article may give you some ideas. Anyways predicting what will happen 162 games into the future will always be difficult.
Hey, Cowboy Sweet N’ Nasty, you’re off base on the whole coffee thing. According to the National Coffee Association, coffee should be served at 185 degrees. That’s definitely hot enough to cause a third degree burn. Mcdonalds has stated that they have never reduced the temperature of their coffee. I believe starbucks sells theirs at 175. She spilled it in her lap with sweatpants on. You could have cooler coffee than that and need skin graphs considering the circumstances.