The Year BaseRuns Failed

Around here, you know that we spend a lot of time working with metrics that attempt to strip noise out of results. Often times, we’re less concerned with what has happened and more concerned with what is going to happen, and these component metrics often do a better job of isolating either a player or a team’s overall contribution to the results, while removing some of the factors that lead to those results but aren’t likely to continue in the future.

At the team level, the most comprehensive component metric we host is called BaseRuns, which evaluates a team’s quality based on all the plays they were involved in, without regard for the sequence in which those events occurred. BaseRuns essentially gives us a context-neutral evaluation of a team’s performance, assuming that the distribution of hits and runs isn’t really something a team has a lot of control over. BaseRuns can be thought of as the spiritual successor to Bill James‘ implementation of the pythagorean theorem to baseball, as pythag strips sequencing out of the conversion of runs to wins, but doesn’t do anything to strip the sequencing effects out of turning specific plays into runs scored and runs allowed.

Historically, BaseRuns has worked really well. For the years we have historical BaseRuns data (2002 to 2014), one standard deviation was right around four wins, and the data appears to be normally distributed; 73% of team-seasons have fallen within one standard deviation, 97% of team-seasons have fallen within two standard deviations, and no team had ever exceeded three standard deviations. There have been years here and there where a team sequenced their way to an extra 11 or 12 wins, but they weren’t very common, and that was usually the only break from the norm in that season.

Until this year. Here is the year by year standard deviation in BaseRuns wins versus actual wins for every year that we have the data.

BaseRunsStDev

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

As you can see, the numbers are pretty close to the four win standard deviation in almost every year, with a slight increase in 2009 (4.76) and then a sharp drop-off in 2010 (2.83) before it re-stabilized right back around 4.00 or so again. Outside of that weird 2010 season where BaseRuns was amazingly prescient — every team was +/- five wins, except the Astros, who were +8 — one standard deviation has been between 3.6 and 4.8 wins in every season, more regularly in that 3.9 to 4.3 range.

Then there’s this year. The current standard deviation in actual wins versus BaseRuns is 5.7 wins, and that’s with only about 75% of the season in the books. Keep in mind, this is not a projection of how many games teams will lose versus their expected record over a full 162 game season; this is how many wins they’ve already deviated from in their first ~125 contests or so. The Royals and Twins have already added 10 wins to their ledger through the power of sequencing, while the Cardinals are nine wins up because of the ordering of their events. On the other end of the spectrum, the A’s have already squandered 13 wins because of their horrific sequencing of events, which would be the largest deviation from BaseRuns of any team in any season since we have the data.

Three teams have outperformed their BaseRuns by 12 wins prior, but that was 12 wins over 162 games. The A’s have underperformed by 13 wins in 126 games; if they sustained this pace over the rest of the year, they’d end up 17 wins off their BaseRuns total by the time the season ended. The Royals and Twins are both on pace to end up at +13 wins, no team has done in the previous 13 years we have BaseRuns data for. The Cardinals would end up at +12, tying the previous record for largest deviation.

If we extrapolate the current deviations from BaseRuns out to 162 games for all teams, then the final season standard deviation for the league would be 7.4 wins, nearly double what it has been in previous years. Of course, the teams that have defied BaseRuns so far probably won’t continue to do so at the same rate over the final six weeks of the season, but even if every team plays exactly as BaseRuns would expect over the rest of the year, we’re still likely to end the year with three teams at at least 10 games off of their BaseRuns record, and the Cardinals could easily make it four. Heck, the Reds, Marlins, and Rangers are all close enough where it wouldn’t be entirely crazy if any of them ended up +10 or -10 by the end of the year, so theoretically, we could have seven teams end the year with double-digit differences between their expected records and their actual records.

In the previous 390 team-seasons that we’ve tracked BaseRuns for, that has happened a grand total of nine times, or about 2.3% of the time. This year, we’re looking at somewhere between 10-20% of teams finishing to the very far edges of the spread. It’s been a weird year, to say the least.

Most likely, this is just the opposite of what happened in 2010, where sequencing didn’t really matter much at all, and the standard deviation dropped to its lowest point before shooting right back up to normal the next year. This is probably just a blip, the kind of thing that happens given enough years, and probably doesn’t mean that teams have figured out how to cluster their good performances together in a sustainable way.

But it’s worth keeping an eye on, at least. While no one has ever really been able to show that organizations can continually build teams that specialize in sequencing skills, it is certainly possible to imagine ways in which those skills could be found in the future. A team that could figure out how to build reliably dominant bullpens, for instance, would very likely beat a context-neutral estimation of their win totals on a regular basis, since the runs relievers allow (or don’t allow) have an outsized impact on win-loss records. But because bullpen performances are highly volatile, it’s a difficult strategy to execute on an annual basis.

We certainly shouldn’t discard BaseRuns as an effective model of team quality simply because this year’s results haven’t conformed to the expectations. Historical precedence still suggests that team’s can mostly only control the quantity of positive and negative events they’re involved in, and not the timing of when those events occur. But the timing can make a big difference, and this year, a lot of teams have sequenced their ways to very different outcomes than expected.





Dave is the Managing Editor of FanGraphs.

101 Comments
Oldest
Newest Most Voted
Damaso
10 years ago

interesting article at espn about the pirates apparently “cultivating the extremes” might shed some light here.

or might not.

Phillies113
10 years ago

“We certainly shouldn’t discard BaseRuns as an effective model of team quality simply because this year’s results haven’t conformed to the expectations.”

Many will almost certainly try.

King of the Byelorussian Square Dancers
10 years ago
Reply to  Phillies113

I did try, but failed.

Byelorussian Square Dancer
10 years ago

I’m being repressed!

Well-Beered Englishman
10 years ago

King of the Byelorussian Square Dancers is back!! Where have you been?!?!

Bart Simpson
10 years ago
Reply to  Phillies113

I can’t promise I’ll try, but I’ll try to try.

MikeS
10 years ago
Reply to  Phillies113

You mean this doesn’t invalidate all of sabermetrics, ever?

dom
10 years ago

are some teams better at defensive positioning when runners are on base? do good base stealers actually cause an increase in fastballs/hittable pitches thrown to hitters behind them? didn’t we just talk about how adding a good hitter to a good offense creates non linear gains in runs scored because offenses perform better with men on base? could any of these help explain the variances?

dl80Member since 2026
10 years ago
Reply to  dom

I would think they might, but why would teams be so much better at exploiting them this year suddenly? What’s changed?

Damaso
10 years ago
Reply to  dl80

could be that there’s simply been a large increase in all teams buying into the extreme tactics. no reason to think the change has to be slow and incremental.

Philip
10 years ago
Reply to  dl80

StatCast now exists in all staudiums. Though I doubt the entirety of baseball is efficient enough to cause this break from the trend.

Gabriel Gutierrez
10 years ago
Reply to  Philip

Maybe not all teams, but just STL, KC, PIT and MIN, who happen to be in the playoff race. A good analysis would be if there’s any correlation between OAK’s woes and the teams that have victimized them. Maybe teams that should’ve won less than BaseRuns says beat the A’s, creating a “double effect”.

Logan
10 years ago
Reply to  Philip

Gabriel,
FWIW, the Royals won 5 out of 6 this year against the A’s while only outscoring them by 5 runs, 23-18. That’s kind of interesting.

evo34Member since 2023
10 years ago
Reply to  Philip

Uh, I’m gonna guess no, considering that half of those four teams aren’t even in the same league as the A’s.

Stinky Pete
10 years ago
Reply to  dl80

I’m probably going to get 50 down votes for this hypothesis, but could the increased deployment of defensive shifts account for some of this variation? Since it’s been demonstrated that shifts are effective in suppressing offense, and teams are less likely to employ a radical shift with runners on base, could that be skewing the sequencing data (more negative outcomes in low leverage situations and vice versa)?

RMD4
10 years ago
Reply to  Stinky Pete

You’re on to something. We’re seeing more of a league-wide “Ryan Howard effect.” He’d consistently have much higher RBI totals than expected because teams couldn’t shift against his batted ball tendencies.

evo34Member since 2023
10 years ago
Reply to  Stinky Pete

These are both excellent points. If it’s happening across baseball, we should see the std. deviation of team runs per game go up.

Brian L
10 years ago
Reply to  dom

Possibly, but none of those are new things this year. They could be factors, and maybe to varying degrees this year vs. prior years, but they have always been factors. So they’re likely just a part of the overall year-to-year randomness / variance.

Eric M. Van
10 years ago
Reply to  dom

If we had this data going back only to 2010, we’d be convinced we were seeing an increasing trend, driven by analytic discoveries of sequencing skills. But of course the actual data has only the tiniest suggestion of this. We’ll know better in another five years, I’d say.

joser
10 years ago
Reply to  dom

didn’t we just talk about how adding a good hitter to a good offense creates non linear gains in runs scored because offenses perform better with men on base?

This is exactly why we don’t project team performance by adding up its players’ individual WAR, but use BaseRuns instead. In theory, BaseRuns captures some of that nonlinearity (at least more so than simply summing WAR). In practice, of course, it may not capture all of it (or all of anything else, which is why it’s imperfect, and always will be — weird shit like this year happens, and that, as they say, is why they play the games).

Avattoir
10 years ago
Reply to  dom

Also, when a team adds a better hitter or two to its order, or the converse, does that have knock-on effects beyond the contiguous positions in the order where the added (or subtracted) hitter is placed. I think it must, because the result will be that Hitter X, say a RFer already in the line-up largely due to his defensive value and hitting a competent 6th relative to that line-up, may now get bumped lower (or higher, depending on his skill set, tho higher seems unlikely for a RFer). So, adding a hitter or two causing that RFer to thereafter bat 7th or 8th (or 9th even, particularly in the DHL), means hitters contiguous in position to the RFer so moved can be expected to benefit from a better hitter in that lower position.

Would not this explain why a SF Giants team with a healthy fit Aoki, Panik and Pence will produce more than expected simply by replacing them with Maxwell, Adrianza and Byrd and their relatively inferior htting stats would predict?

bmarkham
10 years ago
Reply to  dom

“didn’t we just talk about how adding a good hitter to a good offense creates non linear gains in runs scored because offenses perform better with men on base?”

BaseRuns presupposes this. It’s one of the reasons that it more effective than adding up WAR to replacement wins to estimate expected records.

Scott
10 years ago

what is the % chance that these anomolous results extrapolate out at their current pace vs the % chance that these anomolous results regress/normalize over the final 6 weeks in the other direction and become less anomolous?

Scott
10 years ago
Reply to  Dave Cameron

Thanks for the reply Dave.

So if I am interpreting your response correctly you are saying that even if every team played to its exact baseruns expected record the rest of the way the StdDev in the graph above for 2016 would still remain significantly higher than the other years displayed?

I realize it’d be the gambler’s fallacy to expect the anomolous results to over-correct in the other direction but it seems ~6 weeks of more teams playing to their baseruns expected records would at least knock down the current >6% StdDev to some extent and that is still the most likely outcome given the assumption that baseruns are a reliable predictive model.

Lanidrac
10 years ago
Reply to  Scott

However, the standard deviation is based on what is essentially a counting stat. If every team played to their expected Base Runs over the remainder of the season, the Royals and Twins would still have +10 wins over what is expected, the Cardinals at +9, the Athletics at -10, etc., so the standard deviation wouldn’t actually change.

Peter Denton
10 years ago
Reply to  Scott

We would expect some teams to continue to excel (suck) at sequencing, and other teams to switch direction. Assuming no team has optimized sequencing (or is doing something, for some reason, to make for terrible sequencing), we would expect each team to have an equal chance of improving their sequencing gains as worsening their sequencing gains (and same for losses). So of the teams with huge sequencing gains (losses), about half of them will end the season with even higher gains (losses). Of course, there aren’t that many games remaining, so an extreme scenario (all the huge winners/losers winning/losing even more, or all the huge winners/losers losing/winning) is also quite likely.

atgo
10 years ago
Reply to  Dave Cameron

I dont know if that is right Dave.

Baseruns is just a model of true talent. What we should expect is for teams on average to play at their actual true talent in the future, not play at the model of their true talent. The model residuals should get smaller as the sample size gets bigger(if its a good model, which you show).

Nivra
10 years ago
Reply to  atgo

Yes, they should, but he’s not using a scaled stat. He’s using an absolute stat, which wouldn’t scale as you suggest.

Nivra
10 years ago
Reply to  Dave Cameron

Yet, in the article you don’t mention this. Instead, you extrapolate out the anomolous tendencies, assuming that they won’t regress their sequencing to more or less random.

atgo
10 years ago
Reply to  Nivra

im replying to your reply above. im confused by your response. hes looking at residuals over one sample size and comparing them to the residuals over a larger sample size. the greater the sample size the smaller the impact of random variation and thus the smaller the residuals, if its a good model. there is still a lot of random variation in these results! It is not designed as a true talent estimate which would require regression to the mean. Thus, i dont see any reason to expect teams to play at their baseruns estimate after 3/4 of the season for the last 1/4 of the season.

atgo
10 years ago
Reply to  Nivra

oops, i get what youre saying now

RMR
10 years ago

I wonder if the overall low scoring environment but remains of power result in the higher variance. That is, with fewer runs being scored overall, there is greater potential for HRs and extra base hits to have disproportionate relative contributions given good timing.

John
10 years ago
Reply to  RMR

I think that makes intuitive sense, but 2015 has been a higher run-scoring environment than 2013 & 2014 so the data doesn’t necessarily support your theory.

evo34Member since 2023
10 years ago
Reply to  John

I think what RMR is saying is that as the HR/R ratio increases (caused by either runs going down or HR going up), the timing of events becomes more important.

And yes, the HR/R ratio is up this season — more than 11% higher than 2014. It was 0.225 from 2010-2014, vs. 0.236 this season. Clearly not the whole story, but probably part of it.

Taylor
10 years ago
Reply to  RMR

Another way to think about this is that the low scoring environment does not “cause” this to happen (see 2013-14), but that low scoring environments generally are more governed by randomness than are higher run environments. Fewer events means less confidence in all predictions, and therefore wider confidence bands around the point estimate of team performance. Just my $0.02.

MjwW
10 years ago

When will Fangraphs apologize for throwing countless Red Sox articles at us in the offseason? They are a dumpster fire team and I think I speak for many when I say Fangraphs needs to own up to this glaring failure.

Pirates Hurdles
10 years ago
Reply to  MjwW

The funny thing is that it was the Sox and A’s that were most vehemently defended all winter. The Sox have actually performed right at their base run record this year. Of course one season is no reason to throw out the models, but some folks certainly have egg on their face. We really aren’t much closer to figuring out the magic with the Cards, Twins, and Royals this year.

Dick Bremer
10 years ago

As someone who watches the Twins consistently, there are only a few things I can attribute to their success. The line-up has stayed relatively healthy where the primary and secondary contributors have stayed on the field. Perkins was automatic up until the All-Star break. The first wave of fielding prospects such as Rosario, Sano, and Buxton haven’t been overwhelmed by the transition to the MLB. They also seem to have a knack for going cold at the plate, then stringing 3 or 4 hits together in one inning. By any statistical measure, the Twins should not have a record anywhere close to 63-61.

L. Ron Hoyabembe
10 years ago
Reply to  MjwW

Did you bet your kid’s college fund on the Red Sox this year? A lot of smart people were wrong about the Red Sox. No reason to apologize for that.

joser
10 years ago
Reply to  MjwW

You don’t speak for me, schmuck.

Jason B
10 years ago
Reply to  joser

Right?! I loooove when people think they’re owed an apology for an opinion . How dare you have the temerity to not predict the future with 100% accuracy! To believe something different than me!

Look, you can opine that John Jaha was the greatest player ever to grace the field. Sure, you might be batsh!t crazy, but you don’t owe me, or anyone, an apology.

Damaso
10 years ago
Reply to  MjwW

no need for an apology, but there’s a growing pile of evidence suggesting we may need to start re-examining some basic principles here.

would be a bit frustrating if fangraphs once again decided to throw up their hands, cry “baseball”, and not scrutinize their process at all, though.

BipMember since 2016
10 years ago
Reply to  Damaso

“fangraphs” doesn’t have a process. the site uses a bunch of model developed by other people. A big one is ZiPS, which is developed entirely by Szymborski, not by fangraphs. I’m not against improving projections at all, but fangraphs is a site that employs writers, not hardcore statistical analysts and researchers, so let’s direct our scathing criticisms to the right place.

Damaso
10 years ago
Reply to  Bip

hmm, i think you’re underrating what fangraphs has done and why they are so respected and influential.

Eric
10 years ago
Reply to  MjwW

Cmon fdop. This is low, even for you

Damaso
10 years ago
Reply to  Eric

swear its not me.

%
10 years ago
Reply to  MjwW

Maybe next time, just type the comment out, go have a glass of water, walk around outside for a bit, then delete it?

Ernie Camacho
10 years ago

The reference to “dominant bullpens” here and in other discussions about beating BaseRuns seems to miss a major piece. Even dominant pitchers can sequence their outcomes poorly. Cody Allen, for example, had a negative WPA as recently as a couple weeks ago. And the success of the 2012 Orioles and recent Royals teams relied in part on the distribution of outcomes produced by the individual pitchers in those pens, not just the [context-neutral] quality of those pitchers .

I know this was a short post so couldn’t really drill down, but it would be interesting to know how much of the past success of the “dominant bullpen” approach is really attributable to roster construction versus individual “clutch” performance distribution.

BipMember since 2016
10 years ago
Reply to  Ernie Camacho

That distinction is made for a very important reason: A pitcher’s distribution of events is basically not in his control at all. If a guy could choose to not give up hits at the time when hits are most harmful to him, then he just wouldn’t give up hits. Whether a guy can bear down and pitch better when it is more beneficial for him to do so is unclear, but there is little evidence it happens. That’s why baseruns ignores it.

On the other hand, it is entirely within a team’s control when they bring in their best and worst relievers.

Ernie Camacho
10 years ago
Reply to  Bip

Yes, that there is a distinction is my point. The “good bullpen” seasons people point to as examples of roster construction beating BaseRuns (like the 2012 Orioles, discussed in the DC piece linked to in the post) also relied heavily on clutch pitching outcome distribution.

Same thing with the current Royals season. So bravo to DM for assembling a great pen, but without more analysis, we have no idea how much credit for beating BaseRuns goes to him versus the dumb luck of having a clutch score of 1.65.

Ernie Camacho
10 years ago
Reply to  Ernie Camacho

Or to be more specific, all the key Royals relievers individually having a positive clutch score (i.e., it’s not just about when they are put in games; they all optimized their performance even within their own appearances).

BipMember since 2016
10 years ago
Reply to  Ernie Camacho

That’s the question though. Did they optimize their performance within their own appearances? I mean, they “did” in that it happened, but we don’t know if they did it intentionally. If they can’t do that intentionally, it doesn’t make sense to include it in Baseruns, because the point of it is to project using the factors we expect to continue.

Ernie Camacho
10 years ago
Reply to  Ernie Camacho

Bip, that not rally “the question” I care about. I think sequencing is “dumb luck” as I said above, so don’t even consider it an open question. I emphatically don’t want optimal sequencing to be captured by BaseRuns.

I’m simply curious how much a good bullpen can really be counted on to beat BaseRuns, as Dave suggest might be the case. So, for example, an estimate of what the Orioles’ record would have been in 2012 had its excellent bullpen had neutral outcome clustering instead of extremely favorable clustering. How much of its results success was due to player talent versus sequencing?

jerry60555
10 years ago

http://fivethirtyeight.com/features/is-2015-the-year-baseballs-projections-failed/

Dave?did you ever read this article?It said projection system is pretty struggle this year?and one of the most important reason is Luck spike whole MLB this year?and spike more than previous several seasons?do you think “luck”is the main reason to destroy base runs method this year?

jerry60555
10 years ago
Reply to  jerry60555

My cellphone typing is pretty wierd….fix again….sorry

http://fivethirtyeight.com/features/is-2015-the-year-baseballs-projections-failed/

Dave,did you ever read this article?It said projection system is pretty struggle this year,and one of the most important reason is Luck spike whole MLB this year,and spike more than previous several seasons,do you think “luck”is the main reason to destroy base runs method this year?

Umpire Weekend
10 years ago
Reply to  jerry60555

Thanks for fixing. Very easy to understand now.

To Serbian, Vietnamese, French, and Back
10 years ago
Reply to  jerry60555

Mobile operators rather strange … again difficulties. sorry

????://???????????????.???/????????/??-2015-???-????-?????????-???????????-??????/

Dave, I read articles that say that the great projection system is in trouble this year and is one of the most important causes of happiness Spike Major League Baseball all this year and jump to award Last season you think “happiness” is the main reason to destroy the main method begins this year?

BipMember since 2016
10 years ago

oh my i so love the substitution of “happiness” for “luck” (i’m aware that’s what happiness used to mean)

Horseshoe Jones
10 years ago
Reply to  jerry60555

Luck is an overused trope to explain away what people have yet to understand. Luck definitely exists in Baseball, but not at the levels we see expounded upon.

Look more closely at your defensive measures, park factors and rethink regressing to the mean (ie – regress to the familia instead)

Jeurys Familia
10 years ago

Why would people regress to me? So racist.

BipMember since 2016
10 years ago

you have some very specific suggestions without any hint as to why those suggestions would point us towards an over-attribution to luck.

Luck is the lack of an explanation, and it is also the correct explanation when there is no other. So we don’t know for a fact that the Cardinals pitching staff is just getting really lucky, but there is a good chance that is what is happening.

Horseshoe Jones
10 years ago

I don’t really have the space here to expound, but I’m working on a thesis – that if proves true would essentially show that Defense and a better calculated Park Factor could replace upwards of half what is currently attributed to Luck in projection systems and BABIP.. combined with using batter profiles, comped to pitch arsenals

It may not pan out or I may be in over my head with the math.. but I disagree that luck is the correct explanation when there is no other (see: History of Religion v Modern Science).. I think people aren’t giving some athletes and some GM’s enough credit..

Horseshoe Jones
10 years ago

I get that you’re joking and appreciated it, but for those that didn’t get my Familia reference, it is from Tom Tango’s “The Book”

BipMember since 2016
10 years ago

So, luck is a matter of perspective. Nothing is caused by “luck”. you can trace each event in a baseball game to a cause. However, luck, from the perspective of any given player, can be defined as the things that affect the outcome a player is trying to achieve that are not under his control. So, a hitter can control to an extent how hard he hits the ball, but he can’t “control” where his hit goes to the extent that he can prevent the defense from getting it. So, if a hard-hit ball finds a fielder, that is bad luck, even though the directionality of the ball was determined almost entirely by the pitch and how the ball made contact with it.

Similarly, a pitcher can consider his defense to be “luck”, in that it affects his ability to get outs, but he cannot control it, save for perhaps giving a bit of feedback as to where they should play. However, at the team level, defense is not luck at all. The team can intentionally construct a good defense.

So, in order to find out if something is “luck”, you basically just have to find out to what extent the player or team can control it. So, it’s not quite making an argument from ignorance, as my original statement made it sound.

AB
10 years ago

How does BaseRuns perform compared to adding up a team’s total WAR + 43?

doffbhoya
10 years ago

could it be that with fewer runs being scored, leading to more low-scoring games, there is simply a lot of variation in W-L records? teams with good (or bad) bullpens are disproportionately seeing that reflected in their actual reacords.

Lanidrac
10 years ago
Reply to  doffbhoya

That would be a Pythagorean outlier, not a Base Runs outlier. Base Runs predicts how many runs a team will actually score and give up based on individual events, which then indirectly affects the expected W/L record through Pythag, which predicts wins and losses based on the actual number of runs scored and allowed.

Ernie Camacho
10 years ago
Reply to  Lanidrac

Which raises the question: is the breakdown this year between BaseRun wins and actual wins due more to the discrepancy between real and modeled runs scored/allowed or in the way the runs are sequenced to make wins?

Theodore
10 years ago
Reply to  Ernie Camacho

Put another way: shouldn’t the “BaseRuns” effect — the effect of sequencing of hits — on W-L record truly be the difference between the Pythag record and BaseRuns record? By that count, the As are only -3 due to sequencing of hits, and -10 due to sequencing of runs. They’ve also had something of a notoriously bad bullpen, though as I recall the Pythag differential and bullpen performance aren’t all that well connected.

evo34Member since 2023
10 years ago
Reply to  Ernie Camacho

EXACTLY. This article for some reason is ignoring the fact that converting hits to runs and converting runs to wins are two completely independent processes.

How Cameron could take the time to write this article but not address this issue is beyond me.

Jacob
10 years ago

Dave–I haven’t found any study on BaseRuns that passes academic muster. The best description available seems to be from Tangotiger, but it is tough to glean specifics from that article, especially since (I assume) some updates have been made since then. Have a recommendation for a study that compares different models (with cross validation)?

HamelinROY
10 years ago
Reply to  Jacob

Additionally, is it possible for Fangraphs to share their formula they use to calculate it?

joser
10 years ago
Reply to  Jacob

How many studies in all of Sabermetrics pass “academic muster”? Have you asked Tango, or Steve Staude (see his spreadhseet here)?

BipMember since 2016
10 years ago
Reply to  joser

some probably pass muster on methodology, but probably almost none pass on peer review.

Cave Dameron
10 years ago
Reply to  Jacob

I agree I don’t see where it passes any muster. FG most likely runs all sorts of #’s until they find something they apply to prior years and see if there is a pattern, then call it a projection system.

I said it many weeks ago when they put an article on RaseBuns out that it was the Twins awful offensive start that skewed there overall #’s. Plus a lot of the times they had “lucky sequencing” was when they won 12-1, 14-2 etc., not in 3-2 or 4-3 wins.

Stats Guy
10 years ago

Maybe it’s time to add RBI and wins to the Baseruns formula.

Lanidrac
10 years ago
Reply to  Stats Guy

But what about the runs that score without an RBI, such as by error, wild pitch, balk, double play, etc.?

Plus Mathman
10 years ago
Reply to  Stats Guy

Also, what grit coefficient were you working with?

CSW
10 years ago
Reply to  Plus Mathman

#saveFireJoeMorgan

thebot
10 years ago

do you think perhaps you should start using intangibles in the BaseRuns formula, like clubhouse presence and experience?

DavidK
10 years ago

Hey-

Very interesting piece- one question: why do your rankings in terms of overachievement/underachievement seem to differ (pretty significantly) from the ones at thepowerrank – https://thepowerrank.com/cluster-luck/ ? I thought you guys were using the same formulae, but I might be wrong. (Their system seems to have CLE as the major outlier in terms of bad luck, and includes Toronto as a major outlier in terms of good luck, you didn’t mention either of these teams). Not sure if y’all are taking different things into account or something, but would love to know what the differences are in the analysis (powerrank explains theirs here:http://thepowerrank.com/were-the-2011-pittsburgh-pirates-lucky-or-good/)

D

Vince Clortho
10 years ago

After all the wailing and rending of garments from O’s Truthers in these sorts of posts over the years, how cruel it is BaseRuns ‘fails’ in a year they’ve underperformed their pythag by 5 wins. How they have waited for this day!

wiggly
10 years ago
Reply to  Vince Clortho

Are you Vinz’s brother?

Vince Clortho
10 years ago
Reply to  wiggly

Nah coincidence. That guy is Sumerian while our family’s name is bastardized from the original Polish ‘Czlortho’

Mike Green
10 years ago

I think that it is simple random variation.

There are two elements to the conversion from BaseRuns to wins and losses- converting the elements of runs scoring and prevention into actual runs scored and allowed, and converting runs scored and allowed into actual wins. For almost of all of the teams that are problematic, the second half of the process (call it the Pythagorean correlation) is an important part. The exception is the Twins who have scored many more runs than anticipated due to sequencing hits. The distribution of team’s actual records vs. Pythagoras is not particularly unusual. Similarly, the distribution of teams’ Base Runs runs scored and allowed vs. actual runs scored and allowed does not seem to me to be particularly unusual.

What is unusual this year is that both halves of the process have tended in the same direction for the great majority of the teams with the Jays and the Mariners being the exceptions.

evo34Member since 2023
10 years ago
Reply to  Mike Green

Thank you for writing this. I wish FG authors would actually read the comments.

Damaso
10 years ago
Reply to  evo34

agreed.

not only this great comment from mike but in general i’ve found the comments over the last month or so in particular to be treasure troves of insight.

Plucky
10 years ago

I have a conceptual problem with the way “cluster luck” is described and thought of. At the micro level, a pure Markov chain assumption on the team level is a terrible assumption. The biggest problem is on the pitching side. Teams don’t have uniform pitching quality, and teams do control what quality of pitcher is on the mound at any given time. If you look at a whole game in which a team gives up 8 hits, of which 3 were XBH’s in the 8th inning, you can chalk it up to “bad cluster luck” when what actually what happened was the starter brought his A game and the 8th inning guy blow up. “Cluster luck” only works in the context of individual players, not pitching staffs as a whole, and that assumes the pitchers play at a uniform true talent level rather than some fluctuating performance level.

Take the Dodgers this year. They are exactly at their pythag W-L, but 6 games behind their BaseRuns expected record. Is that actually bad luck, or is that just what happens when your pitching staff consists of Zach Greinke, Clayton Kershaw, and… yea. Is there apparent unluckiness with clustering due to actual bad luck? Or is it because BaseRuns implicitly assumes Clayton Kershaw and Brett Anderson will both pitch like =AVERAGE(Kershaw, Anderson)?

The “broken-ness” of BaseRuns this year should not be in reference to higher variance this year. What shows the broken-ness is that the StDev of Actual Wins vs BaseRun expected is larger than that for the simple pythagorean expectation. That is evidence there’s something wrong with BaseRuns. What does the above chart look like when you compare BaseRuns StDev vs Pythag StDev for previous years?

L. Ron Hoyabembe
10 years ago
Reply to  Plucky

Interesting point. I’m not sure if what you’re saying is necessarily a problem with BaseRuns, but it might point to how we are misusing it. I hope somebody who understands BaseRuns more than I do can respond to this.

dcs
10 years ago

I am the creator of BaseRuns. Despite the inflammatory title of this article (at least to me), my ‘prior’ would be that BaseRuns has not failed. Instead, that someone’s implementation and/or interpretation is inadequate. BaseRuns is really just a framework for how runs are scored . Correct implementation/interpretation is all-important.

Famous Mortimer
10 years ago
Reply to  dcs

Isn’t it just an equation you plug the figures into? How can the implementation be wrong?

Paul Clarke
10 years ago

You can have different BaseRuns formulas provided that they follow the overall form:

runs = A*B/(B + C) + D

where A = base runners, B = ‘advancement’ factor, C = outs and D = home runs. A, C and D are reasonably straightforward, but there are different ways of estimating B. See http://www.tangotiger.net/wiki/index.php?title=Base_Runs for details and three different examples of BaseRuns formulas that use different stats.

pft
10 years ago

Back in the day we called sequencing luck. Any evidence that sequencing is a skill? If so, isn’t that the same as clutch, which does not exist in the saberworld? Clutch for most of the population of players is just luck, as are RBI’s (also opportunity).

atgo
10 years ago

ill argue that baseruns actually strips out too much context. Even though it accounts for individual events, it values those events with linear weights! Thus, it presumes a single is worth the same in every environment, which we know is not true. Baseruns therefore removes park effects to a degree and it also removes team defense, both factors teams get to keep going forward.

bmarkham
10 years ago
Reply to  atgo

No it doesn’t, on all accounts.

atgo
10 years ago
Reply to  bmarkham

i dont understand what you’re saying. baseruns uses the linear weight value of a single to score how much a single is worth. that linear weight is the average value of a single in all run environments, from coors to petco, from kershaw to non-kershaw. there is at least one early implementation of baseruns where team specific linear weights are used.

Famous Mortimer
10 years ago

I like Fangraphs, but this is an odd article, with a sense of annoyance that reality hasn’t conformed to the equations you’ve developed to measure it. Perhaps BaseRuns has been lucky by hewing relatively close to reality for the past few years?

Damaso
10 years ago

this is an important perspective to keep in mind for sure. just because something correlates better than others for a few years doesn’t mean it’s necessarily more accurate.

newer metrics are also obviously more susceptible to this kind of recency bias.

vivalajeter
10 years ago

Is it possible that bullpen management – and not just the caliber of bullpen pitchers – has a big influence? And by bullpen management, I also mean whether or not you pull the starter at the right time.

We’ve seen it plenty of times before, where a starting pitcher seems to be doing well until things fall apart quickly. Through 7 innings they’ve given up 5 baserunners with plenty of K’s, but they’re at 105 pitches and gave up 2 hard liners in the 7th inning. The manager has him start the 8th, and he gives up 2 hits and a walk before departing. According to baseruns, that’s bad cluster luck. According to reality, the cluster of hits is because of poor managing.

The Orioles are a team that has done well over the last few years – is Showalter considered a very good manager when it comes to the bullpen? Does he know when to pull the starter, and when to avoid putting relievers in bad situations (such as using a LOOGY too often against righties?).

On the other hand, this wouldn’t explain why things are different this year compared to prior years.

WerthlessMember since 2020
10 years ago

I would expect that the distribution of standard deviations would fluctuate around a mean, but would approach a normal curve itself. Law of large numbers. You’re bound to get years with above or below average deviations, and there is not necessarily a need to attach a narrative (ie. set of reasons why) to this year to year variation.