How Do Prospect Grades Translate to Future Outcomes?

Hello, and welcome to Prospect Week! (Well, closer to Prospect Fortnight — as you can probably tell from the navigation widget above, the fun will continue well into next week, including the launch of our Top 100.) I’m not your regular host – that’d be Eric Longenhagen – but not to worry, you’ll get all the Eric you can handle as he and the team break down all things minor leagues, college baseball, and MLB draft. I’m just here to set the stage, and in support of that goal, I have some research to present on prospect grades and eventual major league equivalency.
When reading coverage of the minor leagues, I often find myself wondering what it all means. The Future Value scale does a great job of capturing the essence of a prospect in a single number, but it doesn’t translate neatly to what you see when you watch a big league game. Craig Edwards previously investigated how prospect grades have translated into surplus value, but I wanted to update things from an on-field value perspective. Rather than look at what it would cost to replace prospect production in free agency, I decided to measure the distribution in potential outcomes at each Future Value tier.
To do that, I first gathered my data. I took our prospect lists from four seasons, 2019-22, and looked at all of the prospects with a grade of 45 FV or higher. I separated them into two groups — hitters and pitchers — then took projections for every player in baseball three years down the line. For example, I paired the 2019 prospect list with 2022 projections and the 2022 prospect list with 2025 projections. In this way, I came up with a future expectation for each player.
I chose to use projections for one key reason: They let us get to an answer more quickly. In Craig’s previous study, he looked at results over the next nine years of major league play. I don’t have that kind of time – I’m trying to use recent prospect grades to get at the way our team analyzes the game today. If I used that methodology, the last year of prospect lists I could use would be 2015, in Kiley McDaniel’s first term as FanGraphs’ prospect analyst.
Another benefit of using projections is that they’re naturally resistant to the sample-size-related issues that always crop up in exercises like this. A few injuries, one weird season, a relatively small prospect cohort, and you could be looking at some strange results. Should we knock a prospect if his playing time got blocked, or if his team gamed his service time? I don’t think so, and projections let us ignore all that. I normalized all batters to a 600 plate appearance projection and all pitchers to a 200 innings pitched projection.
I decided to break future outcomes down into tiers. More specifically, I grouped WAR outcomes as follows. I counted everything below 0.5 WAR per season as a “washout,” including those players who didn’t have major league projections three years later. Given that we project pretty much everyone, that’s mostly players who had either officially retired or never appeared in full-season ball. I graded results between 0.5 and 1.5 WAR as “backup.” I classified seasons between 1.5 and 2.5 WAR as “regular,” as in a major league regular. Finally, 2.5-4 WAR merited an “above average” mark, while 4-plus WAR got a grade of “star.” You could set these breakpoints differently without too much argument from me; they’re just a convenient way of showing the distribution. There’s nothing particularly magical about the cutoff lines, but you have to pick something to display the data, and a simple average of WAR projections probably isn’t right.
With that said, let’s get to the results. My sample included 685 hitters from 45-80 FV. Allowing for some noise at the top end due to small sample size, the distribution looks exactly like you’d hope:
| FV | Washed Out | Backup | Regular | Above Average | Star | Count |
|---|---|---|---|---|---|---|
| 45 | 51% | 25% | 17% | 6% | 1% | 295 |
| 45+ | 52% | 18% | 19% | 11% | 1% | 91 |
| 50 | 23% | 24% | 30% | 21% | 2% | 197 |
| 55 | 17% | 17% | 30% | 31% | 6% | 54 |
| 60 | 14% | 12% | 19% | 38% | 17% | 42 |
| 65 | 0% | 33% | 33% | 0% | 33% | 3 |
| 70 | 0% | 0% | 0% | 0% | 100% | 2 |
| 80 | 0% | 0% | 0% | 0% | 100% | 1 |
Consider the 55 FV line for an explanation. Of the players we graded as 55 FV prospects, 17% look washed three years later – Jeter Downs, a 2020 55 FV, for example. Another 17% have proven to be backup-caliber, like 2022 55 FV Curtis Mead, or 2019 55 FV Taylor Trammell if you don’t think Mead’s trajectory is set just yet. Continuing down the line, 30% look like big league regulars – 2021 55 FV Alek Thomas, perhaps. A full 31% appear to be above-average major league contributors three years later, like 2019 55 FV Sean Murphy or 2021 55 FV Royce Lewis. Finally, 6% project as stars three years later – Jackson Merrill, a 55 FV in 2022, feels appropriate as an example.
Two things immediately jump out to me when looking at this data. First, the “above average” and “star” columns increase at every tier break, and the “washout” column decreases at every tier break. In other words, the better a player’s grade, the more likely they are to be excellent, while the worse their grade, the more likely they are to bust. That’s a great sign for the reliability of our grades; they’re doing what they purport to do, essentially.
Second, each row feels logically consistent. The 45 FV prospects are most likely to bust, next-most-likely to end up as backups, and so on. The 45+ FVs look like the 45 FVs, only with a better top end; their chances of ending up above average are meaningfully better. The 50 FVs are a grab bag; their outcomes vary widely, and plenty of those outcomes involve being a viable major leaguer. By the time you hit the 55 and 60 FV prospects, you’re looking at players who end up as above-average contributors a lot of the time. The gap between 55 and 60 seems clear, too; the 60 FVs are far more likely to turn into stars, more or less. Finally, there are only six data points above 60 FV, so that’s mostly a stab in the dark.
This outcome pleases me greatly. Looking at that chart correlates strongly with how I already perceived the grades. For a refresher, roughly 30 prospects in a given year grade out as a 55 FV or above, give or take a few. Something like three quarters of those tend to be hitters. That means that in a given year, 20-ish prospects look like good bets to deliver average-regular-or-better performance. The rest of the Top 100? They’re riskier, with a greater chance of ending up in a part-time role and a meaningfully lower chance of becoming a star. But don’t mistake likelihood for certainty – plenty of 55 and 60 FVs still end up at or below replacement level, and 45 FVs turn into stars sometimes. Projecting prospect performance is hard!
How should you use this table? I like to think of Future Value in terms of outcome distributions, and I think that this does a good job of it. Should a team prefer to receive two 50 FV prospects in a trade, or a 55 FV and a 45 FV? You can add up the outcome distributions and get an idea of what each combination of prospects looks like. Here are the summed probabilities of those two groups:
| Group | Washed Out | Backup | Regular | Above Average | Star |
|---|---|---|---|---|---|
| Two 50 FVs | 46% | 49% | 60% | 42% | 4% |
| One 55, One 45 | 68% | 42% | 47% | 37% | 6% |
Another way of saying that: If you go with the two-player package that has the 55 and 45 FV prospects, you’re looking at a higher chance of developing a star. You’re also looking at a greater chance of ending up with at least one complete miss, and therefore lower odds of ending up with two contributors. Adding isn’t exactly the right way to handle this, but it’s a good shorthand for quick comparisons. If you want to get more in depth, I built this little calculator, which lets you answer a simple question: For a given set of prospects, what are the odds of ending up with at least X major leaguers of Y quality or better? You can make a copy of this sheet, define X and Y for yourself, and get an answer. In our case, the odds of ending up with at least one above-average player (or better) are 40.7% for the two 50s and 41.4% for the 45/55 split. The odds of ending up with two players who are at least big league regulars? That’d be 28.1% for the two 50 FVs, and 16.1% for the 45/55 pairing. Odds of at least one star? That’s 4% for the two 50 FVs and 6% for the 45/55 group. In other words, the total value is similar, but the shape is meaningfully different.
For example, you’d have to add together a ton of 50 FV prospects to get as high of a chance of finding a star as you would from one 60 FV. On the other hand, if you have three 50 FVs, the odds of ending up with at least a solid contributor are quite high. Meanwhile, even 60 FV prospects end up as backups or worse around a quarter of the time. That description of the relative risks and rewards makes more sense to me than converting players into some nebulous surplus value. Prospects are all about possibility, so representing them that way tracks analytically for me.
Take another look at the beautiful cascade of probabilities in that table of outcomes for hitting prospects, because we’re about to get meaningfully less pretty. Let’s talk about pitching prospects. Here, the outcomes are less predictable:
| FV | Washed Out | Backup | Regular | Above Average | Star | Count |
|---|---|---|---|---|---|---|
| 45 | 53% | 26% | 16% | 5% | 0% | 230 |
| 45+ | 38% | 24% | 25% | 13% | 0% | 68 |
| 50 | 27% | 27% | 24% | 20% | 2% | 96 |
| 55 | 17% | 20% | 37% | 27% | 0% | 30 |
| 60 | 17% | 33% | 25% | 25% | 0% | 12 |
| 65 | 0% | 0% | 0% | 100% | 0% | 1 |
| 70 | 0% | 0% | 100% | 0% | 0% | 1 |
I have tons of takeaways here. First, there are substantially fewer pitching prospects ranked, particularly as 50 FVs and above. Clearly, that’s a good decision by the prospect team, because even the highest-ranked pitchers turn into backups at a reasonable clip. Pitching prospects just turn into major league pitchers in a less predictable way, or so it would appear from the data.
Second, there are fewer stars among the pitchers than the hitters. That’s true if you look at 2025 projections, too. There are only six pitchers projected for 4 WAR or higher, while 42 hitters meet that cutoff. It’s also true if you look at the results on the field in 2024; 36 hitters and 12 pitchers (22 by RA9-WAR) eclipsed the four-win mark. You should feel free to apply some modifiers to your view of pitcher value if you think that WAR treats them differently than hitters, but within the framework, the relative paucity of truly outstanding outcomes is noticeable.
Another thing worth mentioning here is that pitchers don’t develop the same way that hitters do. Sometimes one new pitch or an offseason of velocity training leads to a sudden change in talent level in a way that just doesn’t happen as frequently with hitters. Tarik Skubal was unmemorable in his major league debut (29 starts, with a 4.34 ERA and 5.09 FIP). Then he made just 36 (very good) starts over the next two years due to injuries. Then he was the best pitcher in baseball in 2024. Good luck projecting that trajectory. Perhaps three-year-out windows of pitcher performance just aren’t enough thanks to the way they continue to develop even after reaching the majors.
There’s one other limitation of measuring pitchers this way: I don’t have a good method for dealing with the differential between reliever and starter valuation. Normalizing relievers to 200 innings pitched doesn’t make a ton of sense, but handling them on their own also feels strange, and I don’t have a good way of converting reliever WAR to the backup/regular/star scale that I’m using here. A 3-WAR reliever wouldn’t be an above-average player, they’d be the best reliever in baseball. I settled for putting them up to 200 innings and letting that over-allocaiton of playing time handle the different measures of success. For example, a reliever projected for 3.6 WAR in 200 innings would check in around 1.2 for a full season of bullpen work. That’s a very good relief pitcher projection; only 20 players meet that bar in our 2025 Depth Charts projections.
In other words, the tier names still mostly work for relievers, but you should apply your own relative positional value adjustments just like normal. A star reliever is less valuable than a star outfielder. A star starting pitcher might be more valuable than a star outfielder, depending on the degree of luminosity, but that one’s much closer. This outcome table can guide you in terms of what a player might turn into. It can’t tell you how to value each of those outcomes, because that’s context-specific and open to interpretation.
This study isn’t meant to be the definitive word on what prospects are “worth.” Grades aren’t innate things, they’re just our team’s best attempt at capturing the relative upside and risk of yet-to-debut players. Being a 60 FV prospect doesn’t make you 17% likely to turn into a star; rather, our team is trying to identify players with s relatively good chance of stardom by throwing a big FV on them. And teams aren’t beholden to our grades, either. They might have better (or worse!) internal prospect evaluation systems.
With those caveats in mind, I still find this extremely useful in my own consumption of minor league content. The usual language you hear when people discuss prospect trades – are they on a Top 100, where do they rank on a team list, what grade are they – can feel arcane, impenetrable even. Breaking it down in terms of likelihood of outcome just works better for me, and I hope that it also provides valuable information to you when you’re reading the team’s excellent breakdown of all things prospect-related this week.
Ben is a writer at FanGraphs. He can be found on Bluesky @benclemens.
Absolutely incredible piece, Ben!
Yeah, this is the stuff I come to Fangraphs for. Thanks Ben.
Agreed!
Great stuff. Was wondering Ben how you treated a player during the time period where the FV started below 45 but in the next year or more increased to FV45 or better? Were these “risers” added to data? If not, Trying to populate to see what this group looked Iike vs. initial FV45+ guys might also be interesting. Keep your research coming. Very enjoyable.
I got a late start on the morning (driving back from a weekend ski trip), so expect a lot of answers down the comments, starting here. I counted every grade every year, so some prospects were included multiple times. I didn’t do any further study on trajectory, though, might be interesting further research.
Good stuff & fun article- Couple questions:
Seems like it would have to be Gore, as he’s the only 70FV pitcher in the sample. And if the projections were taken 4 years later instead of 3, then he would’ve been in the above-average group.
Yep, three buckets. And yep, Gore.
Prospects week is here!
Very interesting. One takeaway–from what I recall, Longenhagen has some of the harshest prospect grades in the industry. And yet, for 50 and below, I wonder if they should be shaded down half of a grade–as the median outcome for 50FV hitters seems to be the very low-end of the “regular” group.
Agreed. And also wondering if we should be thinking most about modal and cumulative outcomes. A 55 has at least a 67% chance of being a regular or better vs 46% for a 50.
He didn’t use to (look at his 2017-2018 grades and compare to 2024) , although I find that simply a reflection of how often prospects actually pan out and how other publications may feel obligated to generate excitement (this is obviously speculation, of course).
That seems likely. IIRC, MLB Pipeline gives every player on the top 100 at least a 50, and they probably give out 50s beyond that.
Actually I just checked and they give out 55s to the bottom of the 100! And I don’t think this is because their analysts are inferior (it seems like Jim Callis is considered to be the best in the industry by his peers), so it must be that there’s grade inflation built in.
MLB essentially give everyone 5 more than FG would.
Eric has become much more conservative with grades over time (and to a lesser extent, Kiley McDaniel at ESPN).
He and Kiley did a ton of work here a few years ago (I think some of which wound up in Future Value as well) that established a connection between the WAR distribution and the FV grades. Then they wound up aligning the grades with the distribution. This means that there are some things that rarely happen, like almost no pitchers get an FV65 or FV70, because they are so rare.
Either that or we just suck at identifying them! FV40 on Cole Ragans, for example, could’ve been a FV65 or FV70 in hindsight.
IMO this seems pretty accurate to me. An FV50 is supposed to be a regular, and the fact that it’s in that bucket is reassuring. There’s always going to be some disconnect between the individual grades and the overall distribution.
Part of that is that the FV grade will always be a blunt tool, albeit a necessary one if the exercise involves putting everyone on a uniform scale.
Normally we think of probabilities along a bell curve. But I bet if you sat down with scouts to talk about the range of outcomes for say…2022 Elly de la Cruz, it wouldn’t look anything like a bell curve. With huge bust potential but a near-certainty that if he made his skill-set work, he’d be a star and very little possibility he’d settle in as just ok, it would look more like a Bactrian camel (that’s the two-humped one, right?). For a boom-or-bust guy like that, it’s hard to convey a lot of meaning by assigning a number like 50 or 55FV.
But, of course, we also want to know about those guys. So typically they get lumped in with the guys who have much more traditional quartile/decile projections, and 50FV becomes more of a hedge against the downside than a statement about the most likely/median outcome.
2 things: 1) if there were a problem in terms of definition vs results, wouldn’t that reflect an issue with the individual grades, rather than with the grading system itself? and
2) Um…isn’t that a very good result? A 50FV grade is, roughly speaking, an opinion that a given prospect has a median projection as a regular, someone who produces in the 1.6-2.4 WAR range (Ben used 1.5-2.5, but 1.6-2.4 is how Eric/Kiley originally mapped it out).
If 50% of FV hitters are turning into regulars or better, and 50% aren’t, isn’t that a pretty good indication that Eric’s doing a really good job of assigning those grades? In statistical terms, that means half of the guys are playing at or above their median projections, and half aren’t. I’m not a data scientist, but to my layperson’s understanding when you’re working with models based on probability, about half performing at or above the median and half performing below it is pretty much exactly what you’re going for, no?
I (sincerely) love that Prospect Week has escaped the containment of a single week.
This is a great piece. I think it would overwhelm the analysis, but I’m curious how the added layer of risk plays in… I know FV numbers discount risk, but a risky 55 and a safe 50 presumably have different outcomes, and I’m curious how different teams quantify/qualify this
Maybe it’s just a different way of saying the same thing, but I wonder how age affects things. Does a 23-year-old player with a 55 FV have less variance than a 19-year-old 55 player? I would assume so, but “risk,” despite being far less granular than age, might be a better variable as it subjectively incorporates age, injury, and traits/flaws.
I didn’t try to do this because I think it would make the sample size miniscule, but I completely agree that that would be more interesting. If Eric keeps putting out grades, I’ll keep adding them to this study, and at some point we’ll have enough to do that.
Great article. Those washout rates need to be stapled to every trade article when fans lose their minds over giving up 45/50 FV guys to get an actual major leaguer
Thank you! This is amazing.
How do you think the numbers might change if you stretched it to 5 year forecasts, giving pitchers a better chance to develop?
Also, this might be rephrasing the question asked by another commentator, but did age come up as a factor on “hit” rate? As in the college hitters given a 50 were more likely to end up solid regulars than the younger prospects given the same grade?
Age didn’t come up as a factor, but I think that’s another issue of sample size. I’m going to start extending the horizons every year we get more data and answer your first question; I’m excited to see the results.
TINSTAAPP is greatly misunderstood, but I think the second half of this post explains the concept pretty well.
Especially this line:
Very Interesting. It seems like the last 3 levels of pitchers and hitter have too small of a sample size. Would you consider consolidating the bottom 3? Other wise if next year you get the rare player with an extremely high fv who only becomes a back up or gasp a washout it would skew the data.
Oh yeah, if I were using this data analytically to try to value prospects, I’d be doing some bundling and regressing towards the bottom end. The sizes are just too small to be reliable, which is why I included them. I didn’t want to give the impression that the data contained there was just as weighty as the huge group of 45’s and 50’s.
Am I reading this wrong or could one conclusion be that we are giving too much weight to “deep” systems and not enough to “top heavy” ones?
I had the same thought. The farm system rankings here on fangraphs do address this by giving a lower $ value to each lower FV grade. However, I think 40 and 35 FV prospects are still a bit overvalued. For example, Arizona currently has 20 prospects with a FV of 35+ which equates to $10M in value. A 45+ hitter is valued at $8M. Would any team be willing to give up their 45+ hitter for 20 35+ lottery tickets? I doubt it. I wouldn’t.
For sure teams fall in love with a random 40/35 prospect and request them as a trade throw-in; but as a whole they are not super valuable. For that reason i don’t really think including anything below a 40+ is very helpful in farm rankings.
It seems to me that there is context that applies in the lower end of the FV ratings in a system. In particular, a lot of starters with relief risk end up in the 35+ range (I think off the top of my head). Their WAR ceiling is clearly limited, but they can be the pipeline that a real team needs to staff a playoff quality bullpen?
I think that there are plenty of scenarios where I trade a 45FV (DH only slugger in his late 20s) for 20 35+FV “lottery picks” (reliever risk, skill sets that align with my systems development strengths, etc).
I like the Fangraphs existing system ranking approach. It allows me to see the most data and draw my own conclusions.
I totally agree with your comments from a Fantasy perspective.
You’re arguing that there’s no way 20 35+ FV picks could outvalue the 31% chance 1 45+ FV player has of being a regular or better? Couldn’t the 35+ FV guys average a 2% chance of making regular or better? I think they could. Roster scarcity. At this extreme, does become a problem, but in a frictionless universe….
Yeah exactly – you can only roster so many guys in your farm system and still have them development properly, right? So due to the roster scarcity I think large sums of 35+ prospects begin to lose value.
I think the conclusion is that we have a way to quantify such systems’ quality. (Or rather, these probability distributions are a second way, in addition to surplus value.)
I ring this bell as often as I can. In most years the 50FV tier extends from prospect ~30 or so to about prospect ~130, and I’d bet there are at least another 150 or so 45/45+ guys.
Put another way, the difference between prospect ~35 and prospect ~280 is generally a half-grade, and qualified evaluators can and do have half-grade differences of opinion on guys all the time.
I’d rather have 2 60s or 3 55s and nobody else in the 50FV tier than 6 50s.
Obviously it’s not quite so simple, as there are tiers within tiers, and I bet most people would see a bigger difference than that between a “high 50” and a “low 45.”
But it’s simple enough that you can say pretty comfortably that not a ton of expected future value separates a prospect ranked 80th or so and a prospect ranked 200th (if rankings went that far, which most don’t).
Love this measure of standard deviation. Question – did this table include Wander Franco’s 80 FV ranking?
He’s the one 80 FV batter that made the table I believe. I’m not sure the last time there was an 80 grade prospect before him but my memory is poor.
He was the first one
So sad
What were Puig and Jose Fernandez’s FV’s? I remember them being very high, but maybe not 80.
Very interesting way to look at prospect projections vs MLB player projections of those prospects.
It looks like a 55 FV is roughly 3x more likely to be a starter (or better) than a 45 FV while a 45 FV is 3x more likely to washout. But I’d be very interested in seeing how the numbers workout with a larger sample, or over a longer timeframe.
Perhaps I missed it, but what projection system did you use?
I used Steamer for all of these. When I was doing the study, the initial 2025 ZiPS release wasn’t out yet. I did do some tests of 2022-2024 Steamers vs. ZiPS vs. Depth Charts blend and didn’t find much signal or bias so I stuck with Steamer as the one that had a full release for all the years I was interested in at time of study.
One thing I think is important translating the usual language of prospects is that a Top 100 prospect is a 50 FV or better prospect (and, if referred to in those terms, probably is a 50 FV because people at the top end of the scale tend to be referred to more precisely).
Something’s not right here, or I’m not understanding the methodology. Lux last showed on the 2020 prospect report as a 70 FV. That means we should be looking at his 2023 projections, which say 2 WAR. How is there a 100% star rate on the 70 FV prospects?
Gavin Lux isn’t on the 2020 updated list, which can be found here:
https://www.fangraphs.com/prospects/the-board/2020-in-season-prospect-list
Here’s all the data I used in my study if you’re interested in diving into the specifics more deeply:
https://docs.google.com/spreadsheets/d/114XEGzTwFwHkUzj6rEaUOl42MWFFmZ1NBfDoA03mGZs/edit?usp=sharing
I see, so he was a 60 in the 2019 mid season, moved up to 70 2020 beginning of the season but dropped off the list by mid season so he probably went into your numbers as a 2019 60 (can’t get to that google doc at the moment to validate)
Not sure enough guys follow that particular trajectory to mess with the numbers in terms of evaluating them against the player’s *peak* FV, though I would bet most of them who do are higher FV players.
That makes sense though, thank you for the response!
Yeah, he was a 2019 60, exactly. Yeah, I do think that more data points would be interesting, I’d have to think about how to handle it though. Maybe groups of risers vs. fallers or something?
I’ve wanted this assessment for so long. Thank you for bringing this eval to us fortunate and grateful readers.
Great article.
The one caveat is the lack of data. Not Ben’s problem though going backward a bit more may have smoothed out some numbers. However, at the highest levels you’d still have issues with tiny samples.
I find it mildly interesting that roughly 50% of the 50 FV prospects in this sample became big league regulars or better. (53% of hitters and 46% of pitchers.)
This is the probalem with systems like zips…..50% chance theyre war total is 0. Thats why he overvalues prospects.
FV is a probability distribution so finding that this probability distribution has 50% below the median expectation is not insane at all. Also its not ZIPS. That is an entirely different thing.
I know that. I’m saying zips shouldnt be using reversion to means on propsects. That is antiquated method and why very few are even looking at zips these days outside of casual fans. Teams are way past that method. I don’t know how he’s going to hang 4 war on a AAA guy and not laugh….he need to be using some system like this.
> I don’t know how he’s going to hang 4 war on a AAA guy and not laugh
Wut? Most of the star projections are for players _already_ in the big leagues, but who were prospects 3 years before.
I am curious what the percents are for hitters who rank below 40, or at 40.
I tried to run this and ran into meaningful data cleaning problems. By the time I got into the 40 tier, I was having to dig through the data very extensively to remove players who had retired within three years, or never played anywhere above the DSL, or things like that. I ended up ditching it because it was unwieldy and the hit rates were miniscule. My interpretation is that by the time you get to prospects of that level, the particulars are far more important than the grade. Do you like Caleb Killian’s command a ton? Bump him up. That kind of thing. When you’re in ‘high likelihood of being organizational filler’ territory, it’s more important to look for outliers than to use the mean.
But couldn’t you pretty simply label all the players without projections as “washouts”? I mean, getting injured and retiring is a different kind of washout than becoming an org guy, but it still means they aren’t generating any WAR from now on.
And I don’t know if I agree about the mean value being less relevant when the probability of success is low. It seems like it matters a lot whether it’s 3% or 6%, and if you’re able to find cohorts – like certain sets of tool grades – that meaningfully over or under perform their FV peers, I’d think that would be incredibly helpful.
Last year Cowser and Butler both had 45+ grades. Are they still viewed in that manner? Do either of them now have star or above average futures?
Regarding pitcher grading and outcomes, I frequently think about this off-the-cuff Eric monologue from an EW episode last year:
So do we need a LOOK column on the pitching side of the prospect board?
I’m not sure how religiously updated this page is, but we already do?
I like the methodology. An offshoot of this methodology would be to take the projection at age 27 ir 28, which in many ways is what scouts are projecting.
Great idea and one I’m gonna look into in the coming weeks, that sounds kinda fun.
Prospect week creep = 10 days of prospect week
(I’m here for it) (unlike christmas decorations on nov 1)
Wow, this is amazing. Thank you, Ben! And of course thank you to Eric as well for his incredibly valuable prospect assessments.
Ben, this is amazing work. Kudos!
I may be wrong, but aren’t prospect grades part of the ZIPS projections? Therefore wouldn’t that hurt the data a bit because one variable already affects the other? Am I wrong about this?
I am decently sure that ZIPS is FV-agnostic.
I’d love to know if there are any trends within position groups. Are shortstops and center fielders more likely to live up to their FV than catchers or first basemen? It’s possible the differences are already baked into their FV number, and I realize breaking it up by position makes the sample sizes even smaller.
But it’d be cool to know if, let’s say, it’s easier to identify middle infielders than others, and corner outfielders are a struggle (or whatever it might actually be).
Interesting thought. I remember some years ago Baseball Prospectus was advancing a theory of SBPODE: second base prospects often don’t evolve. The natural following question being whether there’s something special about second base prospects, or whether SBPODE is just a result of looking to closely at PODE.
I’d bet different franchises are better or worse at certain positions rather than positions themselves. Although, minor league starting shortstops often end up almost anywhere on the field in the end.
This is fascinating to look at, but it’s very unclear how much signal there is in all the noise. The sample sizes are small, the projections are projections and not actual measured performance any more than the prospect grades… it’d be interesting to revisit this same pool of players in, say, 5 years and look at what happened to see if anything meaningful can be prised out of it. As it is, it’s hard to look at that data and conclude anything at all.
In that case, stay tuned for:
“How’s My Driving: 2018 Top 100 Audit”
coming up later in the week!
Based on my proprietary projection system, that article grades out as exactly what I’d have predicted based on this one.
I know you run into sample size issues pretty quickly, but I’d be curious to know splits – at least within the FV50 and lower tiers – by max level, age, projected position, and underlying tool grades, among other things.
I’d think you’d be able to populate one dimensional splits, even if full xtabs are a nonstarter.
Sorry, not sorry, for the snark, but maybe the top-100 list would be a little better if you released it at the trade deadline next August?
FanGraphs’ way of doing things is a known quantity. You have several other good choices if this one doesn’t suffice.
There’s a major update released before the trade deadline
This was a great article. Thank you. I wonder if you could do the reverse…take the current major leaguers in the buckets (or use guys with 2-4 years experience) and run the distribution of prospect grades. I know you’ll have multiple prospect grades for many players, but you’ll figure it out….
Forrest Whitley’s prospect outcome is “All-Star?” What?
I’m not sure where you get that – here’s a link to the data:
https://docs.google.com/spreadsheets/d/114XEGzTwFwHkUzj6rEaUOl42MWFFmZ1NBfDoA03mGZs/edit?usp=sharing
From my quick checking, Whitley was a 60 in 2019 and projected for 1.1 WAR/200 in 2022. Then he was a 55 in 2020 and projected for 0-ish WAR/200 in 2023. He wasn’t on either of the 2021 or 2022 lists.
Just looking at the Board, Whitley was a 65 in the 2019 report (60 was the 2019 updated) and the only 65 pitcher in the sample is listed as an All-Star. Is there a different 65 pitcher I’m missing?
It’s 2022 Eury Perez
I’m confused then — Whitley is listed as a 65 on the 2019 report — why do the outcomes only include Perez?
Great work, Ben. I wish FanGraphs made season opening projections from past years available to subscribers.
They are available! We put them up as of January 28th:
https://blogs.fangraphs.com/instagraphs/introducing-new-steamer-split-projections-and-more/
woo hoo!
Would love to see this experiment run by GM/team too. I intuitively _know_ Ben Cherington has had higher rates of washout hitters than most other GMs, for example.
Who the hell was an 80 grade prospect???
Wander Franco.
Nicely done. Wondering if next time you could look at how indiv skill grades (raw/game power, present and future, etc.) correlate to MLB success.
This is excellent, Ben!
“depending on the degree of luminosity” is an absolutely wonderful turn of phrase!
Some prospects are a 50 due to high probability of a modest floor and are lumped in with 50’s with much higher ceiling and higher risk. If we are using this for fantasy, then historically, which type of 50 is best?
I’d love to see this methodology applied to different ranking systems — is Law, BP, Fangraphs, etc. better at predicting future performance?
Is it intentional that the headings for pitchers and hitters are different?
I tried to take this and combine with the Farm system rankings to see which teams would have what expected count in each category (only including 45 and up). I just multiplied the outcome % by the count in each category (e.g. 50 FV Bat) for each team and summed them up. Sorry if the formatting sucks.
Caveat: I’m just doing this for fun and realize that SSS would really throw off the top tiers. But I like the idea here that instead of just a straight ranking you could show the likelihoods of different types of players for a farm system. Food for thought.
Long live Prospect Fortnight
Wasn’t Moncada a fv70? I do not consider him a star. So he does not fit your data. Unless one season of 5 WAR in 9 seasons is all you need using this model. he has 13 war in those 9 seasons. How am I wrong?
oh i guess he is outside the data sample.
Echoing other posts, this is excellent! One thing I’m curious about is the predictiveness of other tool grades, both for overall future value and for the tool in question. Put another way, which tool grade has the highest correlation with future value (e.g., power for hitter, fastball or command for pitcher)? And how predictive are power grades for future HR or ISO, and speed grades for SB?
Great article, Ben.
Really fascinating article Ben! One note – the columns are labeled differently for Position Players and Pitchers. Is there a reason for that (I’m assuming just a change during editing that got missed)
Yeah I changed the nomenclature around a little bit on my spreadsheets while compiling the data and grabbed one of them right in here. The break points are still the same, and I’ve updated the names to match each other.
Ben, this is way more entertaining of an audit than the ones at my corporate job.
Really excellent piece and extremely useful to the whole community.
One question i’m curious about: What % of mlb players that have made their debuts over the past x seasons have actually even had prospect projections? And what percentage falls within each prospect grade (or no prospect grade)? From that, then potentially doing a similar kind of analysis as you did above with WAR to then break that out further to see the actual value of that MLB time?
I’m really just curious about how many mlb players of recent never had prospect projections :).
Thanks a bunch Ben, really enjoyed the article!
This is great work. Prospect grades are evaluations, not height measurements, so a 2024 55+ might be different than a 2020 55+, as the evaluator’s skills may “improve” and the grades become better predictors. Breaking down by year may throw too much variance due to smaller sample sizes, so I don’t think it is possible to tease out any more nuance doing so.
You make a really interesting point about how difficult it is to evaluate starters vs. relievers, but I disagree with how you handled the relievers. Relievers have far less value than starters do, which is why they are paid far less and most organizations will put considerable effort into making their best pitchers starters before resorting to making them relievers. For them to top out somewhere around average regular makes perfect sense.
The idea of ‘normalizing’ relievers to 200 IP is especially problematic. First, most relievers have demonstrated that, due to either stuff or durability, they cannot handle a full-time role. Second, relievers have leverage baked into their WAR. When normalize their WAR to 200 IP, you’re not just giving them credit for innings they are almost certainly incapable of pitching, you’re giving them triple credit for leverage shouldn’t be included. (The 200 IP marker is similarly, though far less, problematic for starters.)
In the end, the result is distortion of their value as prospects compared to other pitchers. It assumes teams would be just as happy with a 60 FV starter turning into a 1.3 WAR reliever as they would be with him turning into a 4 WAR starter.
who was the 80 grade prospect?
This is a great piece. Interesting just for the sake of baseball and I’m going to be using this calculator when I think about trades in my dynasty league! I’ve always felt folks in my league over-value prospects and this kind of backs that up.
Great foundational peice to build on analyses for years to come.