What the headline measures
The production headline is the v30 multivariate brms model run with cmdstanr MCMC. For every draft pick from 1980–2025, it conditions on draft slot, position, draft year, and the player’s observed career window before estimating whether the GM’s portfolio repeatedly beat expectation.
This is a drafted-player outcome estimate that bundles evaluation and development. It is not a franchise-wins model, a coaching rating, or a claim that one person made every decision in a front office.
Seven outcomes, three equal voices
The active v30 drafted-player outcome set groups games played, starts, and position-normalized snap share as Opportunity; trade-credit on-team AV and AV/snap quality where observable as Quality; and Pro Bowls plus AP1 All-Pro seasons as Recognition. The model combines outcomes with precision weighting inside each block, then gives each block equal weight so playing time does not drown out value or honors.
Why the public score is conservative
The raw posterior is not the public answer. Good GM applies a maturity layer based on capped player-years: portfolios with less time to mature are pulled toward zero and widened before publication. The Big Board’s score is the resulting maturity-adjusted composite θ; every interval is a maturity-adjusted 95% credible interval.
Verdict labels are relative tiers in the fitted GM pool, not guarantees. An asterisk signals direction evidence; the interval bar shows the remaining uncertainty more directly.
The current release now has a dated validation suite covering posterior-predictive behavior, five-year estimate stability, and maturity-setting sensitivity. The PPC finds material zero-rate and tail misfit, so the evidence is published as a limitation and method-review input, not a pass badge.
The record vs. the verdict
Every GM profile shows two dots on one axis: an open dot for his record alone and a filled dot for the published maturity-adjusted composite. The record-only estimate is a descriptive projection of raw pick outcomes, not a score, rating, index, or ranking. It takes era-normalized on-team games and AV (within-draft-year z-scores), draft capital (1/√pick), and Pro Bowl credit per pick, then maps those through one fixed regression per release onto the composite scale. The calibration coefficients, fit quality, and source-leaderboard checksum ship in the release data so the projection is reproducible.
The bar between the dots is the adjustment the model makes beyond the record: partial pooling with era and draft-capital context, plus the maturity discount. It runs in both directions. Thin tenures with dire raw numbers get pulled toward league average because a handful of picks proves little, and one legendary afternoon gets the same treatment. For most GMs the dots nearly coincide: the model mostly follows the record and intervenes only where evidence is thin.
Why not simply plot the raw per-pick numbers? Uncalibrated z-units would not share the composite’s scale, so the gap would conflate units with signal. And comparing only the raw posterior with the published composite would miss the pooling story entirely, because a rookie class’s raw posterior is already pooled. The fitted-projection choice keeps one honest axis.
People, GM-team careers, and support views
The release contains 190 people in 209 GM-team career rows. The Big Board ranks each person once while preserving a representative team context. The GM profiles use per-pick residuals to explain individual outcomes, and team pages show historical evidence trajectories. Those are support views, not separate headline ratings.
The unified-theta Stan route remains experimental/comparison work. It is not the source of the public leaderboard.
What it leaves out
Cap accounting is too indirect for this attribution question, and Hall of Fame recognition arrives too late to evaluate modern front offices. Full-career and any-team numbers can appear as player or draft context, but the headline only credits the model’s drafted-player outcome definition.