← Methodology MODEL CHECKS 2026-07-30-v30-validation

What the model gets right (and misses)

Three checks ask different questions. None is a single pass/fail grade, and none turns a support estimate into observed truth.

01

Read the checks separately

Posterior predictive check
Can the active v30 likelihood generate data shaped like the observations it was fitted to?
Five-year stability screen
How much do earlier raw and maturity-adjusted estimates move when compared with a later support-route estimate?
Sensitivity grid
How much does the public ordering change across the tested maturity-layer settings?

02

The likelihood reproduces broad means, but misses zeros and tails

Using 1000 posterior draws randomly selected with a fixed seed (and fixed simulation seeds) from the reviewed deep v30 fit, 158 of 346 screened summaries (45.7%) landed outside their replicated 95% ranges. The flags concentrate in exact-zero rates and upper tails for games, starts, AV, and honors.

This is material model-misfit evidence: Gaussian likelihoods for nonnegative career totals cannot reproduce exact point masses at zero, sometimes generate impossible negative totals, and understate several observed tails. It does not by itself show how much GM rankings would move under another likelihood; that requires a reviewed method comparison and a new fit.

Scope boundary: this PPC checks the raw seven-response likelihood. It does not evaluate the public maturity layer or the maturity-adjusted 95% credible intervals.

PPC flags by active response
ResponseFamilyObserved rowsFlagged cells
Games playedgaussian10,82936 / 50
Games startedgaussian10,82940 / 50
Pro Bowlspoisson10,82924 / 50
AP1 All-Pro seasonspoisson10,82916 / 50
Position-normalized snapsgaussian3,9864 / 48
On-team AVgaussian10,82934 / 50
AV/snap qualitygaussian2,7034 / 48

Flags are per-cell diagnostics, not a family-wise pass/fail test; some flags are expected when many correlated summaries are screened.

03

The public layer is conservative, not empirically calibrated here

Across 3,759 overlapping five-year comparisons from 170 GMs, a later raw support-route point fell inside the earlier raw interval 99.0% of the time and inside the earlier public interval 100.0% of the time. The public layer widened intervals, while its mean was slightly farther from the later point (0.174 versus 0.161 θ).

Those repeated comparisons are not independent. The percentages are stability rates, not coverage of true GM skill. The later target is another model estimate from the support-only trajectory route, whose fixed-effect prior and per-GM applicability logic are not identical to the headline route.

04

Rank order is stable across the tested maturity settings

The active maturity formula reconstructs the checked-in public score and interval to within 4.4e-16. Across 36 combinations of player-year cap, prior strength, and prior uncertainty, there are 12 distinct point-ranking transforms (prior uncertainty changes intervals, not point ranks). Rank correlation never fell below 0.998. Top-quintile overlap stayed at or above 94.7%, bottom-quintile overlap at or above 94.7%, and the largest individual movement was 15 places.

That supports robustness to the tested publication-layer knobs. It does not test alternate outcome families, priors inside the brms model, block definitions, attribution rules, or the experimental latent route.

05

Release interpretation

The evidence supports a cautious reading: the maturity layer is mechanically reproducible and its ordering is insensitive to the tested settings, but the production likelihood has visible distributional misfit. The current evidence does not authorize a headline-model change, a new release status, or a claim of empirical calibration.

The public bounds remain the release’s maturity-adjusted 95% credible intervals; technically, they are moment-matched uncertainty ranges and do not carry an empirical coverage guarantee. A future method review should compare zero-aware and heavier-tailed families on held-out or cross-validated performance before any refit is considered for promotion.

Download the evidence

Every file belongs to the dated check run and is fingerprinted by its archive inventory.