The cleanest test in the residual audit is this: residual honors, conditioned on the position-and-pick baseline, against drafting team. If the model is well-specified, this should be flat. Two GMs drafting the same caliber of player, in the same era, into the same role, should produce the same residual on average.
It is not flat. The team-level effect on Pro Bowl residuals is small but emphatically significant: η² = 0.0071, F = 2.51 across 32 teams, p < 1e-5. All-Pro residuals show the same pattern at lower magnitude (η² = 0.0053, p = 0.003). This survived the Phase 1.5 calibration cleanly; it's not a leakage artifact.
The teams
Sorted by mean honors residual (the average of rz_pb and rz_ap), 1980–2025:
| Top 10 | rz_pb | rz_ap | mean |
|---|---|---|---|
| BAL | +0.17 | +0.16 | +0.16 |
| KC | +0.15 | +0.11 | +0.13 |
| DAL | +0.14 | +0.09 | +0.12 |
| SF | +0.07 | +0.14 | +0.11 |
| PIT | +0.07 | +0.09 | +0.08 |
| MIN | +0.07 | +0.06 | +0.07 |
| PHI | +0.07 | +0.04 | +0.05 |
| DEN | +0.05 | +0.03 | +0.04 |
| TEN | +0.04 | +0.04 | +0.04 |
| MIA | +0.02 | +0.02 | +0.02 |
| Bottom 10 | rz_pb | rz_ap | mean |
|---|---|---|---|
| JAX | −0.13 | −0.09 | −0.11 |
| NYJ | −0.11 | −0.08 | −0.10 |
| CIN | −0.09 | −0.10 | −0.09 |
| ATL | −0.07 | −0.11 | −0.09 |
| ARI | −0.07 | −0.08 | −0.08 |
| LV | −0.11 | −0.03 | −0.07 |
| HOU | −0.07 | −0.06 | −0.06 |
| BUF | −0.03 | −0.08 | −0.06 |
| LAR | −0.10 | +0.01 | −0.05 |
| CAR | −0.06 | −0.00 | −0.03 |
The spread is roughly a quarter of a standard deviation between top and bottom: not enormous in absolute terms, but real, persistent, and orthogonal to the actual on-field performance the model is already conditioning on.
Not market size, exactly
The market-size hypothesis predicts that the largest media markets attract more honors votes. The data partly fits: Dallas, San Francisco, Philadelphia, Pittsburgh are top-half. But Kansas City sits at #2, Tennessee at #9, Minnesota at #6; none are big-market teams. Meanwhile the New York Jets are second from the bottom; the Bears (a top-five media market by population) sit just outside the top-10. The pattern is closer to sustained franchise quality than market size: the teams in playoff contention more often, more decades.
A cleaner reading: voters watch the same forty-some primetime games per year, and players on perennial winners show up in those windows more than equally-talented players on bottom-feeders. A first-team All-Pro vote is a top-five-at-position award; if you are a top-five edge rusher having a quiet year on a 5-12 team, you might not make it. Same stat line on the Ravens, you might.
Why we preserve it instead of fixing it
The three-category triage is: ARTIFACT (suppress), ERA-ECONOMIC (normalize), REAL-STRUCTURAL (preserve). This trend is REAL-STRUCTURAL. It's not a model bug; it's a real fact about how honors get distributed. We don't add a "team prestige" covariate to the baseline because doing so would launder credible team-level effects out of the composite. If a GM systematically puts players in environments where they win honors, that's part of the GM's job.
The implication for the leaderboard: a couple of percent of a Ravens or Chiefs GM's composite θ is, mechanically, "drafted into a context that votes well for honors." We are explicitly not removing that.