Backtested against real fantasy outcomes, per position, per year.
Earned by math. Verified by outcomes.
Every season the engine reruns as a walk-forward backtest — train on year Y, grade the preseason projections against what actually happened in year Y+1. The numbers below are the result. They don't get adjusted toward a target. Outliers don't get excluded. The cohort doesn't get curated.
An honest note on leakage. Engine parameters and rookie baselines are not frozen per reference year, so these columns are next-season-out-of-sample for the predictions but not a fully leakage-free holdout. For a clean out-of-sample read on the rookie model specifically, our leave-one-class-out estimate is MAE ≈ 3.46 (vs 3.27 in-sample — a +0.19 optimism gap), still well below the 4.71 of the prior approach. The 2026 season is reserved as a pre-registered, untouched holdout for prospective scoring.
How to read the table. r is correlation between projection and actual. Closer to 1.0 is better. MAE is mean absolute error in fantasy points per game. Lower is better. bias is the average over/under by position. Positive means the engine projected high. The volume-weighted variants downweight small-sample players, and that's the version that matters for projection trust.
| Position | 2021 | 2022 | 2023 | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| n | r | MAE | bias | n | r | MAE | bias | n | r | MAE | bias | |
| QB | 34 | 0.637 | 5.47 | 4.24 | 31 | 0.466 | 6.20 | 4.40 | 34 | 0.699 | 4.58 | 3.28 |
Volume-weighted: 2021: MAEw 5.16 · biasw 3.82 · 2022: MAEw 5.46 · biasw 3.95 · 2023: MAEw 4.79 · biasw 1.65 | ||||||||||||
| RB | 70 | 0.590 | 4.18 | 3.46 | 70 | 0.641 | 3.40 | 2.62 | 73 | 0.696 | 3.56 | 2.78 |
Volume-weighted: 2021: MAEw 2.85 · biasw 1.41 · 2022: MAEw 3.01 · biasw 0.40 · 2023: MAEw 2.25 · biasw 0.15 | ||||||||||||
| WR | 119 | 0.760 | 3.20 | 2.86 | 123 | 0.673 | 2.80 | 1.86 | 113 | 0.734 | 2.79 | 2.11 |
Volume-weighted: 2021: MAEw 2.66 · biasw 1.92 · 2022: MAEw 2.56 · biasw 0.80 · 2023: MAEw 2.30 · biasw 1.02 | ||||||||||||
| TE | 61 | 0.686 | 2.50 | 1.51 | 59 | 0.730 | 2.17 | 1.15 | 63 | 0.676 | 2.36 | 1.24 |
Volume-weighted: 2021: MAEw 1.99 · biasw 0.50 · 2022: MAEw 2.11 · biasw 0.30 · 2023: MAEw 2.13 · biasw 0.20 | ||||||||||||
Publishing accuracy receipts means publishing the misses too. These are the calibration gaps we know about, what causes them, and what's planned.
Overnight sprint combined ship (Ryan-directed acceleration; evidence _sprint_v1131/, combined A/B tune affected n=384 dMAE -0.1014 CI [-0.1893,-0.0187], pooled -0.0181 CI excl zero, validation no-harm; ablations replicate workstream receipts, interaction sub-additive). G-38 cameo-mover gate (xfp_base.volume_reproject.qb_min_ref_sf 0.35): the Bridgewater-29.6 class was the G-31 mover volume-reproject applying a cameo-derived factor (clipped at cap 2.0) to a starter-upgraded rate; inflating factors now require ref-season sf >= 0.35, docks kept, DISC-R8 guard remains backstop. G-35 qb_starter_status_prior (post-shrinkage): non-established QBs blended toward measured class realizations (backup class projected 15.4 vs realized 9.1); tune affected dMAE -1.68 CI [-2.46,-0.92]; validation neutral no-harm (disclosed); QB12->24 rate spread 4.12->5.25. G-34 y2_rate_uplift: Y2 veterans +1.4/+0.9/+0.6/+0.5 ppg (QB/RB/TE/WR, 0.5x measured cohort bias, ungated by round); tune affected n=280 CI excl zero, WR slice CI excl zero; bust-slice cost +0.16/+0.30 disclosed; Y3 = receipted null. G-41 reverse-spike relief ships DARK (_enabled false, 0=legacy proven): tune null-ledgered, recent-era revisit signal logged. G-40 two_way_position_overrides to config (meta + stats rows; map == legacy hardcode, zero board effect). G-43 publication stat lines (redraft.publication_stat_lines): per-component lines derived with the same recency weights, scaled to reconcile exactly (stat_line_scale audit), ranking_sanity reconciliation gate tol 0.15 ppg; rookies ppg/total-only v1; publish_redraft hook config-conditional. G-44 marketing gate fail-closed (missing/unreadable/malformed backtest_results.json -> BLOCK; explicit kill switch, distinct check id). New goals G-47/G-48/G-49; G-45 keep-verdict diagnosis and G-46 targets-basis diagnosis on file.