Backtested against real fantasy outcomes, per position, per year.
Earned by math. Verified by outcomes.
Every season the engine reruns as a walk-forward backtest — train on year Y, grade the preseason projections against what actually happened in year Y+1. The numbers below are the result. They don't get adjusted toward a target. Outliers don't get excluded. The cohort doesn't get curated.
An honest note on leakage. Engine parameters and rookie baselines are not frozen per reference year, so these columns are next-season-out-of-sample for the predictions but not a fully leakage-free holdout. For a clean out-of-sample read on the rookie model specifically, our leave-one-class-out estimate is MAE ≈ 3.46 (vs 3.27 in-sample — a +0.19 optimism gap), still well below the 4.71 of the prior approach. The 2026 season is reserved as a pre-registered, untouched holdout for prospective scoring.
How to read the table. r is correlation between projection and actual. Closer to 1.0 is better. MAE is mean absolute error in fantasy points per game. Lower is better. bias is the average over/under by position. Positive means the engine projected high. The volume-weighted variants downweight small-sample players, and that's the version that matters for projection trust.
Every player the engine ranked who took a snap is scored, and a player's season points are divided by the FULL SEASON rather than by the games he happened to be available for. That matches the projection, which is already discounted for expected availability.
Prediction: dynasty_ppg ·
Actual: points / games_in_season ·
Minimum games to be scored: 1
Break in series (2026-08-09). Figures published before this date scored an availability-discounted projection against points-per-game-PLAYED. That mismatch made the engine look systematically pessimistic at every position. Both sides now include availability, and the reported bias fell accordingly — the model did not change, the measurement did. Do not compare these numbers with earlier ones as though the method were the same.
These figures describe a single dynasty league
configuration — the one backtest.py builds. They are not
a measurement of every board SignalTuned serves; see the note below the
redraft table.
These are next-season figures, and that is
not a measurement of lifetime value. Each reference year is scored
against the season immediately after it, on a per-game rate
(dynasty_ppg). The dynasty board itself is ordered by
discounted value across a multi-year horizon
(ltv_discounted) — a different quantity over a longer
window. A player we correctly identify as a five-year asset who has a quiet
next season counts against us here; an aging player with one good year left
counts for us. Read this table as "how well did we call next season", which
is a fair question and not the one dynasty is for.
The horizon measurement now exists, and it says
something different. Scoring ltv_discounted against
realized discounted value gives rank correlations of 0.64–0.72 across
seven reference years — better ordering than this table shows. But it
also shows the horizon projections run HIGH: on the one reference year whose
seven seasons have finished (2018, n=813), running backs are over-projected
by 41% and tight ends by 34% over the horizon. Those two results are not
comparable to the numbers in this table — different predictor, and a
cohort that keeps players who left the league at zero instead of dropping
them, which makes ordering look easier. Calibration rests on a single
cohort until the engine emits value at multiple horizons.
The cohort these figures are computed on. Across 9 reference year(s), 7,361 board rows were scored against realized seasons and 4,447 matched (60%). Of those, 569 are rookies whose board id is synthetic — a player with no NFL snap has no league id yet — and had to be resolved by name. Before 2026-08-11 those rows were dropped silently, which made these figures better than they should have been, because rookies are the hardest players on the board to project. Restoring them lowered the rank correlations. Unmatched: 2,455 with no scored season (did not play, or missed the games filter), 454 unresolved rookie ids (never appeared in the scored season), 5 listed at a different position in the stats. The remaining gap is players who did not play, and that is a survivorship filter we have not fixed yet: this table still grades us mostly on players who stayed on the field.
| Position | 2021 | 2022 | 2023 | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| n | r | MAE | bias | n | r | MAE | bias | n | r | MAE | bias | |
| QB | 81 | 0.779 | 3.22 | -1.39 | 75 | 0.631 | 4.56 | -1.49 | 73 | 0.792 | 3.73 | -1.76 |
Volume-weighted: 2021: MAEw 3.99 · biasw -2.05 · 2022: MAEw 5.13 · biasw -2.41 · 2023: MAEw 4.64 · biasw -2.89 | ||||||||||||
| RB | 137 | 0.598 | 3.21 | -0.99 | 126 | 0.646 | 2.93 | -0.88 | 119 | 0.645 | 2.89 | -0.39 |
Volume-weighted: 2021: MAEw 4.20 · biasw -3.40 · 2022: MAEw 4.38 · biasw -3.33 · 2023: MAEw 3.66 · biasw -2.94 | ||||||||||||
| WR | 198 | 0.710 | 2.43 | -0.98 | 193 | 0.728 | 2.28 | -0.65 | 203 | 0.735 | 2.30 | -0.64 |
Volume-weighted: 2021: MAEw 3.15 · biasw -2.29 · 2022: MAEw 2.83 · biasw -1.81 · 2023: MAEw 2.84 · biasw -1.79 | ||||||||||||
| TE | 111 | 0.776 | 1.75 | -0.68 | 103 | 0.751 | 1.90 | -0.81 | 107 | 0.763 | 1.87 | -0.71 |
Volume-weighted: 2021: MAEw 2.42 · biasw -1.99 · 2022: MAEw 2.61 · biasw -2.01 · 2023: MAEw 2.50 · biasw -1.87 | ||||||||||||
Scoring format: half_ppr_te_premium ·
reference years 2015-2024 · generated 2026-04-26T12:07:16.489384
| Position | mean r | mean MAE | mean bias | top-N hit rate | Years |
|---|---|---|---|---|---|
| QB | 0.484 | 2.967 | 0.239 | 0.658 | 10 |
| RB | 0.684 | 2.893 | -0.123 | 0.692 | 10 |
| WR | 0.743 | 2.269 | -0.343 | 0.663 | 10 |
| TE | 0.747 | 1.980 | -0.089 | 0.642 | 10 |
What is still not measured. SignalTuned serves eight board types — dynasty and redraft, each superflex and 1QB, each with and without TE premium. The dynasty figures above come from one dynasty league configuration; these redraft figures come from one scoring format. The other six are served but not yet measured, and no number on this page should be read as covering them.
Publishing accuracy receipts means publishing the misses too. These are the calibration gaps we know about, what causes them, and what's planned.
Decision engine v3.4 (2026-08-28, mock 10, G-490/G-491). Mock 10 was played 100% engine-obedient; total leak vs engine-optimal was ~3-4 board points concentrated in two picks with one root cause. G-490: (a) an open-slot row's ladder gone-class is now MIN(own raw survival read, position clock) instead of clock substitution - substitution let a 98%-safe Quentin Johnston inherit the dying WR clock (top WR dying but same-team-discounted) and become THE PICK over Kyle Pitts at 5.05, which cascaded into a TE-tier loss (Pitts went 6.06, Warren 5.07); Lawrence stays demoted, mock-8 Nix stays urgent; (b) the position clock and per-position VONA anchor on the best row per position by OUR BOARD (blended_rank_score), not the first row in the game-theory re-rank, which need-weighting can reorder (3.05: all-WR want-set with Breece Hall priced off the board); (c) the hero-card pass-reason branches on roster fit before survival - 'Goff is 79% taken-by-then, so take Shakir' stated a dying player's doom as a reason to pass him; the actual reason was a filled QB slot. G-491: momentum above 1.0 decays toward 1.0 by the fraction of seats that can still START the position - Caleb Williams was priced 30%-to-survive during a QB run that was already finished (zero of the next nine picks took a QB).
A cell is flagged when every reference year agrees on the direction of the error. Published to make the self-correction loop visible — including the part that argues against taking any single row too seriously.
Read the denominator first. Of 46 cells tested, 2 cleared the agreement bar, and roughly 1.6 would clear it by chance alone. Agreement also gets EASIER as evidence gets thinner — fewer signals means fewer chances to disagree — so the flags concentrate in the sparsest cells rather than the best-measured ones. That is a property of the test, not a finding about those ages.
What this does and does not support. 2 flags against 1.6 expected is more than chance would produce. It is not support for any single coefficient: correcting for 46 comparisons puts the bar at P=0.00109, and the best-supported row here is P=0.250. So the rows below are evidence, not a list of changes to make. They do not auto-apply; each is reviewed against sample size and stability before anything ships.
And a reason to distrust the direction. Every position under-predicts, with bias at 82–92% of MAE — nearly all the error is one constant offset rather than scatter, which a calibrated model does not do. Accuracy is measured only on players who cleared a games-played floor of 8, which across 2015–2025 excludes 41% of everyone who took a snap — and 54% of quarterbacks. Busts and the injured leave the pool; the survivors, who outperform the cohort actually ranked, are all that is scored. Measured directly, that selection accounts for about 88% of the apparent pessimism at QB and roughly half of it at the skill positions. That alone would make the engine look pessimistic everywhere, and these age flags may be that same offset re-expressed per age bucket. Being tracked as G-347; until it is resolved, read the direction of these rows as unproven.