SignalTuned SignalTuned

Rankings tuned to the signal, not the noise.

Need help? Sign in

Model Card

Every engine version. Every change. Every accuracy review.

Current production version: v1.6.175. 200 releases on record.

Every change documented. Every version archived.

SignalTuned doesn't ship "the latest rankings." It ships a specific, versioned engine with every parameter spelled out. When a parameter changes, the version increments and the change shows up here with the reason. When the engine runs an accuracy review against real outcomes, the results show up here too.

Every ranking the engine produces carries a manifest_hash tied to one of these versions. Click the hash on any rankings page to see which engine version generated that output.

v1.6.172

2026-08-28

Decision engine v3.4 (2026-08-28, mock 10, G-490/G-491). Mock 10 was played 100% engine-obedient; total leak vs engine-optimal was ~3-4 board points concentrated in two picks with one root cause. G-490: (a) an open-slot row's ladder gone-class is now MIN(own raw survival read, position clock) instead of clock substitution - substitution let a 98%-safe Quentin Johnston inherit the dying WR clock (top WR dying but same-team-discounted) and become THE PICK over Kyle Pitts at 5.05, which cascaded into a TE-tier loss (Pitts went 6.06, Warren 5.07); Lawrence stays demoted, mock-8 Nix stays urgent; (b) the position clock and per-position VONA anchor on the best row per position by OUR BOARD (blended_rank_score), not the first row in the game-theory re-rank, which need-weighting can reorder (3.05: all-WR want-set with Breece Hall priced off the board); (c) the hero-card pass-reason branches on roster fit before survival - 'Goff is 79% taken-by-then, so take Shakir' stated a dying player's doom as a reason to pass him; the actual reason was a filled QB slot. G-491: momentum above 1.0 decays toward 1.0 by the fraction of seats that can still START the position - Caleb Williams was priced 30%-to-survive during a QB run that was already finished (zero of the next nine picks took a QB).

v1.6.173

2026-08-28

G-492 (verification mock 2026-08-28, 3.05 + 5.05): (a) multi-open-slot positions now wait-deflate on the same market-survival factor as single-open ones, on a higher floor multi_open_wait_floor=0.75 - a two-open WR shelf whose board-best was 78%-safe monopolized the want-set over a dying board-#1 Breece Hall at an open RB slot; (b) the G-490 position-clock anchor now applies the G-486 same-team discount when choosing each position's best row (market-faller waiver mirrored), so a discounted teammate cannot keep his position's need urgent while the selector refuses to draft him - which is how a 98%-safe Quentin Johnston became THE PICK at 5.05.

v1.6.174

2026-08-28

G-493 (mock 12 2026-08-28, 3.05): the want-set ladder's final board-order tiebreak now applies the G-486 same-team discount (board_rank_score x same-team multiplier; market fallers keep 1.0). The sel bucket already carried the discount, but a bucket tie - the ladder's own definition of close - fell through to raw board order, letting a same-team Jameson Williams lead Rice/Higgins at 3.05 with ASB rostered and no market fall. No new knobs.

v1.6.171

2026-08-28

Decision engine v3.3 (2026-08-28, mock 9, G-486/G-488/G-489). G-486: same-team WR/TE target-share discount - ASB went 1.05 and Jameson Williams (DET) was THE PICK at 6.05; two premium WRs on one NFL team cap each other's ceiling, so a candidate WR/TE sharing a team with a rostered WR/TE takes a need discount (same_team_penalty) scaled by the rostered player's draft capital ((7-round)/6), faded after same_team_fade_round; a discount, never a gate - a real market fall punches through; QB stacks and RB handcuffs untouched. G-488: ladder urgency reworked twice over - (a) a row filling an OPEN dedicated slot carries its POSITION's clock (gone-prob of the position's best option, from _gone_pos) instead of its own: a dying 5th-QB Lawrence (12% survival) was THE PICK over a safe Nix/Goff/Stroud shelf at 9.05, while mock-8 Nix (shelf top dying) stays urgent; depth rows keep their personal raw read (Pierce over a 70%-safe Addison); (b) the primary sort key is a BINARY likelier-gone-than-not class on the RAW (pre-demand-deflation) read - among already-dying players fine survival ordering is noise (QJ 32% vs Addison 9% cannot both be kept: the better one leads), and G-485's quarter-buckets over deflated urgency compressed dying-vs-safe into one bucket. G-489: recs now serve wait_gone (raw wait-horizon gone-prob); the hero card's survival claims read it instead of the on-clock-zeroed denial_prob, which had the card telling Ryan a 32%-survival QJ had a 100% chance to still be there.

v1.6.170

2026-08-27

Decision engine v3.2 (2026-08-27, mock 8, G-485). Mock 8, 10.05: THE PICK was Tyrone Tracy Jr. (98% survives-to-next-pick, +43 vs crowd) over Bo Nix, whom the market model gave 29% to survive - the room took Nix three picks later and Ryan had to override a round after that to rescue Goff, the last startable QB. Three coupled fixes: (a) the G-481 wait factor's windowed-need seats term is replaced by market survival of the position's top option - rooms draft by ADP, not open slots (second QBs went all draft), so a single open slot defers only while its best option is likely to STILL BE THERE; (b) the G-474b zero-demand hard zero on decision urgency is removed - G-471's demand_floor deflation (raised 0.2 -> 0.45) keeps opponent-need information without silencing the market; (c) want-set sequencing is perishability-first (urgency quarter-buckets, then value x need 0.1-buckets, then our board) - every set member is a player we intend to take with remaining picks, so who dies first orders them; both-gone/both-safe tie through to value then board (G-480b preserved).

v1.6.169

2026-08-27

Decision engine v3.1 (2026-08-27, mock 7, G-481..G-483). G-481: the v168 demand/supply factor counted supply inside a value band of the position's top anchor - dense-anchor positions (RB/WR) read as bottomless supply and had their need crushed ~3x while sparse-top TE/QB kept full need: the exact INVERSE of the wait edge (McBride as THE PICK at 1.05 against his own 66% survives-to-next-pick read; Hurts at 3.05; all-TE/all-QB also-fits; Judkins over a 0-of-3-WR hole again). Replaced with a per-position VONA wait factor: a SINGLE open dedicated slot defers (down to wait_floor) only when the position's best available loses ~nothing to waiting (vona/vona_ref, scaled by windowed demand seats/2); MULTIPLE open dedicated slots are never wait-discounted - a 0-of-3 WR hole is fill pressure by definition. Urgency coverage-decay removed (same band pathology - it zeroed decision_urgency board-wide in prod); sequencing buckets urgency by quarters so sub-0.25 urgency ties break on OUR BOARD. demand_supply_floor/gain retained but unused. G-482 (adp_engine): parse_sleeper_board stamped K/DEF position=None, so the G-458 K/DEF market-row gate matched nothing and endgame fired into a pool with zero K/DEF rows in every mock - the actual root of the three-mock K/DEF failure; positions now preserved (DST normalized to DEF). G-483 (war-room.css): own filled cells brightened (38%/20% mix + ring) - they read DIMMER than opponents' picks.

v1.6.168

2026-08-26

Decision engine v3 (G-477..G-480, Belichick mock 6). G-477 ROOT CAUSE: redraft cache path served boards with no LeagueConfig -> _league_slot_counts empty -> endgame K/DEF fill, demand urgency, TE flex discount and bench coverage ALL silently disabled on the live mock path (harness passed: it passes roster_slots explicitly). Slots now resolve from the config on the cache path and from Sleeper draft settings slots_* engine-side. G-478: slot-aware need - open dedicated slots at full need (mild open-count x pressure scaling), dedicated-full = flex depth, need x demand/supply factor (wide-band startable pool) so low-demand positions defer at SELECTION. G-479: bench need = gap to depth targets (RB/WR starters+flex+2; TE/QB starters only) - inverts the req/rostered formula that made TE2 the 'thinnest' spot. G-480: want-set ladder - value x need bucket, then perishability bucket, then OUR BOARD; tier urgency no longer multiplies. Serve-time only; backtests unaffected.

v1.6.167

2026-08-26

Decision engine v2 (mock 5, G-473..G-476). SELECT-THEN-SEQUENCE (G-475): THE PICK's candidate set is the top-N rows by OUR value x need, N = remaining discretionary picks (picks left minus unfilled required K/DEF); urgency can only sequence within the set, never promote into it (Worthy - market's #1 remaining, our rank 16 lower, value gain -6.7 - was THE PICK at 12.05 purely on 'won't last'). TWO-PHASE NEED: starter phase keeps flex-aware need with TE's flex share discounted in non-premium leagues once the dedicated TE slot fills (G-473 - Pitts as TE2 'open starter slot' at 6.05); bench phase switches to per-position depth COVERAGE, required starters / rostered, never pooled (G-476 - RB 2-of-2 was being TAXED for the WR overflow sharing its flex pool while Ryan sat one injury from an empty lineup slot). TIER URGENCY (G-474): urgency = windowed demand seats vs tier supply (players within supply_band of the position's best remaining value) - five QBs for five needy seats means waiting is free; ZERO windowed demand defers the position outright (G-474b - 'nobody else needs one... this is where we get the edge by pushing qb not just taking one'); player-level denial only breaks ties within a tier. Endgame diagnostics now count fill candidates per open required slot (the mock-4/5 TE-over-empty-DEF failures produced nothing on the wire to debug). Hero copy gated: 'unlikely to last' only when decision_urgency backs it, with the percentage shown (Stafford was fronted by a sentence the engine had measured false). Serve-time only; cached boards and walk-forward artifacts unaffected beyond the config-hash restamp.

v1.6.166

2026-08-26

Mock-4 live fixes (G-470/G-471). G-470: the wait horizon. On the clock picks_until_me==0, so every P(gone before my pick) was computed over ZERO picks - trivially ~0 board-wide - flattening the G-466 urgency to its floor and making VONA read 'waiting is free' at the exact moment of decision; THE PICK fell back to need-adjusted raw value (live: Quentin Johnston, market-priced 4-5 rounds later, recommended at 6.05). The horizon now falls through to the distance to my NEXT owned pick (owned_cells, else snake/reversal math); auctions keep 0 - no slot order, waiting genuinely free. Applied to the decision urgency AND VONA (test_g197 on-clock contract rewritten accordingly). G-471: demand-aware urgency. The denial model priced demand off league-wide ADP and never looked at the rosters actually picking in between - Bo Nix was THE PICK at 9.05 with 'the market is about to take him' while all eleven other teams already had their QB. The urgency denial is now multiplied by the fraction of distinct windowed seats with an open slot the position can start in (dedicated; flex share RB/WR/TE; SUPER_FLEX for QB), floored at decision.demand_floor (0.2) for off-need stashing. Scoped to the DECISION urgency; displayed denial chips keep the pure market read. New wire field recommendations[].decision_urgency; hero copy claims 'unlikely to last' only when the demand-aware number backs it.

v1.6.165

2026-08-26

Mock-3 batch A+C (G-466..G-469). G-466: THE PICK becomes a DECISION - server stamps decision_score = anchor(board value percentile) x need_mult x (survival_floor + (1-survival_floor) x P(gone before my next pick)); the SPA hero orders by it instead of raw value (which had recommended TE2 at 4.05 and TE3 at 14.05 as depth adds while WR/K/DEF starter slots sat empty, and 3-round reaches with zero wait-awareness - Ryan: 'best available and likeliness to go anytime soon ... thats where the edge is'). Endgame roster-completion: when remaining picks <= unfilled required dedicated starters, rows filling one (K/DEF included, market-ordered) dominate the decision order - K/DEF carry no engine score so nothing else could ever surface them as THE PICK. Signal Board stays a pure value board; hero copy explains when the decision pick differs from the raw-value #1 and why. Config: draft_board.decision {enabled, survival_floor 0.55, endgame_fill} + schema. G-467: the contention-window strategy lens is disabled for near-redraft keeper leagues (derived/explicit horizon <= 2) exactly like redraft - Auto had resolved to 'Build' in a keeps-2 league and the youth lens ordered a rookie TE over a higher-scoring vet WR. G-468: The Room's subtitle and status dot now name the actual pricing source from vibe_source (Sleeper ADP market vs VibeRank crowd) instead of hardcoded VibeRank copy. G-469: DEF position color moved from #B08968 to #7A5238 - the first brown sat too close to the TE orange on the draft grid. UI-only changes carry no ranking effect; decision layer affects live-draft recommendations only (no cached boards, no backtests touched by design - engine_config hash moves, so walk-forward artifacts restamped).

v1.6.164

2026-08-25

Belichick mock-draft QA batch (G-459..G-465). G-459: max_keepers added to _config_fingerprint + _H2_DRIFT_FIELDS and threaded through the boot-warm reconstructor (data_status), the admin fp-repair endpoint (caught by the armed G-363 parity guard) and the manual request models; keeper-horizon resolution stamped into rankings meta (meta.keeper_horizon) with a logger.warning on refusal. Root cause: the boot warm rebuilt keeper configs without max_keepers, the G-392 derivation refused silently, and the full 7-season dynasty board was cached under the same fingerprint the correct config computes - reconnect/roster-sync/drift all left it standing. G-460: build_live_board now passes top_n=len(available) to rank_picks (default-20 cap starved the surfaced payload; the per-position backfill was dead code; client Flex filter showed 1 player). G-461: surplus QB in a 1QB league (no SUPER_FLEX seat) drops straight to the need floor instead of the graduated curve (0.65 QB2 was the least-taxed depth on the board; live all-QB Signal Board rounds 8-12). G-462: war-room ADP column resolution passes keeper through (resolver maps keeper to the redraft tier itself) and reads rec pts from the resolved LeagueConfig scoring before the live probe (mocks carry no league_id; full-PPR keeper league was priced off ADP dynasty half ppr). G-463: market/crowd vibe stamps now sync into the ranked copies in the same response (Room ALL-tab warming-up flap, one-poll-stale VIBE, K/DEF pinned to top) and K/DEF market rows carry synthetic unique player_ids. Scarcity: value_over_replacement_ranking.descarcity_horizon_factor flipped TRUE (G-391 double-count removal, measured McBride TE #5->#12 on the shallow-bench sibling league); baseline/elasticity recalibration stays parked pending harness extension. engine_config.json rewritten ensure_ascii=True (fixes test_g244 non-ASCII bytes). UI (no engine effect): K/DEF position color tokens, strategy-lens sort-basis note, Auto lens shows its resolved strategy.

v1.6.163

2026-08-22

G-451: name-join fallback + ADP read-time guard, bundled so one walk-forward rebuild covers both. (1) adp_engine.match_to_rankings and market_layer._build_reconciliation gain two rescue stages consulted ONLY after the existing match logic misses: a position-scoped name_utils.name_key lookup (resolves Kenny/Kenneth Gainwell, Matt/Matthew Hibner - Sleeper publishes short forms that appear in NO players.csv name field, so Jaccard reads 1/3 and SequenceMatcher ~0.83 vs the 0.85 gate) and a (last-name, position, team) unique-triple index (resolves De'Zhaun-Ryan vs De'Zhaun Stribling); ambiguous keys refuse by construction and free agents never participate. The war room already used name_key (G-307 family) and is untouched. name_utils gains elijah->eli (Eli Raridon, players.csv first_name Elijah). (2) adp_engine.SLEEPER_BOARD_MAX_ADP = 700.0: _is_board_null_adp rejected only >= 999, so Sleeper's 700.0 ranking-horizon clamp (38 cells >= 700.0 on the 2026-08-22 export, max 700.9) ingested as real ADP; the refresh's build-time clamp protects only boards that build produces, a user's scoped upload never passes through it. validate_adp_board.py now imports the ceiling from the engine so gate and runtime cannot drift. Measured stake: Kenny Gainwell (RB TB, dynasty-SF ADP 139.8) missing from EDGE since at least 2026-08-11.

v1.6.161

2026-08-18

G-437d: three defects of one class - each resolves to NOTHING rather than erroring, which is why none was noticed. (1) POST /v1/admin/news-events validated player_name and took player_id on faith; it accepted the literal string '<the right id>', returned status ok, and wrote an event that could never match a player. Now checked via news_events.player_exists against the same crowd_player_ratings table resolve_player reads, 404 on unknown. (2) _apply_young_qb_ceiling's docstring claimed dual-threat QBs are excluded while archetypes_eligible lists dual_threat and every QB in the ref2024 top five IS dual_threat. (3) That whole method is inert under the current config - the Pattern 13b guard short-circuits it whenever qb_use_v2_ltv is true, and ceiling_bonus reads exactly 1.000 on all 5,790 rows across 7 ref years and every position - so nothing tests it and flipping the flag off re-arms a 1.55x multiplier. Both notes now describe the code. Deliberately NOT changed: the eligibility list or the flag; that is a modelling ruling with an accuracy consequence, and this is documentation plus validation.

v1.6.160

2026-08-18

G-437c: G-437B COULD NOT REACH THE PLAYER IT WAS BUILT FOR. Reported durations were gated behind three guards that belong to the team_change CALIBRATION, not to games-available ARITHMETIC: the dynasty-only framework check, an is_rookie skip at the top of the event loop, and the fact that RedraftEngine never calls the layer at all. Jordyn Tyson - the first real use, logged as event 47 with 'out 2 months' - is a 2026 rookie, so the event was accepted and could never apply. Fix: parse_duration_games + news_duration_correction move to availability_layer (the shared layer G-243 established for exactly this, so redraft does not grow a second copy that drifts - this repo already carries two rookie-id conventions from that failure mode); the is_rookie skip narrows to the team_change branch; RedraftEngine calls the shared helper on redraft_proj immediately after its live-availability block, inside the non-snapshot else so backtests stay inert. The distinction now stated in code: a calibration fitted on next-season PPG for veteran movers has no rookie or redraft meaning; (season_games - games)/season_games has both. Regression: test_g437b_duration_reaches_the_player.py, whose rookie and redraft cases were written to FAIL against G-437b.

v1.6.159

2026-08-18

G-437b: A REPORTED INJURY DURATION COULD NOT MOVE A PROJECTION. The news ledger already accepted event_type='injury' and the news/situation layer already converted a duration into a year-0 factor, but only for suspensions -- an injury event produced a NEWS flag and no price movement. This opens that gate (news_situation.injury, ships DISABLED) and adds parse_injury_duration_games, which reads games/weeks/months and takes the pessimistic end of a range, refusing anything unreadable so an unparseable note stays flag-only exactly as today. THE ORDERING BUG THIS AVOIDS: _apply_live_injury_adjustment (step 2h) has already multiplied projected_season_ppg by the severity tier before this layer runs at 2h-bis, and the LTV horizon loop multiplies both stamped columns, so stamping the duration factor directly would double-charge -- Tyson at 8 weeks and Questionable would price 0.529 x 0.92 = 0.487. The stamped value is therefore a CORRECTION RATIO, target / live_injury_year_0_factor, making the product equal the reported duration. Authoritative because a reported prognosis beats the cohort average G-437 measured; same reasoning as the suspension branch skipping players the Sleeper feed already marks suspended. Players zeroed by season-ending IR are skipped. Snapshot-guarded via the existing FF_SNAPSHOT_AS_OF_SEASON early return, so walk-forward caches stay byte-identical.

v1.6.158

2026-08-18

G-437: THE LIVE INJURY TIERS WERE CALIBRATED AGAINST NOTHING. engine_config.live_injury.severity year_0 factors moved questionable 0.98->0.92, doubtful 0.95->0.68, out 0.92->0.62. Measured from nflverse weekly injury reports joined to nflverse snap counts, 2016-2025 (sample starts after the league eliminated Probable and redefined Questionable for 2016 - published play rate moved 53% to 75%, so 2015 is a different designation wearing the same word; including it moves the values by 0.003): for each designation-week, the share of the player's remaining regular-season team games in which he took an offensive snap, divided by the same figure for undesignated established starters observed in the SAME season-week. Relative because year_0_factor multiplies a per-game rate. Established = offensive snaps in >=3 of the previous 4 team games; both arms pass the same gate. Results: questionable 0.919 CI [0.903,0.934] n=1766; questionable 0.922 CI [0.908,0.938] n=1633; doubtful 0.683 CI [0.629,0.738] n=161; out 0.616 CI [0.592,0.640] n=941. All three CIs exclude the shipping value; zero contradicting position slices; out too generous in 11 of 11 seasons. SANITY ARM: players ON the injury report carrying no game status measure 0.988 CI [0.974,1.001], so the metric separates injury from roster churn rather than measuring churn. TWO EARLIER PASSES WERE WRONG AND ARE KEPT AS RECEIPTS: grading against injury_spells is invalid because build_injury_recovery_data sets OUT_STAT={out,doubtful} and so excludes Questionable by construction; and inferring appearance from the presence of a weekly stat row put UNDESIGNATED starters at 0.695, worse than designated ones, because a healthy WR targeted zero times has no row. Snap counts fixed it. WHY NO A/B: G-252 makes this layer return early under FF_SNAPSHOT_AS_OF_SEASON, so it is invisible to every backtest and no walk-forward arm can grade it; the caches nonetheless rebuild because G-182 fingerprints the whole config file. That rebuild is itself the G-252 conformance test - historical ranked columns MUST NOT move. KNOWN INCOMPLETE: `out` carries a 0.302 calendar-week gradient (0.728 wk4-6 -> 0.426 wk13+) that a constant cannot express; it ships anyway because it is strictly closer than 0.92 in every week bucket. Week-awareness for `out` is the follow-on and is due before week 1. Spell age also carries signal for `out` (0.747 fresh -> 0.536 at 5+ weeks designated) but is confounded with calendar week and was NOT disentangled; status_since already exists (G-239 stamps it) and nothing in the pricing path reads it.

v1.6.156

2026-08-15

G-405: A BETTER OFFENSE WAS LOWERING A REDRAFT DRAFT SCORE. Ruled and enabled by Ryan 2026-08-15 -- the THIRD application of his 2026-08-14 law that a status change must never move a player the wrong direction, after FA status (2026-08-14) and availability monotonicity (G-397, v1.6.150). THE DEFECT, derived from redraft_engine.py rather than argued: val is computed at [3/8] as (redraft_proj - replacement_ppg); _apply_team_context then MUTATES redraft_proj (team_context.py:385) and re-clamps ceiling but NOT floor; val is never recomputed. draft_score is val*sv + cw*ceiling_premium - fw*floor_penalty + opp - aging, where ceiling_premium = (ceiling - redraft_proj).clip(0) and floor_penalty = (redraft_proj - floor).clip(0). So a boost of +D contributed NOTHING through val (stale), -cw*D through ceiling_premium and -fw*D through floor_penalty. At the live weights cw=0.3 and fw=0.1 the net was -0.4*D: a player moving to a BETTER offense ranked WORSE for it and one moving to a worse offense ranked better, while the user-visible 'Team context (offense change)' chip reported context_adj_pct with the OPPOSITE sign. Bounded by max_adj_pct=0.1, so a SIGN error rather than a magnitude one, on the board most users actually draft from. MEASURED ON THE LIVE 736-ROW BOARD, flag off vs on: 115 team-context movers, 113 draft_score changes, ALL 113 now agreeing with the chip, ZERO contradicting it, and ZERO non-movers touched. 294 rank changes are positional displacement from those 113 moving, not independent movement. Named cases: Deebo Samuel +10.0% adj, rank 54 -> 45, score +2.17 -> +2.76; Cedrick Wilson Jr. +10.0%, 141 -> 123, -0.92 -> -0.50; Andy Dalton +8.1%, 59 -> 49; Tutu Atwell -8.8%, 95 -> 115, +0.04 -> -0.37; Deven Thompkins -9.7%, 157 -> 167. THE REPLACEMENT BAR IS HELD, DELIBERATELY, and that was a measured choice rather than a stylistic one. The obvious implementation re-calls _compute_val, which ALSO re-derives replacement levels from post-context projections; measured, that moved RB 4.35 -> 4.40 and WR 4.96 -> 4.86, and because draft_score multiplies val by a PER-PLAYER availability factor a uniform bar shift reorders players who never changed teams -- 522 scores and 306 ranks moved for a fix aimed at 115 players. Holding the bar keeps the change attributable to exactly the players team context touched, per G-243's scope rule that the boundary is what makes the A/B readable. Whether the bar SHOULD reflect team context is a real and separate question, filed rather than smuggled in. A PREDICTION THE MEASUREMENT CORRECTED: the algebra says the fix only restores the law when sv > cw+fw = 0.4, so I predicted ~9 movers would still contradict the chip at the 0.30 availability floor. ZERO did. The arithmetic is still right and is pinned in tests; the population that would hit it is empty on today's board, most likely because G-397's positive_val_only forces sv=1.0 whenever val<=0 and the movers near the availability floor are also near val=0. Both facts are recorded rather than one rounded away. WHY A RULING AND NOT A HARNESS: no A/B can say whether a better offense SHOULD improve a draft ranking -- it can only report how many players move and by how much. Same substitution as the FA law and G-397, and deliberate. Shipped default-OFF first, measured, then flipped. TESTS: tests/test_g405_team_context_sign.py, 22 tests, including the exact sv = cw+fw boundary, the residual's arithmetic, and a guard that the replacement bar is not re-derived.

v1.6.155

2026-08-15

G-403: SLEEPER IS NOW REFRESHED DURING THE DAY. NO RANKED-OUTPUT DELTA -- this changes how often the box does I/O, never what any board says. This is the piece that actually moves the measured 16h15m, and the last item on the critical path for Ryan's D1 ('within the hour, tighter on game days'). THE GAP. The Sleeper table was force-pulled at exactly TWO moments: 02:00 (_roster_refresh_core) and 02:30 (crowd reseed). Every other reader passes force=False, and because the 02:00 pull resets the sidecar manifest, the 24h TTL never elapsed during the day -- so no request-path or roster-sync read ever refetched. A Sunday inactive announced ~11:30 ET reached the cache at the Mon 02:00 pull and reached boards at the Mon 03:45 warm: 16h15m typical, ~25h for a change landing just after 02:30. SHIPPED, FOUR PARTS. (1) TWO NEW JOBS: sleeper_refresh_hourly at :00 and sleeper_refresh_gameday on Sundays 10-23 every 15 min, both self-skipping outside the in-season month window {8,9,10,11,12,1} -- August included, because the draft and the Waller signing are both August events. Both do I/O ONLY and never engine work, which is what stops a tighter cadence from saturating the single recompute lane. (2) THE TTL IS A BACKSTOP, NOT THE REFRESH, and that distinction is the design. load_sleeper_players' own docstring says request-path reads keep force=False 'so a board view never triggers a 50MB download' -- true only while the cache is fresh, so cutting the TTL to chase freshness would have moved a ~50MB inline download onto board builds in a single-worker app. In-season it drops to 6h: tight enough that a DEAD scheduler self-heals in hours instead of a day, loose enough that it never fires while the jobs work, and well inside the 48h at which player_status stops trusting depth and injury. Refresh, then backstop, then refuse. (3) THE resolve(force=True) IN THE JOB IS LOAD-BEARING, and missing it would have made the whole chain silently inert: player_status memoises on the cache file's MTIME and peek_fingerprint answers ONLY from that memo, returning None ('cannot judge', so serve) on a miss. A fresh pull changes the mtime and leaves the memo keyed to the old one, so without re-resolving in the job, peek would return None until some board build happened to repopulate it and the G-399/G-401 drift gate would never fire. (4) THE PER-LEAGUE REBUILD COOLDOWN, shipped in the SAME change rather than after it. With injury in the fingerprint (G-401) and a 15-minute game-day pull, drift is near-continuous across 604 currently-flagged players; engine runs are serialised at ~35-105s apiece, so an unbounded enqueue would queue every league on every read and leave user-initiated roster syncs waiting behind drift work. Default 900s, matched to the game-day pull interval on purpose -- anything tighter buys nothing because the data underneath has not moved again. Throttled reads still SERVE and still flag; only the enqueue is suppressed, and rankings.read.status_drift_throttled makes the throttled:enqueued ratio visible before anyone retunes it. ALSO CLOSED: scripts/daily_data_refresh.py's 'invalidation' was a Path.touch() running in a SUBPROCESS, structurally unable to reach this process's caches; the parent now calls data_loader.invalidate_player_data() itself after the child succeeds. OPERATIONALLY: sleeper_refresh.dry_run=true logs what each firing WOULD pull, including current cache age, without fetching -- intended to be watched across one game-day window before going live. enabled=false restores the pre-G-403 cadence exactly. A NOTE ON HOW THE JOB REGISTRATION NEARLY SHIPPED BROKEN: JOB_REGISTRY is built at module-load time ABOVE where run_sleeper_refresh is defined, so registering beside the function raised NameError at import and would have taken the entire scheduler down on boot -- and py_compile passed it. Caught by importing the module rather than compiling it, which is now part of the verification.

v1.6.154

2026-08-15

G-401: THE DRIFT GATE COULD NOT SEE A SUNDAY INACTIVE. Self-correction on v1.6.152, same day. NO RANKED-OUTPUT DELTA -- the fingerprint is provenance, not a price. THE DEFECT. G-399 shipped a status-drift gate whose entire purpose is to stop serving a board that predates a live-status change, and it compares meta.player_status.fingerprint. That fingerprint hashed TEAM and DEPTH_CHART_ORDER only. A Sunday inactive changes neither: the player stays on his team, in his depth slot, and flips injury_status to 'Out'. So the gate built to catch Sunday inactives would have slept through every one of them, and the 16h15m latency finding would have been 'closed' by a mechanism structurally unable to observe the event. The fingerprint was not wrong -- it was built for G-182 roster drift and its docstring said 'roster' the whole time. G-399 reused it for a second purpose without re-deriving what that purpose required. MEASURED against the live 12,210-player Sleeper cache, 518 of which carry a designation (Questionable 294, NA 89, PUP 73, IR 49, Out 6, Sus 3, COV 2, DNR 2): ruling one depth-1 starter Out left the old recipe BYTE-IDENTICAL at ad54f68c124c1087, and moves the new one 0cbacb461b315bf3 -> 4f1573c9e036ec3f. SHIPPED: injury_status and injury_body_part join team and depth in the recipe. Both move a PRICE -- injury_status selects the severity tier, injury_body_part drives career-bender escalation (ACL / Achilles / neck / head / spine) -- which is the inclusion rule, not 'everything the feed carries'. practice_participation is DELIBERATELY EXCLUDED and a test documents why: availability_layer stamps it in _TEXT_COLUMNS for display and never reads it into _FACTOR_COLUMNS, so it moves no number while changing most weekdays for most listed players; including it would multiply rebuilds for nothing. The recipe now reads from snapshot.raw, defensively, so a gsis present in `teams` without a raw record cannot take down a board build. COST: changing the recipe invalidates every cached board exactly once, on the deploy that ships it -- which is free, because the caches are in-process and already empty after a deploy. There is no ongoing cost increase: Sleeper still refreshes only at 02:00 and 02:30, so the fingerprint still moves twice a day. That changes when the in-season cadence lands, and a per-league rebuild cooldown belongs in the same change. TESTS: tests/test_g401_injury_in_fingerprint.py, 16 tests, including one per live designation, the Questionable->Out transition, the Knee vs 'Knee - ACL' body-part case, no-regression on signings and depth moves, and a recipe test that fails if any field which moves a factor is missing from the hash. THE LESSON IS NOT 'ADD A FIELD'. A value built for one purpose was reused for a second without re-deriving its requirements, and nothing in the type system, the tests or the docstring objected. The recipe test exists so the next reuse has to state what it needs.

v1.6.153

2026-08-15

G-400: FOUR PLACES THE PLATFORM REPORTED SUCCESS WHILE DOING NOTHING. NO RANKED-OUTPUT DELTA on any live board -- one of the four changes what a BACKTEST produces, and that is the point. One failure mode, four instances, none of which could ever have surfaced as an error. (1) THE REDRAFT HARNESS HAS BEEN SCORING AN EMPTY BOARD. redraft_engine.run()'s snapshot guard read `return self._finish_live_availability_noop(...)` where it meant `proj_df = ...` -- and it is in the body of run(), not a helper. So under FF_SNAPSHOT_AS_OF_SEASON, run() returned a DataFrame and stages [3/8] through [8/8] never executed: no VAL, no ceiling/floor, no opportunity overlay, no expected_games, no draft_score, no ADP, no tiers, no _build_output. It survived because the consumer, validate_redraft_board_outcome.py:295 -> _board_df, does `result.get('rankings') or {}` -- and on a DataFrame .get() is a COLUMN lookup, so it missed, returned None, the loop never ran, and an empty board came back with no exception. Every reference year scored nothing and the run reported success. Fixed to an if/else so the live layer is still skipped on historical boards (the G-252 guarantee) while the rest of the pipeline runs. _board_df now RAISES on a non-dict result, so the class cannot recur silently. Production was never affected: the env var is not set on the serve path. Dynasty walk-forward was never affected either -- backtest.py mentions redraft exactly once, in a comment, and never builds a redraft board -- so no dynasty verdict is retracted by this. (2) ONE MISSING CSV KILLED EVERY CRON JOB. api.startup_event's step-4 gate returns when player_stats_season.csv / players.csv / engine_config.json are absent, logging 'Skipping startup rankings'. start_scheduler() ran BELOW that return, and looked like module-level code because a column-0 '# ENDPOINTS' banner sat between them -- Python emits no DEDENT for comment-only lines, so those statements were still inside the function and the banner disguised it perfectly. Consequence: a missing data file silently disabled the 02:00 Sleeper pull, the 03:45 league warm, the daily data refresh and the 5-minute cache reaper for the life of the container, and the scheduler is exactly what REPAIRS a missing-data boot (it carries a make-up data refresh), so the one condition that most needed it was the one that disabled it. Scheduler moved to step 3.8, above the gate; the gate's log is now an ERROR that states what it actually costs. A test fails if indented code ever follows that banner again. (3) THE STATUS RESOLVER LOGGED NOTHING, EVER. player_status.resolve() takes `log=print` in its signature and never calls it -- zero occurrences in the body. Every failure path only appended to snapshot.warnings. So a total loss of every live signal (no live fetch AND no cache file) produced no stdout line, no log record and no metric: the FA discount, buried depth, the QB depth gate, live injury and the mover predicate all silently stopped firing, and the only trace was meta.player_status.warnings on a payload nobody reads until something already looks wrong. This module's own header guarantee 5 -- 'Degradation is LOUD' -- was true of the snapshot object and false of the process running it. (4) THE 48h EXPIRY WAS EQUALLY SILENT. Two consecutive failed 02:00 pulls put the Sleeper cache past 48h; usable() flips False for depth AND injury; every board then prices injured and buried players as fully healthy and available -- which from outside is indistinguishable from a quiet week. Both now route through one _report() helper called once per real resolution (never on a memo hit, so a stale cache pages once per resolution rather than once per board), emitting logger.error plus counters player_status.unavailable and player_status.signal_expired. Those counters, and boot.warm_skipped_missing_files / boot.scheduler_start_failed, are registered in obs_metrics.WATCHDOG_COUNTER_THRESHOLDS at 1, so the hourly metrics watchdog escalates them to Sentry -- incrementing a counter nothing watches would be the same silence in a new costume. TESTS: tests/test_g400_silent_success.py pins each fix AND the structural trap that hid it, including an AST assertion that the snapshot guard contains no Return and a scan that no indented code follows the ENDPOINTS banner. THE THROUGH-LINE: in all four cases the code read correctly. Only running it, or diffing what it produced, exposed the defect -- which is now the fourth, fifth, sixth and seventh finding in this body of work that inspection would have missed.

v1.6.152

2026-08-15

G-399: A SLEEPER CHANGE NOW REACHES A CACHED BOARD. NO RANKED-OUTPUT DELTA -- this changes WHEN a board is recomputed, never what it says. THE DEFECT. _config_fingerprint (league_config_service.py:186) hashes league_id, format, num_teams, superflex, te_premium, scoring, roster_slots, my_user_id, keeper_horizon_years and status -- ten league-SHAPE fields and NOTHING about the data. So new Sleeper data has never changed a cache key and has never, by itself, caused a miss. The SPA's fast path (GET /v1/rankings?config_fp=...) then returns cached['result'] with no age check of any kind: BoundedTTLCache.get (cache_store.py:98-103) only does move_to_end + return, and the reaper judges on the value's write-time timestamp, which a read never refreshes. Net: the only things that have ever replaced a live board are the 03:45 nightly force-warm, a roster sync, and the reaper dropping the entry past 7200s. MEASURED PATH for a Sunday inactive announced ~11:30 ET: into sleeper_players_cache.json at the Mon 02:00 forced pull -- the only forced pulls are 02:00 and 02:30, every other reader passes force=False, and 02:00 resets the manifest so the 24h TTL never elapses during the day -- then onto boards at the Mon 03:45 nightly warm. 16h15m typical; ~25h for a change landing just after 02:30. THE FIX WAS ALREADY HALF-BUILT. Every board already stamps the status fingerprint it was built with, at meta.player_status.fingerprint, for exactly this purpose ('two boards with the same fingerprint saw the same roster'). Nothing ever compared it. This compares it. WHY THE CACHE KEY WAS NOT TOUCHED. The obvious fix -- add a data component to _config_fingerprint -- is wrong and would have made staleness PERMANENT: the fp is an opaque md5 persisted on the leagues DB row, the SPA asks with that stored fp, and a data-dependent fp orphans it the moment Sleeper moves, leaving the SPA requesting an fp that still resolves to the OLD board, forever. It would also arm the G-363 parity guard, which intersects _config_fingerprint's field names with from_manual_input's parameters and would then require every leagues-row reconstructor to pass a value that is not a league attribute. The fingerprint is untouched; the comparison is on the READ side against the stamp the board already carries. WHY IT SERVES RATHER THAN REFUSES. The 02:00 pull moves the fingerprint for EVERY league at once. Refusing on drift would send the first read of each of ~58 leagues into a rebuild, and engine runs are serialised on max_workers=1 at ~35-105s apiece -- hours of queue, triggered by whoever opens the app at 02:15. So a drifted board is SERVED immediately, stamped meta.status_drift, and its canonical single-flight warm enqueued through the existing AUD-11 rescue path including that path's ownership check. The next poll gets the fresh board. Same stale-while-revalidate shape rankings_pipeline already uses on TTL expiry, keyed on data instead of on the clock. COST CONTROL. player_status.peek_fingerprint() answers ONLY from the memo resolve() already fills, keyed identically on the Sleeper cache mtime, and returns None on a cold memo or under FF_SNAPSHOT_AS_OF_SEASON. It never calls load_sleeper_players, which re-downloads the ~50MB Sleeper table inline once the cache is past its 24h TTL -- correct on a board build, unacceptable on a request path in a single-worker app. Every ambiguity resolves to 'serve': None from the peek, no stamp on the board, or a board built with no usable live signal all mean 'cannot judge', never 'changed'. A false positive costs a needless rebuild; a false negative costs one more poll interval on a board that was already stale. OBSERVABILITY BEFORE ENFORCEMENT: rankings.read.status_drift and rankings.read.status_drift_warm ship with it, and status_freshness_gate.enqueue_rebuild=false retains the counters while removing all added load. NOT DONE HERE, FILED: the redraft read path (cache_resolvers._resolve_redraft_result) carries the identical hole, and this closes the detection gap rather than the 16h refresh cadence itself -- the Sleeper TTL and the game-day pull window are separate work.

v1.6.151

2026-08-15

G-398: close the redraft measurement gap, including one this session CREATED. RANKED-OUTPUT NEUTRAL on redraft (measured). THE SELF-INFLICTED HALF. v1.6.149 closed 'a gate that fires and records nothing' for dynasty. v1.6.150 then landed the DISC-Q2 QB depth gate on redraft -- and stamped no per-player flag, creating a fresh instance of the same defect on the other board, in the same session. Caught by diffing the week0b and week0c freezes rather than by reading the diff. THE INSTRUMENT WAS ALSO WRONG. freeze_prospective_board carried ONE global gate map and applied it to all three formats. Redraft runs a different layer stack -- no stale decay, no committee detection, no pre-LTV structural ceiling, team_context in place of news_situation -- so the single map made the redraft freeze ADVERTISE TWELVE gates it could not grade, while silently omitting TWO it could: season_availability_factor (the RDR-AVAIL-2 term the G-397 law had just changed, 656 players) and context_adj_pct (team_context, 117 players). An instrument that overstates its own coverage is worse than one with a known gap, because the gap is at least visible. Gate maps are now per-format, and every freeze prints an OVERCLAIM line naming any gate the board advertises but cannot grade. SHIPPED: redraft now stamps injury_discount, injury_pattern, injury_risk (it built the profiles and discarded everything but two reporting fields) plus qb_depth_order and is_qb_depth_capped. Five gates that fired blind on the board most users actually draft from. REMAINING, HONESTLY SCOPED: team_context records magnitude but not direction, and redraft's FA discount runs a different path than dynasty's graded by_pos_age schedule so there is no single factor to stamp -- that one is engine math, not plumbing, and is filed rather than faked.

v1.6.150

2026-08-15

G-397: BECOMING LESS AVAILABLE MUST NEVER RAISE A DRAFT SCORE. Ryan's ruling, 2026-08-15 -- the mirror of his 2026-08-14 law that signing with an NFL team must never LOWER a ranking, and ruled the same way and for the same reason. THE DEFECT. redraft's RDR-AVAIL-2 season-value factor multiplies `val`, which is value ABOVE REPLACEMENT and therefore NEGATIVE below it. Scaling a negative number by an availability factor < 1 makes it LESS negative, so a below-replacement player ranked HIGHER for being unavailable. Found while measuring a separate coherence fix, not by reading code: capping 23 backup QBs at 2.5 expected games moved ALL 23 UP and NONE down -- Kirk Cousins draft_score -3.280 -> -0.470 and rank 62 -> 23, Justin Fields -1.400 -> +0.090, crossing from below replacement to ABOVE it purely by being ruled unavailable. The behaviour was DELIBERATE and documented in the source ('a below-replacement player who misses time costs you FEWER deficit-weeks'). That is a sound LINEUP argument -- if you are forced to start him, his absence costs less -- and an unsound DRAFT-BOARD one, because nobody is forced to start him: you start replacement instead, so his season value over replacement is ~0 either way. WHY THIS WAS RULED, NOT A/B'd. No existing harness can grade it. scripts/validate_redraft_season_value.py tests whether expected-games beats flat-games at predicting season TOTALS and never touches val or draft_score. The question is semantic -- what a draft board MEANS -- not calibrational, which is exactly why the FA case was ruled rather than measured. SHIPPED: redraft.season_value_expected_games.positive_val_only = true (schema-registered). The availability factor applies to POSITIVE val only. Minimal change that enforces the law without inventing a penalty the data has not earned. MEASURED on the live redraft board: 563 of 750 rows move, ALL DOWN, ZERO UP -- the monotonicity property holds by construction, since the guard can only remove an upward distortion. The sub-replacement tail stops being compressed toward zero by unavailability (Collin Johnson -2.22 -> -7.68). QB top-12 ordering byte-identical; the single top-36 'mover' is Joe Burrow at -0.06 with NO rank change. UNBLOCKED AND LANDED IN THE SAME CHANGE: dynasty/redraft QB availability parity. redraft's _enrich_expected_games called build_injury_index WITHOUT qb_depth_chart, so the DISC-Q2 cap fired on dynasty only and 35 of 77 QBs carried two different expected_games across formats (Will Levis 2.5 vs 11.3, Mac Jones 2.5 vs 11.3, J.J. McCarthy 2.5 vs 10.0) while ZERO non-QB skill players diverged. The Waller defect class one level up: two subsystems holding contradictory beliefs about one player. Redraft now builds the map from the SAME status resolver with the SAME G-387 snapshot guard. Divergence measured 35 -> 0. Backup QBs now FALL where the parity fix alone would have raised them: Sam Howell 23 -> 44, Anthony Richardson 24 -> 45, Nick Mullens 30 -> 51. Parity was deliberately HELD on its own earlier the same day precisely because shipping it without the guard would have made the board worse for 23 quarterbacks; a test now fails if the guard is ever reverted while parity remains. DYNASTY AND KEEPER ARE UNTOUCHED by both halves.

v1.6.149

2026-08-15

G-396: four gates fired and left no per-player trace. RANKED-OUTPUT NEUTRAL (measured: 0 of 837 dynasty values changed, 0 ranks changed) -- this is instrumentation, not math. It exists because G-390's prospective freeze could not grade what the engine does not record. (1) mover_volume_factor arrived NULL on every dynasty payload row while the engine logged '[volume_reproject] re-projected 111 mover weighted_ppg'. Root cause: metrics_xfp.apply_volume_reproject multiplied weighted_ppg IN PLACE and returned only a count -- the dynasty payload has emitted the key since 2026-08-14 and nothing ever wrote the column. redraft_engine has its own writer and already stamped it. Now stamped at the point of application: 105 dynasty rows carry a real factor. (2) structural_ceiling_factor (207 players, mean 0.846), (3) stale_decay_factor (120 players, ladder 0.12/0.34/0.36) -- both computed per player and dropped before the payload, so a boolean said a gate fired but never by how much. (4) qb_depth_order / is_qb_depth_capped -- the DISC-Q2 depth gate emitted NO per-player flag at all; 51 board QBs now carry one. WHY THIS MATTERS FOR THE SEASON: six live-ingestion gates are snapshot-guarded, so no walk-forward run can grade them and no live board is reproducible after the fact (G-390). The only grading path is a pre-season freeze scored against outcomes, and a gate with no per-player factor can be scored in aggregate only -- 'the board was N% accurate' rather than 'this gate earned its keep'. Re-freeze before Week 1 to capture all four. HELD, NOT SHIPPED (filed G-397): dynasty and redraft disagree about expected_games for 35 of 77 quarterbacks (Will Levis 2.5 vs 11.3 games, Mac Jones 2.5 vs 11.3, J.J. McCarthy 2.5 vs 10.0; ZERO non-QB skill players diverge). Cause: redraft's _enrich_expected_games calls build_injury_index without qb_depth_chart, so the DISC-Q2 cap fires on dynasty only. The parity fix was written and MEASURED, and it is blocked: redraft's RDR-AVAIL-2 season-value factor multiplies VAL, which is NEGATIVE below replacement, so cutting a backup QB's availability RAISES his draft_score. All 23 affected QBs moved UP, none down -- Kirk Cousins +2.81 and 39 rank spots, Justin Fields -1.400 -> +0.090, crossing from below replacement to above it by being ruled unavailable. The behaviour is deliberate and documented in the source ('a below-replacement player who misses time costs you FEWER deficit-weeks'), which is a sound LINEUP argument and an unsound DRAFT-BOARD one. Resolve the negative-VAL treatment first, then land parity. || COMPLETES THE v1.6.148 RECORD (added same day, after the fact). v1.6.148 shipped on a binding SHIP from scripts/validate_availability_layer.py, and that verdict carried NO RANK-METRIC CHECK -- rank_delta_ci was passed as None, which eval_lib.verdict_block treats as passing. CLAUDE.md defines SHIP as 'dMAE CI95 excluding zero AND no rank-metric regression' (audit D5), so the verdict was incomplete against the house bar. The check has now been run. scripts/live_ab._set_gate was extended to accept DOTTED PATHS so nested gates can reach the canonical harness at all -- before this it could only flip top-level keys, which is why most of the injury_model block, redraft.* and xfp_base.* had never been A/B'd through it. RESULT, board-level walk-forward, ref2022/2023/2024, OFF=ungated vs ON=max_seasons 2, sentinel FIRED on all three (102/124/102 rows): TIE at all four positions, NO-SHIP. QB dMAE -0.2572 CI95 [-5.9283, +5.3516]; RB -0.4341 [-1.7550, +0.8487]; WR -0.5904 [-1.7617, +0.6550]; TE -0.3586 [-1.4331, +0.6883]. Every point estimate is an IMPROVEMENT and every CI crosses zero. Crucially, the rank check -- the specific thing that was missing -- PASSES at all four: Dro QB -0.0571 [-0.1217, +0.0021], RB +0.0081, WR -0.0024, TE -0.0058, no regression beyond CI anywhere. Per-season direction WR 3/3 and TE 3/3 improve. READ IT HONESTLY: this is NOT independent board-level confirmation. It is absence of harm plus a directionally consistent point estimate at n=59-216 players, against the layer-level test's 773. A single multiplier among ~20 layers, moving 104 of 837 values, cannot move a board-level CI at that power -- which is itself the finding: the board-level harness can only ever confirm LARGE changes, and individual gates need the layer-level matched-control design. The change stands on the layer-level evidence (SHIP in and out of sample, every position and history-depth slice guard non-positive, 21/21 and 8/8 seasons) plus the qualitative ghost-class defect it closes, NOT on this run. ALSO NOTE the advisory survivor-ppg view degrades in every season and position (dMAE +0.04 to +0.69, bias more negative). That is expected and arguably confirmatory rather than contrary: the change charges low-availability players more, and a survivor-conditional target scores only players who went on to PLAY. It is precisely the target eval_lib demotes to advisory (audit D6/A2) because it drops the cohort the gate is about. TWO OTHER SELF-CHECKS RUN THE SAME DAY, both passing: a PLACEBO on the layer-level design (split the control cell, score half B against half A's yardstick -- 0.9981 CI95 [0.9565, 1.0386], 1.0 inside; the originally-cited 'control lands on 1.000' figure was near-TAUTOLOGICAL and is retracted as evidence), and HORIZON SENSITIVITY (SHIP at H=1, 2, 3 and 5 in both windows, effect growing with horizon, so the 3-season choice was conservative rather than load-bearing).

v1.6.148

2026-08-15

G-388: the availability layer's first A/B, and the one thing it found. injury_model.healthy_pattern_bypass (DISC-Q1f) restored the legacy alpha-power discount for any player the pattern classifier called 'healthy' whose availability_scaling_v2 discount fell below 0.90. Its own config note calls it 'surgically narrow' and names ONE case: Daniels, 17/17 rookie + 7/17 sophomore -- two seasons of history. It was not narrow. _classify_injury_pattern only returns 'chronic' at 2+ seasons under 0.30 availability (~under five games), so a player who plays six or seven games EVERY year never trips it, stays 'healthy' forever, and the bypass hands him back ~0.79-0.88 -- the legacy 0.792 floor that availability_scaling_v2's own note says it exists to escape ('the legacy floor=0.792 trap that made Wentz/Rivers/Martinez fantasy-relevant'). Twenty players on the current pool were in that state, by name: Deshaun Watson (6, 6, 7 games -> 0.820), Kendre Miller (7, 6, 7 -> 0.792), Xavier Gipson (14, 7, 6 -> 0.792), Zay Jones (9, 6, 6 -> 0.843). DISC-Q1f had quietly re-opened the exact ghost-player class DISC-Q1 closed, for anyone whose absences were chronic but never quite season-ending. SHIPPED: healthy_pattern_bypass.max_seasons = 2 -- the bypass fires only for the 1-2 season cohort it was written for. EVIDENCE: scripts/validate_availability_layer.py, new, the first A/B this layer has ever had. Matched-control by design after the two 2026-08-14 self-corrections: realized 3-season discounted TOTAL-POINTS retention (unconditional) divided by a did-nothing control matched on position x age band x prior-rate tier, so natural age decline (control retention 0.747 / 0.590 / 0.437 by band), mean reversion and league-average games missed all divide out. The control cell is defined by STATE (avg_avail < position_baseline, the branch condition itself), not by the multiplier under test, so no arm can move its own yardstick. 1,959 treated anchors / 773 players, anchors 2002-2022, 2026 hard-refused. VERDICT: binding SHIP in BOTH windows -- dMAE -0.0396 CI95 [-0.0479, -0.0313] full, -0.0467 [-0.0584, -0.0350] on anchors after 2014 -- every position AND history-depth slice guard non-positive, 21/21 and 8/8 seasons agreeing in direction. WHY GATED AND NOT REMOVED: removing the bypass outright scores marginally better on the aggregate (-0.0415) and is what the aggregate alone would have shipped, but its own motivating cohort refuses to resolve -- hist=1 is n=49 in sample / n=23 out of sample and FLIPS SIGN between windows, so outright removal fails its own binding slice guard out of sample. The change stops where the evidence does. Daniels 0.859 -> 0.859, Hampton 0.858 -> 0.858, Nabers and Waller untouched. CROSS-CHECK: experiments/g152_availability_layer_sizing.py re-run under both configs reproduces its published 2026-07-25 values to four significant figures under the old config (RB withheld 0.5117 vs 0.512, WR 0.3408 vs 0.341, TE 0.3859 vs 0.386), then reports withholding rising RB 0.512->0.944, WR 0.341->0.568, TE 0.386->0.604 with every position STILL UNDER_PRICED -- two instruments built for different questions agreeing on sign, and no overshoot. NEGATIVE RESULTS KEPT: availability_scaling_v2 itself is correct and load-bearing (disabling it is a decisive DEFEND, +0.0528 [+0.0427, +0.0628], 0/21 seasons improve); scaling_floor 0.15 -- the one knob whose note openly admits to a judgement call -- is INERT (0.15 -> 0.30 moves dMAE +0.0001 [-0.0003, +0.0005]); a 12-cell per-position schedule buys nothing over one flag; min_sample_gate alone is a no-op (-0.0002 [-0.0012, +0.0007]) and stays OFF, which IS the binding-CI A/B its config note has been waiting for. BOARD IMPACT: dynasty 104 of 837 values move, max 75 rank spots, 3 positional top-24 swaps; keeper 98 of 837, max 79; redraft ZERO (redraft consumes expected_games / availability_rate, never injury_discount, so this layer has never reached it). LOGGED NOT FIXED: (1) reported availability_rate and applied discount disagree for the moderate-single-miss cohort -- expected_games_recovery_floor lifts the REPORTED field to G-30's 0.726 while the discount keeps the backward average (Evans 0.709, G. Wilson 0.606, Aiyuk 0.599); the ungated bypass was incidentally papering over it. Not reconciled here because G-30's number is a ONE-season availability rate and this study's target is three-season control-relative points retention. (2) G-152's aged-TE accounting moves from net +0.040 [-0.381, +0.451] to +0.198 [-0.235, +0.623] -- still straddling zero, no double-count demonstrated, but the first cell to re-measure if availability tightens further; visible as three deep TEs newly on the zero floor, an artifact of a SUBTRACTIVE flat charge sitting downstream of MULTIPLICATIVE discounts. Receipt: docs/RESULT_G388_availability_layer_ab_2026-08-15.md. Tests: tests/test_g388_availability_bypass_gate.py.

v1.6.147

2026-08-14

v1.6.146

2026-08-14

v1.6.145

2026-08-14

v1.6.144

2026-08-06

G-308 Buy Window roster context. The recommendation surface now reads the manager's own positional grades before proposing buys, so a position already graded elite is no longer surfaced as a buy target - the Buy Window was proposing five QBs into a QB room graded A. Reuses _positional_grades(); adds a surplus-grade constant and an explicit empty shape. The ranked board is unchanged: this filters the recommendation surface, not the rankings. Version bumped because the recommendation payload's CONTENTS change.

v1.6.131

2026-07-13

Overnight sprint combined ship (Ryan-directed acceleration; evidence _sprint_v1131/, combined A/B tune affected n=384 dMAE -0.1014 CI [-0.1893,-0.0187], pooled -0.0181 CI excl zero, validation no-harm; ablations replicate workstream receipts, interaction sub-additive). G-38 cameo-mover gate (xfp_base.volume_reproject.qb_min_ref_sf 0.35): the Bridgewater-29.6 class was the G-31 mover volume-reproject applying a cameo-derived factor (clipped at cap 2.0) to a starter-upgraded rate; inflating factors now require ref-season sf >= 0.35, docks kept, DISC-R8 guard remains backstop. G-35 qb_starter_status_prior (post-shrinkage): non-established QBs blended toward measured class realizations (backup class projected 15.4 vs realized 9.1); tune affected dMAE -1.68 CI [-2.46,-0.92]; validation neutral no-harm (disclosed); QB12->24 rate spread 4.12->5.25. G-34 y2_rate_uplift: Y2 veterans +1.4/+0.9/+0.6/+0.5 ppg (QB/RB/TE/WR, 0.5x measured cohort bias, ungated by round); tune affected n=280 CI excl zero, WR slice CI excl zero; bust-slice cost +0.16/+0.30 disclosed; Y3 = receipted null. G-41 reverse-spike relief ships DARK (_enabled false, 0=legacy proven): tune null-ledgered, recent-era revisit signal logged. G-40 two_way_position_overrides to config (meta + stats rows; map == legacy hardcode, zero board effect). G-43 publication stat lines (redraft.publication_stat_lines): per-component lines derived with the same recency weights, scaled to reconcile exactly (stat_line_scale audit), ranking_sanity reconciliation gate tol 0.15 ppg; rookies ppg/total-only v1; publish_redraft hook config-conditional. G-44 marketing gate fail-closed (missing/unreadable/malformed backtest_results.json -> BLOCK; explicit kill switch, distinct check id). New goals G-47/G-48/G-49; G-45 keep-verdict diagnosis and G-46 targets-basis diagnosis on file.

v1.6.130

2026-07-12

EYEBALL-4 final polish pass (Ryan-directed). F1: redraft_rookie_model.draft_picks_2026 rebuilt from rookie_prospects.csv draft data -- the hand-populated list had 15 players (zero TEs; the TE list existed only as an empty stray key at the wrong nesting level, schema-required but never read) and missed Kenyon Sadiq (TE R1-16), Makai Lemon (WR R1-20, KTC WR21), Zachariah Branch (WR R3-79), Nicholas Singleton (RB R5-165); three tier errors fixed (KC Concepcion R2/40 -> R1/24, Taylen Green and Kaytron Allen R4 -> R6). Now 72 drafted skill players, names/teams reconciled to players.csv (RDR-9a 0 unmatched). F2: superflex_qb_replacement 20 -> 24 (SF boards were near-clones of 1QB: QB replacement ppg moved only 17.39 -> 16.17; 12-team SF starts ~24 QBs). Zero 1QB/publication impact by construction; QB-rate compression noted to G-35. Companion (non-config): routers/seo.py chrome class fix (st-public-header did not exist in public-pages.css -> unstyled banner), Jayden Daniels bounded reviewed exception (standard coherence, dynasty-side investigation queued G-39), test_factcheck_blog unique-surname fixture Bowers -> Egbuka (players.csv now has six Bowers).

v1.6.129

2026-07-12

DISC-R8-GUARD (Ryan bug-hunt notes): the shrinkage elite skip/taper now require an EARNED rate -- some season among the player's three most recent with >= 6 games at >= 0.75x his current rate. Kills the cameo class the skip was never meant to protect (Bridgewater 29.6 ppg on 3 games -> board QB5; Minshew 23.7 published QB29; Dalton 21.4) while sparing earned elites with a short injury year (Nabers/CMC/Rice/Breece-2022 verified untouched). A/B on the v5 fidelity harness: tune affected dMAE -0.63 CI [-1.34,-0.07], pooled -0.008 CI [-0.017,-0.001]; validation affected -1.62 CI [-2.53,-0.04]; all exclude zero. Companion harness fix RDR-FIDELITY-6: redraft_backtest loader now ingests REG rows only (REG+POST/POST rows entered as duplicate seasons for 94.6% of player-seasons; engine was always REG-only via data_loader). Receipts: data/disc_r8_guard_final_ab.json, data/disc_r8_guard_live_impact.json.

v1.6.128

2026-07-12

RETROACTIVE ENTRY (missed at ship, added 2026-07-12 evening): shrinkage session -- G-28 redraft.y2_capital_prior (Y2 R1-R2 QB/TE shrink toward capital-conditional priors, monotone-favorable relief valve; Y2 QB bias +4.15->+2.98, TE +2.42->+2.00; Daniels QB15->QB8) + wide DISC-R8 taper (taper_start_frac 0.8->0.0 all formats; tune pooled dMAE -0.0116 CI excl zero) + G-31 mover volume-reprojection redraft parity (xfp_base.volume_reproject.redraft_enabled; Wan'Dale WR22->WR32, Nacua WR2 held) + G-33 redraft.spike_regression (WF-AGE-5 port, rf=0.35 RB-only; spike cohort dMAE -0.181 CI excl zero; Javonte 13.18->11.76, Kyren held). Receipts in data/g28_*, g31_*, g33_*; record RUNSHEET_shrinkage_session_v1128_2026-07-12.md.

v1.6.127

2026-07-12

G-30 forward-looking expected_games: acute-recovery players floored at the acute_event_taper's calibrated forward rates (Nabers eg 6.6->11.1); NEW moderate-single-miss branch floored at 0.726 (cohort n=151: Daniels 9.3->12.3, G.Wilson, LaPorta, Bucky, Aiyuk). First-season/recurrent missers held with cohort receipts (realize 0.528/0.560). Reporting fields only; LTV discount path untouched. EYEBALL-2 macro theme 1.

v1.6.126

2026-07-11

MOVE-1 volume_reproject re-enabled (mover predicate -> team_codes.is_real_team_change: vocab-normalized, FA excluded; mover audit 190->110 real; LTV harness TIE-no-harm, Phase 1 Y+1 proxy binding) + age_projection_discount re-calibrated via 11-arm grid (RB [email protected], WR [email protected], TE [email protected], QB [email protected], max 0.35; 30+ bias ~0 all positions, MAE 2.434->2.410). EYEBALL-1 themes 2+3. Receipts: volume_reproject_mover_audit / age_discount_grid / y2_cohort_study 2026-07-11.

v1.6.125+redraft-ppos

2026-07-11

Per-position recency weight optimization (S-3). 11 weight changes. Scores: QB=0.3207, RB=0.4286, WR=0.4766, TE=0.4643.

v1.6.125+redraft-mc.standard

2026-07-11

Redraft Monte Carlo optimization (S-2) — format: standard. 7 parameter changes. Best composite score: 0.4168.

v1.6.125+redraft-ppos

2026-07-11

Per-position recency weight optimization (S-3). 9 weight changes. Scores: QB=0.3172, RB=0.4318, WR=0.4955, TE=0.4772.

v1.6.125+redraft-mc.full_ppr_te_premium

2026-07-11

Redraft Monte Carlo optimization (S-2) — format: full_ppr_te_premium. 7 parameter changes. Best composite score: 0.4237.

v1.6.125+redraft-ppos

2026-07-11

Per-position recency weight optimization (S-3). 12 weight changes. Scores: QB=0.3207, RB=0.4316, WR=0.4955, TE=0.4770.

v1.6.125+redraft-mc.full_ppr

2026-07-11

Redraft Monte Carlo optimization (S-2) — format: full_ppr. 7 parameter changes. Best composite score: 0.4246.

v1.6.125+redraft-ppos

2026-07-11

Per-position recency weight optimization (S-3). 10 weight changes. Scores: QB=0.3190, RB=0.4304, WR=0.4892, TE=0.4742.

v1.6.125+redraft-mc.half_ppr

2026-07-11

Redraft Monte Carlo optimization (S-2) — format: half_ppr. 7 parameter changes. Best composite score: 0.4218.

v1.6.125+redraft-ppos

2026-07-11

Per-position recency weight optimization (S-3). 12 weight changes. Scores: QB=0.3184, RB=0.4304, WR=0.4892, TE=0.4772.

v1.6.125+redraft-mc.half_ppr_te_premium

2026-07-11

Redraft Monte Carlo optimization (S-2) — format: half_ppr_te_premium. 6 parameter changes. Best composite score: 0.4223.

v1.6.116

2026-06-25

RK-MEAS Phase 0: removed combine_delta from the empirical rookie model (rookie_model.combine_delta_enabled true->false). Rookie projected_peak_ppg now = bucket_median + college_delta + trajectory. EVAL-1-gated LOCO OOS A/B (scripts/validate_combine_delta_removal_eval1.py): paired delta-MAE -0.135 CI95 [-0.207, -0.063], rank delta-rho +0.023 [+0.007, +0.040], no per-position regression (QB -0.374 / RB -0.155 / TE -0.089 / WR -0.057), 9/11 seasons improve. The combine signal was net-harmful OOS; removal is simpler and more accurate. Holdout-clean (<=2025 fit, 2026 refused). Spec: docs/SPEC_rookie_dnp_college_weighting_2026-06-25.md.

v1.6.114+redraft-ppos

2026-06-21

Per-position recency weight optimization (S-3). 11 weight changes. Scores: QB=0.3261, RB=0.4369, WR=0.4924, TE=0.4792.

v1.6.114+redraft-ppos

2026-06-21

Per-position recency weight optimization (S-3). 12 weight changes. Scores: QB=0.3292, RB=0.4351, WR=0.4860, TE=0.4758.

v1.6.114+redraft-ppos

2026-06-21

Per-position recency weight optimization (S-3). 12 weight changes. Scores: QB=0.3308, RB=0.4330, WR=0.4736, TE=0.4662.

v1.6.114+redraft-ppos

2026-06-21

Per-position recency weight optimization (S-3). 11 weight changes. Scores: QB=0.3261, RB=0.4369, WR=0.4924, TE=0.4792.

v1.6.114+redraft-ppos

2026-06-21

Per-position recency weight optimization (S-3). 12 weight changes. Scores: QB=0.3292, RB=0.4351, WR=0.4860, TE=0.4758.

v1.6.114+redraft-ppos

2026-06-21

Per-position recency weight optimization (S-3). 12 weight changes. Scores: QB=0.3308, RB=0.4330, WR=0.4736, TE=0.4662.

v1.6.117

2026-06-21

Per-position recency weight optimization (S-3). 11 weight changes. Scores: QB=0.3261, RB=0.4369, WR=0.4924, TE=0.4792.

v1.6.116

2026-06-21

Per-position recency weight optimization (S-3). 12 weight changes. Scores: QB=0.3292, RB=0.4351, WR=0.4860, TE=0.4758.

v1.6.115

2026-06-21

Per-position recency weight optimization (S-3). 12 weight changes. Scores: QB=0.3308, RB=0.4330, WR=0.4736, TE=0.4662.

v1.6.113

2026-06-17

v1.6.111

2026-06-06

v1.6.110

2026-06-06

v1.6.107

2026-06-05

v1.6.106

2026-06-05

v1.6.105

2026-06-05

v1.6.104

2026-06-04

v1.6.103

2026-06-04

v1.6.102

2026-06-03

v1.6.101

2026-06-02

STAGED (flag-OFF) the QB shrinkage elite-skip — primary lever for the +2.50 PPG QB backtest bias. Ported the redraft DISC-R8 elite-skip to dynasty: metrics.apply_shrinkage now accepts config.shrinkage.elite_skip {enabled:false, positions:[QB], elite_mult:1.0} and, when ON, skips Bayesian shrinkage for a thin-sample player whose weighted_ppg > elite_mult x positional prior (a demonstrated above-prior producer is durable, not a one-year wonder dragged toward a sub-replacement prior). Root cause per docs/SPEC_qb_shrinkage_fix_2026-06-02.md: the dominant of two shrinkage ops; young producing QBs (Daniels class) are the affected set. FLAG STAYS OFF: the candidate's QB CI crosses zero on the thin cohort (n~33), so ship only after 2026 actuals tighten it AND a clean full-engine backtest clears (3 prior QB attempts false-SHIP'd on proxies). Flag-OFF = no ranking change; walk-forward caches re-stamp identical (the proof). Files: metrics.py, engine_config.json, engine_config.schema.json.

v1.6.100

2026-06-02

v1.6.99

2026-06-01

v1.6.98

2026-06-01

v1.6.97

2026-05-30

v1.6.96

2026-05-30

v1.6.95

2026-05-28

v1.6.94

2026-05-28

v1.6.93

2026-05-28

v1.6.92

2026-05-28

v1.6.91

2026-05-27

v1.6.90

2026-05-27

v1.6.89

2026-05-27

v1.6.88

2026-05-27

v1.6.87

2026-05-27

v1.6.86

2026-05-27

v1.6.85

2026-05-27

v1.6.84

2026-05-26

v1.6.83

2026-05-26

v1.6.82

2026-05-26

v1.6.78

2026-05-26

v1.6.77

2026-05-26

v1.6.76

2026-05-26

v1.6.75

2026-05-26

v1.6.74

2026-05-26

v1.6.73

2026-05-26

v1.6.72

2026-05-26

v1.6.71

2026-05-26

v1.6.70

2026-05-26

v1.6.69

2026-05-26

v1.6.68

2026-05-26

v1.6.67

2026-05-26

v1.6.66

2026-05-26

v1.6.65

2026-05-26

v1.6.64

2026-05-26

v1.6.63

2026-05-25

v1.6.62

2026-05-24

v1.6.61

2026-05-24

v1.6.60

2026-05-24

v1.6.59

2026-05-24

v1.6.58

2026-05-24

v1.6.57

2026-05-23

v1.6.56

2026-05-23

v1.6.55

2026-05-22

v1.6.54

2026-05-22

v1.6.53

2026-05-22

v1.6.52

2026-05-22

v1.6.51

2026-05-22

v1.6.50

2026-05-22

v1.6.49

2026-05-22

v1.6.48

2026-05-22

v1.6.47

2026-05-22

v1.6.46

2026-05-21

v1.6.45

2026-05-21

v1.6.44

2026-05-20

v1.6.43

2026-05-20

v1.6.42

2026-05-20

v1.6.41

2026-05-20

v1.6.40

2026-05-19

v1.6.39

2026-05-19

v1.6.38

2026-05-19

v1.6.37

2026-05-18

v1.6.36

2026-05-18

v1.6.35

2026-05-18

v1.6.34

2026-05-18

v1.6.33

2026-05-18

v1.6.32

2026-05-18

v1.6.31

2026-05-18

v1.6.29

2026-05-18

v1.6.28

2026-05-18

v1.6.27

2026-05-18

v1.6.26

2026-05-18

v1.6.25

2026-05-17

v1.6.24

2026-05-17

v1.6.23

2026-05-17

v1.6.22

2026-05-17

v1.6.21

2026-05-17

v1.6.20

2026-05-17

v1.6.19

2026-05-13

v1.6.18

2026-05-13

v1.6.17

2026-05-13

v1.6.16

2026-05-13

v1.6.15

2026-05-13

v1.6.14

2026-05-13

v1.6.13

2026-05-13

v1.6.12

2026-05-12

v1.6.11

2026-05-12

v1.6.10

2026-05-12

v1.6.10

2026-05-12

v1.6.9

2026-05-12

v1.6.8

2026-05-12

v1.6.7

2026-05-12

Sprint C accuracy ship: WR age curve flattened to 1.0 across all ages (REBUILT_WR_CURVE in age_curves_rebuilt.py). Rationale: AS-7 ablation (Session 93) showed disabling the WR age curve improves WR MAE_dyn by 0.43 ppg/year; AS-8 minimal-engine head-to-head confirmed flat-curve naive baseline beats production WR MAE_dyn by 0.44. Sprint C experiments with curve reshape (AS-2-informed lifts; naive piecewise shape) both WORSENED WR MAE (4.27, 4.91 vs baseline 2.89 on 2019). Root cause: the LTV horizon math (7-season weighted average of weighted_ppg × age_factor) interacts non-linearly with non-flat values — only TRUE flat (age_factor=1.0 at every age) reproduces AS-7's improvement. Walk-forward AB 2019-2024 aggregate vs v1.6.6: QB MAE_dyn 3.891 → 3.690 (-0.20); RB 3.216 → 3.085 (-0.13); WR 2.702 → 2.268 (-0.43); TE 2.332 → 2.205 (-0.13). vs naive baseline (AS-8): QB now beats naive by 0.74; TE by 0.05; WR ties within 0.005; RB still loses by 0.16 (next sprint target). Trade-off disclosure: flat WR curve removes dynasty age-discrimination from WR dynasty_ppg — a 30-year-old WR with weighted_ppg=10 is valued the same as a 25-year-old at 10. This is intentionally optimizing for year-1 projection MAE; multi-year dynasty value differentiation is preserved in ltv_discounted (sum, not average) and in archetype modifiers. Old REBUILT_WR_CURVE values preserved in age_curves_rebuilt.py comment block for rollback. No other curves touched.

v1.6.6

2026-05-12

Session 93 Option-C ramp moderation. v1.6.5 walk-forward AB returned numbers identical to v1.6.4 across all four positions to two decimal places (QB MAE 3.89 / bias +2.58 / r 0.458; TE r 0.658). Diagnosis update: the QB age curve smoothing in v1.6.4 / rollback in v1.6.5 had no engine-level effect — the actual driver of the v1.6.4 regression (vs v1.6.3 baseline QB bias +2.35, TE r 0.719) was the rookie_development_ramp.QB change from [0.88, 0.94, 1.0] to [0.50, 0.70, 0.85, 1.0]. The mechanism cascade is ramp -> rookie projections -> P13/VOS positional priors -> TE r drop. Surgical fix: bracket the ramp magnitude to find the smallest step back from [0.50, ...] that still flips the Maye/Mendoza/Simpson inversion. Sandbox bracket sweep: [0.50, 0.70, 0.85, 1.0] gives Maye QB5 / Mendoza QB9 (4-rank gap, baseline); [0.55, 0.75, 0.88, 1.0] gives Maye QB5 / Mendoza QB7 (2-rank gap, robust); [0.60, 0.78, 0.90, 1.0] gives Maye QB5 / Mendoza QB6 (1-rank gap, thin); [0.65, 0.80, 0.90, 1.0] inverts back to Mendoza QB4 / Maye QB6. v1.6.6 selects [0.55, 0.75, 0.88, 1.0] — smallest step back with a robust safety margin on the inversion fix. Predicted effect on walk-forward AB: partial recovery of QB bias toward +2.35 and TE r toward 0.719, since the ramp magnitude on year-0 (the dominant lever) is 10% less aggressive vs v1.6.5. Pending: walk-forward AB validation. QB curve stays at restored trim-25% values (no rollback of v1.6.5's curve restore).

v1.6.5

2026-05-12

Session 93 surgical rollback of v1.6.4 QB age curve smoothing. Walk-forward AB on v1.6.4 regressed two of the four positions: QB bias drifted further from zero (+2.35 -> +2.57 wPPG, wrong direction; decision rule was 'bias improves toward 0' -> FAILED) and TE r dropped -0.061 (0.719 -> 0.658), with TE hit@top-12 dropping -8pp (62% -> 54%) despite TE curves / ramp / config block being untouched in v1.6.4 (suspected P13 / VOS cascade from QB-side positional priors shifting). WR untouched (curves not modified). RB MAE +0.14 worse, bias slightly improved. Diagnosis: the smoothed QB curve lifted peak-age values (27: 0.880->0.97, 28: 0.935->0.99), which are the ages most QB-seasons sit in across the walk-forward cohort, so cohort-aggregate projection went UP and over-projection bias got WORSE. The rookie ramp change (rookie_development_ramp.QB = [0.50, 0.70, 0.85, 1.0]) was the surgical lever that actually flipped the Maye / Mendoza / Simpson rookie inversion -- hand-calc confirmed Mendoza year-0 multiplier 0.88->0.50 = -43% year-0 PPG hit did the bulk of the LTV drop. v1.6.5 restores REBUILT_QB_CURVE in age_curves_rebuilt.py to the original trim-25% values and KEEPS rookie_development_ramp.QB at [0.50, 0.70, 0.85, 1.0]. Predicted result: Maye lands QB5-7 (smaller relative gap to Mendoza than v1.6.4 showed, but inversion still flipped vs v1.6.3 baseline). v1.6.4 entry marked superseded_by=v1.6.5 for audit trail; not deleted. Methodology note: this ship violated the AB-harness-before-wire-in gating rule formalized 2026-05-11 -- the one-player smoke check on Maye placement was not a population-level accuracy validation. Rule re-affirmed: any age curve / ramp / projection-parameter change MUST run through walk_forward_backtest before wire-in.

v1.6.4

2026-05-11

Session 92 rookie ranking inversion fix. Two related changes addressing Drake Maye / Mendoza / Simpson QB ranking inversion: (1) QB age curve smoothing in age_curves_rebuilt.py REBUILT_QB_CURVE. Pre-smooth values had 4 monotonicity violations (22>23, 26>27, 32<33, and 36->37->38->39 ascending from 0.771 to 0.951 — survivor-bias inversion implying a 39-year-old QB produces 95% of peak). Smoothed values enforce strict monotonicity ascending 22->29 and descending 30->41, with late-career magnitudes aligned to CLAUDE.md framework ("Decline gradual; effectively done Late 30s"). Concretely: 23: 0.773->0.84, 27: 0.880->0.97, 32: 0.894->0.93, 36: 0.771->0.68, 38: 0.933->0.46, 39: 0.951->0.32, 40: 0.848->0.20. (2) Rookie development ramp for QB steepened from [0.88, 0.94, 1.0] to [0.50, 0.70, 0.85, 1.0]. Old ramp projected a #1-pick QB at 17-18 ppg in year 1, ~equal to a 2nd-year vet's weighted_ppg. New ramp aligns year-0 projection with empirical year-1 outcomes (Mahomes was a backup year 1; Allen/Hurts/Lawrence/Burrow/Stroud/Daniels/Williams all had material year-1 growing pains). Combined effect on top-15 dynasty QB: Maye 24 lifts QB7->QB3; Mendoza 22 (#1 pick proj) drops QB2->QB5; Simpson 22 (#13 pick proj) drops QB6->QB14. Late-career QBs (Wilson 37, Rodgers 42, Cousins 38) drop out of top-25 entirely. RB/WR/TE curves untouched. Pending: walk-forward AB validation.

v1.6.3

2026-05-10

Sprint C bias_correction calibration shipped. Per-position bias_w (actual - weighted_ppg) measured on 2021-2023 multi-year backtest post-Pattern-13b-flip and post-12-age-curve-tweaks. All four positions show direction-consistent bias across all 3 ref years. Applied with 0.30 conservative blend (matches accuracy_review.blend_factor). QB: 0.0 -> +0.62 (was systematically under). RB: 0.0 -> -0.54 (was systematically over — likely committee volume discount under-corrects). WR: 0.0 -> -0.12 (small over). TE: 0.0 -> -0.03 (basically calibrated; shipped as documentation). Affects projected_season_ppg column only — dynasty_ppg / ltv_discounted unchanged. Dynasty-layer over-discount on RB/WR/TE (separate concern: dynasty_ppg under-projects despite weighted_ppg being calibrated/over) tracked as follow-up task.

v1.6.2

2026-05-10

Sprint C consensus age-curve tweaks applied (12 ages across RB/WR/TE). All 12 ages flagged as 'too_pessimistic' by 3/3 reference years 2021-2023 in the multi-year backtest run on 2026-05-09. RB ages 25/29/31, WR ages 25/33/34, TE ages 26/27/30/32/33/34. Updates use the standard blend_factor=0.30 conservative weight from the existing accuracy_review pipeline. QB had 0 consensus signals; QB calibration deferred to a post-Pattern-13b-flip re-run.

v1.7.0

2026-05-09

Pattern 13b/c flip ON. qb_use_v2_ltv: false -> true after ab_harness validation on 2021-2024 snapshots. Verdict SHIP: 3/4 seasons pass kill criterion (dMAE <= -0.10). 2022: -0.80, 2023: -0.51, 2024: -0.47, 2021: +0.40 (single regression dominated by 3 wins). Net win directly attacks the QB +2.18 bias_w from Sprint C 2015-2024 backtest. Session 77 _apply_young_qb_ceiling guard prevents the double-boost interaction with young_qb_ceiling that originally rejected Pattern 13b on 2026-05-05. v2 path uses age_curves_v2.dynasty_lifetime_value_v2 (relative-factor formula) for QB position only; other positions still use v1 absolute or te_use_v2_ltv path.

future_pick_discount_market_align_2026-05-09

lasso_residual_te_2026-05-09

future_pick_discount_2026-05-09

lift_signals_round3_2026-05-09

lift_signals_round2_2026-05-09

lift_signals_enabled_2026-05-09

v1.6.1

2026-05-02

v1.5.1

2026-04-28

Task #28: added edge_score block for Personalized Edge Score

v1.4.9

2026-04-26

Redraft improvement wave: (1) Age curve rates doubled based on LOO backtest evidence. (2) QB starter-upgrade detection block added. (3) Opportunity blend config block added.

v1.4.8

2026-04-26

Per-position recency weight optimization (S-3). 6 weight changes. Scores: QB=0.3425, RB=0.4340, WR=0.4466, TE=0.4463.

v1.4.7

2026-04-26

Redraft Monte Carlo optimization (S-2) — format: standard. 7 parameter changes. Best composite score: 0.4139.

v1.4.6

2026-04-26

Per-position recency weight optimization (S-3). 12 weight changes. Scores: QB=0.3363, RB=0.4305, WR=0.4626, TE=0.4669.

v1.4.5

2026-04-26

Redraft Monte Carlo optimization (S-2) — format: full_ppr_te_premium. 6 parameter changes. Best composite score: 0.4199.

v1.4.4

2026-04-26

Per-position recency weight optimization (S-3). 6 weight changes. Scores: QB=0.3401, RB=0.4322, WR=0.4568, TE=0.4594.

v1.4.3

2026-04-26

Redraft Monte Carlo optimization (S-2) — format: half_ppr. 7 parameter changes. Best composite score: 0.4181.

v1.4.2

2026-04-26

Per-position recency weight optimization (S-3). 12 weight changes. Scores: QB=0.3403, RB=0.4305, WR=0.4626, TE=0.4649.

v1.4.1

2026-04-26

Redraft Monte Carlo optimization (S-2) — format: full_ppr. 7 parameter changes. Best composite score: 0.4205.

v1.4.0

2026-04-26

Per-format parameter architecture. Format-sensitive redraft params (recency_weights, td_regression, shrinkage, position_recency_weights) moved under redraft.formats.{format_key}. Engine resolves format from LeagueConfig.scoring_format at runtime. half_ppr_te_premium is fully optimized (v1.3.5); other formats seeded and pending MC runs.

v1.3.5

2026-04-25

Per-position recency weight optimization (S-3). 12 weight changes. Scores: QB=0.3397, RB=0.4324, WR=0.4568, TE=0.4653.

v1.3.1

2026-04-25

Redraft Monte Carlo optimization (S-2). 7 parameter changes. Best composite score: 0.4196.

v1.3.0

2026-04-25

v1.2.5

2026-04-25

v1.2.2

2026-04-25

YoY production delta as trajectory momentum signal for RB and WR. production_delta = ppg_most_recent_full_season - ppg_prior_full_season computed in weighted_ppg_for_player() (metrics.py) from full_records sorted by season desc. Exposed as production_delta column in enrich_production_metrics(). New _apply_trajectory_momentum_bonus() method in rankings_engine.py — parallel to _apply_young_qb_ceiling(), applied immediately after. Eligible: RB and WR, ascending/early_peak trajectory, ≥4.0 weighted_ppg, ≥2 full seasons. Ascending (delta ≥ +1.5 ppg): bonus_pct = delta × 0.012, capped at +18% dynasty_ppg. Falling (delta ≤ -1.5 ppg): penalty_pct = |delta| × 0.008, capped at -12% dynasty_ppg. Config-driven in trajectory_momentum block. momentum_bonus_applied column added for accuracy_review tracking.

v1.2.3

2026-04-25

Per-player LTV uncertainty bands (P25/P75/CV). compute_player_ltv_bands() added to monte_carlo.py — multiplicative noise model: production_factor × availability_factor × age_curve_factor, each drawn from a clamped Gaussian. 1000 simulations per player, pure NumPy (no engine re-run). Production sigma scales with seasons_included (1 season: 0.28, 4+: 0.08). Injury sigma scales with injury_risk tier (low_risk: 0.04, high_risk: 0.18). Age curve sigma fixed at 0.08 (population mean has ~8% individual stdev). CV = (P75-P25)/median — risk metric. risk_profile: safe_floor (CV<0.15), balanced (0.15-0.28), boom_bust (CV≥0.28). ltv_p25/ltv_p75/ltv_cv/risk_profile added to all rankings output rows. Trade analyzer: risk_profile surfaces in _player_summary() and _build_narrative() flags mismatched risk profiles between trade sides. Config-driven in player_ltv_bands block.

v1.2.4

2026-04-25

Position-specific recency weights. Bug fix: global recency_weights from engine_config.json were NOT flowing to weighted_ppg_for_player() in metrics.py — hardcoded module constant (3.0/2.0/1.0) was used instead of Monte Carlo-optimized values (2.966/0.9/0.279). Fixed by threading recency_weights param through weighted_ppg_for_player() -> enrich_production_metrics() -> build_player_rankings_base() -> _build_base() in rankings_engine.py. position_recency_weights added to engine_config.json: RB (4.5/0.65/0.15 — heavy current-season, least autocorrelated), QB (2.2/1.3/0.60 — flatter, most autocorrelated), WR (3.5/0.85/0.22 — moderate), TE (2.6/1.1/0.40 — sticky once established). Per-position override picked in enrich_production_metrics() before calling weighted_ppg_for_player(). Falls back to global weights if position absent. ModelParameters.position_recency_weights loaded from config. Pre-backtest calibration — accuracy_review.py will refine annually.

v1.2.1

2026-04-25

College school quality multiplier in rookie_model. Dominator rating at Alabama (P4_strong, ×1.0) ≠ NDSU (FCS, ×0.78). 5-tier system: P4_strong (1.0), P4_avg (0.95), G5 (0.87), FCS (0.78), unrated (0.90, default). Multiplier applied to raw college production stats BEFORE baseline subtraction in project_peak_ppg() — so a 35% target share at FCS becomes effectively 27.3% before the 20% baseline is subtracted, compressing the contribution from 15pp to 7.3pp. Applies to: college_target_share, college_rush_share, college_completion_pct, college_pass_ypa. Config-driven in rookie_model.school_quality. School lookup is case-insensitive; handles nflverse abbreviated formats (Penn St., Ohio St., North Dakota St.). school_name parameter added to project_peak_ppg(); get_rookie_rows() passes p.get('school'). school_quality_mult and school_quality_tier exposed in output rows.

v1.1.8

2026-04-24

Bayesian shrinkage on weighted_ppg for thin-sample players. Formula: w = n/(n+k), shrunk_ppg = w*player_ppg + (1-w)*position_prior. k=3.0, min_seasons_no_shrink=4. Prior = mean weighted_ppg of above-replacement players at each position. Only applied to above-replacement players (below-replacement are not shrunk toward starter mean). Rookies (seasons_included=0 or is_rookie=True) are always skipped. Key corrections: McCaffrey 21.4->16.3 (1 qualifying season / injury history), Daniels 20.9->19.0, Nabers 14.6->11.4, Bowers 15.0->13.0, Gibbs 18.5->14.9. 81 above-replacement players affected out of 825. Mean delta: +0.32 ppg. raw_weighted_ppg column preserved for transparency.

v1.1.9

2026-04-24

Rookie development ramp in LTV calculation. Adds per-year additional multipliers applied ON TOP OF the age curve factor for is_rookie=True rows in dynasty_lifetime_value(). Corrects survivor bias: age-21 WR factor (0.72) was derived from all 21-year-olds including experienced ones; true year-1 rookies produce ~55-65% of career peak (implied factor ~0.61), not 72%. Ramps: WR [0.85, 0.92, 1.0], TE [0.80, 0.88, 0.95, 1.0], RB [0.92, 0.97, 1.0], QB [0.88, 0.94, 1.0]. Config-driven in rookie_development_ramp block. Implementation: _load_development_ramps() in age_curves.py; dynasty_lifetime_value() accepts development_ramp parameter; _ltv_row() in enrich_age_curves() detects is_rookie and passes position ramp. No effect on non-rookie players.

v1.2.0

2026-04-25

Injury type granularity: structural ceiling discount applied to weighted_ppg BEFORE LTV calculation. Distinct from existing injury_discount (post-LTV availability factor). Structural injuries (seasons with <5 games = ACL proxy) now permanently bend the peak ceiling, propagating through the full 7-season LTV projection. New InjuryProfile fields: structural_ceiling_factor, n_structural_injuries, healthy_seasons_since_structural. Formula: base_ceiling[pos] - age_penalty(age_at_injury - cliff_age) + recovery_credit(healthy_seasons_since) × compound_multiplier^(n_structural-1). Config block: injury_model.structural_injury in engine_config.json. Pipeline: rankings_engine._apply_structural_ceiling() runs after rookie injection, before enrich_age_curves(). Availability discount (injury_model.alpha/floor) unchanged and still applied post-LTV.

v1.1.7

2026-04-24

Rookie model: 3-fix calibration. Fix 1 — base_ceiling raised via bias calibration against 2019-2023 classes (n=284): QB 12.0→16.09 (+34%), RB 11.0→13.66 (+24%), WR 10.0→11.94 (+19%), TE 8.0→7.63 (-5%). Fix 2 — elite athleticism breakout multiplier: SS≥elite_threshold +15-20%, SS≥mod_threshold +7-10%, applied pre-pick_factor; TE thresholds raised to 116/110 (vs 110/105) to avoid over-projection. Fix 3 — QB draft round starter bonus: R1 +7.0, R2 +3.5, R3 +1.5 ppg (pre-mult additive). All breakout/starter-bonus thresholds now config-driven in engine_config.json rookie_model block.

v1.1.6

2026-04-21

Superflex QB dynasty corrections. (1) young_qb_ceiling.min_weighted_ppg: 10.0 → 4.0 — was blocking ascending QBs with limited samples (Maye 7.49, Williams 8.24, Stroud 8.11). (2) young_qb_ceiling.bonus_per_year_to_peak: 0.0362 → 0.18 — previous value produced ~11% max bonus for a 23yo QB; new value produces ~30-55% for ascending QBs 3 years from peak. (3) young_qb_ceiling.cap: 1.1 → 1.55 — allows meaningful upside for franchise QBs with thin sample. (4) superflex_qb_dynasty_premium multiplier: 1.12 — applies to all QB dynasty_ppg in superflex dynasty leagues to reflect structural scarcity (20 starter slots vs ~16 reliable starters). Known remaining issue: 30-32yo QBs (Mayfield, Goff, Murray) still overranked relative to PP due to flat QB age curve 31-34 in engine_config; will partially self-correct with 2025 data load. (5) young_qb_ceiling.archetypes_eligible: added dual_threat — Maye and Williams were miscategorized as dual_threat based on rookie scrambling stats, blocking their ceiling bonus. Trajectory filter (ascending/early_peak) already prevents veteran dual-threat QBs like Jackson from qualifying. (6) young_qb_ceiling.trajectories_eligible: [ascending, early_peak] → [ascending] only. early_peak included established starters (Daniels 25.7, Purdy 26.7, Nix 26.5) who were being over-boosted. Developmental ceiling bonus should only apply to true ascending QBs with thin samples.

v1.1.5

2026-04-21

Reference season advanced to 2025 after full 2025 season data loaded via Sleeper fetch. Recency weight index 0 now maps to 2025. Former-starter discount now correctly identifies QBs who backed up in 2025. ACTIVE_THRESHOLD unchanged at 2024 (broad pool). Backtest re-run and Monte Carlo optimization pending.

v1.1.4

2026-04-20

Monte Carlo optimization (S-1). 32 parameter changes. Best composite score: 0.4029.

v1.1.3

2026-04-20

Multi-year backtest parameter update (20 reference years). 35 parameter changes applied.

v1.1.2

2026-04-19

Multi-year backtest parameter update (reference years: 2019/2020/2021). 20 parameter changes applied.

v1.1.1

2026-04-19

Multi-year backtest parameter update (reference years: 2019/2020/2021). 17 parameter changes applied.

v1.1.0

2026-04-19

Backtest-driven corrections using 2019 reference year (2015–2019 training, 2020–2024 evaluation, n=320 players). Four targeted fixes: (1) RB age curve recalibrated at ages 26–33 — original cliff was too steep, causing systematic -1.9 to -3.8 ppg underranking of 26–30 year old RBs; ages 26/27/28/29 factors raised from 0.94/0.83/0.70/0.55 to 0.97/0.92/0.84/0.72. (2) RB peak window extended from [22,26] to [22,28] — data shows productive prime runs through age 28. (3) QB pocket_passer post_peak_modifier raised from 1.10 to 1.15 — veteran pocket QBs (Brady, Brees, Rodgers, Roethlisberger) were underranked because age cliff was too aggressive after 35; base QB curve also flattened at 36–39. (4) confidence_flags.age_uncertainty_rb_over lowered from 28 to 27 — flagging uncertainty one year earlier to match actual production cliff. Overall backtest: r=0.769 overall, RB bias improved by recalibration. WR model was most accurate (r=0.754, bias +0.18). QB had highest variance (stdev 4.92) driven by backup QBs with inflated starter-season wPPG.

Accuracy review
reference_year2019
training_years2015-2019
evaluation_years2020-2024
n_players320
overall_r0.769
overall_mae2.27
overall_bias-0.43
tier_exact_hit_rate0.375
tier_within_1_hit_rate0.803
PositionnrbiasMAE
QB460.491-0.43.59
RB870.732-1.212.52
WR1210.7540.181.89
TE660.697-0.561.7

v1.0.0

2026-04-18

Initial parameters. Baseline from POC build. All values theory-grounded, not yet back-tested against actual outcomes. Includes: young_qb_ceiling block (ascending pocket QB bonus), trade_model block (trade analyzer constants).