Every engine version. Every change. Every accuracy review.
Current production version: v1.6.175.
200 releases on record.
Every change documented. Every version archived.
SignalTuned doesn't ship "the latest rankings." It ships a specific, versioned engine
with every parameter spelled out. When a parameter changes, the version increments and the
change shows up here with the reason. When the engine runs an accuracy review against
real outcomes, the results show up here too.
Every ranking the engine produces carries a manifest_hash tied to one of
these versions. Click the hash on any rankings page to see which engine version
generated that output.
Decision engine v3.4 (2026-08-28, mock 10, G-490/G-491). Mock 10 was played 100% engine-obedient; total leak vs engine-optimal was ~3-4 board points concentrated in two picks with one root cause. G-490: (a) an open-slot row's ladder gone-class is now MIN(own raw survival read, position clock) instead of clock substitution - substitution let a 98%-safe Quentin Johnston inherit the dying WR clock (top WR dying but same-team-discounted) and become THE PICK over Kyle Pitts at 5.05, which cascaded into a TE-tier loss (Pitts went 6.06, Warren 5.07); Lawrence stays demoted, mock-8 Nix stays urgent; (b) the position clock and per-position VONA anchor on the best row per position by OUR BOARD (blended_rank_score), not the first row in the game-theory re-rank, which need-weighting can reorder (3.05: all-WR want-set with Breece Hall priced off the board); (c) the hero-card pass-reason branches on roster fit before survival - 'Goff is 79% taken-by-then, so take Shakir' stated a dying player's doom as a reason to pass him; the actual reason was a filled QB slot. G-491: momentum above 1.0 decays toward 1.0 by the fraction of seats that can still START the position - Caleb Williams was priced 30%-to-survive during a QB run that was already finished (zero of the next nine picks took a QB).
G-492 (verification mock 2026-08-28, 3.05 + 5.05): (a) multi-open-slot positions now wait-deflate on the same market-survival factor as single-open ones, on a higher floor multi_open_wait_floor=0.75 - a two-open WR shelf whose board-best was 78%-safe monopolized the want-set over a dying board-#1 Breece Hall at an open RB slot; (b) the G-490 position-clock anchor now applies the G-486 same-team discount when choosing each position's best row (market-faller waiver mirrored), so a discounted teammate cannot keep his position's need urgent while the selector refuses to draft him - which is how a 98%-safe Quentin Johnston became THE PICK at 5.05.
G-493 (mock 12 2026-08-28, 3.05): the want-set ladder's final board-order tiebreak now applies the G-486 same-team discount (board_rank_score x same-team multiplier; market fallers keep 1.0). The sel bucket already carried the discount, but a bucket tie - the ladder's own definition of close - fell through to raw board order, letting a same-team Jameson Williams lead Rice/Higgins at 3.05 with ASB rostered and no market fall. No new knobs.
Decision engine v3.3 (2026-08-28, mock 9, G-486/G-488/G-489). G-486: same-team WR/TE target-share discount - ASB went 1.05 and Jameson Williams (DET) was THE PICK at 6.05; two premium WRs on one NFL team cap each other's ceiling, so a candidate WR/TE sharing a team with a rostered WR/TE takes a need discount (same_team_penalty) scaled by the rostered player's draft capital ((7-round)/6), faded after same_team_fade_round; a discount, never a gate - a real market fall punches through; QB stacks and RB handcuffs untouched. G-488: ladder urgency reworked twice over - (a) a row filling an OPEN dedicated slot carries its POSITION's clock (gone-prob of the position's best option, from _gone_pos) instead of its own: a dying 5th-QB Lawrence (12% survival) was THE PICK over a safe Nix/Goff/Stroud shelf at 9.05, while mock-8 Nix (shelf top dying) stays urgent; depth rows keep their personal raw read (Pierce over a 70%-safe Addison); (b) the primary sort key is a BINARY likelier-gone-than-not class on the RAW (pre-demand-deflation) read - among already-dying players fine survival ordering is noise (QJ 32% vs Addison 9% cannot both be kept: the better one leads), and G-485's quarter-buckets over deflated urgency compressed dying-vs-safe into one bucket. G-489: recs now serve wait_gone (raw wait-horizon gone-prob); the hero card's survival claims read it instead of the on-clock-zeroed denial_prob, which had the card telling Ryan a 32%-survival QJ had a 100% chance to still be there.
Decision engine v3.2 (2026-08-27, mock 8, G-485). Mock 8, 10.05: THE PICK was Tyrone Tracy Jr. (98% survives-to-next-pick, +43 vs crowd) over Bo Nix, whom the market model gave 29% to survive - the room took Nix three picks later and Ryan had to override a round after that to rescue Goff, the last startable QB. Three coupled fixes: (a) the G-481 wait factor's windowed-need seats term is replaced by market survival of the position's top option - rooms draft by ADP, not open slots (second QBs went all draft), so a single open slot defers only while its best option is likely to STILL BE THERE; (b) the G-474b zero-demand hard zero on decision urgency is removed - G-471's demand_floor deflation (raised 0.2 -> 0.45) keeps opponent-need information without silencing the market; (c) want-set sequencing is perishability-first (urgency quarter-buckets, then value x need 0.1-buckets, then our board) - every set member is a player we intend to take with remaining picks, so who dies first orders them; both-gone/both-safe tie through to value then board (G-480b preserved).
Decision engine v3.1 (2026-08-27, mock 7, G-481..G-483). G-481: the v168 demand/supply factor counted supply inside a value band of the position's top anchor - dense-anchor positions (RB/WR) read as bottomless supply and had their need crushed ~3x while sparse-top TE/QB kept full need: the exact INVERSE of the wait edge (McBride as THE PICK at 1.05 against his own 66% survives-to-next-pick read; Hurts at 3.05; all-TE/all-QB also-fits; Judkins over a 0-of-3-WR hole again). Replaced with a per-position VONA wait factor: a SINGLE open dedicated slot defers (down to wait_floor) only when the position's best available loses ~nothing to waiting (vona/vona_ref, scaled by windowed demand seats/2); MULTIPLE open dedicated slots are never wait-discounted - a 0-of-3 WR hole is fill pressure by definition. Urgency coverage-decay removed (same band pathology - it zeroed decision_urgency board-wide in prod); sequencing buckets urgency by quarters so sub-0.25 urgency ties break on OUR BOARD. demand_supply_floor/gain retained but unused. G-482 (adp_engine): parse_sleeper_board stamped K/DEF position=None, so the G-458 K/DEF market-row gate matched nothing and endgame fired into a pool with zero K/DEF rows in every mock - the actual root of the three-mock K/DEF failure; positions now preserved (DST normalized to DEF). G-483 (war-room.css): own filled cells brightened (38%/20% mix + ring) - they read DIMMER than opponents' picks.
Decision engine v3 (G-477..G-480, Belichick mock 6). G-477 ROOT CAUSE: redraft cache path served boards with no LeagueConfig -> _league_slot_counts empty -> endgame K/DEF fill, demand urgency, TE flex discount and bench coverage ALL silently disabled on the live mock path (harness passed: it passes roster_slots explicitly). Slots now resolve from the config on the cache path and from Sleeper draft settings slots_* engine-side. G-478: slot-aware need - open dedicated slots at full need (mild open-count x pressure scaling), dedicated-full = flex depth, need x demand/supply factor (wide-band startable pool) so low-demand positions defer at SELECTION. G-479: bench need = gap to depth targets (RB/WR starters+flex+2; TE/QB starters only) - inverts the req/rostered formula that made TE2 the 'thinnest' spot. G-480: want-set ladder - value x need bucket, then perishability bucket, then OUR BOARD; tier urgency no longer multiplies. Serve-time only; backtests unaffected.
Decision engine v2 (mock 5, G-473..G-476). SELECT-THEN-SEQUENCE (G-475): THE PICK's candidate set is the top-N rows by OUR value x need, N = remaining discretionary picks (picks left minus unfilled required K/DEF); urgency can only sequence within the set, never promote into it (Worthy - market's #1 remaining, our rank 16 lower, value gain -6.7 - was THE PICK at 12.05 purely on 'won't last'). TWO-PHASE NEED: starter phase keeps flex-aware need with TE's flex share discounted in non-premium leagues once the dedicated TE slot fills (G-473 - Pitts as TE2 'open starter slot' at 6.05); bench phase switches to per-position depth COVERAGE, required starters / rostered, never pooled (G-476 - RB 2-of-2 was being TAXED for the WR overflow sharing its flex pool while Ryan sat one injury from an empty lineup slot). TIER URGENCY (G-474): urgency = windowed demand seats vs tier supply (players within supply_band of the position's best remaining value) - five QBs for five needy seats means waiting is free; ZERO windowed demand defers the position outright (G-474b - 'nobody else needs one... this is where we get the edge by pushing qb not just taking one'); player-level denial only breaks ties within a tier. Endgame diagnostics now count fill candidates per open required slot (the mock-4/5 TE-over-empty-DEF failures produced nothing on the wire to debug). Hero copy gated: 'unlikely to last' only when decision_urgency backs it, with the percentage shown (Stafford was fronted by a sentence the engine had measured false). Serve-time only; cached boards and walk-forward artifacts unaffected beyond the config-hash restamp.
Mock-4 live fixes (G-470/G-471). G-470: the wait horizon. On the clock picks_until_me==0, so every P(gone before my pick) was computed over ZERO picks - trivially ~0 board-wide - flattening the G-466 urgency to its floor and making VONA read 'waiting is free' at the exact moment of decision; THE PICK fell back to need-adjusted raw value (live: Quentin Johnston, market-priced 4-5 rounds later, recommended at 6.05). The horizon now falls through to the distance to my NEXT owned pick (owned_cells, else snake/reversal math); auctions keep 0 - no slot order, waiting genuinely free. Applied to the decision urgency AND VONA (test_g197 on-clock contract rewritten accordingly). G-471: demand-aware urgency. The denial model priced demand off league-wide ADP and never looked at the rosters actually picking in between - Bo Nix was THE PICK at 9.05 with 'the market is about to take him' while all eleven other teams already had their QB. The urgency denial is now multiplied by the fraction of distinct windowed seats with an open slot the position can start in (dedicated; flex share RB/WR/TE; SUPER_FLEX for QB), floored at decision.demand_floor (0.2) for off-need stashing. Scoped to the DECISION urgency; displayed denial chips keep the pure market read. New wire field recommendations[].decision_urgency; hero copy claims 'unlikely to last' only when the demand-aware number backs it.
Mock-3 batch A+C (G-466..G-469). G-466: THE PICK becomes a DECISION - server stamps decision_score = anchor(board value percentile) x need_mult x (survival_floor + (1-survival_floor) x P(gone before my next pick)); the SPA hero orders by it instead of raw value (which had recommended TE2 at 4.05 and TE3 at 14.05 as depth adds while WR/K/DEF starter slots sat empty, and 3-round reaches with zero wait-awareness - Ryan: 'best available and likeliness to go anytime soon ... thats where the edge is'). Endgame roster-completion: when remaining picks <= unfilled required dedicated starters, rows filling one (K/DEF included, market-ordered) dominate the decision order - K/DEF carry no engine score so nothing else could ever surface them as THE PICK. Signal Board stays a pure value board; hero copy explains when the decision pick differs from the raw-value #1 and why. Config: draft_board.decision {enabled, survival_floor 0.55, endgame_fill} + schema. G-467: the contention-window strategy lens is disabled for near-redraft keeper leagues (derived/explicit horizon <= 2) exactly like redraft - Auto had resolved to 'Build' in a keeps-2 league and the youth lens ordered a rookie TE over a higher-scoring vet WR. G-468: The Room's subtitle and status dot now name the actual pricing source from vibe_source (Sleeper ADP market vs VibeRank crowd) instead of hardcoded VibeRank copy. G-469: DEF position color moved from #B08968 to #7A5238 - the first brown sat too close to the TE orange on the draft grid. UI-only changes carry no ranking effect; decision layer affects live-draft recommendations only (no cached boards, no backtests touched by design - engine_config hash moves, so walk-forward artifacts restamped).
Belichick mock-draft QA batch (G-459..G-465). G-459: max_keepers added to _config_fingerprint + _H2_DRIFT_FIELDS and threaded through the boot-warm reconstructor (data_status), the admin fp-repair endpoint (caught by the armed G-363 parity guard) and the manual request models; keeper-horizon resolution stamped into rankings meta (meta.keeper_horizon) with a logger.warning on refusal. Root cause: the boot warm rebuilt keeper configs without max_keepers, the G-392 derivation refused silently, and the full 7-season dynasty board was cached under the same fingerprint the correct config computes - reconnect/roster-sync/drift all left it standing. G-460: build_live_board now passes top_n=len(available) to rank_picks (default-20 cap starved the surfaced payload; the per-position backfill was dead code; client Flex filter showed 1 player). G-461: surplus QB in a 1QB league (no SUPER_FLEX seat) drops straight to the need floor instead of the graduated curve (0.65 QB2 was the least-taxed depth on the board; live all-QB Signal Board rounds 8-12). G-462: war-room ADP column resolution passes keeper through (resolver maps keeper to the redraft tier itself) and reads rec pts from the resolved LeagueConfig scoring before the live probe (mocks carry no league_id; full-PPR keeper league was priced off ADP dynasty half ppr). G-463: market/crowd vibe stamps now sync into the ranked copies in the same response (Room ALL-tab warming-up flap, one-poll-stale VIBE, K/DEF pinned to top) and K/DEF market rows carry synthetic unique player_ids. Scarcity: value_over_replacement_ranking.descarcity_horizon_factor flipped TRUE (G-391 double-count removal, measured McBride TE #5->#12 on the shallow-bench sibling league); baseline/elasticity recalibration stays parked pending harness extension. engine_config.json rewritten ensure_ascii=True (fixes test_g244 non-ASCII bytes). UI (no engine effect): K/DEF position color tokens, strategy-lens sort-basis note, Auto lens shows its resolved strategy.
G-451: name-join fallback + ADP read-time guard, bundled so one walk-forward rebuild covers both. (1) adp_engine.match_to_rankings and market_layer._build_reconciliation gain two rescue stages consulted ONLY after the existing match logic misses: a position-scoped name_utils.name_key lookup (resolves Kenny/Kenneth Gainwell, Matt/Matthew Hibner - Sleeper publishes short forms that appear in NO players.csv name field, so Jaccard reads 1/3 and SequenceMatcher ~0.83 vs the 0.85 gate) and a (last-name, position, team) unique-triple index (resolves De'Zhaun-Ryan vs De'Zhaun Stribling); ambiguous keys refuse by construction and free agents never participate. The war room already used name_key (G-307 family) and is untouched. name_utils gains elijah->eli (Eli Raridon, players.csv first_name Elijah). (2) adp_engine.SLEEPER_BOARD_MAX_ADP = 700.0: _is_board_null_adp rejected only >= 999, so Sleeper's 700.0 ranking-horizon clamp (38 cells >= 700.0 on the 2026-08-22 export, max 700.9) ingested as real ADP; the refresh's build-time clamp protects only boards that build produces, a user's scoped upload never passes through it. validate_adp_board.py now imports the ceiling from the engine so gate and runtime cannot drift. Measured stake: Kenny Gainwell (RB TB, dynasty-SF ADP 139.8) missing from EDGE since at least 2026-08-11.
G-437d: three defects of one class - each resolves to NOTHING rather than erroring, which is why none was noticed. (1) POST /v1/admin/news-events validated player_name and took player_id on faith; it accepted the literal string '<the right id>', returned status ok, and wrote an event that could never match a player. Now checked via news_events.player_exists against the same crowd_player_ratings table resolve_player reads, 404 on unknown. (2) _apply_young_qb_ceiling's docstring claimed dual-threat QBs are excluded while archetypes_eligible lists dual_threat and every QB in the ref2024 top five IS dual_threat. (3) That whole method is inert under the current config - the Pattern 13b guard short-circuits it whenever qb_use_v2_ltv is true, and ceiling_bonus reads exactly 1.000 on all 5,790 rows across 7 ref years and every position - so nothing tests it and flipping the flag off re-arms a 1.55x multiplier. Both notes now describe the code. Deliberately NOT changed: the eligibility list or the flag; that is a modelling ruling with an accuracy consequence, and this is documentation plus validation.
G-437c: G-437B COULD NOT REACH THE PLAYER IT WAS BUILT FOR. Reported durations were gated behind three guards that belong to the team_change CALIBRATION, not to games-available ARITHMETIC: the dynasty-only framework check, an is_rookie skip at the top of the event loop, and the fact that RedraftEngine never calls the layer at all. Jordyn Tyson - the first real use, logged as event 47 with 'out 2 months' - is a 2026 rookie, so the event was accepted and could never apply. Fix: parse_duration_games + news_duration_correction move to availability_layer (the shared layer G-243 established for exactly this, so redraft does not grow a second copy that drifts - this repo already carries two rookie-id conventions from that failure mode); the is_rookie skip narrows to the team_change branch; RedraftEngine calls the shared helper on redraft_proj immediately after its live-availability block, inside the non-snapshot else so backtests stay inert. The distinction now stated in code: a calibration fitted on next-season PPG for veteran movers has no rookie or redraft meaning; (season_games - games)/season_games has both. Regression: test_g437b_duration_reaches_the_player.py, whose rookie and redraft cases were written to FAIL against G-437b.
G-437b: A REPORTED INJURY DURATION COULD NOT MOVE A PROJECTION. The news ledger already accepted event_type='injury' and the news/situation layer already converted a duration into a year-0 factor, but only for suspensions -- an injury event produced a NEWS flag and no price movement. This opens that gate (news_situation.injury, ships DISABLED) and adds parse_injury_duration_games, which reads games/weeks/months and takes the pessimistic end of a range, refusing anything unreadable so an unparseable note stays flag-only exactly as today. THE ORDERING BUG THIS AVOIDS: _apply_live_injury_adjustment (step 2h) has already multiplied projected_season_ppg by the severity tier before this layer runs at 2h-bis, and the LTV horizon loop multiplies both stamped columns, so stamping the duration factor directly would double-charge -- Tyson at 8 weeks and Questionable would price 0.529 x 0.92 = 0.487. The stamped value is therefore a CORRECTION RATIO, target / live_injury_year_0_factor, making the product equal the reported duration. Authoritative because a reported prognosis beats the cohort average G-437 measured; same reasoning as the suspension branch skipping players the Sleeper feed already marks suspended. Players zeroed by season-ending IR are skipped. Snapshot-guarded via the existing FF_SNAPSHOT_AS_OF_SEASON early return, so walk-forward caches stay byte-identical.
G-437: THE LIVE INJURY TIERS WERE CALIBRATED AGAINST NOTHING. engine_config.live_injury.severity year_0 factors moved questionable 0.98->0.92, doubtful 0.95->0.68, out 0.92->0.62. Measured from nflverse weekly injury reports joined to nflverse snap counts, 2016-2025 (sample starts after the league eliminated Probable and redefined Questionable for 2016 - published play rate moved 53% to 75%, so 2015 is a different designation wearing the same word; including it moves the values by 0.003): for each designation-week, the share of the player's remaining regular-season team games in which he took an offensive snap, divided by the same figure for undesignated established starters observed in the SAME season-week. Relative because year_0_factor multiplies a per-game rate. Established = offensive snaps in >=3 of the previous 4 team games; both arms pass the same gate. Results: questionable 0.919 CI [0.903,0.934] n=1766; questionable 0.922 CI [0.908,0.938] n=1633; doubtful 0.683 CI [0.629,0.738] n=161; out 0.616 CI [0.592,0.640] n=941. All three CIs exclude the shipping value; zero contradicting position slices; out too generous in 11 of 11 seasons. SANITY ARM: players ON the injury report carrying no game status measure 0.988 CI [0.974,1.001], so the metric separates injury from roster churn rather than measuring churn. TWO EARLIER PASSES WERE WRONG AND ARE KEPT AS RECEIPTS: grading against injury_spells is invalid because build_injury_recovery_data sets OUT_STAT={out,doubtful} and so excludes Questionable by construction; and inferring appearance from the presence of a weekly stat row put UNDESIGNATED starters at 0.695, worse than designated ones, because a healthy WR targeted zero times has no row. Snap counts fixed it. WHY NO A/B: G-252 makes this layer return early under FF_SNAPSHOT_AS_OF_SEASON, so it is invisible to every backtest and no walk-forward arm can grade it; the caches nonetheless rebuild because G-182 fingerprints the whole config file. That rebuild is itself the G-252 conformance test - historical ranked columns MUST NOT move. KNOWN INCOMPLETE: `out` carries a 0.302 calendar-week gradient (0.728 wk4-6 -> 0.426 wk13+) that a constant cannot express; it ships anyway because it is strictly closer than 0.92 in every week bucket. Week-awareness for `out` is the follow-on and is due before week 1. Spell age also carries signal for `out` (0.747 fresh -> 0.536 at 5+ weeks designated) but is confounded with calendar week and was NOT disentangled; status_since already exists (G-239 stamps it) and nothing in the pricing path reads it.
G-405: A BETTER OFFENSE WAS LOWERING A REDRAFT DRAFT SCORE. Ruled and enabled by Ryan 2026-08-15 -- the THIRD application of his 2026-08-14 law that a status change must never move a player the wrong direction, after FA status (2026-08-14) and availability monotonicity (G-397, v1.6.150). THE DEFECT, derived from redraft_engine.py rather than argued: val is computed at [3/8] as (redraft_proj - replacement_ppg); _apply_team_context then MUTATES redraft_proj (team_context.py:385) and re-clamps ceiling but NOT floor; val is never recomputed. draft_score is val*sv + cw*ceiling_premium - fw*floor_penalty + opp - aging, where ceiling_premium = (ceiling - redraft_proj).clip(0) and floor_penalty = (redraft_proj - floor).clip(0). So a boost of +D contributed NOTHING through val (stale), -cw*D through ceiling_premium and -fw*D through floor_penalty. At the live weights cw=0.3 and fw=0.1 the net was -0.4*D: a player moving to a BETTER offense ranked WORSE for it and one moving to a worse offense ranked better, while the user-visible 'Team context (offense change)' chip reported context_adj_pct with the OPPOSITE sign. Bounded by max_adj_pct=0.1, so a SIGN error rather than a magnitude one, on the board most users actually draft from. MEASURED ON THE LIVE 736-ROW BOARD, flag off vs on: 115 team-context movers, 113 draft_score changes, ALL 113 now agreeing with the chip, ZERO contradicting it, and ZERO non-movers touched. 294 rank changes are positional displacement from those 113 moving, not independent movement. Named cases: Deebo Samuel +10.0% adj, rank 54 -> 45, score +2.17 -> +2.76; Cedrick Wilson Jr. +10.0%, 141 -> 123, -0.92 -> -0.50; Andy Dalton +8.1%, 59 -> 49; Tutu Atwell -8.8%, 95 -> 115, +0.04 -> -0.37; Deven Thompkins -9.7%, 157 -> 167. THE REPLACEMENT BAR IS HELD, DELIBERATELY, and that was a measured choice rather than a stylistic one. The obvious implementation re-calls _compute_val, which ALSO re-derives replacement levels from post-context projections; measured, that moved RB 4.35 -> 4.40 and WR 4.96 -> 4.86, and because draft_score multiplies val by a PER-PLAYER availability factor a uniform bar shift reorders players who never changed teams -- 522 scores and 306 ranks moved for a fix aimed at 115 players. Holding the bar keeps the change attributable to exactly the players team context touched, per G-243's scope rule that the boundary is what makes the A/B readable. Whether the bar SHOULD reflect team context is a real and separate question, filed rather than smuggled in. A PREDICTION THE MEASUREMENT CORRECTED: the algebra says the fix only restores the law when sv > cw+fw = 0.4, so I predicted ~9 movers would still contradict the chip at the 0.30 availability floor. ZERO did. The arithmetic is still right and is pinned in tests; the population that would hit it is empty on today's board, most likely because G-397's positive_val_only forces sv=1.0 whenever val<=0 and the movers near the availability floor are also near val=0. Both facts are recorded rather than one rounded away. WHY A RULING AND NOT A HARNESS: no A/B can say whether a better offense SHOULD improve a draft ranking -- it can only report how many players move and by how much. Same substitution as the FA law and G-397, and deliberate. Shipped default-OFF first, measured, then flipped. TESTS: tests/test_g405_team_context_sign.py, 22 tests, including the exact sv = cw+fw boundary, the residual's arithmetic, and a guard that the replacement bar is not re-derived.
G-403: SLEEPER IS NOW REFRESHED DURING THE DAY. NO RANKED-OUTPUT DELTA -- this changes how often the box does I/O, never what any board says. This is the piece that actually moves the measured 16h15m, and the last item on the critical path for Ryan's D1 ('within the hour, tighter on game days'). THE GAP. The Sleeper table was force-pulled at exactly TWO moments: 02:00 (_roster_refresh_core) and 02:30 (crowd reseed). Every other reader passes force=False, and because the 02:00 pull resets the sidecar manifest, the 24h TTL never elapsed during the day -- so no request-path or roster-sync read ever refetched. A Sunday inactive announced ~11:30 ET reached the cache at the Mon 02:00 pull and reached boards at the Mon 03:45 warm: 16h15m typical, ~25h for a change landing just after 02:30. SHIPPED, FOUR PARTS. (1) TWO NEW JOBS: sleeper_refresh_hourly at :00 and sleeper_refresh_gameday on Sundays 10-23 every 15 min, both self-skipping outside the in-season month window {8,9,10,11,12,1} -- August included, because the draft and the Waller signing are both August events. Both do I/O ONLY and never engine work, which is what stops a tighter cadence from saturating the single recompute lane. (2) THE TTL IS A BACKSTOP, NOT THE REFRESH, and that distinction is the design. load_sleeper_players' own docstring says request-path reads keep force=False 'so a board view never triggers a 50MB download' -- true only while the cache is fresh, so cutting the TTL to chase freshness would have moved a ~50MB inline download onto board builds in a single-worker app. In-season it drops to 6h: tight enough that a DEAD scheduler self-heals in hours instead of a day, loose enough that it never fires while the jobs work, and well inside the 48h at which player_status stops trusting depth and injury. Refresh, then backstop, then refuse. (3) THE resolve(force=True) IN THE JOB IS LOAD-BEARING, and missing it would have made the whole chain silently inert: player_status memoises on the cache file's MTIME and peek_fingerprint answers ONLY from that memo, returning None ('cannot judge', so serve) on a miss. A fresh pull changes the mtime and leaves the memo keyed to the old one, so without re-resolving in the job, peek would return None until some board build happened to repopulate it and the G-399/G-401 drift gate would never fire. (4) THE PER-LEAGUE REBUILD COOLDOWN, shipped in the SAME change rather than after it. With injury in the fingerprint (G-401) and a 15-minute game-day pull, drift is near-continuous across 604 currently-flagged players; engine runs are serialised at ~35-105s apiece, so an unbounded enqueue would queue every league on every read and leave user-initiated roster syncs waiting behind drift work. Default 900s, matched to the game-day pull interval on purpose -- anything tighter buys nothing because the data underneath has not moved again. Throttled reads still SERVE and still flag; only the enqueue is suppressed, and rankings.read.status_drift_throttled makes the throttled:enqueued ratio visible before anyone retunes it. ALSO CLOSED: scripts/daily_data_refresh.py's 'invalidation' was a Path.touch() running in a SUBPROCESS, structurally unable to reach this process's caches; the parent now calls data_loader.invalidate_player_data() itself after the child succeeds. OPERATIONALLY: sleeper_refresh.dry_run=true logs what each firing WOULD pull, including current cache age, without fetching -- intended to be watched across one game-day window before going live. enabled=false restores the pre-G-403 cadence exactly. A NOTE ON HOW THE JOB REGISTRATION NEARLY SHIPPED BROKEN: JOB_REGISTRY is built at module-load time ABOVE where run_sleeper_refresh is defined, so registering beside the function raised NameError at import and would have taken the entire scheduler down on boot -- and py_compile passed it. Caught by importing the module rather than compiling it, which is now part of the verification.
G-401: THE DRIFT GATE COULD NOT SEE A SUNDAY INACTIVE. Self-correction on v1.6.152, same day. NO RANKED-OUTPUT DELTA -- the fingerprint is provenance, not a price. THE DEFECT. G-399 shipped a status-drift gate whose entire purpose is to stop serving a board that predates a live-status change, and it compares meta.player_status.fingerprint. That fingerprint hashed TEAM and DEPTH_CHART_ORDER only. A Sunday inactive changes neither: the player stays on his team, in his depth slot, and flips injury_status to 'Out'. So the gate built to catch Sunday inactives would have slept through every one of them, and the 16h15m latency finding would have been 'closed' by a mechanism structurally unable to observe the event. The fingerprint was not wrong -- it was built for G-182 roster drift and its docstring said 'roster' the whole time. G-399 reused it for a second purpose without re-deriving what that purpose required. MEASURED against the live 12,210-player Sleeper cache, 518 of which carry a designation (Questionable 294, NA 89, PUP 73, IR 49, Out 6, Sus 3, COV 2, DNR 2): ruling one depth-1 starter Out left the old recipe BYTE-IDENTICAL at ad54f68c124c1087, and moves the new one 0cbacb461b315bf3 -> 4f1573c9e036ec3f. SHIPPED: injury_status and injury_body_part join team and depth in the recipe. Both move a PRICE -- injury_status selects the severity tier, injury_body_part drives career-bender escalation (ACL / Achilles / neck / head / spine) -- which is the inclusion rule, not 'everything the feed carries'. practice_participation is DELIBERATELY EXCLUDED and a test documents why: availability_layer stamps it in _TEXT_COLUMNS for display and never reads it into _FACTOR_COLUMNS, so it moves no number while changing most weekdays for most listed players; including it would multiply rebuilds for nothing. The recipe now reads from snapshot.raw, defensively, so a gsis present in `teams` without a raw record cannot take down a board build. COST: changing the recipe invalidates every cached board exactly once, on the deploy that ships it -- which is free, because the caches are in-process and already empty after a deploy. There is no ongoing cost increase: Sleeper still refreshes only at 02:00 and 02:30, so the fingerprint still moves twice a day. That changes when the in-season cadence lands, and a per-league rebuild cooldown belongs in the same change. TESTS: tests/test_g401_injury_in_fingerprint.py, 16 tests, including one per live designation, the Questionable->Out transition, the Knee vs 'Knee - ACL' body-part case, no-regression on signings and depth moves, and a recipe test that fails if any field which moves a factor is missing from the hash. THE LESSON IS NOT 'ADD A FIELD'. A value built for one purpose was reused for a second without re-deriving its requirements, and nothing in the type system, the tests or the docstring objected. The recipe test exists so the next reuse has to state what it needs.
G-400: FOUR PLACES THE PLATFORM REPORTED SUCCESS WHILE DOING NOTHING. NO RANKED-OUTPUT DELTA on any live board -- one of the four changes what a BACKTEST produces, and that is the point. One failure mode, four instances, none of which could ever have surfaced as an error. (1) THE REDRAFT HARNESS HAS BEEN SCORING AN EMPTY BOARD. redraft_engine.run()'s snapshot guard read `return self._finish_live_availability_noop(...)` where it meant `proj_df = ...` -- and it is in the body of run(), not a helper. So under FF_SNAPSHOT_AS_OF_SEASON, run() returned a DataFrame and stages [3/8] through [8/8] never executed: no VAL, no ceiling/floor, no opportunity overlay, no expected_games, no draft_score, no ADP, no tiers, no _build_output. It survived because the consumer, validate_redraft_board_outcome.py:295 -> _board_df, does `result.get('rankings') or {}` -- and on a DataFrame .get() is a COLUMN lookup, so it missed, returned None, the loop never ran, and an empty board came back with no exception. Every reference year scored nothing and the run reported success. Fixed to an if/else so the live layer is still skipped on historical boards (the G-252 guarantee) while the rest of the pipeline runs. _board_df now RAISES on a non-dict result, so the class cannot recur silently. Production was never affected: the env var is not set on the serve path. Dynasty walk-forward was never affected either -- backtest.py mentions redraft exactly once, in a comment, and never builds a redraft board -- so no dynasty verdict is retracted by this. (2) ONE MISSING CSV KILLED EVERY CRON JOB. api.startup_event's step-4 gate returns when player_stats_season.csv / players.csv / engine_config.json are absent, logging 'Skipping startup rankings'. start_scheduler() ran BELOW that return, and looked like module-level code because a column-0 '# ENDPOINTS' banner sat between them -- Python emits no DEDENT for comment-only lines, so those statements were still inside the function and the banner disguised it perfectly. Consequence: a missing data file silently disabled the 02:00 Sleeper pull, the 03:45 league warm, the daily data refresh and the 5-minute cache reaper for the life of the container, and the scheduler is exactly what REPAIRS a missing-data boot (it carries a make-up data refresh), so the one condition that most needed it was the one that disabled it. Scheduler moved to step 3.8, above the gate; the gate's log is now an ERROR that states what it actually costs. A test fails if indented code ever follows that banner again. (3) THE STATUS RESOLVER LOGGED NOTHING, EVER. player_status.resolve() takes `log=print` in its signature and never calls it -- zero occurrences in the body. Every failure path only appended to snapshot.warnings. So a total loss of every live signal (no live fetch AND no cache file) produced no stdout line, no log record and no metric: the FA discount, buried depth, the QB depth gate, live injury and the mover predicate all silently stopped firing, and the only trace was meta.player_status.warnings on a payload nobody reads until something already looks wrong. This module's own header guarantee 5 -- 'Degradation is LOUD' -- was true of the snapshot object and false of the process running it. (4) THE 48h EXPIRY WAS EQUALLY SILENT. Two consecutive failed 02:00 pulls put the Sleeper cache past 48h; usable() flips False for depth AND injury; every board then prices injured and buried players as fully healthy and available -- which from outside is indistinguishable from a quiet week. Both now route through one _report() helper called once per real resolution (never on a memo hit, so a stale cache pages once per resolution rather than once per board), emitting logger.error plus counters player_status.unavailable and player_status.signal_expired. Those counters, and boot.warm_skipped_missing_files / boot.scheduler_start_failed, are registered in obs_metrics.WATCHDOG_COUNTER_THRESHOLDS at 1, so the hourly metrics watchdog escalates them to Sentry -- incrementing a counter nothing watches would be the same silence in a new costume. TESTS: tests/test_g400_silent_success.py pins each fix AND the structural trap that hid it, including an AST assertion that the snapshot guard contains no Return and a scan that no indented code follows the ENDPOINTS banner. THE THROUGH-LINE: in all four cases the code read correctly. Only running it, or diffing what it produced, exposed the defect -- which is now the fourth, fifth, sixth and seventh finding in this body of work that inspection would have missed.
G-399: A SLEEPER CHANGE NOW REACHES A CACHED BOARD. NO RANKED-OUTPUT DELTA -- this changes WHEN a board is recomputed, never what it says. THE DEFECT. _config_fingerprint (league_config_service.py:186) hashes league_id, format, num_teams, superflex, te_premium, scoring, roster_slots, my_user_id, keeper_horizon_years and status -- ten league-SHAPE fields and NOTHING about the data. So new Sleeper data has never changed a cache key and has never, by itself, caused a miss. The SPA's fast path (GET /v1/rankings?config_fp=...) then returns cached['result'] with no age check of any kind: BoundedTTLCache.get (cache_store.py:98-103) only does move_to_end + return, and the reaper judges on the value's write-time timestamp, which a read never refreshes. Net: the only things that have ever replaced a live board are the 03:45 nightly force-warm, a roster sync, and the reaper dropping the entry past 7200s. MEASURED PATH for a Sunday inactive announced ~11:30 ET: into sleeper_players_cache.json at the Mon 02:00 forced pull -- the only forced pulls are 02:00 and 02:30, every other reader passes force=False, and 02:00 resets the manifest so the 24h TTL never elapses during the day -- then onto boards at the Mon 03:45 nightly warm. 16h15m typical; ~25h for a change landing just after 02:30. THE FIX WAS ALREADY HALF-BUILT. Every board already stamps the status fingerprint it was built with, at meta.player_status.fingerprint, for exactly this purpose ('two boards with the same fingerprint saw the same roster'). Nothing ever compared it. This compares it. WHY THE CACHE KEY WAS NOT TOUCHED. The obvious fix -- add a data component to _config_fingerprint -- is wrong and would have made staleness PERMANENT: the fp is an opaque md5 persisted on the leagues DB row, the SPA asks with that stored fp, and a data-dependent fp orphans it the moment Sleeper moves, leaving the SPA requesting an fp that still resolves to the OLD board, forever. It would also arm the G-363 parity guard, which intersects _config_fingerprint's field names with from_manual_input's parameters and would then require every leagues-row reconstructor to pass a value that is not a league attribute. The fingerprint is untouched; the comparison is on the READ side against the stamp the board already carries. WHY IT SERVES RATHER THAN REFUSES. The 02:00 pull moves the fingerprint for EVERY league at once. Refusing on drift would send the first read of each of ~58 leagues into a rebuild, and engine runs are serialised on max_workers=1 at ~35-105s apiece -- hours of queue, triggered by whoever opens the app at 02:15. So a drifted board is SERVED immediately, stamped meta.status_drift, and its canonical single-flight warm enqueued through the existing AUD-11 rescue path including that path's ownership check. The next poll gets the fresh board. Same stale-while-revalidate shape rankings_pipeline already uses on TTL expiry, keyed on data instead of on the clock. COST CONTROL. player_status.peek_fingerprint() answers ONLY from the memo resolve() already fills, keyed identically on the Sleeper cache mtime, and returns None on a cold memo or under FF_SNAPSHOT_AS_OF_SEASON. It never calls load_sleeper_players, which re-downloads the ~50MB Sleeper table inline once the cache is past its 24h TTL -- correct on a board build, unacceptable on a request path in a single-worker app. Every ambiguity resolves to 'serve': None from the peek, no stamp on the board, or a board built with no usable live signal all mean 'cannot judge', never 'changed'. A false positive costs a needless rebuild; a false negative costs one more poll interval on a board that was already stale. OBSERVABILITY BEFORE ENFORCEMENT: rankings.read.status_drift and rankings.read.status_drift_warm ship with it, and status_freshness_gate.enqueue_rebuild=false retains the counters while removing all added load. NOT DONE HERE, FILED: the redraft read path (cache_resolvers._resolve_redraft_result) carries the identical hole, and this closes the detection gap rather than the 16h refresh cadence itself -- the Sleeper TTL and the game-day pull window are separate work.
G-398: close the redraft measurement gap, including one this session CREATED. RANKED-OUTPUT NEUTRAL on redraft (measured). THE SELF-INFLICTED HALF. v1.6.149 closed 'a gate that fires and records nothing' for dynasty. v1.6.150 then landed the DISC-Q2 QB depth gate on redraft -- and stamped no per-player flag, creating a fresh instance of the same defect on the other board, in the same session. Caught by diffing the week0b and week0c freezes rather than by reading the diff. THE INSTRUMENT WAS ALSO WRONG. freeze_prospective_board carried ONE global gate map and applied it to all three formats. Redraft runs a different layer stack -- no stale decay, no committee detection, no pre-LTV structural ceiling, team_context in place of news_situation -- so the single map made the redraft freeze ADVERTISE TWELVE gates it could not grade, while silently omitting TWO it could: season_availability_factor (the RDR-AVAIL-2 term the G-397 law had just changed, 656 players) and context_adj_pct (team_context, 117 players). An instrument that overstates its own coverage is worse than one with a known gap, because the gap is at least visible. Gate maps are now per-format, and every freeze prints an OVERCLAIM line naming any gate the board advertises but cannot grade. SHIPPED: redraft now stamps injury_discount, injury_pattern, injury_risk (it built the profiles and discarded everything but two reporting fields) plus qb_depth_order and is_qb_depth_capped. Five gates that fired blind on the board most users actually draft from. REMAINING, HONESTLY SCOPED: team_context records magnitude but not direction, and redraft's FA discount runs a different path than dynasty's graded by_pos_age schedule so there is no single factor to stamp -- that one is engine math, not plumbing, and is filed rather than faked.
G-397: BECOMING LESS AVAILABLE MUST NEVER RAISE A DRAFT SCORE. Ryan's ruling, 2026-08-15 -- the mirror of his 2026-08-14 law that signing with an NFL team must never LOWER a ranking, and ruled the same way and for the same reason. THE DEFECT. redraft's RDR-AVAIL-2 season-value factor multiplies `val`, which is value ABOVE REPLACEMENT and therefore NEGATIVE below it. Scaling a negative number by an availability factor < 1 makes it LESS negative, so a below-replacement player ranked HIGHER for being unavailable. Found while measuring a separate coherence fix, not by reading code: capping 23 backup QBs at 2.5 expected games moved ALL 23 UP and NONE down -- Kirk Cousins draft_score -3.280 -> -0.470 and rank 62 -> 23, Justin Fields -1.400 -> +0.090, crossing from below replacement to ABOVE it purely by being ruled unavailable. The behaviour was DELIBERATE and documented in the source ('a below-replacement player who misses time costs you FEWER deficit-weeks'). That is a sound LINEUP argument -- if you are forced to start him, his absence costs less -- and an unsound DRAFT-BOARD one, because nobody is forced to start him: you start replacement instead, so his season value over replacement is ~0 either way. WHY THIS WAS RULED, NOT A/B'd. No existing harness can grade it. scripts/validate_redraft_season_value.py tests whether expected-games beats flat-games at predicting season TOTALS and never touches val or draft_score. The question is semantic -- what a draft board MEANS -- not calibrational, which is exactly why the FA case was ruled rather than measured. SHIPPED: redraft.season_value_expected_games.positive_val_only = true (schema-registered). The availability factor applies to POSITIVE val only. Minimal change that enforces the law without inventing a penalty the data has not earned. MEASURED on the live redraft board: 563 of 750 rows move, ALL DOWN, ZERO UP -- the monotonicity property holds by construction, since the guard can only remove an upward distortion. The sub-replacement tail stops being compressed toward zero by unavailability (Collin Johnson -2.22 -> -7.68). QB top-12 ordering byte-identical; the single top-36 'mover' is Joe Burrow at -0.06 with NO rank change. UNBLOCKED AND LANDED IN THE SAME CHANGE: dynasty/redraft QB availability parity. redraft's _enrich_expected_games called build_injury_index WITHOUT qb_depth_chart, so the DISC-Q2 cap fired on dynasty only and 35 of 77 QBs carried two different expected_games across formats (Will Levis 2.5 vs 11.3, Mac Jones 2.5 vs 11.3, J.J. McCarthy 2.5 vs 10.0) while ZERO non-QB skill players diverged. The Waller defect class one level up: two subsystems holding contradictory beliefs about one player. Redraft now builds the map from the SAME status resolver with the SAME G-387 snapshot guard. Divergence measured 35 -> 0. Backup QBs now FALL where the parity fix alone would have raised them: Sam Howell 23 -> 44, Anthony Richardson 24 -> 45, Nick Mullens 30 -> 51. Parity was deliberately HELD on its own earlier the same day precisely because shipping it without the guard would have made the board worse for 23 quarterbacks; a test now fails if the guard is ever reverted while parity remains. DYNASTY AND KEEPER ARE UNTOUCHED by both halves.
G-396: four gates fired and left no per-player trace. RANKED-OUTPUT NEUTRAL (measured: 0 of 837 dynasty values changed, 0 ranks changed) -- this is instrumentation, not math. It exists because G-390's prospective freeze could not grade what the engine does not record. (1) mover_volume_factor arrived NULL on every dynasty payload row while the engine logged '[volume_reproject] re-projected 111 mover weighted_ppg'. Root cause: metrics_xfp.apply_volume_reproject multiplied weighted_ppg IN PLACE and returned only a count -- the dynasty payload has emitted the key since 2026-08-14 and nothing ever wrote the column. redraft_engine has its own writer and already stamped it. Now stamped at the point of application: 105 dynasty rows carry a real factor. (2) structural_ceiling_factor (207 players, mean 0.846), (3) stale_decay_factor (120 players, ladder 0.12/0.34/0.36) -- both computed per player and dropped before the payload, so a boolean said a gate fired but never by how much. (4) qb_depth_order / is_qb_depth_capped -- the DISC-Q2 depth gate emitted NO per-player flag at all; 51 board QBs now carry one. WHY THIS MATTERS FOR THE SEASON: six live-ingestion gates are snapshot-guarded, so no walk-forward run can grade them and no live board is reproducible after the fact (G-390). The only grading path is a pre-season freeze scored against outcomes, and a gate with no per-player factor can be scored in aggregate only -- 'the board was N% accurate' rather than 'this gate earned its keep'. Re-freeze before Week 1 to capture all four. HELD, NOT SHIPPED (filed G-397): dynasty and redraft disagree about expected_games for 35 of 77 quarterbacks (Will Levis 2.5 vs 11.3 games, Mac Jones 2.5 vs 11.3, J.J. McCarthy 2.5 vs 10.0; ZERO non-QB skill players diverge). Cause: redraft's _enrich_expected_games calls build_injury_index without qb_depth_chart, so the DISC-Q2 cap fires on dynasty only. The parity fix was written and MEASURED, and it is blocked: redraft's RDR-AVAIL-2 season-value factor multiplies VAL, which is NEGATIVE below replacement, so cutting a backup QB's availability RAISES his draft_score. All 23 affected QBs moved UP, none down -- Kirk Cousins +2.81 and 39 rank spots, Justin Fields -1.400 -> +0.090, crossing from below replacement to above it by being ruled unavailable. The behaviour is deliberate and documented in the source ('a below-replacement player who misses time costs you FEWER deficit-weeks'), which is a sound LINEUP argument and an unsound DRAFT-BOARD one. Resolve the negative-VAL treatment first, then land parity. || COMPLETES THE v1.6.148 RECORD (added same day, after the fact). v1.6.148 shipped on a binding SHIP from scripts/validate_availability_layer.py, and that verdict carried NO RANK-METRIC CHECK -- rank_delta_ci was passed as None, which eval_lib.verdict_block treats as passing. CLAUDE.md defines SHIP as 'dMAE CI95 excluding zero AND no rank-metric regression' (audit D5), so the verdict was incomplete against the house bar. The check has now been run. scripts/live_ab._set_gate was extended to accept DOTTED PATHS so nested gates can reach the canonical harness at all -- before this it could only flip top-level keys, which is why most of the injury_model block, redraft.* and xfp_base.* had never been A/B'd through it. RESULT, board-level walk-forward, ref2022/2023/2024, OFF=ungated vs ON=max_seasons 2, sentinel FIRED on all three (102/124/102 rows): TIE at all four positions, NO-SHIP. QB dMAE -0.2572 CI95 [-5.9283, +5.3516]; RB -0.4341 [-1.7550, +0.8487]; WR -0.5904 [-1.7617, +0.6550]; TE -0.3586 [-1.4331, +0.6883]. Every point estimate is an IMPROVEMENT and every CI crosses zero. Crucially, the rank check -- the specific thing that was missing -- PASSES at all four: Dro QB -0.0571 [-0.1217, +0.0021], RB +0.0081, WR -0.0024, TE -0.0058, no regression beyond CI anywhere. Per-season direction WR 3/3 and TE 3/3 improve. READ IT HONESTLY: this is NOT independent board-level confirmation. It is absence of harm plus a directionally consistent point estimate at n=59-216 players, against the layer-level test's 773. A single multiplier among ~20 layers, moving 104 of 837 values, cannot move a board-level CI at that power -- which is itself the finding: the board-level harness can only ever confirm LARGE changes, and individual gates need the layer-level matched-control design. The change stands on the layer-level evidence (SHIP in and out of sample, every position and history-depth slice guard non-positive, 21/21 and 8/8 seasons) plus the qualitative ghost-class defect it closes, NOT on this run. ALSO NOTE the advisory survivor-ppg view degrades in every season and position (dMAE +0.04 to +0.69, bias more negative). That is expected and arguably confirmatory rather than contrary: the change charges low-availability players more, and a survivor-conditional target scores only players who went on to PLAY. It is precisely the target eval_lib demotes to advisory (audit D6/A2) because it drops the cohort the gate is about. TWO OTHER SELF-CHECKS RUN THE SAME DAY, both passing: a PLACEBO on the layer-level design (split the control cell, score half B against half A's yardstick -- 0.9981 CI95 [0.9565, 1.0386], 1.0 inside; the originally-cited 'control lands on 1.000' figure was near-TAUTOLOGICAL and is retracted as evidence), and HORIZON SENSITIVITY (SHIP at H=1, 2, 3 and 5 in both windows, effect growing with horizon, so the 3-season choice was conservative rather than load-bearing).
G-388: the availability layer's first A/B, and the one thing it found. injury_model.healthy_pattern_bypass (DISC-Q1f) restored the legacy alpha-power discount for any player the pattern classifier called 'healthy' whose availability_scaling_v2 discount fell below 0.90. Its own config note calls it 'surgically narrow' and names ONE case: Daniels, 17/17 rookie + 7/17 sophomore -- two seasons of history. It was not narrow. _classify_injury_pattern only returns 'chronic' at 2+ seasons under 0.30 availability (~under five games), so a player who plays six or seven games EVERY year never trips it, stays 'healthy' forever, and the bypass hands him back ~0.79-0.88 -- the legacy 0.792 floor that availability_scaling_v2's own note says it exists to escape ('the legacy floor=0.792 trap that made Wentz/Rivers/Martinez fantasy-relevant'). Twenty players on the current pool were in that state, by name: Deshaun Watson (6, 6, 7 games -> 0.820), Kendre Miller (7, 6, 7 -> 0.792), Xavier Gipson (14, 7, 6 -> 0.792), Zay Jones (9, 6, 6 -> 0.843). DISC-Q1f had quietly re-opened the exact ghost-player class DISC-Q1 closed, for anyone whose absences were chronic but never quite season-ending. SHIPPED: healthy_pattern_bypass.max_seasons = 2 -- the bypass fires only for the 1-2 season cohort it was written for. EVIDENCE: scripts/validate_availability_layer.py, new, the first A/B this layer has ever had. Matched-control by design after the two 2026-08-14 self-corrections: realized 3-season discounted TOTAL-POINTS retention (unconditional) divided by a did-nothing control matched on position x age band x prior-rate tier, so natural age decline (control retention 0.747 / 0.590 / 0.437 by band), mean reversion and league-average games missed all divide out. The control cell is defined by STATE (avg_avail < position_baseline, the branch condition itself), not by the multiplier under test, so no arm can move its own yardstick. 1,959 treated anchors / 773 players, anchors 2002-2022, 2026 hard-refused. VERDICT: binding SHIP in BOTH windows -- dMAE -0.0396 CI95 [-0.0479, -0.0313] full, -0.0467 [-0.0584, -0.0350] on anchors after 2014 -- every position AND history-depth slice guard non-positive, 21/21 and 8/8 seasons agreeing in direction. WHY GATED AND NOT REMOVED: removing the bypass outright scores marginally better on the aggregate (-0.0415) and is what the aggregate alone would have shipped, but its own motivating cohort refuses to resolve -- hist=1 is n=49 in sample / n=23 out of sample and FLIPS SIGN between windows, so outright removal fails its own binding slice guard out of sample. The change stops where the evidence does. Daniels 0.859 -> 0.859, Hampton 0.858 -> 0.858, Nabers and Waller untouched. CROSS-CHECK: experiments/g152_availability_layer_sizing.py re-run under both configs reproduces its published 2026-07-25 values to four significant figures under the old config (RB withheld 0.5117 vs 0.512, WR 0.3408 vs 0.341, TE 0.3859 vs 0.386), then reports withholding rising RB 0.512->0.944, WR 0.341->0.568, TE 0.386->0.604 with every position STILL UNDER_PRICED -- two instruments built for different questions agreeing on sign, and no overshoot. NEGATIVE RESULTS KEPT: availability_scaling_v2 itself is correct and load-bearing (disabling it is a decisive DEFEND, +0.0528 [+0.0427, +0.0628], 0/21 seasons improve); scaling_floor 0.15 -- the one knob whose note openly admits to a judgement call -- is INERT (0.15 -> 0.30 moves dMAE +0.0001 [-0.0003, +0.0005]); a 12-cell per-position schedule buys nothing over one flag; min_sample_gate alone is a no-op (-0.0002 [-0.0012, +0.0007]) and stays OFF, which IS the binding-CI A/B its config note has been waiting for. BOARD IMPACT: dynasty 104 of 837 values move, max 75 rank spots, 3 positional top-24 swaps; keeper 98 of 837, max 79; redraft ZERO (redraft consumes expected_games / availability_rate, never injury_discount, so this layer has never reached it). LOGGED NOT FIXED: (1) reported availability_rate and applied discount disagree for the moderate-single-miss cohort -- expected_games_recovery_floor lifts the REPORTED field to G-30's 0.726 while the discount keeps the backward average (Evans 0.709, G. Wilson 0.606, Aiyuk 0.599); the ungated bypass was incidentally papering over it. Not reconciled here because G-30's number is a ONE-season availability rate and this study's target is three-season control-relative points retention. (2) G-152's aged-TE accounting moves from net +0.040 [-0.381, +0.451] to +0.198 [-0.235, +0.623] -- still straddling zero, no double-count demonstrated, but the first cell to re-measure if availability tightens further; visible as three deep TEs newly on the zero floor, an artifact of a SUBTRACTIVE flat charge sitting downstream of MULTIPLICATIVE discounts. Receipt: docs/RESULT_G388_availability_layer_ab_2026-08-15.md. Tests: tests/test_g388_availability_bypass_gate.py.
SECOND SELF-CORRECTION IN ONE DAY. v1.6.146 shipped stale_player_decay.by_age_band = 0.36/0.19/0.08. That derivation labelled a player stale using the Y1 OUTCOME and then scored Y1..Y1+2, so one of three target years was a guaranteed ZERO by construction of the label -- information the engine cannot have when it ranks. Caught by re-running the cohort the way the engine actually sees the world.
THE ENGINE FIRES THIS DECAY ON AN OBSERVED FACT: no row in the latest COMPLETED season. The honest analogue observes the miss and scores the seasons AFTER it, against a control anchored identically so it carries the same one-year gap in its projection window. Measured n=612 stale / 5,197 control, 2000-2021, unconditional target, 2026 hard-refused, cluster-bootstrapped by player: <26 0.361 [0.246,0.491]; 26-29 0.343 [0.243,0.469]; 30+ 0.116 [0.075,0.163]. The shipped ladder was BELOW the CI on two of three bands -- roughly 2x too harsh. Corrected to 0.36 / 0.34 / 0.12.
FA GATE REMOVED from the stale ladder. v1.6.146 fired it only when the player was also unsigned, on the reasoning that the joint state was what had been measured. With the anchor corrected, the measurable cohort is ALL players who missed a season -- rostered and unsigned alike -- and it cannot be split by historical FA status, which is not observable (the same proxy limit validate_redraft_fa_discount documents). Gating a general-population estimate to a subpopulation is the same error class as fitting on a mis-specified target, so the ladder now applies to every stale player. Consequence: stale-but-rostered players move from the flat 0.50 to the measured ladder, which is why the board delta is dominated by rostered players who missed 2025 (Brandon Aiyuk WR107->WR144, Tank Dell WR114->WR148, K.J. Osborn, Trey Palmer).
REVISES AN EARLIER CLAIM. v1.6.146's note said audit A10's stack direction was wrong -- that the live stack was TOO SOFT. That reading came from the contaminated cohort. On the corrected one the ladder I had shipped was TOO HARSH. A10's concern about compounding correlated evidence is UNRESOLVED, not refuted.
OPEN AND DELIBERATELY NOT GUESSED: how stale_player_decay should COMPOUND with fa_status_discount, and with structural_ceiling / live_injury, for a player carrying several at once. Brandon Aiyuk is the live named case -- an ACL absence priced by the stale ladder, the structural ceiling and the live-injury layer simultaneously. The joint cells cannot be identified from history, so the multipliers simply multiply, as before. Flagged for the prospective harness rather than tuned.
REDRAFT FA DISCOUNT: MEASURED, NO CHANGE NEEDED. scripts/validate_redraft_fa_discount.py has the same three method defects (no matched control, pooled subtypes, and i.i.d. row resampling rather than the cluster bootstrap eval_lib requires -- audit D4 says that makes CIs too narrow). Re-measured on the corrected matched-control design at horizon 1: every late_sign cell's live multiplier sits INSIDE the measured CI (QB .85/.80 vs .81/.73; RB .85/.76/.65 vs .69/.83/.71; WR .85/.80/.75 vs .69/.80/.80; TE .85/.85/.65 vs .79/.84/.75). The aging leak that broke the dynasty fit is small over a single season because the did-nothing control retains ~1.0. The redraft schedule stands as-is; its VERDICT should still be re-derived on a cluster bootstrap before being cited.
SELF-CORRECTION. v1.6.145 shipped fa_status_discount.by_pos_age fitted on RAW retention (RB .44/.41/.19, WR .44/.40/.20, TE .50/.45/.25) with a binding SHIP in and out of sample. The verdict was real; the TARGET was mis-specified, and a correct A/B on the wrong target still ships the wrong number. Found by building the stack harness and looking at the CONTROL cell.
DOUBLE-COUNT 1 -- AGING. Players who changed nothing (same team, full season) realize 1.06 / 0.87 / 0.73 by age band. A player who stayed put still loses a quarter of his rate by 30+, and enrich_age_curves already prices that on every player before these discounts are reached. Fitting a status multiplier on raw retention charges the age curve twice. Divided by a matched control the FA penalty is roughly FLAT (0.55-0.78); the steep 0.44->0.19 gradient was aging, not free agency. This also corrects how audit A10 was read: its '~0.7 under 26, ~0.15 over 30' describes RAW retention, not the marginal cost of being unsigned.
DOUBLE-COUNT 2 -- STALENESS. The FA cohort pools two subtypes the engine treats differently: late_sign (unsigned but PLAYED) and y1_missing (unsigned AND missed the season). Fitting one factor on the pool drags it toward y1_missing, and stale_player_decay then multiplies 0.50 on top of exactly those players -- the stale penalty charged twice to one group and once to a group that never earned it. Against control the two cells are 0.63/0.72/0.61 vs 0.23/0.14/0.05: not one population.
CORRECTED DESIGN: matched-control, subtype-separated -- the same mover-vs-matched-control shape team_change_deltas.json already uses, which should have been the precedent the first time. fa_status_discount is now derived from the late_sign cell relative to a did-nothing control matched on (position, age band): RB .59/.68/.57, WR .65/.67/.55, TE .67/.69/.70. Binding verdict on the FA+stale pair via eval_lib.verdict_block: SHIP, dMAE -0.0247 CI95 [-0.0352,-0.0146]; out-of-sample (derive y1<=2014, score y1>2014) SHIP, dMAE -0.0299 CI95 [-0.0441,-0.0157]. QB again held at the flat default: in-sample slice CI crosses zero (-0.0295 [-0.0656,+0.0074]). Measured QB is 0.78/0.74, so the retained 0.50 is CONSERVATIVE for quarterbacks, not lenient.
stale_player_decay GRADED and SCOPED. The flat 0.50 had a cohort of one player (Joe Mixon, named in its own docstring) and moves 117 players. New by_age_band 0.36 / 0.19 / 0.08 is derived so that FA x stale reproduces the measured joint state. It fires ONLY when the player is also unsigned -- the exact cell the CI was computed on. 'Stale but still rostered' (missed the season on IR, under contract) never appears in the observable cohort, so he keeps the flat 0.50 rather than being priced off an extrapolation. CORRECTS AUDIT A10'S DIRECTION: A10 predicted the 0.5x0.5 stack OVER-penalizes correlated states; measured, the live stack is TOO SOFT for older players -- it applied 0.12 to a 30+ stale-and-unsigned player who realizes 0.05.
Board delta vs the pre-2026-08-14 baseline: 403 ranks move, mean 7.1, max 96, ZERO positional top-24 disturbed. Direction is now coherent: unsigned-but-played players are treated far more gently than the morning values did (Waller TE152 -> TE87, not TE87-harsh), while stale-AND-unsigned players fall correctly (Diontae Johnson WR194 -> WR290, Amari Cooper WR193 -> WR289, Joe Mixon RB100 -> RB158 -- the canonical stale case finally priced).
UNSCOPED DEFAULT BOARD (G-BOARD-1). Two stacked bugs. (a) LeagueConfig.from_league_info_json passed _detect_sleeper_format the bare metadata dict, so settings.type was never consulted and its legacy string scan fell through to `return "dynasty"`. Belichick Cup carries settings.type=0 (redraft) and a metadata dict with no type string, so production served /health default_board {league_name: 'Belichick Cup', format: 'dynasty', rec: 1.0} -- the dynasty LTV framework on a redraft league's scoring, to all 27 _build_league_config(None) consumers. (b) The deeper flaw that misread was hiding: the unscoped default resolved to whichever league happened to sit in league_info.json on that container -- a PRIVATE league's scoring and roster shape served to anonymous callers, and to crowd_seed.seed_from_engine, which seeds the VibeRank BASE PLAYER POOL. Fixing (a) alone is not enough and the measurement is why: honestly resolved Belichick is redraft, and a redraft base board carries 751 players against dynasty's 837 -- the crowd pool would silently lose 86 dynasty-only names. New schema-registered default_board.source selects an EXPLICIT reference: 'house_basis' (default) = LeagueConfig.dynasdeez_nutz(), the basis rookie_empirical.py's baselines are denominated in, so INV-30's rookie format-scale is a no-op by construction and no rookie is silently rescaled on the anonymous board; 'league_info' keeps the legacy path reachable, not recommended. League-SCOPED requests are untouched. Verified: crowd base pool back to 837, framework dynasty_ltv, no warnings.
TESTS. tests/test_default_board_resolution.py (7) pins settings.type precedence, the legacy metadata fallback, that default_board is schema-registered and settable, that it defaults to the house basis, and that the house basis is the widest format and the rookie calibration basis. tests/test_status_coherence.py gains two: status multipliers must be measured against a control (pins the corrected range so nobody re-fits on raw retention), and the stale ladder must stay scoped to the state that was measured.
NEW HARNESS scripts/validate_discount_stack.py: measures the three observable cells (neither / FA only / FA+stale) against a matched control and compares each to the multiplier the live config actually applies, so a compounding error is visible instead of inferred.
ROOT CAUSE (the Waller regression). The FA CREDIT read sleeper_players_cache.json::team while the signing DEBITS (MOVE-1 mover volume, redraft team_context) and the displayed team read players.csv::latest_team. Two files, two jobs, two clocks -- so between them the engine held two contradictory beliefs about one player. Measured on Darren Waller, positional TE rank, four league formats: sleeper=FA/csv=FA TE152 (true free agent); sleeper=CAR/csv=FA TE65 (the BEST row on the board, still labelled 'FA'); sleeper=CAR/csv=CAR TE72 (correct); sleeper=FA/csv=CAR TE162 (WORSE than being a free agent). Ryan saw a player sign with Carolina and drop, and asked how a player with no team could outrank one with a team. The answer: he was not being priced as a free agent -- the credit had already lifted while the label had not.
NEW player_status.py. One load per board build (memoised on cache mtime), one fallback rule (live fetch -> cache file -> unavailable), one freshness policy per signal class, one snapshot-mode rule, a fingerprint over the resolved state, and LOUD degradation. Consumers rewired: dynasty fa_status_discount, dynasty buried_depth_discount, the QB depth-chart gate, availability_layer live injury (both engines), redraft FA + depth maps, metrics_xfp mover predicate, team_context. VERIFIED behaviour-identical on a fresh cache: 0 of 837 dynasty rows differ from the pre-refactor board.
CLOSES THE REDRAFT SNAPSHOT-GUARD GAP. redraft_engine's FA and buried-depth paths had NO FF_SNAPSHOT_AS_OF_SEASON check while scripts/validate_redraft_board_outcome.py sets it -- so every redraft walk-forward board was built with 2026 free agency and 2026 depth charts applied to a historical season, and the redraft FA-discount SHIP verdict was measured on those boards. Same class as G-385/G-387; the 2026-08-11 finding flagged redraft as 'worth the same check' and it was never done. The resolver returns empty under snapshot mode, closing both at once. That verdict now needs a re-run.
FRESHNESS + FALLBACK PARITY. Three consumers of one cache had three policies: fa_status_discount fell back to the cache file and applied with NO age check; buried_depth refused past 48h; the QB depth gate had NEITHER a fallback NOR an age check and silently skipped an ~85% availability cap on ~75 QBs on any load error (observed live: 'Depth-chart gate: sleeper load failed (ProxyError) - skipped' while the FA discount kept firing on a 74-hour-old cache in the same run). live_injury had neither either. All four now share status_resolution: team never expires (season-stable), depth and injury expire at 48h, and every skip is recorded on board meta rather than lost to stdout.
team_context DEFECT-1 fix. It used `old != new` rather than is_real_team_change, and team_norm('FA') returns the STRING 'FA' -- notna() and != 'KC' -- so an unsigned player was flagged team_change=True with new_team='FA', which is not a key in _offense_scores and therefore scored the neutral 0.50 default: a better offensive context than roughly 25 real teams. Waller measured +1.0% while unsigned and -4.1% once he signed with Carolina (CAR offense 0.247). Being a free agent was worth more than having a team. Board delta: 192 -> 125 team-changers, 269 redraft ranks move, only 2 inside a positional top-24; every large mover is an unsigned veteran losing a boost he was never owed.
team_context and status_resolution REGISTERED in engine_config.json + engine_config.schema.json. team_context had no config block and no schema entry under additionalProperties:false, so context_adj_alpha=0.20, max_adj_pct=0.10 and opp_signal_discount_changer=0.50 were UNSETTABLE and pinned at hardcoded defaults -- the same read_unschemad class the 2026-06-06 sweep caught for committee_detection.weighted_ppg_factor and missed here (its receiver names are `raw`/`tc`, outside that sweep's cfg/config/params match). Registered at the code defaults, so behaviour-identical. They remain UNVALIDATED and are surfaced to users as the 'Team context (offense change)' chip.
AGE-GRADED DYNASTY FA DISCOUNT (SHIP). The flat 0.50 shipped v1.6.52 with no harness, no cohort and no CI, under a retroactive version-history stub, justified only by analogy to redraft. NEW scripts/validate_dynasty_fa_discount.py measures realized 3-season discounted retention (the quantity discount_factor actually predicts) on the same FA-proxy cohort validate_redraft_fa_discount uses, with an UNCONDITIONAL target -- a season not played scores 0, not NaN, which is the discipline SPEC_stack1 found validate_acute_injury_taper had violated. n=1401 pairs / 1058 players, 2000-2023, 2026 hard-refused. Realized: RB 30+ 0.195 [0.146,0.248], WR 30+ 0.199 [0.163,0.235], TE 30+ 0.253 [0.191,0.320]; every under-26 cell contains 0.50. Binding verdict via eval_lib.verdict_block: SHIP, dMAE -0.0652 CI95 [-0.0730,-0.0572]; out-of-sample (derive y1<=2014, score y1>2014) SHIP, dMAE -0.0542 CI95 [-0.0672,-0.0416]. This CORRECTS audit A10: its '~0.7 under 26' is NOT supported, its '~0.15 over 30' essentially is. Live board delta: 341 ranks move, mean 9.1, max 78, ZERO positional top-24 disturbed; the movers are aged unsigned veterans (Keenan Allen, Tyreek Hill, DeAndre Hopkins, Amari Cooper). 70 of the 126 discounted players were 30+, i.e. the majority of the cohort was over-valued by roughly 2.5x.
QB IS DELIBERATELY EXCLUDED from the graded schedule and keeps the flat 0.50. Its out-of-sample slice CI crosses zero (-0.0148 [-0.0593,+0.0293]) -- a TIE -- and this repo does not ship QB on a proxy after three prior QB false-SHIPs. QB re-enters when the 2026 cohort plays out.
PROVENANCE ON THE ROW. is_nfl_fa, is_buried_depth, fa_status_factor, mover_volume_factor and news_year_0_factor now reach both payloads, and board meta carries a player_status block (source, as_of, age_hours, signals_usable, fingerprint, warnings). Diagnosing the original regression took a day because projection_observed reads 10.02 both in the correct signed state and in the desync state that ranked him 90 spots lower.
TESTS. NEW tests/test_status_coherence.py (8) encodes Ryan's law -- signing must never lower a ranking below the unsigned value, in any format -- plus the freshness split, the loud-degradation contract and the fingerprint property. tests/test_g385 and tests/test_g387 re-pointed from source-inspecting nine inline guard copies to pinning the single seam, and strengthened with a test that every live consumer reads the same snapshot. 149 root tests pass; tests/ green except the g182/g217 freshness-fingerprint family, which is the expected consequence of an engine-source change and needs the re-stamp ritual.
KEEPER HORIZON ENABLED (keeper_ratio_horizon, G-392). Ryan was asked for a keeper horizon in years and answered with the rule instead: 'it depends on the percent of keepers to your total roster -- the lower the percent the more like redraft it is and closer to year 0, the higher the percent the closer to dynasty it is.' That is verbatim the curve the block already implemented, effective = round(1 + ratio * (horizon-1)) on ratio = max_keepers/total_roster_spots, so the gap was never the formula -- it was that no A/B could grade it, and validate_keeper_horizon_relevance.py says in its own words that none can: 'That is a definition-of-value question and no A/B can grade it.' The owner has now defined the value. WHAT IT FIXES: keeper routes to RankingsEngine, so a keeper league was ranked over the full 7-season dynasty horizon where year 0 is only 21.3% of ltv_discounted (horizon 7, discount_rate 0.863). Every year-0 signal was diluted ~2.5x -- injuries, the FA discount, the news year-0 factor. Measured before: Kittle, Charbonnet and Pearsall moved IDENTICALLY in keeper and dynasty, to the rank, on the same designations. Measured after at 3-of-16 keepers (ratio 0.188 -> 2-year horizon): live-injury retention Kittle 63% -> 33%, Charbonnet 78% -> 54%, Pearsall 77% -> 45%; Waller TE72 -> TE56 (a 33-year-old TE is worth more when the horizon stops asking about years he will not play). Curve verified monotone: 1-of-16 -> 1yr, 3-of-16 -> 2yr, 8-of-16 -> 4yr, 12-of-16 -> 6yr, 16-of-16 -> 7yr (no-op). An explicit keeper_horizon_years on the league still wins. Dynasty and redraft boards are untouched -- the derivation only runs on keeper format.
KEEPER RATIO PLUMBED END TO END (2026-08-14). Enabling keeper_ratio_horizon was not enough: keeper_ratio = max_keepers / total_roster_spots is a RATIO computed per league, but max_keepers reached a LeagueConfig on exactly ONE construction path -- from_league_info_json, which reads a local file. from_platform_api (the documented primary consumer entry point, 'user enters a league ID -> this method auto-configures the engine') never set it, from_manual_input had no parameter for it, the leagues table had no column for it, and the DB-row reconstructor therefore could not restore it. The derivation refuses rather than guessing when max_keepers is absent, so the feature would have worked on Ryan's local file and for no real user, silently leaving every connected keeper league on the 7-season dynasty horizon. Plumbed: platform_connectors Sleeper adapter reads settings.max_keepers; LeagueConfig.from_manual_input takes and stores it; leagues.max_keepers column + idempotent ALTER; upsert_league persists it with COALESCE so a connect that omits it cannot clobber a stored value; routers/leagues.connect_league sends it; league_config_service restores it on cold-cache rebuild. Verified end to end across seven league shapes, DB round-trip included: 1-of-16 ratio 0.062 -> 1yr, 3-of-16 0.188 -> 2yr, 4-of-10 0.400 -> 3yr, 8-of-16 0.500 -> 4yr, 12-of-16 0.750 -> 6yr, 4-of-25 0.160 -> 2yr, 16-of-16 1.000 -> 7yr (no-op). Bench size matters and is handled: 4-of-25 lands at 2 years while 4-of-10 lands at 3.
G-308 Buy Window roster context. The recommendation surface now reads the manager's own positional grades before proposing buys, so a position already graded elite is no longer surfaced as a buy target - the Buy Window was proposing five QBs into a QB room graded A. Reuses _positional_grades(); adds a surplus-grade constant and an explicit empty shape. The ranked board is unchanged: this filters the recommendation surface, not the rankings. Version bumped because the recommendation payload's CONTENTS change.
EYEBALL-4 final polish pass (Ryan-directed). F1: redraft_rookie_model.draft_picks_2026 rebuilt from rookie_prospects.csv draft data -- the hand-populated list had 15 players (zero TEs; the TE list existed only as an empty stray key at the wrong nesting level, schema-required but never read) and missed Kenyon Sadiq (TE R1-16), Makai Lemon (WR R1-20, KTC WR21), Zachariah Branch (WR R3-79), Nicholas Singleton (RB R5-165); three tier errors fixed (KC Concepcion R2/40 -> R1/24, Taylen Green and Kaytron Allen R4 -> R6). Now 72 drafted skill players, names/teams reconciled to players.csv (RDR-9a 0 unmatched). F2: superflex_qb_replacement 20 -> 24 (SF boards were near-clones of 1QB: QB replacement ppg moved only 17.39 -> 16.17; 12-team SF starts ~24 QBs). Zero 1QB/publication impact by construction; QB-rate compression noted to G-35. Companion (non-config): routers/seo.py chrome class fix (st-public-header did not exist in public-pages.css -> unstyled banner), Jayden Daniels bounded reviewed exception (standard coherence, dynasty-side investigation queued G-39), test_factcheck_blog unique-surname fixture Bowers -> Egbuka (players.csv now has six Bowers).
DISC-R8-GUARD (Ryan bug-hunt notes): the shrinkage elite skip/taper now require an EARNED rate -- some season among the player's three most recent with >= 6 games at >= 0.75x his current rate. Kills the cameo class the skip was never meant to protect (Bridgewater 29.6 ppg on 3 games -> board QB5; Minshew 23.7 published QB29; Dalton 21.4) while sparing earned elites with a short injury year (Nabers/CMC/Rice/Breece-2022 verified untouched). A/B on the v5 fidelity harness: tune affected dMAE -0.63 CI [-1.34,-0.07], pooled -0.008 CI [-0.017,-0.001]; validation affected -1.62 CI [-2.53,-0.04]; all exclude zero. Companion harness fix RDR-FIDELITY-6: redraft_backtest loader now ingests REG rows only (REG+POST/POST rows entered as duplicate seasons for 94.6% of player-seasons; engine was always REG-only via data_loader). Receipts: data/disc_r8_guard_final_ab.json, data/disc_r8_guard_live_impact.json.
PICK-VALUATION P2 FLIP: draft_pick_model.slot_values rebuilt from the engine's own rookie EV (System C, user-scored) and reconciled to dynasty_ppg via the rookie->LTV age-curve path (pick_value_rookie_ltv; parity-proven == on-board rookie dynasty_ppg). Replaces the frozen 2018-2022 realized-PPG curve. 1.01 ~18.3->~18.1 (unit preserved); SF QB-premium slots +1-2; deep-4th tail discounted. Cohort 2015-2024. Holdout: 2026 refused. Spec docs/SPEC_pick_valuation_P2_unit_reconciliation_2026-06-17.md.
Walk-forward caches are CONTENT-UNAFFECTED (slot_values feeds the trade/pick engines, not the dynasty rankings build) — re-stamp them to this version with scripts/rebuild_walk_forward_caches.py --rebuild --force (byte-identical content) so the freshness gate passes; no behavior rebuild needed.
DYNASTY H3b buried-depth-chart discount SHIPPED (ports the redraft H3b gate; audit P1 'port the three redraft-only gates to dynasty'). Rostered RB/WR/TE listed at Sleeper depth_chart_order >= 4 (4th-string or deeper) get weighted_ppg x 0.90 (10% cut) before LTV; QB excluded (DISC-Q2 handles QB depth). New buried_depth_discount config block + schema entry (schema entry added up front to avoid the v1.6.110 te_use_v2_ltv dead-flag trap), _apply_buried_depth_discount in dynasty_engine.py wired as step 2g-ter after fa_status_discount.
SNAPSHOT-GUARDED (returns early when FF_SNAPSHOT_AS_OF_SEASON is set): depth_chart_order is a point-in-time live signal, so it never runs in backtests/walk-forward and the caches are byte-identical content (re-stamped to v1.6.111, not behavior-rebuilt). Diverges from the unguarded fa_status_discount sibling on purpose: team affiliation is season-stable, depth-chart order is not.
VALIDATION (live-board OFF/ON, scripts/validate_buried_depth_dynasty.py; NOT ab_harness-gateable since snapshot-guarded): cohort = 185 deep backups (highest rank 122, Parker Washington), 0 discounted players in the top 60 (no startable asset touched), named cases Aiyuk #242->263 + Tank Dell #263->277 discounted as expected. Factor held at 0.90 (vs the 0.85 proposal) because June depth charts are offseason-provisional and a few high-capital ascending sophomores (Troy Franklin, Jack Bech, Kayshon Boutte whom the engine's own YPRR tier rates elite) sit in the cohort.
SCOPE NOTE: the other two 'redraft-only gates' from the backlog were NOT ported — H7 missed-year-0 is already covered (and stronger, 0.50 vs 0.20) by _apply_stale_player_decay; DISC-R4 young-rushing-QB is a QB lever deferred to the Sept 2026 QB pass.
TE LTV formula: te_use_v2_ltv -> false (TE switched from v2 relative-factor LTV back to v1 absolute). Audit A4/A11 + DIAG-1 pt3 found dynasty LTV lost to raw weighted_ppg 4/4 ref years at TE, the position running v2 unconditionally; v2's relative-factor flattens the vet-heavy TE field.
ACCEPTED TRADEOFF: ab_harness has no named-case guard; v2's only virtue was protecting ascending elite TEs (Bowers 15.31->13.87, Kelce class), which v1 demotes. Shipped per the field-accuracy gain (Ryan's call).
PREREQUISITE FIX: te_use_v2_ltv was read in age_curves._ltv_row but ABSENT from engine_config.schema.json; under additionalProperties:false any attempt to set it invalidated the config and _ltv_row's try/except silently fell back to the hardcoded v2 default — the gate was unsettable. Added to schema this session; the ab_harness sentinel caught it (zero rows differed until the schema fix). Setting it false is the behavior change; walk-forward caches REBUILT (TE ltv_discounted changes), hence this version bump.
PHASE 4 MODIFIER RETIREMENTS (routing memo final two; verdicts from scripts/validate_modifier_routing.py, the eval_lib-gated re-test built this session — both TIE, route-only-on-SHIP rule applies; 2 rows in docs/experiments_ledger.jsonl):
trajectory_momentum RETIRED via enabled=false, knobs kept: ltv-routing re-test TIE (headline Δrank-MAE -0.036 CI95 [-0.082, +0.011] on unconditional multi-year realized total points, touched slice -0.31, 4/7 ref years favorable — direction positive but CI crosses zero). Re-testable at a future recal; production_delta stays live as a descriptive signal (recommendations narrative, explain_engine MOMENTUM).
contract_year BOOST retired, FLAG kept: re-test TIE with the touched slice wrong-direction (+0.16 Δrank-MAE on only 49 touched top-N rows across ref 2018-2024). _apply_contract_year_modifier now flag-only (contract_year_boost stamped 0.0 for payload compat); boost_pct/_boost_note removed from config+schema; ltv_explainer + app.js LAYER_INFO boost entries retired; app.js signal chip reads the flag. Extension-event machinery (news layer Class B) stays — flag + extension are display-layer together per memo §2.
Both modifiers multiplied dynasty_ppg only (dead on production rank, DIAG-3), so walk-forward RANKS are unchanged by construction — but dynasty_ppg display values move where the boosts fired, hence cache REBUILD (not re-stamp) + this version bump.
DEAD-MODIFIER RETIREMENT (modifier routing memo, approved by Ryan 2026-06-05; audit A5 + DIAG-3). Four dynasty_ppg-only modifiers that never reached production rank (rank sorts on ltv_discounted; pre-retirement kill-test: scheme_fingerprint OFF -> ref2024 rebuild -> RANK DIFFS: 0) removed as behavior-preserving cleanup:
scheme_fingerprint RETIRED OUTRIGHT: _apply_scheme_fingerprint_modifier deleted, config block + schema removed, scheme_modifier payload field + ltv_explainer entry + app.js LAYER_INFO entry removed. Descriptive scheme_label/scheme_pass_rate columns from metrics.py survive. Audit A9: stale-equilibrium signal; vacated-opportunity feature is the planned replacement.
snap_opportunity BOOST retired, FLAG kept: _apply_snap_opportunity_boost deleted; suppression_boost_pct removed from config+schema. snap_opportunity_flag (metrics.py) unchanged on payloads/UI.
qb_dependency BOOST retired, FLAG kept: poor-CPOE suppression boost (Step 3) deleted from _apply_qb_dependency_flag; cpoe_poor_threshold/suppression_boost/min_targets_for_boost removed from config+schema. qb_dependency_risk flag + ltv_p25 widening remain live.
yprr modifier retired, tier label kept: _apply_yprr_modifier renamed _classify_yprr_tier (descriptive elite/below_avg only); elite_boost_pct/below_boost_pct removed from config+schema. Re-introduction as engine math requires re-SHIP on ltv rank after the harness target re-alignment.
young_qb_ceiling DEFERRED to Sept 2026 QB pass; trajectory_momentum + contract_year re-tests gated on harness target re-alignment (memo sequencing rule). No engine math changed on any live path — walk-forward ranks byte-identical by construction.
GHOST-PLAYER HARD FILTER (engine audit 2026-06-05 A10 + DIAG-3 confirmed bug; docs/ENGINE_AUDIT_2026-06-05.md). New ghost_player_filter config block + dynasty_engine._apply_ghost_player_filter (step 2g-pre, before stale decay): drops any non-rookie whose last played season is >= max_seasons_missing (2) behind the dataset's latest season. Root cause: the walk-forward runner injects ALL stats <= ref_year, so decades-retired players (Aikman/Marino/McNair/Couch/Harrington at ref2024 QB ranks 43-53, LTV 17-21) entered the pool and the stacked soft discounts (stale 0.5 x FA 0.5 = 0.25) could not remove them. Live path was already protected by load_active_players thresholds (code-verified) - this is primarily an eval-universe + FA-pool correctness fix. Deliberately NOT snapshot-guarded: walk-forward caches CHANGE and must be REBUILT (--rebuild --force), not re-stamped. Division of labor preserved: 1 missing season = stale_player_decay 0.5x; 2+ missing seasons = dropped.
A8 DUPLICATE-DEF FIX (monte_carlo.py): deleted the dead first compute_player_ltv_bands definition (~line 1178, config-driven model) that was shadowed by the live second definition (~line 1397, hand-set constants). No behavior change - the live model is unchanged. player_ltv_bands config _note updated to state honestly that the block is NOT read by the live band model (constants hardcoded in monte_carlo.py); wiring config -> live model deferred to the D9 band-coverage calibration, which will refit the constants anyway. Unblocks the D9 coverage script (it now measures the model production actually serves).
2026 DECLARED A LOCKED HOLDOUT (audit D2): no 2026 stat rows may feed any calibration/fit/tuning script. 2025 is burnt (rows already feed calibration). Declaration + seven stale-doc corrections landed in CLAUDE.md.
NEWS/SITUATION LAYER PHASE 2 APPLY + Class B math + arbitrage-surface news wiring (spec sections 4 + 5b; the engine now ACTS on real-world events, not just receives them).
(1) TEAM-CHANGE MULTIPLIER LIVE: dynasty_engine._apply_news_situation_adjustment (step 2h-bis) applies the calibrated team_change_deltas.json multiplier to players carrying an active trade/signing ledger event. Year-0-ONLY via the DISC-Q4 channel: projected_season_ppg + LTV horizon year 0 (news_year_0_factor, threaded in age_curves + the age-aware proj recompute, NaN-safe); years 1-6 stay at baseline because the calibration measured next-season PPG. Role tier = weighted_ppg rank within position (QB 12/24, RB 24/48, WR 36/72, TE 12/24 = calibration cells). Decay-to-1.0 linearly as post-event-season games accrue (decay_games=12) + hard apply_days=330 window. Confidence downgraded one tier while active. Rookies skipped; snapshot-guarded (FF_SNAPSHOT_AS_OF_SEASON) so backtests/walk-forward stay byte-identical -> caches re-stamped only.
(2) TWO-STAGE CELL GATING: calibration reliability was necessary but NOT sufficient. New scripts/validate_team_change_apply.py (mover-cohort walk-forward gate: applies multipliers to REAL engine projections ref 2018-2024, scores vs realized outcomes, paired bootstrap CI + A.J. Brown named-case). All-4-reliable-cells run: TIE (dMAE -0.044, CI crosses; RB.starter HURTS +1.02 MAE - the committee/stale-decay/RB-cliff machinery already prices RB moves, the raw multiplier double-counts; WR.starter TIE +0.10). SHIPPED CONFIG = cells allowlist [QB.mid, WR.mid]: n=38, MAE 3.832 -> 3.371, dMAE -0.461, CI95 [-0.844, -0.060] all-negative, bias +1.33 -> -0.07, A.J. Brown WR7 -> WR7 PASS (his WR.starter cell excluded - marquee consensus untouched) => VERDICT: SHIP.
(3) SUSPENSION MATH (Class B, locked): games-available arithmetic, factor = (17 - games)/17 on the same year-0 channel. Games parsed from event severity ('6') or note ('6 games'); unparseable -> flag-only. Skips players Sleeper already marks suspended (no double-count). Arithmetic, not modeling - no calibration gate.
(4) CONTRACT-EXTENSION RESET (Class B, locked): new 'extension' event type; _apply_contract_year_modifier suppresses the draft-round contract-year boost for players with an extension event within extension_suppress_years=3 (Drake London case - the heuristic would boost a player who just got paid). Snapshot-guarded; news_events.extended_player_ids.
(5) ARBITRAGE-SURFACE NEWS WIRING (the accepted Phase 0/1 gap, closed): news_events.arbitrage_suppression_flags (never-raise) consumed by arbitrage_engine.compute_signals (direction -> news_flagged + news payload; recs Buy-Low/Sell-High + Sell/Buy Windows inherit + new news_suppressed list), trade_simulator._build_hype_index (flagged players excluded from hype -> neutral tilt in Trade Finder + Trade Central match finder + ARBITRAGE ANGLES), market_inefficiency_engine.detect (consensus_direction -> news_flagged + NEWS note; undervalued/overvalued filters drop them). 'extension' events do NOT suppress (informational, not repricing).
(6) scheduler.run_daily_news_detect: trade/signing events now wipe + re-warm the rankings cache when team_change.enabled (board math moves same-morning, not at next organic rebuild).
News/situation layer Phase 1 (docs/SPEC_news_situation_layer_2026-06-04.md): the engine now RECEIVES real-world events. (1) news_detect.py diffs the Sleeper player feed daily against a stored snapshot (news_detect_state) and writes trade/signing/released/retirement events into the player_news_events ledger (source=sleeper_diff; engine-universe filtered; 3-day dedupe; cold-start snapshots silently). Detection rules calibrated on the live A.J. Brown trade + Russell Wilson retirement: team change -> trade/signing/released; retirement = active flag dropping on a team-less player (Sleeper does not mark retirement at event time, so manual admin events remain the fast path).
(2) RETIREMENT EXCLUSION in data_loader.load_active_players: players with a retirement event are dropped from the LIVE engine pool (dynasty + redraft boards via load_active_stats). Guarded: FF_SNAPSHOT_AS_OF_SEASON set (backtests/walk-forward) -> exclusion skipped, historical output byte-identical; ledger failure -> pool unfiltered. Un-retire = DELETE the ledger event.
(3) scheduler.run_daily_news_detect (03:15 ET daily + JOB_REGISTRY manual trigger): runs the diff; any retirement event wipes + re-warms the rankings cache so the exclusion bites the same morning.
Rankings content changes ONLY when a retirement event exists (today: Russell Wilson drops from live boards). Backtest/walk-forward output unchanged by the snapshot guard -> caches re-stamped to v1.6.103. Phase 2 (calibrated team-change multipliers) staged separately. Files: news_detect.py (new), news_events.py, data_loader.py, scheduler.py, engine_config.json.
MARKET ARBITRAGE RELIT on the first-party VibeRank crowd source. market_layer.enabled=true with NEW primary_source='crowd': market_layer.fetch() now reads the per-format VibeRank crowd board (crowd_store.resolve_fmt joins the league's dynasty/superflex/te_premium flags to its board) instead of third-party data — exact player_id reconciliation, 0..9999 crowd values, ATTRIBUTION_CROWD. dynasty_engine._load_market_values() gains the same crowd branch (rankings-row arbitrage_score/signal + league-card edge callout relight); legacy ktc-gated DB path retained but dark (ktc.enabled stays false; do not flip — its market_values table holds stale third-party rows).
Rank-based Signal Gap calibration (crowd_store.get_rankings): `signal` is now derived from gap_ranks = crowd_rank - engine_rank vs an adaptive threshold (max(4, 15% of the better rank); 2x = gap_big), replacing the value-scale subtraction that inflated mid-board gaps (market_seed is linear-in-rank, engine_score follows the value curve). Value-scale signal_gap stays on the payload as a diagnostic.
NEW: My Board (PureSignal, login-required) — GET /v1/crowd/my-board replays a user's logged crowd_votes through a fresh per-user aggregator (market_seed priors, full weight, lighter k=2 display shrink) and returns their personal board with vs_crowd / vs_engine rank diffs within the voted subset. /v1/crowd/* otherwise stays public.
Engine math (LTV / weighted_ppg / ordering) untouched — market layer remains diagnostic-only, so rankings CONTENT is unchanged; walk-forward caches re-stamped to v1.6.102 to satisfy the freshness gate. Version bump required so the config edit survives the start.sh image-vs-volume tie (v1.6.40/84/91/98 class). Files: crowd_store.py, crowd_rankings.py (none), routers/crowd.py, market_layer.py, dynasty_engine.py, routers/trade.py, routers/recommendations.py, static/crowd.html, engine_config.json, engine_config.schema.json.
STAGED (flag-OFF) the QB shrinkage elite-skip — primary lever for the +2.50 PPG QB backtest bias. Ported the redraft DISC-R8 elite-skip to dynasty: metrics.apply_shrinkage now accepts config.shrinkage.elite_skip {enabled:false, positions:[QB], elite_mult:1.0} and, when ON, skips Bayesian shrinkage for a thin-sample player whose weighted_ppg > elite_mult x positional prior (a demonstrated above-prior producer is durable, not a one-year wonder dragged toward a sub-replacement prior). Root cause per docs/SPEC_qb_shrinkage_fix_2026-06-02.md: the dominant of two shrinkage ops; young producing QBs (Daniels class) are the affected set. FLAG STAYS OFF: the candidate's QB CI crosses zero on the thin cohort (n~33), so ship only after 2026 actuals tighten it AND a clean full-engine backtest clears (3 prior QB attempts false-SHIP'd on proxies). Flag-OFF = no ranking change; walk-forward caches re-stamp identical (the proof). Files: metrics.py, engine_config.json, engine_config.schema.json.
Rookie calibration refreshed to the 2015-2024 cohort (was 2015-2023; refresh_rookie_baselines.py --cohort-end 2024) — folds 2024 draft-class outcomes in via 2025 season data; regenerates rookie_pick_bucket_baselines/combine/college deltas + historical_rookie_outcomes.parquet. Walk-forward caches rebuilt (2018-2024 ref years); full-engine rookie accuracy healthy (r 0.64-0.80, |bias_peak|<0.55, no regression).
Breakout-age trajectory layer REACTIVATED for live classes. breakout_age + senior_year_decline were 0% populated for 2024-2026 (static CSV column, no derivation pipeline); built scripts/derive_breakout_age.py (multi-season cfbfastR dominator -> age at first >=0.20 season; RB yards_per_team_play>=1.2; validated 95%+ vs historical column) + wired into refresh_rookies_post_draft.py; backfilled 2024-2026 to ~93%.
Trajectory betas RECALIBRATED (rookie_empirical._apply_beta_shrinkage_trajectory). The layer was the ONLY rookie delta layer not shrinking per-bucket betas toward the position fallback (combine + college both do), so small-n cells ran at full strength and produced board-breaking swings (WR R3 beta_senior_decline=+2.0, wrong-signed, drove +5 ppg onto low-production WRs). Fix: breakout_age now uses the robust NEGATIVE position fallback for all buckets (WR -0.52 / RB -0.31 / TE -0.50); senior_year_decline disabled (its fallbacks were wrong-signed). Surfaced by scripts/breakout_age_board_diff.py before ship.
SHIP gate = BIAS gate (scripts/ab_trajectory_breakout.py, paired ON-vs-OFF projection vs realized peak, cohort 2015-2024 n=572). MAE a well-powered TIE (agg dMAE -0.001, CI [-0.027,+0.027]). Breakout age reliably DE-BIASES rookie under-projection, concentrated in WR (OFF bias +0.576 [+0.099,+1.059] -> ON +0.409; paired shift CI [-0.21,-0.12] RELIABLE). RB/TE neutral. A centering win, not a discrimination win.
LTV discount_rate config corrected 0.863 -> 0.92 (documentation alignment, NO ranking change). DEAD KNOB diagnosed 2026-06-01: enrich_age_curves omits discount_rate from _ltv_kwargs, so the engine uses the module default age_curves.DISCOUNT_RATE=0.92 and NEVER the config 0.863 (proven by config sweep giving identical output + code trace). Reset config to 0.92 to match reality + added _discount_rate_note marking it dormant. Walk-forward caches re-stamp IDENTICAL (dead knob -> rankings provably unchanged; the identical rebuild is the proof). Version bump only so the config edit survives the start.sh image-vs-volume tie.
Outcome A/B (scripts/validate_discount_rate_change.py): the discount rate is a WEAK lever for RB-vs-WR ordering (flat at every rate 0.84-0.99) and the overall-Spearman signal that favored gentler rates is confounded by an undiscounted realized target (even perfect projections would show it), so the rate is not cleanly calibratable from outcomes. 0.92 retained; RB-vs-WR-via-discount hypothesis CLOSED. Open follow-up: wire the knob if a deliberate rate change is ever desired (needs full SHIP).
Market arbitrage PULLED from the platform: market_layer.enabled set false (ktc.enabled already false). Both market-data entry points now dark — the rankings/recs path (ktc-gated) and the trade-analyzer + market_inefficiency path (market_layer-gated, previously live on stale FantasyPros static CSVs). All downstream UI surfaces hide gracefully.
No engine math change: market consensus was always diagnostic-only (never modified LTV / weighted_ppg / rankings ordering). Walk-forward cache CONTENT is unchanged; caches re-stamped to v1.6.98 only to satisfy the freshness gate. Version bump required so the engine_config.json edit survives the start.sh image-vs-volume tie (same class as v1.6.40/84/91).
Rationale: the only IP-clean reachable market source (FantasyPros) is stale expert opinion, not true crowd sentiment; Sleeper has no public ADP endpoint; true crowd trade-value data is ToS-blocked (KTC/FantasyCalc class). Pivot to building our own crowd-sourced signal — see BACKLOG P1. Scaffolding kept dormant for relight.
value_over_replacement_ranking ENABLED (QB over-valuation fix). Cross-position Overall sort now blends raw-LTV pct with value-over-replacement pct (dynasty_vos scaled to lifetime), applied to rookies too. Format-conditional weight: 1QB vor_weight=0.7, SF vor_weight_superflex=0.3.
Validated by scripts/validate_vor_ranking.py (walk-forward Spearman vs realized value-over-replacement, 2018-2023): 1QB +0.0213 at w=0.7 (clears +0.010 bar; every weight beats pure-LTV, monotonic to +0.0348 at pure VOR). SF a wash (+0.0067 best at w=0.3); SF set to the mild 0.3 to demote replacement-level rookie QBs while preserving elite QBs (Allen #1, Lamar #6).
Effect: 1QB board now RB/WR-led (Gibbs/Bijan/Achane/Chase/Nacua top 5); Ty Simpson rookie #2 -> overall #195 (1QB) / #50 (SF). ltv_discounted magnitude unchanged; only the Overall ordering (blended_rank_score) changes.
Item A (PlayerProfiler Benchmark Round 3): rookie-year injury victims (acute_single_event, no prior healthy season) were misclassified chronic in the injury_discount path because _enrich_injury did not pass age_lookup into build_injury_index, so current_age was None and the acute branch (age<=25) never fired. Fixed: _enrich_injury now builds and passes age_lookup (parity with _apply_structural_ceiling).
Two-branch EMPIRICAL acute discount (both validated by scripts/validate_acute_injury_taper.py against realized forward availability; the 0.955 recovery curve over-projects BOTH cohorts and is retired for them, kept only as null-fallback). UNPROVEN (no prior healthy season, --cohort unproven n=28): realizes ~0.50; ship unproven_rookie_discount=0.50 (MAE 0.203 vs chronic 0.359, bias +0.004 CI contains 0). PROVEN (>=1 healthy season, --cohort proven n=27): realizes ~0.586; ship proven_recovery_discount=0.65 (MAE 0.234 vs chronic 0.431, bias +0.064 CI contains 0). Chronic under-projects both (bias -0.32/-0.40, CI excludes 0); 0.955 over-projects both (+0.46/+0.37).
Walk-forward cache rebuild REQUIRED before deploy (injury_discount affects dynasty rankings). Run scripts/rebuild_walk_forward_caches.py --rebuild --force locally and recommit the 7 ref-year JSONs.
Session 157 ext bug fix: keeper_horizon_years override was being silently swallowed for young players because age_horizon_caps cap=7 (default-horizon no-op marker) was binding via min(base_horizon, cap). Browser-validated symptom: setting Unlimited renewal (10) on Mystic produced byte-identical LTVs because every QB under age 32 had QB cap=7 from the smallest-key-ge-age lookup, and min(10, 7) clamped horizon to 7. The override silently failed.
Fix in age_curves.py::enrich_age_curves compose-effective-horizon logic: cap binds only when cap < HORIZON (real aged-player truncation). When cap >= HORIZON the cap value is a no-op marker for young players and base_horizon passes through. Aged-player caps (e.g., QB age 38 -> cap 4) still bind even under high base_horizon (Goff stays capped at 4 years regardless of league keeper window).
horizon=3 already worked correctly pre-fix because cap=7 > base=3, and min(3, 7) = 3 (base wins). horizon=10 was the actual broken case: cap=7 < base=10, min(10, 7) = 7. The fix preserves correct horizon=3 behavior and unblocks horizon=10.
Engine math change: 1-line condition added in age_curves.py compose-horizon block. _meta.version v1.6.94 -> v1.6.95. Walk-forward cache rebuild NOT required (same reason as v1.6.94 ship: baselines are format=dynasty, override path inert).
Session 157 ext (BACKLOG Item 10 close): per-league LTV horizon override for keeper-format leagues. New LeagueConfig.keeper_horizon_years field (Optional[int], None default) drives age_curves.enrich_age_curves(base_horizon=N) when set. Sleeper API does not expose this field; user enters their league's effective keeper window via the new Keeper rules modal in the UI.
Architectural rationale: Mystic and other keeper leagues had Rankings tab running dynasty 7-year LTV math regardless of actual keeper window. For unlimited-renewal keeper leagues (Mystic-style: pick N keepers per year, no max term) the 7-year approximation is fine. For finite-window keeper leagues (e.g. 2-year max), 7-year LTV materially over-projects tail production for aging veterans. The user-input field bridges the spectrum.
UI dropdown options: 1-9 years (literal max keeper window) + 'Unlimited renewal' (stores 10 internally, captures ~85% of true-infinity effect without going far past calibration regime). Default empty = engine default of 7. Round-trip: stored value pre-populates the dropdown on next open.
Engine math: when config.format=='keeper' and keeper_horizon_years is set, _apply_framework passes base_horizon kwarg to enrich_age_curves. age_horizon_caps still bind for aged players (effective_horizon = min(base_horizon, age_cap)). Goff at age 31 still gets capped to 7 by the 32-key cap; Maye at age 24 gets the full 10. No change to any other engine layer (Bayesian shrinkage, archetype mods, breakout lifts, scarcity multipliers).
API surface: POST /v1/leagues/{league_id}/keeper-settings accepts {keeper_horizon_years: 1-12 or null}, validates, updates DB, recomputes config_fp, pops the stale cache entry, and reruns rankings under the new horizon. 422 on non-keeper-format leagues (since the field has no semantic meaning for dynasty/redraft).
DB: keeper_horizon_years INTEGER NULL column added to leagues table via boot migration (idempotent ALTER). _config_fingerprint includes the field so a flip invalidates the rankings cache. _H2_DRIFT_FIELDS includes the field for the H2 drift detection on roster-sync.
Filed follow-up (deferred): paired-observation harness for horizon calibration. The 7-year base value was set early in the engine's history without strict empirical validation against long-horizon outcomes. Building one requires 6+ years of held-out actuals per player (i.e. 2018 cohort tested through 2024). Not blocking; the override gives users the lever immediately, math underneath remains the validated dynasty pipeline.
Walk-forward cache rebuild: NOT required for this ship. The 7 historical caches are stamped at the engine version and produced against Dynasdeez baseline (format=dynasty). The keeper_horizon override only fires when config.format=='keeper' AND keeper_horizon_years is set; baseline cache generation skips both gates. Math byte-identical for the cache rebuild path.
Session 157 (PP-Round-2 Item 5 ship): surgical bust-filter applied to rookie_pick_bucket_baselines.json on 8 cells: RB.R4_5, RB.R6_7_UDFA, WR.R3, WR.R4_5, WR.R6_7_UDFA, TE.R3, TE.R4_5, TE.R6_7_UDFA. Closes the late-round-bucket-median pollution class surfaced by the PP-Round-2 benchmark (engine ranked 5 named UDFA/Day-3 RBs (LeQuint Allen, Tahj Brooks, Kaytron Allen, Demond Claiborne, Ollie Gordon) 250-330 spots LOWER than PP because R6_7_UDFA RB median was 0.00, dragged to zero by 57.7% bust rate).
Filter definition: drop players with n_seasons_played == 0 (career-bust cohort, never had a season with 8+ NFL games) before computing median/mean/quantiles. bust_rate + n preserved on full cohort so downstream consumers can reconstruct the prior on roster-failure for prospects who haven't yet made an NFL team.
Preserved cells (explicitly NOT updated): RB.R2 + RB.R3 + WR.R2 (walk-forward-validated in Session 111 DISC-Q5-followup); RB.R1_top + RB.R1_late + WR.R1_top + WR.R1_late + WR.R2 + TE.R1_top + TE.R1_late + TE.R2 (career_avg-tuned in v1.6.15/v1.6.16 RESEARCH-11b, not peak_ppg); ALL QB cells (filter collapses QB late-round sample sizes from n=14-30 to n=2-10; remaining cohort is dominated by starter-caliber outliers, current median 0.00 for QB R3+ is defensible).
Combine + college delta JSONs (rookie_combine_deltas.json, rookie_college_deltas.json) unchanged. Those regressions fit (peak_ppg - bucket_median) against z-scored measurables on the full cohort; refitting under the bust filter would alter their semantics (busted players' combine + college data is real signal about what HASN'T predicted success).
Engine script change: scripts/refresh_rookie_baselines.py extended with --exclude-busts-positions CLI flag (default empty). Future annual rebuilds can pass --exclude-busts-positions RB,WR,TE to apply the filter automatically; QB intentionally excluded. compute_baselines() signature extended to accept exclude_busts_for: Optional[set].
Named-case smoke: all 5 PP-Round-2 RBs are R6_7_UDFA bucket. Their projected_peak_ppg lifts by +4.67 ppg (the bucket median shift). Combine + college deltas unchanged. Engine ranks should move 100-200 spots toward PP without becoming consensus-tuned.
Filed follow-up: bust-decomposition refactor (1-2 day structural fix). Store conditional median + bust probability separately per cell; engine layer interpolates based on live roster signal. Cleaner long-term solution than the surgical filter ship. No paired-observation harness exists for rookie projection accuracy yet - validation is via backtest accuracy + named-case sanity check.
Engine math change: bucket_median lift in 8 cells. _meta.version v1.6.92 -> v1.6.93. Walk-forward cache rebuild required on next push (math touches projected_peak_ppg for partial-volume rookies in affected buckets).
Session 157 (PP-Round-2 Item 2 ship): scarcity_elasticity_per_position bumped uniformly 0.4 -> 1.0 across QB/RB/WR/TE. Closes the 1QB QB over-ranking class surfaced by the PP-Round-2 benchmark (Mystic Dynasty top-14: 11 QBs vs PP's 2).
Harness verdict (scripts/validate_league_aware_ltv_decision_quality.py at elasticity 1.0, per_player_scaling=false): SHIP, aggregate Spearman delta = +0.0726 across 5 configs x 6 ref years. SF 10t baseline no-op at +0.0000 (invariant preserved). SF 12t worst at -0.0003 (well within DEFEND tolerance of -0.020). 1QB configs gain +0.116 to +0.127 Spearman (cross-config decision quality, top-150 scope).
Per-player scaling infrastructure shipped Session 157 Item 2 build (engine_config.league_aware_ltv.per_player_scaling block + dynasty_engine._apply_league_scarcity_to_ltv per-player branch + harness extension). Sweep verdict: DEFEND across all elasticity x anchor combinations. pos_top_vos lost ~50% of OFF improvement uniformly. pos_p75_vos was effectively equivalent to OFF (most top-150 players sit above p75 -> share clamps to 1.0). Architectural hypothesis from Item 1 was wrong-headed; harness is truth. Flag stays OFF in production.
Floor at min_multiplier=0.60 is now binding for 1QB QB (raw mult = (10/24)^1.0 = 0.417 clamps to 0.60). Future elasticity tuning beyond 1.0 requires a min_multiplier change. Filed as a follow-up.
Locked discipline: any future change to scarcity_elasticity_per_position OR per_player_scaling must re-pass scripts/validate_league_aware_ltv_decision_quality.py with SHIP verdict on the cross-config decision-quality cohort + SF baseline no-op invariant intact.
Docs-only version bump to force start.sh to pick up the v1.6.62-v1.6.84 version_history backfill committed earlier this session.
Root cause: start.sh's image-vs-volume comparison uses sort -V on _meta.version. When the image and volume share a version string (both v1.6.90), the 'keep volume' branch wins and content-only edits to engine_config.json silently get reverted on every deploy.
Same bug class as v1.6.40 (Session H DEGRADED hotfix) and v1.6.84 (Session 152 schema-validation outage). Engine math change: zero. The bump exists solely to make image > volume so start.sh copies the freshly-committed config into /app/data/.
New config knob engine_config.auction.positional_allocation_exponent (default 0.7) introduced in auction_engine.compute_auction_values. Applies a power-mean transform on the (player_ltv - replacement_ltv) terms in Step 1 (positional budget allocation), compressing steep-distribution positions (RB top-5 spike) and boosting flat-curve ones (QB clusters tightly).
Wired into routers/draft.py POST /v1/auction/values: reads engine_config.auction.positional_allocation_exponent at request time and passes to compute_auction_values; defensively defaults to 1.0 on config read failure.
Sweep harness scripts/validate_auction_allocation_exponent.py confirms directional movement: 1QB top QB price $28->$36 (+30%), 1QB QB allocation 5.2%->6.7%, SF top QB price $41->$44, SF QB allocation 17.5%->19.1%. RB share compresses from 46.1%->41.5% (1QB) and 40.0%->35.8% (SF). WR + TE shift modestly (no position breaks).
Magnitude-bounded: even at aggressive exp=0.5 (sweep extreme), SF QB allocation only reaches 19.8% vs BACKLOG-cited 25-30% target. Structural fix for full closure filed as follow-up - needs real-auction-price discipline data from MFL/FantasyPros to validate a more aggressive default or a complementary mechanism (e.g. qb_sf_bonus restore alongside the exponent).
New config block engine_config.wr_sophomore_upside (default enabled=false). When true, applies a years-to-peak multiplicative bonus on dynasty_ppg for ascending or early_peak WRs in R1_top/R1_late/R2 with 1-3 NFL seasons completed and weighted_ppg >= 8.0. Mirrors young_qb_ceiling architecture.
New method dynasty_engine._apply_wr_sophomore_upside (~115 LOC) wired into rankings pipeline at step 3c-bis (after _apply_young_qb_ceiling, before _apply_trajectory_momentum_bonus).
Schema entry added with all 9 properties + required list.
A/B harness scripts/validate_wr_sophomore_upside.py (320 LOC) - paired-observation MAE on n=110 historical WR Y1-Y3 ascending high-pedigree pairs (2016-2024).
VERDICT at default params (bonus_per_year=0.12, cap=1.30): DEFEND. Agg delta +0.605, CI95 [+0.210, +1.004] all-positive (candidate REGRESSES). R1_top + R2 + Y2 + early_peak slices all break at CI-significant level.
Parameter sweep across 9 (bonus_per_year, cap) variants (0.025-0.12 x 1.08-1.30): every variant TIE or DEFEND. No calibration earns SHIP.
Diagnostic: cohort is heterogeneous - half ascending high-pedigree WRs break out (JSN 22, Jefferson 21, A.J. Brown 20, Jaylen Waddle 22, JuJu 18), half regress to mean or get injured (Odell 16, Allen Robinson 16, Chase 23, Aiyuk 21, DK 21, A.J. Brown 21). The bonus mechanism stamps every eligible WR uniformly, can't distinguish breakouts from busts, and pulls predictions further from actuals on the bust class. Baseline (weighted_ppg x age_curve) is already MAE-tight at 2.40 ppg on this cohort.
Per locked discipline [[feedback_calibration_via_harness_only]] + [[feedback_consensus_is_signal_not_target]]: harness is the truth, engine math is statistically defensible. PlayerProfiler benchmark gap (JSN +43, Nacua +24, Nabers +119, etc.) reflects consensus heuristics that don't predict Y+1 outcomes better than our model.
Flag stays OFF in production. Infrastructure ships flag-gated as a documented 'tested, did not validate' result, mirroring Pattern 4 (v1.6.87) playbook. Future operators can re-test with different mechanism designs - e.g., add a target_share trajectory gate to distinguish role-trajectory ascenders from injury/role-disruption regressors, or use a per-player breakout-probability signal.
New config block engine_config.thin_sample_regression with dedup_with_unified_projection (default now true).
When true, _apply_thin_sample_regression in dynasty_engine.py excludes rows where apply_unified_projection (DISC-1, Session 110) already fired with w_data < 1.0. Eliminates a small-sample double-counting bug class: a 1-season non-rookie with a draft_pick was previously shrunk twice (once toward rookie_empirical bucket prior, then again toward live position median), landing partial-volume Y1 RBs ~30% below their bucket floor.
Hampton named case: weighted_ppg 9.28 -> 11.90 (raw 13.08, prior 11.40 for R1_late pick=22, w_data=0.30); committee discount layer unchanged.
A/B harness scripts/validate_thin_sample_dedup.py SHIP verdict on n=571 historical Y0->Y+1 partial-volume rookies (2014-2023). Agg MAE delta -0.182 CI95 [-0.354, -0.011] all-negative. WR n=230 SHIP-clean CI [-0.398, -0.068]. RB n=153 directionally consistent (-0.184) but CI crosses zero. No position regresses at CI-significant level. Per [[feedback_per_bucket_decomposition]]: agg + at least one position CI-clean negative + no CI-significant positive on any position satisfied.
Locked discipline: any future change to thin_sample_regression.dedup_with_unified_projection must re-pass scripts/validate_thin_sample_dedup.py with SHIP verdict.
Pattern 4 infrastructure (Session 154 ext): added engine_config.json::qb_nfl_viability_confidence block + dynasty_engine._apply_qb_viability_penalty() method + scripts/validate_qb_viability_penalty.py SHIP-gate harness. Default enabled=false; engine math BYTE-IDENTICAL to v1.6.86 until harness clears.
Math design: multiplies ltv_discounted for QBs by piecewise-linear interpolation over career pass attempts. Breakpoints [0, 100, 300, 500, 1000] -> penalties [0.55, 0.65, 0.80, 0.92, 1.00]. QBs with <4 career starts always get penalty_at_breakpoints[0] regardless of attempts. Applied AFTER enrich_age_curves and BEFORE _apply_young_qb_ceiling so the ramp + ceiling bonus stack on top of a dampened LTV.
Target cohort: McCarthy (3 career attempts, our v1.6.86 rank 50, PP 237) -> ~0.55x LTV penalty. Ty Simpson (0 attempts, college rookie) -> 0.55x. Drake Maye (~570 attempts post-Y1) -> ~0.93x, minimal change. Mahomes/Allen/Burrow class (5000+ attempts) -> 1.00x, no change.
SHIP gate: scripts/validate_qb_viability_penalty.py — paired-observation MAE on QB rookies/sophomores with <500 career attempts 2015-2024. Validates penalty closes Y+1 over-projection on McCarthy-class without regressing Mahomes-class.
Session 154 ext: Pattern 1 SHIP. Flipped engine_config.json::league_aware_ltv.enabled from false to true. The per-league scarcity multiplier (built Session 152, math-validated Session 153) now flows into ltv_discounted across the full ranking pipeline.
SHIP gate: scripts/validate_league_aware_ltv_decision_quality.py returned SHIP at default elasticity 0.40 + default scope top-150. Aggregate Spearman delta = +0.0312 across 5 configs x 6 ref years (2018-2023), 3x the +0.010 SHIP threshold. SF 10-team baseline = +0.0000 (no-op by design); SF 12t = +0.0006 (neutral); 1QB 10t = +0.0498; 1QB 12t = +0.0529; 1QB 14t = +0.0525. All three 1QB configs flipped from NEGATIVE Spearman OFF (worse than random vs true VOR) to POSITIVE Spearman ON. Top-50 hit-rate gain for 1QB users: +1.8pp consistently.
Trigger context: 2026-05-27 PlayerProfiler benchmark exposed that our engine moved players a mean of only 10.9 ranks between SF and 1QB configs vs PP’s 21.3 ranks; QBs shifted -4.0 in our engine (slightly UP in 1QB) vs PP’s +27.8 (DOWN in 1QB). Tua/Geno/Mac/Flacco class were rank-flat across configs in our engine, badly mis-ranked in 1QB. Pattern 1 closes most of that gap.
Production math change: 1QB league users will see top QBs drop ~50-80 ranks (multiplier ~0.76x at elasticity 0.40 for 1QB 12-team). SF baseline users see no change. SF 12-team / 14-team users see slight differentiation. TE Premium and scoring-format multipliers unchanged (separate code path).
Walk-forward caches must be regenerated against v1.6.86 because the engine math is changing for non-baseline configs. The harness verdict itself was computed using the cached LTV values + a post-hoc multiplier (no engine re-run needed for the verdict); but the production rankings now flow the multiplier through compute_replacement_from_slots -> _apply_league_scarcity_to_ltv at the dynasty pipeline step 5.
New SHIP-gate harness: scripts/validate_league_aware_ltv_decision_quality.py (471 LOC, supports --sweep-elasticity, --top-spearman-n, --full-spearman, --elasticity overrides). Default invocation reflects the decision-quality SHIP gate; --full-spearman emits the noisier all-player Spearman for debugging.
Session 154: closed the Session 151 P2 architectural fix. engine_config.json::age_curves is now the single source of truth for QB/RB/WR/TE age curves AND the new QB_POCKET / QB_DUAL archetype-split sub-curves.
engine_config.json::age_curves block rewritten with canonical REBUILT_* values that previously lived in age_curves_rebuilt.py. The pre-existing dead config block contained junk values (ages up to 56 for QB, 51 for RB, etc.) carried forward from an earlier auto-generated calibration pass; replaced with the actual production curves.
engine_config.schema.json::age_curves relaxed from rigid age-by-age enumeration with additionalProperties:false to patternProperties on integer-string ages. Junk ages 41/45/47/56 (QB), 42/46/48/50/51 (RB), 46/49/54 (WR), 47/49 (TE) removed from required-list. QB_POCKET + QB_DUAL added as optional position blocks.
age_curves.py refactored: module-level _CURVES, _QB_CURVE, _RB_BASE_CURVE, _WR_CURVE, _TE_CURVE, _QB_POCKET_CURVE, _QB_DUAL_CURVE now sourced from engine_config.json at module import via new _load_curves_from_config() + _load_qb_archetype_curves() helpers. Hardcoded literals preserved as defensive fallback when config load fails. FF_USE_REBUILT_CURVES env var deprecated (no-op).
age_curves_rebuilt.py converted to a thin compatibility shim: REBUILT_QB_CURVE, REBUILT_RB_BASE_CURVE, REBUILT_WR_CURVE, REBUILT_TE_CURVE, REBUILT_QB_POCKET_CURVE, REBUILT_QB_DUAL_CURVE, REBUILT_CURVES all now read from engine_config.json. Preserves backward compatibility for 7 downstream scripts (validate_curve_change, audit_calibrations, validate_archetype_mods, validate_qb_dual_curve_modern, validate_pwopr_lift, wf_age_7_ab_revalidation, recalibrate_age_curves) without touching their imports.
validate_curve_change.py::current_curve() now reads from engine_config.json directly (not via the shim). Closes the Session 151 bug class permanently: harness baseline = production source by construction.
Engine math: BYTE-IDENTICAL to v1.6.84 production. Same numeric values flow into _CURVES via the new load path. No predictions move. Walk-forward caches re-stamped to v1.6.85 metadata only (no recompute needed since math is unchanged).
Session 152: hotfix for a DEGRADED-mode outage. v1.6.83 added a required league_aware_ltv block to engine_config.schema.json without bumping the engine version past the prior volume-persisted config; start.sh's restore-from-volume path overwrote the new config with the old one, and schema validation failed at boot.
Two-part fix: (a) bump engine version to v1.6.84 so start.sh stops overwriting the new config, (b) make league_aware_ltv optional in the schema so an older volume-persisted config doesn't fail validation.
Engine v1.6.83 -> v1.6.84. Math unchanged. Filed as feedback discipline: new required schema blocks stay optional until the engine version is bumped past the volume-persisted version.
engine_config.json::age_curves was found to be dead config: loaded by dynasty_engine.py:241 but never consumed for actual curve lookups. Live production path is age_curves.py::age_curve_factor -> _CURVES[pos] -> age_curves_rebuilt.py::REBUILT_*_CURVE.
Earlier Session 151 attempted to ship 18 'SHIP'd age-curve calibration updates (v1.6.79-82). Re-tested against the live REBUILT_* baseline after reverting validate_curve_change.py to its original age_curves_rebuilt.py source: all 4 position bundles returned DEFEND or TIE. The REBUILT_* production curves are MAE-defensible. Reverted engine_config.json::age_curves to v1.6.78 values + bumped _meta.version to v1.6.83 to supersede phantom v1.6.79-82 commits.
Engine math: IDENTICAL to v1.6.78 production behavior. No curve values changed in the live production path. No predictions move. backtest_results.json metrics (MAE/r/bias) hold byte-identical between v1.6.78 and v1.6.83.
Filed P2 architectural fix: consolidate the dual config so age_curves.py reads from engine_config.json::age_curves via RankingsEngine.params, deprecating age_curves_rebuilt.py. Eliminates the bug class.
Also shipped: Task #19 last hole closed (/v1/trade/ai-recommendations now scopes verdict math to caller's league config); live overlay surfaces draft_type so auction users see 'Auction' label instead of misleading 'picks until me: 0'.
Catch-up note: version_history entries for v1.6.62 through v1.6.82 are missing (BACKLOG.md is the system-of-record for those sessions; consolidation into version_history filed as P3 hygiene).
Session 151: 18-of-18 age-curve calibration recommendations from the Session 150 backtest SHIPPED via validate_curve_change.py harness through versions v1.6.79, v1.6.80, v1.6.81, v1.6.82 (incremental version bumps in one commit).
Walk-forward caches refreshed to v1.6.82 after recalibration. RB/WR/TE age-curve values written to engine_config.json::age_curves block.
Subsequently reverted in v1.6.83 after the dual-config drift was discovered: engine_config.json::age_curves was loaded by dynasty_engine.py:241 but never consumed for actual curve lookups; production used age_curves_rebuilt.py::REBUILT_* values. Engine math change: zero in production despite the calibration ship.
Frontend Position Counts widget migrated to canonical position_health from the API payload. Removes the last consumer with its own bespoke thin/surplus computation.
Engine v1.6.77 -> v1.6.78. Closes the v1.6.76 unified-foundation migration arc.
Trade Targets migrated to the canonical _position_health foundation from v1.6.76. -19 LOC; replaced site-local thin/surplus synthesis with health[pos]['signal'] read.
Unified _position_health foundation. Three drifting roster-strength views (_weak_positions PPG-median based, _positional_grades A-F tier quality, _compute_position_targets count vs proportional) consolidated into one canonical RosterAnalyzer._position_health method.
Returns a per-position dict {count, target, lo, hi, grade, weak_by_ppg, count_status, signal} where signal synthesizes to 'thin' / 'surplus' / 'ok'. Surfaced on full_analysis() return as position_health so SPA + downstream consumers read one canonical view.
FA opportunities migrated as smoke proof. Engine v1.6.75 -> v1.6.76.
Apostrophe-class sibling fixes: League Overview and Trade Central had their own normalize() paths that needed the same apostrophe-preserving treatment as v1.6.74.
Closes the apostrophe-name-fallback bug class across all SPA surfaces.
Apostrophe name-fallback fix: Tre' Harris and D'Andre Swift class of names with embedded apostrophes were failing the name-match fallback path in _build_player_payload because the normalize() pass didn't preserve curly + straight apostrophe variants.
Unit tests added (tests/test_apostrophe_name_fallback.py) to lock the behavior. Class lockdown locks v1.6.74 against regression.
Silent-except annotation sweep across routers/. Every silent-exception site got an explicit comment explaining what failure mode it absorbs and why, plus a logger.debug breadcrumb so the silent path is observable in production.
Pair with the boot-time data-sources check (v1.6.72) so silent-degrade bug class becomes a runbook concept rather than a hidden failure mode.
Silent-degrade observability: boot-time data-sources check runs at app startup, validates that all CSV + parquet dependencies are present and parseable, and logs a one-line STATUS marker per source.
Closes the bug class where market_layer ran silently no-op for 30 days (Sessions 118-148) because FantasyPros CSVs were gitignored and never reached the Docker image. Now any missing source surfaces in Railway logs within 5 seconds of boot.
Trade Targets surface count-aware weak/surplus signals: position labeled 'thin' only when count below low watermark OR PPG-median below threshold; labeled 'surplus' when count exceeds high watermark.
Closes the inverse of the v1.6.70 fix: previously, a roster's 'trade target' suggestions could mark a position as 'thin' purely on PPG-median while the user already had 7 of that position.
FA opportunities now respect position surplus. A roster carrying 8 WRs no longer surfaces additional WR FA pickups as 'opportunities' just because they show positive LTV-delta against the open pool.
Surplus detection feeds off the per-position weak_by_ppg + count_status path that becomes the unified _position_health foundation in v1.6.76.
Trajectory bucketing fix: collapsed 6 classifier labels (ascending / early_peak / late_peak / plateau / declining / unknown) into 4 display buckets so /recs and /under-the-hood narratives match the actual classifier output without spurious bucket churn between sessions.
POS_IDEAL now derives dynamically from league config + roster size instead of a hardcoded table. Lineup slots, FLEX count, SUPER_FLEX presence, and bench depth all feed the per-position ideal count used by the recommendations engine.
Closes a class of bug where 12-team SF leagues received the same position-targets as 10-team 1QB leagues. Affects FA opportunities, Trade Targets, and roster-grade prose.
Session 149: 5-site sweep closing the last-write-wins bug class across keeper.py (/v1/keeper/optimize, /v1/keeper/evaluate, /v1/keeper/inflation), redraft.py (/v1/redraft/edge), and meta.py cache-status age.
All five sites now scope to caller config_fp with logger.warning on fallback labeled KEEPER_<NAME>_LAST_WRITE_WINS_FALLBACK / REDRAFT_EDGE_LAST_WRITE_WINS_FALLBACK per the Session 148 silent-degrade discipline.
Engine v1.6.66 -> v1.6.67. SPA already passed config_fp on /v1/keeper/optimize so no frontend change required.
Session 149: /v1/redraft/trade scoped to caller config_fp. Sister fix to v1.6.64 closing the same last-write-wins bug class for the redraft engine surface.
Session 149: Recommendations Buy Low / Sell High lists now sort by engine rank instead of arbitrage magnitude alone, and filter by team-needs signal so the targets list reflects roster fit rather than pure market dislocation.
Engine v1.6.64 -> v1.6.65. No math change; surface logic only.
Session 149: Task #19 primary close. /v1/trade and /v1/trade/ai-read now scoped to caller config_fp so trade analysis runs against the user's actual league config, not whichever was last written to the rankings cache.
Defense-in-depth: TRADE_*_DEFAULT_CONFIG_FALLBACK warnings fire on the silent-degrade path so future cache-source drift surfaces in logs instead of silently producing wrong answers.
Engine v1.6.63 -> v1.6.64. Math unchanged on a correctly-scoped request; production behavior now matches the league the user is looking at.
Session 148: market arbitrage shipped end-to-end across Trade Central + Recommendations with LTV-delta math (not just rank-delta).
arbitrage_engine.compute_signals now produces LTV-delta magnitudes alongside rank deltas. Engine LTV at market_rank becomes the implied_market_ltv; engine_ltv - implied_market_ltv = ltv_delta in actual LTV units, sorted by magnitude.
Trade Central surfaces per-player BUY/SELL chips with LTV delta on each give/receive row, plus trade-level market_verdict (roughly fair / steal / overpay), arbitrage_edge_narrative when engine and market diverge meaningfully, and an arbitrage_summary.net_ltv_edge roll-up. Recommendations: trade_targets arbitrage weight bumped 0.20 -> 0.30; buy_low / sell_high entries carry ltv_delta + engine_ltv + implied_market_ltv.
Session 142: scorched-earth removal of Best Ball as a supported league format. The format was experimental Tier-3 work (Sessions 138-141) that never reached UAT-clean status; complexity-to-payoff ratio failed an IP-risk + maintenance review on 2026-05-24.
Session 141 closes the v1.6.60 follow-up. After DISC-6-recal (v1.6.60) closed only ~1% of the +2.65 PPG QB backtest under-projection bias on the weighted_ppg signal layer, and dynasty shrinkage harness DEFENDED (3 candidates all TIE/regression), the next-tier investigation surfaced REBUILT_QB_DUAL_CURVE in age_curves_rebuilt.py as the dominant amplifier.
Original Session 94 v1.6.12 dual_threat curve assumed Vick/Cunningham-class collapse past 28: 27:0.95, 28:0.85, 29:0.72, 30:0.58, 33:0.22. Compounded across the 7-yr LTV horizon, this produced ~64% LTV haircut on Josh-Allen-class age-29 dual QBs. Modern dual-threat cohort (Hurts 27, Burrow 28, Lamar 28, Allen 29, Wilson 31-33) doesn't decline this fast — they extend the peak into the early 30s.
Built scripts/validate_qb_dual_curve_modern.py (paired-observation MAE A/B on dual-threat QB year-pairs, n=32 across 2018-2024 ref years; matches established harness pattern). 5 candidate curves tested, all SHIP per discipline. Selected flat_then_slow_decline (tightest CI; preserves dual-archetype character vs deleting it entirely): agg MAE 4.144 -> 3.345 delta=-0.799 CI95 [-1.729, -0.045] all-negative. Bias closure: -3.075 -> -2.036 PPG (closes 34% of the dual-cohort under-projection).
New curve shape: flat at peak through age 30 (27:1.000, 28:0.98, 29:0.95, 30:0.91), then gentle decline (31:0.85, 32:0.78, 33:0.70, 34:0.60, 35:0.50). Matches empirical performance of the modern dual-threat QB cohort.
Engine math: age_curves_rebuilt.REBUILT_QB_DUAL_CURVE replaced with the new shape. No other engine code changes. Pocket curve unchanged. Schema unchanged (curves are Python constants, not config-driven).
Sample affected players (single-step Y+1 projection): Josh Allen age 29 base 17.24 -> cand 22.74 vs actual 24.04 (error -5.51); Russell Wilson age 33 base 4.48 -> cand 14.24 vs actual 17.34 (error -9.77); Lamar Jackson age 28 base 18.57 -> cand 21.41 vs actual 17.60 (error +2.84 over). Aggregate net win.
Cascades across the 7-yr LTV horizon — a 5+ PPG single-step projection improvement on Allen at 29 compounds into a ~20% LTV improvement for him.
Session 141 backtest regen against v1.6.59 surfaced QB MAE +14% (test 2024) / +41% (test 2023) regression on same test data vs v1.6.19 baseline; QB systematic bias +2.65 PPG under-projection (filed as P2 in BACKLOG). Root-cause hypothesis: DISC-6 (Session 109, v1.6.21) QB recency weights calibrated to nearly-flat 0.5/0.5/0.3/0.3/0.3 on a 6-yr-back n=199 veteran-dominant cohort underweights the modern breakout-young-QB phenomenon (Hurts/Allen/Burrow/Jackson/Daniels/Stroud Y0->Y1 ascensions).
Built scripts/validate_recency_weights_modern.py (paired-observation MAE A/B isolated to metrics.py::weighted_ppg_for_player; no confounders from shrinkage / age curve / LTV downstream). Cohort: 2018-2024 ref years, n=190 QB year-pairs. Bias decomposition at baseline confirmed Session 141 hypothesis at cleanest signal level: breakout-young (age<=25) bias -1.35 PPG (n=77), plateau-vet (age>=28) bias +0.61 PPG (n=86) - same direction the backtest surfaced, ~5x smaller magnitude (full pipeline amplifies the signal-level error downstream).
Grid search (16 single-profile candidates + 9 age-conditional candidates): single-profile candidates show breakout-young improvement CI all-negative across the board, but mid-career n=27 regression eats aggregate gains; no single-profile candidate earns SHIP (best agg CI crosses zero). Age-conditional candidates with cut25 leave mid/vet untouched and earn SHIP at multiple points along the young-tilt curve.
Selected: YOUNG=1.2/0.7/0.4/0.2/0.1 cut25 (best point estimate among CI-clean candidates). Verdict: agg MAE 2.935->2.873 delta=-0.062 CI95 [-0.111,-0.018] all-negative. Breakout-young delta=-0.153 CI95 [-0.274,-0.048] all-negative. Mid-career (n=27) and vet (n=86) cohorts unchanged by construction (cut25). Bias closure: breakout-young -1.35->-1.19 (12% closure); plateau-vet bias unchanged.
Engine math: position_recency_weights.QB is now a structured object with _age_conditional=true, age_cutoff=25, young_weights and old_weights sub-dicts. metrics.py _resolve_recency_weights_for_player helper resolves per-player based on age at reference_season (sourced from age_<season>/age_at_season/birth_date columns already present in stats_df via data_loader). Defensive fallback: missing/zero age uses old_weights so unknown players are never silently treated as breakout-young. engine_config.schema.json QB sub-schema is now oneOf [legacy_flat, age_conditional]; back-compat preserved (any flat QB profile still validates).
Path forward for RB/WR/TE: same harness extensible to other positions; current calibrations DEFEND in this session by lack of audit evidence for the same modern-cohort regression. Re-test all four positions annually as cohort grows.
Session 140 v1.6.58 shipped DISC-R2 with age_min=30 across all four positions per BACKLOG spec. Variant testing during the SHIP exercise showed age-28 cliff RB/WR (TE/QB unchanged at 30) won bigger aggregate MAE. Filed as follow-up.
Harness scripts/validate_redraft_fa_discount.py re-run on the position-specific candidate (RB/WR age_min 28, TE/QB age_min 30): RB agg MAE 4.014->3.814 delta=-0.200 CI95 [-0.265,-0.136]. WR agg 3.744->3.659 delta=-0.085 CI [-0.111,-0.060]. TE + QB TIE by construction (no change for those positions). Young <30 slice (new 28-29 cohort): RB delta=-0.295 CI [-0.390,-0.202], WR delta=-0.132 CI [-0.171,-0.094]. No slice regression.
Over-discount risk check on late_sign 28-29 subset (the cohort that actually played - validating bigger discount not driven purely by y1_missing actual=0 cases): RB n=60 delta=-0.627 CI [-1.005,-0.247] all-negative, WR n=106 delta=-0.246 CI [-0.378,-0.117] all-negative. Bigger discount fits both subtypes.
Discipline rationale for stopping at age-28 vs age-27 (which won even bigger Δ but pushes onto truly prime-age 27 FAs at WR per CLAUDE.md age curves): consensus is signal, never a tuning target; 2-year drop from spec defensible, 3-year drop risks over-fit to y1_missing dominance.
Engine math: no code change needed (the dict-form resolver shipped in v1.6.58 already supports any age_min). engine_config.json only: RB age_min 30->28, WR age_min 30->28. Schema unchanged.
Locked discipline: any future change to fa_status_discount age_min must re-pass scripts/validate_redraft_fa_discount.py with SHIP verdict + late_sign subtype check (the over-discount risk cohort that doesn't show up in agg).
Session 137 deferred DISC-R2 citing 'no historical FA-roster snapshot data to harness against.' Session 140 closed the data blocker by building a proxy cohort from player_stats_season.csv: a player is 'FA-going-into-Y1' if EITHER no Y1 record exists (FA bust, ppg1=0) OR team_Y0 != team_Y1 AND games_Y1 < 14 (late-sign approximation). Cohort: 1526 pairs total (681 y1_missing + 845 late_sign).
Descriptive cut confirmed the hypothesis: aged 30+ FAs regress -2.91 to -4.87 PPG Y0->Y1 vs same-team aged players regressing -0.68 to -1.83 PPG. Excess regression: RB -1.84, WR -2.08, TE -2.23, QB -3.64 PPG attributable to FA status.
Built scripts/validate_redraft_fa_discount.py (paired-observation MAE A/B). Accepts scalar (legacy) OR dict {default, by_pos_age} candidate.
Verdict: SHIP across all four positions. QB agg MAE 6.129->5.976 delta=-0.153 CI95 [-0.222,-0.082]. RB agg 4.301->4.014 delta=-0.287 CI [-0.362,-0.213]. WR agg 3.922->3.744 delta=-0.178 CI [-0.212,-0.144]. TE agg 3.169->2.856 delta=-0.313 CI [-0.402,-0.223]. Aged-30+ slice MAE delta ranges -0.284 (QB) to -0.893 (RB), all CIs all-negative. Young-slice delta=0 by construction. Late-sign aged-30+ subset wins decisively (RB delta=-0.437 CI [-0.791,-0.080], WR delta=-0.320 CI [-0.470,-0.170], TE delta=-0.225, QB delta=-0.092).
Robustness check on adjacent variants: conservative (RT 25, W 20, Q 18) wins smaller; aggressive (RT 45, W 35, Q 30) wins agg but breaks late-sign RB at CI level (CI [-0.937, +0.107] crosses zero). Age-28 cliff variant wins more aggregate MAE but is out of BACKLOG spec - filed as follow-up.
Engine math: redraft.fa_status_discount config field changed from scalar to dict shape. redraft_engine.py __init__ accepts both forms (fa_status_discount_raw stores original, scalar surface kept for gate check). Per-player application site resolves discount via cfg lookup: dict -> by_pos_age[pos].discount if age >= age_min, else default; scalar fallback unchanged.
engine_config.schema.json updated: redraft.fa_status_discount now oneOf [number, object{default, by_pos_age}].
Locked discipline: any future change to fa_status_discount params must re-pass scripts/validate_redraft_fa_discount.py with SHIP verdict on all four positions.
Session 137 post-ship re-benchmark identified Brock Bowers engine #54 vs consensus #18 — TE2 in consensus undervalued by engine. Root cause: Bowers has seasons_data=2 (Y1 rookie 17g + Y2 12g), both elite-PPG, but Bayesian shrinkage applies uniformly to all seasons_data<3 cohorts, pulling his raw 12.6 PPG toward TE prior 5.93 -> 8.9 PPG. Engine cannot distinguish 'elite back-to-back producer' from 'thin-sample one-year wonder.'
Descriptive cut on n=1198 Y2-Y3 transitions (2015-2024): for raw_pred > 1.5 * position_prior cohort, NOT shrinking beats shrinking by Δ -0.7 to -1.1 MAE per position. TE elite delta=-1.10, RB elite delta=-0.69, WR elite delta=-0.50. Non-elite cohort unaffected (engine should still shrink them).
Built scripts/validate_redraft_y2_shrinkage.py — paired-observation MAE A/B with shrinkage-skip when raw_ppg > elite_skip_mult * prior.
Verdict at elite_skip_mult=1.5: SHIP. Agg MAE 2.970 -> 2.813, delta=-0.157, CI95 [-0.241, -0.074] all-negative AND clears strict SHIP_THRESHOLD. Elite cohort (n=298) delta=-0.633, CI [-0.961, -0.306]. TE elite (n=45) delta=-1.102 CI [-1.799, -0.403]. WR elite (n=189) delta=-0.502 CI [-0.867, -0.135]. RB elite (n=64) delta=-0.692 CI [-1.627, +0.268]. Non-elite cohort (n=900): delta=0.000 (perfect guardrail by design).
Engine math: new redraft.shrinkage.elite_skip_mult parameter (default 1.5). In _apply_shrinkage's per-player loop, skip downward-shrinkage if ppg > elite_skip_mult * effective_prior. Otherwise continue as before. Applies AFTER DISC-R4 QB rushing-prior override.
Locked discipline: any future change to shrinkage_elite_skip_mult must re-pass scripts/validate_redraft_y2_shrinkage.py with SHIP verdict.
Session 137 final tally bumped: 5 engine ships (R3-WR + R5 + R4 + R7 + R8) + 3 defends (R1, R3-TE TIE, R6) + 1 deferred (R2 no data) + 7 paired-observation harnesses live in scripts/.
Session 137 benchmark surfaced Rashee Rice engine #100-112 vs consensus #21. Root cause: engine's min_games_full_weight=10 hard threshold puts his 2025 8-game elite-PPG season (18.76 PPR) in partial_records, which are then DISCARDED entirely because his 2023 16-game season exists as a full record (line 818: `active = full_records if full_records else partial_records`). Effective seasons_data=1 -> Bayesian shrinkage at w=0.4 pulls his 2023 PPG of 10.81 toward WR prior of 6.46, projecting 7.94 PPG.
Descriptive cut on n=6590 year-pairs (2015-2024): games0=8-9 MAE 3.00 ≈ games0=10-12 MAE 3.01 (statistically identical). 8-9 game PPG is as predictive of year+1 as 10+ game PPG. The 10-game threshold is over-restrictive.
Built scripts/validate_redraft_min_games.py - paired-observation MAE A/B for the threshold change. Test: predict year+1 PPG using engine recency-weighted multi-season blend under threshold=10 vs threshold=8. Compare MAE on n=5495 total pairs, n=450 affected pairs (where threshold change matters).
Engine math: redraft.min_games_full_weight 10 -> 8 in engine_config. Affects 510 year-pairs of 8-9 game seasons (typically injury returns) that now count as full_records instead of being discarded. Rashee Rice's 2025 season now contributes to his projection; expected ranking move from #100 up to ~#60-80 range.
Top-of-position consensus check: no top-3 player at any position has a games=8-9 season in their recency window currently, so this change doesn't disrupt elite rankings. The change affects mid-tier WRs (Rice), injury-return RBs, backup-to-starter QBs.
Locked discipline: any future change to min_games_full_weight must re-pass scripts/validate_redraft_min_games.py with SHIP verdict.
Session 137 redraft benchmark surfaced young rushing-QB cluster (Maye, Daniels, Dart, Caleb) all ranking 50-70 places below consensus; Stroud (pocket passer) matched consensus as built-in control. Engine internals: shrinkage_factor=0.333 on Maye pulled him toward the QB population mean of 14.96 PPG.
Descriptive on n=650 QB year-pair transitions (2018-2024, games>=8): young (age<=25) low-rush (<15 ypg) Δ=+0.41 PPG Y0->Y1 (n=110), young mid-rush (15-30) Δ=+0.30 (n=59), young high-rush (>=30) Δ=+1.03 (n=37). Full-cohort elite-rush (>=50 ypg) Δ=-2.25 (n=15) warns against uncritical rushing-up prior — older rushing QBs decline.
Built scripts/validate_redraft_qb_rushing_prior.py — paired-observation MAE A/B. Bayesian shrinkage formula: weight = n/(n + k*pos_k_mult); pred = w*ppg0 + (1-w)*prior. Candidate replaces position-wide prior with rush_prior_ppg for QBs satisfying both (rush_ypg0 >= threshold) AND (age0 <= rush_age_max).
Tested 7 variants. Selected Variant 5 (rush_prior=18.0, threshold=25 ypg, age_max=25): agg MAE 2.769->2.741, delta=-0.028, CI95 [-0.048, -0.009] all-negative. Pocket low-rush slice unchanged (0.000 — Stroud-style control). Mid-rush slice delta=-0.055 CI [-0.104, -0.010]. High-rush 30-50 ypg slice delta=-0.253 CI [-0.426, -0.084]. Young high-rush target cohort delta=-0.283 CI [-0.549, -0.010] n=37. Older rushing QBs (age>=27) locked out at delta=0 by age_max gate. Elite-rush (>=50) slice unchanged direction-wise (+0.113 but CI [-0.265, +0.465] crosses zero, n=15).
Engine math: new redraft.qb_rushing_prior config block + per-player qb_rush_prior_map built inside _apply_shrinkage. The shrinkage target switches from uniform position prior to rush_prior_ppg for indexed rows; pocket QBs and older QBs keep the uniform prior (current behavior).
Feature flag _enabled=true. Schema entry under redraft.properties. _apply_shrinkage signature extended to accept Optional stats_df for latest-season rush_ypg lookup. Discipline pattern locked: any future change to qb_rushing_prior params must re-pass validate_redraft_qb_rushing_prior.py with SHIP verdict.
Three engine ships in one night via the harness discipline: v1.6.53 (DISC-R3 WR age), v1.6.54 (DISC-R5 aging-WOPR), v1.6.55 (DISC-R4 QB rushing-prior). Plus one DEFEND (DISC-R1 TEP-gating) and one TIE (DISC-R3 TE).
Session 137 redraft benchmark surfaced WR 28-29 cluster meanDelta=-63.8 ranks + WR wopr>=0.4 + age>=28 cluster meanDelta=-58.3. Hypothesized root cause: engine's opportunity_premium boosts a 32-yr-old high-WOPR WR the same as a 24-yr-old.
Descriptive cut on n=1721 WR year-pair transitions (2018-2024): high-WOPR (>=0.40) regression -0.51 PPG/yr at age 22-26 vs -2.02 PPG/yr at age 30+. Hypothesis empirically supported.
Built scripts/validate_redraft_opportunity_aging.py — paired-observation MAE A/B that subtracts penalty = alpha * max(0, wopr - wopr_anchor) * decay_age(age) from ppg0 prediction. Three variants tested.
Variant 1 (alpha=4.0, age_floor=27, age_cap=32): agg MAE 2.956->2.905, delta=-0.051, CI [-0.064, -0.037] all-negative. Young 22-26 flat. Prime 27-29 delta=-0.055 CI [-0.075, -0.036]. Old 30+ delta=-0.296 CI [-0.394, -0.198]. High-wopr old 30+ delta=-0.370. Verdict: SHIP.
Variant 2 (gentler alpha=2.0): agg delta=-0.028 also SHIP. Variant 3 (alpha=3.0 age_floor=28): agg delta=-0.029 also SHIP. All three variants SHIP cleanly; selected Variant 1 for biggest MAE improvement + match to empirical age curve.
Engine math: new redraft.opportunity_aging_penalty block + _aging_wopr_penalty helper in _compute_draft_score. Penalty subtracted from draft_score (not redraft_proj — pure ranking adjustment, preserves projection transparency). New per-player column aging_wopr_penalty surfaced in output.
Feature flag _enabled=true (defaults to live). Schema entry in engine_config.schema.json under redraft.properties allows the new block. Discipline pattern locked: any future change to opportunity_aging_penalty params must re-pass validate_redraft_opportunity_aging.py.
Session 137 full benchmark of redraft engine top 200 vs PlayerProfiler TEP + FantasyPros standard surfaced WR 28-29 cluster meanDelta=-63.8 ranks (engine over-rates) and WR 30+ cluster meanDelta=-56.2. Hypothesized root cause: WR age cliff_age=29 + rate=0.030/yr too gentle.
Built scripts/validate_redraft_age_curve.py — paired-observation MAE A/B for redraft age_projection_discount changes. Mirrors scripts/validate_curve_change.py (dynasty) but targets the single-season redraft discount.
Harness verdict on n=2303 WR year-pair transitions (games>=8): agg MAE 2.914 -> 2.876, delta=-0.038, CI95 [-0.062, -0.015] all-negative. Prime 27-29 slice MAE 2.799 -> 2.739, delta=-0.060, CI [-0.102, -0.019]. Old 30+ slice MAE 2.854 -> 2.756, delta=-0.098, CI [-0.194, -0.004]. Young 22-26 slice flat at 3.010 (no regression). Verdict: SHIP. CI upper bound on aggregate just misses strict SHIP_THRESHOLD of -0.020 by 0.005; cleanly directional with n=2303 and all-negative CI.
TE candidate also tested (cliff_age 30 -> 28, rate 0.022 -> 0.045). agg MAE 2.177 -> 2.154, delta=-0.023. CI95 [-0.051, +0.004] crosses zero. Verdict TIE per locked Session 112 discipline (CI upper bound must clear threshold). TE unchanged this session; deferred to future iteration with either more cohort years or a less-aggressive candidate (cliff_age 28 rate 0.035).
RB and QB blocks unchanged (not part of DISC-R3 scope this session).
Discipline pattern locked: any future change to redraft.age_projection_discount must pass scripts/validate_redraft_age_curve.py with SHIP verdict + non-regressing slices.
The 26-entry historical version_history backfill (v1.6.20-v1.6.46) shipped at v1.6.50 but did not actually deploy to prod. Root cause: start.sh:82 image-vs-volume compare only prefers the image-shipped engine_config.json when the version is STRICTLY NEWER. When versions are equal (both v1.6.50), the volume's old file wins and the new image-side file gets clobbered on every boot.
Bumping _meta.version to v1.6.51 triggers the image-prefer path, propagating the backfilled version_history entries to the live container.
Filed P3 follow-up: extend start.sh to also prefer image when file hashes differ at equal versions, so future docs-only engine_config changes propagate without a fake version bump.
Extend the H3a Sleeper-cache load to also build gsis_id -> depth_chart_order map for skill positions.
Apply multiplicative discount when rostered RB/WR/TE has depth_chart_order >= buried_depth_threshold (default 4).
New tunable knobs redraft.buried_depth_threshold (default 4) and redraft.buried_depth_discount (default 0.10).
QB intentionally skipped: dynasty engine has its own QB-specific gate (DISC-Q2 via avg_avail cap); redraft QB handling deferred to H2 follow-up.
Threshold=4 calibration: order 1-2 = real roles, order 3 = rotational players with real snaps (Christian Watson GB), order 4+ = genuinely buried (Aiyuk #4 SWR SF, Tank Dell #6 RWR HOU).
Build gsis_id -> team map from Sleeper cache (sleeper_players_cache.json) before per-player loop, applying same DISC-Q2 pattern from dynasty.
Apply multiplicative discount when player has team=None (FA-status).
Two-stage cache load with stale-cache fallback: load_sleeper_players() first (respects TTL), then direct cache-file read if that fails. Production-resilient when network is briefly down + TTL expired.
New tunable knob redraft.fa_status_discount, default 0.15.
Compounds independently with H7: Mixon (FA + missed 2025) and Amari Cooper (FA + missed 2025) get both discounts stacked.
Detect when a player has no record for the reference season in stats data (e.g., Brandon Aiyuk: 2020-2024 in data, 2025 missing entirely due to ACL recovery).
Apply multiplicative discount to weighted_proj when missed_year_0 detected.
New tunable knob redraft.missed_year_0_discount, default 0.20.
First iteration used `active` (post-filtered) for detection and falsely triggered on 237 partial-season-2025 players (Tyreek Hill 4g, Burrow injured, Diggs 17g). Fixed by checking SOURCE group instead. Second iteration: 140 legitimate fires, all spot-verified.
Capped ceiling at proj * (1 + ceiling_max_multiplier) in redraft_engine._compute_ceiling_floor.
New tunable knob redraft.ceiling_max_multiplier, default 0.50.
Identified during 2026-05-22 redraft engine evaluation. Mechanism: ceiling = max(historical_PPGs) produced unrealistic ceilings for depth players with old fluke career years; the 0.30 ceiling_weight then promoted them into top 200 via ceiling_premium contribution.
Cap fires surgically on targeted cases: Evan Engram ceiling 10.19 -> 7.48, Cole Kmet 9.64 -> 6.96, David Njoku 10.23 -> 10.06. Cap does NOT fire on McBride (14.88 < 19.62), Kittle (13.17 < 14.57), Nacua (19.41 < 24.80), Bijan, Kelce, Aiyuk, Jeudy.
Full-engine A/B: Engram #129 -> #152 (drop 23), Cole Kmet #149 -> #174 (drop 25). Top-3 preserved at all four positions.
sessionStorage is scoped per-tab. When Clerk's Google SSO opens the OAuth handshake in a popup or new tab (common with .clerk.accounts.dev redirects), the redirect-back-to-/ can land on a different tab than the one that originally wrote the intent.
New tab has its own empty sessionStorage, _checkPendingCheckout finds nothing, user lands on / instead of Stripe Checkout.
Fix: switch pending_checkout intent storage from sessionStorage to localStorage with 10-min expiry.
Consumed by app.js::ensureIdentified, not login.html (cross-tab survival).
Plus discover: rushing-QB Y0-to-Y+1 persistence (H-34, adds_information) cherry-picked + blog post live at /lab/rushing-qb-y0-100-att-y1-persistence.
Clerk's Google SSO completes off-site and redirects directly to the configured After-sign-in URL (typically /), bypassing /login entirely.
That meant login.html's _resolveRedirectTarget never runs and the sessionStorage pending_checkout intent was never consumed. User landed on / instead of Stripe Checkout.
Fix: app.js ensureIdentified() calls _checkPendingCheckout() after a real user_id resolves. If sessionStorage has a fresh (<10 min) pending_checkout entry, the watcher replays the billing call.
Closes the SSO-bypass-of-checkout-intent failure mode.
SPA: 167 em dashes in app.js + 111 in index.html replaced via context-aware Python sweep, comments preserved, 0 syntax breaks.
Pure Signal pricing page (static/pricing.html, 471 lines): three tier cards (Free / Pure Signal annual+monthly / Lifetime founder-only), 35-row comparison matrix, 4-card hygiene commitments grid, 7-question FAQ.
discover.md SKILL locked 'no em dashes anywhere' + 'specific examples sparingly' + calibration paragraph as north star.
NOTE: This commit also contained the FUSE-truncation outage that hit production for 55 minutes (Session 128 P0). Three additional files (pricing.html, methodology.html, blog md) were also silently truncated by the same FUSE write. Recovery shipped Session 128.
Harness verdict via scripts/validate_pwopr_lift.py: SHIP for WR at threshold 0.10 + lift -0.10 (aggregate MAE delta -0.025, CI -0.040 to -0.011 fully below zero, slice MAE -0.47 ppg/pair, n=121 pairs). TE remains TIE.
Engine integration: metrics_pwopr_lift.py + engine_config.pwopr_regression_lift block (default _enabled: false). Flag-gated; production rankings don't shift until operator flips after named-case review.
FINAL DISPOSITION: persistence filter + magnitude-scaled lift_bands built and tested; no configuration earns SHIP verdict that respects named-case discipline. PWOPR remains descriptive-only signal (chips + badge + filter pill).
Plus market layer revival: FP static CSV + Buy Low / Sell High in Recommendations.
Surfaces nflverse-native WOPR (1.5*target_share + 0.7*air_yards_share, Hermsmeyer 2017) as a Custom Layer enrichment column on every player payload.
Backfilled custom_layer_per_player.parquet for 2018-2025 via join from player_stats_season.csv (no nflverse re-pull required).
Provably no math change for predictions: WOPR was already flowing into the TE LASSO model (coef 0.3512, blend 0.1) via base_df; CL cache overlay rewrites the same numerical value.
Validated via paired-observation A/B harness (scripts/validate_injury_discount_change.py): SHIP verdict with QB best lift +18.1 LTV (Lamar class), WR +13.4, RB +4.1, TE 0.0.
Prime cohort stable, top-3 consensus stable across all 4 positions.
Healthy-classified players with availability_rate in [0.75, 0.90] who previously took a 19% LTV penalty from Q1 availability scaling now get the alpha-power treatment restored.
Plus fourth A/B validator shipped: scripts/validate_shrinkage_anchor_change.py.
Original DISC-8-followup QB SHIP verdict driven by 12 boundary pairs out of 225 total -- the other 213 contribute 0 delta because the mod term cancels when both pair ages are on the same side of peak_end. Aggregate '-0.027' was 12 * -0.494 / 225.
Bootstrap 95% CI on the 12 boundary pairs: [-1.249, +0.276]. CI crosses zero.
Per the discipline, no flip on noisy 12-pair evidence. Reverting archetype_v2.enabled to false (v1.6.33 -> v1.6.34).
Harness v2 changes: bootstrap CI required for all verdicts; aggregate point estimates without slice-level support flagged as suspicious.
CI failed on v1.6.29 push because audit found 5 dead blocks under CI conditions (152 tracked .py files) but local run found 4 (169 files including untracked).
Root cause: scripts/ablation_study.py + trade_simulator.py present locally but absent from origin/main -- only consumers of engine_config.rookie_model and trade_simulator.
Two fixes: acknowledge rookie_model + trade_simulator as known-dead until those scripts ship, OR ship the scripts. Chose acknowledgment.
v1.6.30 was an attempted intermediate fix that got squashed/rolled back; skipped in version sequence.
Walk-forward 2020-2024 across n=388 predictions: RB R3 bucket showed -5.81 avg bias with 5/5 years signed agreement (all 5 over-projected). Survivor-bias artifact -- cohort median inflated by 4-5 breakout backs (Aaron Jones / Tony Pollard class).
Replaced median 11.36 -> 8.50 (lower of full-cohort and recent-5y medians). Widened q25/q75 (1.64, 12.96) to surface bimodal distribution.
Q4 deploy revealed Nabers's SF270 was actually driven by the STRUCTURAL injury layer reading 1 ACL season as 0.39 avg_avail and applying it as permanent flat multiplier on ltv_discounted. Same bug class as Q4, different code path.
Pre-Q4 _apply_live_injury_adjustment dampened weighted_ppg by year_0_factor (career_bender 0.30) which propagated through all 7 LTV horizon seasons. Nabers ACL crushed to 29% of healthy LTV.
Post-Q4: weighted_ppg un-dampened. New live_injury_year_0_factor param on dynasty_lifetime_value(_v2) applies year-0-only inside horizon loop. Years 2-6 use un-injured baseline.
projected_season_ppg still dampened directly so current-year UI reflects injury. Sprint B v2 RB recompute multiplies by y0.
Nabers-class unit math recovers 2.91x (29 to 85 percent of healthy). v2 path 2.87x.
Below-baseline players get injury_discount = max(scaling_floor, avg_avail) direct linear scaling, bypassing the legacy floor=0.792 trap. Above-baseline keeps relative**alpha treatment.
Ghost-QBs dropped: Wentz SF53 to 193, Martinez 77 to 392, Slovis 87 to 415, Trey Lance 88 to 318, Mac Jones 48 to 82, Russell Wilson 34 to 60.
Elite WRs/RBs lifted organically: Nico Collins 83 to 66, CeeDee Lamb 74 to 55, Gibbs 21 to 15. Top-5 QBs held (Allen 1, Hurts 2, Mahomes 3, Maye 5).
Bundled banner UI fix: sticky offsets on tab nav + filter bar moved to .sticky-below-app-header / .sticky-below-tab-nav classes (post-banner-bump corrections).
Feature-flagged via injury_model.availability_scaling_v2.enabled.
New unified_projection.py with compute_prior + compute_posterior + apply_unified_projection.
Replaces the bipartite pipeline (rookie-only empirical model or vet-only observed PPG) with a Bayesian blend that shrinks observed PPG toward the empirical prior until career_games exceeds shrinkage_anchor[pos].
Added regression_factor_by_position config support to production_regression layer for future per-position tuning.
H-10 (TD regression) cohort EMPIRICALLY VALIDATED: 78% of 10+ TD scorers regress Y+1, mean -4.0 TDs (WR -4.4, RB -3.7, TE -4.5).
BUT generic production_regression layer applied to WR/TE failed to move backtest bias as expected (WR went from -0.02 to +0.14 even at low factor=0.05). Generic layer triggers on PPG/TD-vs-career thresholds; for WR/TE the career-mean baseline interacts differently than for RB.
Decision: keep positions=[RB] only. The per-position factor scaffolding stays as foundation for future dedicated TD-haircut layer that targets only the top-TD cohort (>= 10 TDs) with empirically-calibrated 0.32 retention factor.
Increased production_regression.regression_factor from 0.35 to 0.50. Apex Fantasy Leagues research suggested 0.50-0.70 range for one-season-only cohort.
Walk-forward validated 2021-2023 (3 ref years): RB bias improved on EVERY year (closer to zero). Average bias improvement: +0.12 ppg per year. Correlation (r) preserved across all years (small noise: -0.016 / +0.005 / -0.003).
Per-year bias: 2021 -0.86 -> -0.74, 2022 -0.42 -> -0.28, 2023 -0.17 -> -0.08. All three move strictly toward zero.
Per-year MAE essentially unchanged within walk-forward noise band. Net effect: bias correction without MAE regression.
Pre-fix: WR/RB/TE rookie buckets stored median peak season ppg, but engine consumed values as year-by-year weighted_ppg input. Same structural error as v1.6.15 QB fix.
Empirical cohort analysis 2015-2023: peak vs career-avg gap is 2.7-4.4 ppg per LTV year across every R1_top / R1_late / R2 bucket for all 3 positions.
Pre-fix: rookie_pick_bucket_baselines.json QB R1_top median_peak_ppg=18.84. Engine consumed this as weighted_ppg input -> Ty Simpson (R1_top rookie) projected at 18.45 ppg every year of LTV horizon despite never playing an NFL snap.
Empirical cohort analysis (n=25 R1_top QBs drafted 2015-2023, Half-PPR Dynasdeez scoring): median PEAK season ppg = 18.50, but median CAREER-AVG ppg = 15.30. The 3.2 ppg gap is the structural over-projection — peak is a single-season high, career-avg is the sustainable rate.
Pre-fix: REBUILT_RB_BASE_CURVE table ended at age 34 (value 0.709). _base_factor fallback in age_curves.py inherited curve[34]=0.709 for all ages 35+ — so a 39yo RB scored 71% of peak production in the LTV horizon.
Result: aging RBs like Derrick Henry (age 32.7, weighted_ppg 22.71) cleared LTV 95.2 — above elite young WRs (Ja'Marr Chase 90.5) and tied with rookie QB bucket-baseline projections.
Fix: extended curve to {35: 0.500, 36: 0.300, 37: 0.150, 38: 0.080, 39: 0.040, 40: 0.020}. Values cross MIN_FACTOR_FLOOR=0.10 at age 38, truncating the LTV horizon early (matches reality — 38yo starting RBs are statistical anomalies).
Methodology grounding: Apex Fantasy Leagues research found ZERO qualifying RB-seasons age 33+ since 2000. The pre-fix plateau was unjustified extrapolation — not empirical.
Flipped cross_position_rank_blend from ltv_weight=0.5 / vos_weight=0.5 to ltv_weight=1.0 / vos_weight=0.0 (pure LTV).
Root cause: 50/50 blend was calibrated against KTC + FantasyCalc external consensus (4-source validation). Those sources were removed for IP reasons in Sessions 89-90 (de-risk sprint). The blend weighting became orphaned tuning targeting a calibration source that no longer exists in the engine.
Symptom Ryan caught at Session 94 wrap: Josh Allen (LTV 119, age 30) ranked BELOW Davante Adams (LTV 92, age 33) in Overall tab while within-position rankings sort by pure LTV. Cross-position blend was producing inconsistent ordering vs the within-position view.
Methodology alignment: pure LTV matches the locked feedback memory 'LTV is the primary ranking metric — not dynasty_ppg' and makes Overall consistent with within-position ordering.
Walk-forward MAE unchanged — this is a presentation-layer reorder, not a projection change. No retest needed.
Added REBUILT_QB_POCKET_CURVE + REBUILT_QB_DUAL_CURVE to age_curves_rebuilt.py grounded in Apex Fantasy Leagues research (158 qualifying QB-seasons, 4.5-year peak-age gap between archetypes).
Modified age_curves.py::age_curve_factor to dispatch to archetype-specific curve when FF_QB_ARCH_SPLIT_CURVES env=1 + QB position + dual_threat/pocket_passer archetype.
Added engine_config.qb_archetype_split_curves.enabled flag (default OFF).
Pocket curve: peaks age 29-31 at 1.0; holds 0.94+ through age 34; gradual decline thereafter.
Dual-threat curve: peaks age 24-26 at 1.0; cliffs 0.85 at age 28; drops to 0.22 by age 33 (matches Apex's 'only 4.1% of dual-threat seasons age 33+' finding).
Validation surface: enable env + run wf_age_2_runbook. Engine smoke + QB diagnostic should show shape difference for breakout QBs (Lamar, Hurts, Burrow) vs aging pocket passers (Stafford, Brady late-career).
New _apply_production_regression() in dynasty_engine.py — multiplicative blend of weighted_ppg toward career-mean PPG for RBs with 3+ prior seasons AND Y-1 spike (PPG >+1.0 OR TDs >+2.0).
Triggers ~30-50 RBs/year. Mean-reversion is the dominant structural failure WF-AGE-4 identified.
Added projected_season_ppg.age_aware per-position config (default OFF for QB/WR/TE, ON for RB).
When ON, projected_season_ppg = weighted_ppg * age_factor + bias_correction[pos], replacing legacy age-blind formula.
Recalibrated bias_correction[RB] from -0.54 to -0.20 to reflect residual after age_factor does the decline-curve work.
Engine change in dynasty_engine.py runs after _apply_framework so age_factor is available; downstream injury/ceiling/momentum multipliers operate on the corrected baseline.
Sprint C accuracy ship: WR age curve flattened to 1.0 across all ages (REBUILT_WR_CURVE in age_curves_rebuilt.py). Rationale: AS-7 ablation (Session 93) showed disabling the WR age curve improves WR MAE_dyn by 0.43 ppg/year; AS-8 minimal-engine head-to-head confirmed flat-curve naive baseline beats production WR MAE_dyn by 0.44. Sprint C experiments with curve reshape (AS-2-informed lifts; naive piecewise shape) both WORSENED WR MAE (4.27, 4.91 vs baseline 2.89 on 2019). Root cause: the LTV horizon math (7-season weighted average of weighted_ppg × age_factor) interacts non-linearly with non-flat values — only TRUE flat (age_factor=1.0 at every age) reproduces AS-7's improvement. Walk-forward AB 2019-2024 aggregate vs v1.6.6: QB MAE_dyn 3.891 → 3.690 (-0.20); RB 3.216 → 3.085 (-0.13); WR 2.702 → 2.268 (-0.43); TE 2.332 → 2.205 (-0.13). vs naive baseline (AS-8): QB now beats naive by 0.74; TE by 0.05; WR ties within 0.005; RB still loses by 0.16 (next sprint target). Trade-off disclosure: flat WR curve removes dynasty age-discrimination from WR dynasty_ppg — a 30-year-old WR with weighted_ppg=10 is valued the same as a 25-year-old at 10. This is intentionally optimizing for year-1 projection MAE; multi-year dynasty value differentiation is preserved in ltv_discounted (sum, not average) and in archetype modifiers. Old REBUILT_WR_CURVE values preserved in age_curves_rebuilt.py comment block for rollback. No other curves touched.
Session 93 Option-C ramp moderation. v1.6.5 walk-forward AB returned numbers identical to v1.6.4 across all four positions to two decimal places (QB MAE 3.89 / bias +2.58 / r 0.458; TE r 0.658). Diagnosis update: the QB age curve smoothing in v1.6.4 / rollback in v1.6.5 had no engine-level effect — the actual driver of the v1.6.4 regression (vs v1.6.3 baseline QB bias +2.35, TE r 0.719) was the rookie_development_ramp.QB change from [0.88, 0.94, 1.0] to [0.50, 0.70, 0.85, 1.0]. The mechanism cascade is ramp -> rookie projections -> P13/VOS positional priors -> TE r drop. Surgical fix: bracket the ramp magnitude to find the smallest step back from [0.50, ...] that still flips the Maye/Mendoza/Simpson inversion. Sandbox bracket sweep: [0.50, 0.70, 0.85, 1.0] gives Maye QB5 / Mendoza QB9 (4-rank gap, baseline); [0.55, 0.75, 0.88, 1.0] gives Maye QB5 / Mendoza QB7 (2-rank gap, robust); [0.60, 0.78, 0.90, 1.0] gives Maye QB5 / Mendoza QB6 (1-rank gap, thin); [0.65, 0.80, 0.90, 1.0] inverts back to Mendoza QB4 / Maye QB6. v1.6.6 selects [0.55, 0.75, 0.88, 1.0] — smallest step back with a robust safety margin on the inversion fix. Predicted effect on walk-forward AB: partial recovery of QB bias toward +2.35 and TE r toward 0.719, since the ramp magnitude on year-0 (the dominant lever) is 10% less aggressive vs v1.6.5. Pending: walk-forward AB validation. QB curve stays at restored trim-25% values (no rollback of v1.6.5's curve restore).
Session 93 surgical rollback of v1.6.4 QB age curve smoothing. Walk-forward AB on v1.6.4 regressed two of the four positions: QB bias drifted further from zero (+2.35 -> +2.57 wPPG, wrong direction; decision rule was 'bias improves toward 0' -> FAILED) and TE r dropped -0.061 (0.719 -> 0.658), with TE hit@top-12 dropping -8pp (62% -> 54%) despite TE curves / ramp / config block being untouched in v1.6.4 (suspected P13 / VOS cascade from QB-side positional priors shifting). WR untouched (curves not modified). RB MAE +0.14 worse, bias slightly improved. Diagnosis: the smoothed QB curve lifted peak-age values (27: 0.880->0.97, 28: 0.935->0.99), which are the ages most QB-seasons sit in across the walk-forward cohort, so cohort-aggregate projection went UP and over-projection bias got WORSE. The rookie ramp change (rookie_development_ramp.QB = [0.50, 0.70, 0.85, 1.0]) was the surgical lever that actually flipped the Maye / Mendoza / Simpson rookie inversion -- hand-calc confirmed Mendoza year-0 multiplier 0.88->0.50 = -43% year-0 PPG hit did the bulk of the LTV drop. v1.6.5 restores REBUILT_QB_CURVE in age_curves_rebuilt.py to the original trim-25% values and KEEPS rookie_development_ramp.QB at [0.50, 0.70, 0.85, 1.0]. Predicted result: Maye lands QB5-7 (smaller relative gap to Mendoza than v1.6.4 showed, but inversion still flipped vs v1.6.3 baseline). v1.6.4 entry marked superseded_by=v1.6.5 for audit trail; not deleted. Methodology note: this ship violated the AB-harness-before-wire-in gating rule formalized 2026-05-11 -- the one-player smoke check on Maye placement was not a population-level accuracy validation. Rule re-affirmed: any age curve / ramp / projection-parameter change MUST run through walk_forward_backtest before wire-in.
Session 92 rookie ranking inversion fix. Two related changes addressing Drake Maye / Mendoza / Simpson QB ranking inversion: (1) QB age curve smoothing in age_curves_rebuilt.py REBUILT_QB_CURVE. Pre-smooth values had 4 monotonicity violations (22>23, 26>27, 32<33, and 36->37->38->39 ascending from 0.771 to 0.951 — survivor-bias inversion implying a 39-year-old QB produces 95% of peak). Smoothed values enforce strict monotonicity ascending 22->29 and descending 30->41, with late-career magnitudes aligned to CLAUDE.md framework ("Decline gradual; effectively done Late 30s"). Concretely: 23: 0.773->0.84, 27: 0.880->0.97, 32: 0.894->0.93, 36: 0.771->0.68, 38: 0.933->0.46, 39: 0.951->0.32, 40: 0.848->0.20. (2) Rookie development ramp for QB steepened from [0.88, 0.94, 1.0] to [0.50, 0.70, 0.85, 1.0]. Old ramp projected a #1-pick QB at 17-18 ppg in year 1, ~equal to a 2nd-year vet's weighted_ppg. New ramp aligns year-0 projection with empirical year-1 outcomes (Mahomes was a backup year 1; Allen/Hurts/Lawrence/Burrow/Stroud/Daniels/Williams all had material year-1 growing pains). Combined effect on top-15 dynasty QB: Maye 24 lifts QB7->QB3; Mendoza 22 (#1 pick proj) drops QB2->QB5; Simpson 22 (#13 pick proj) drops QB6->QB14. Late-career QBs (Wilson 37, Rodgers 42, Cousins 38) drop out of top-25 entirely. RB/WR/TE curves untouched. Pending: walk-forward AB validation.
Sprint C consensus age-curve tweaks applied (12 ages across RB/WR/TE). All 12 ages flagged as 'too_pessimistic' by 3/3 reference years 2021-2023 in the multi-year backtest run on 2026-05-09. RB ages 25/29/31, WR ages 25/33/34, TE ages 26/27/30/32/33/34. Updates use the standard blend_factor=0.30 conservative weight from the existing accuracy_review pipeline. QB had 0 consensus signals; QB calibration deferred to a post-Pattern-13b-flip re-run.
Pattern 13b/c flip ON. qb_use_v2_ltv: false -> true after ab_harness validation on 2021-2024 snapshots. Verdict SHIP: 3/4 seasons pass kill criterion (dMAE <= -0.10). 2022: -0.80, 2023: -0.51, 2024: -0.47, 2021: +0.40 (single regression dominated by 3 wins). Net win directly attacks the QB +2.18 bias_w from Sprint C 2015-2024 backtest. Session 77 _apply_young_qb_ceiling guard prevents the double-boost interaction with young_qb_ceiling that originally rejected Pattern 13b on 2026-05-05. v2 path uses age_curves_v2.dynasty_lifetime_value_v2 (relative-factor formula) for QB position only; other positions still use v1 absolute or te_use_v2_ltv path.
Per-format parameter architecture. Format-sensitive redraft params (recency_weights, td_regression, shrinkage, position_recency_weights) moved under redraft.formats.{format_key}. Engine resolves format from LeagueConfig.scoring_format at runtime. half_ppr_te_premium is fully optimized (v1.3.5); other formats seeded and pending MC runs.
YoY production delta as trajectory momentum signal for RB and WR. production_delta = ppg_most_recent_full_season - ppg_prior_full_season computed in weighted_ppg_for_player() (metrics.py) from full_records sorted by season desc. Exposed as production_delta column in enrich_production_metrics(). New _apply_trajectory_momentum_bonus() method in rankings_engine.py — parallel to _apply_young_qb_ceiling(), applied immediately after. Eligible: RB and WR, ascending/early_peak trajectory, ≥4.0 weighted_ppg, ≥2 full seasons. Ascending (delta ≥ +1.5 ppg): bonus_pct = delta × 0.012, capped at +18% dynasty_ppg. Falling (delta ≤ -1.5 ppg): penalty_pct = |delta| × 0.008, capped at -12% dynasty_ppg. Config-driven in trajectory_momentum block. momentum_bonus_applied column added for accuracy_review tracking.
Position-specific recency weights. Bug fix: global recency_weights from engine_config.json were NOT flowing to weighted_ppg_for_player() in metrics.py — hardcoded module constant (3.0/2.0/1.0) was used instead of Monte Carlo-optimized values (2.966/0.9/0.279). Fixed by threading recency_weights param through weighted_ppg_for_player() -> enrich_production_metrics() -> build_player_rankings_base() -> _build_base() in rankings_engine.py. position_recency_weights added to engine_config.json: RB (4.5/0.65/0.15 — heavy current-season, least autocorrelated), QB (2.2/1.3/0.60 — flatter, most autocorrelated), WR (3.5/0.85/0.22 — moderate), TE (2.6/1.1/0.40 — sticky once established). Per-position override picked in enrich_production_metrics() before calling weighted_ppg_for_player(). Falls back to global weights if position absent. ModelParameters.position_recency_weights loaded from config. Pre-backtest calibration — accuracy_review.py will refine annually.
College school quality multiplier in rookie_model. Dominator rating at Alabama (P4_strong, ×1.0) ≠ NDSU (FCS, ×0.78). 5-tier system: P4_strong (1.0), P4_avg (0.95), G5 (0.87), FCS (0.78), unrated (0.90, default). Multiplier applied to raw college production stats BEFORE baseline subtraction in project_peak_ppg() — so a 35% target share at FCS becomes effectively 27.3% before the 20% baseline is subtracted, compressing the contribution from 15pp to 7.3pp. Applies to: college_target_share, college_rush_share, college_completion_pct, college_pass_ypa. Config-driven in rookie_model.school_quality. School lookup is case-insensitive; handles nflverse abbreviated formats (Penn St., Ohio St., North Dakota St.). school_name parameter added to project_peak_ppg(); get_rookie_rows() passes p.get('school'). school_quality_mult and school_quality_tier exposed in output rows.
Bayesian shrinkage on weighted_ppg for thin-sample players. Formula: w = n/(n+k), shrunk_ppg = w*player_ppg + (1-w)*position_prior. k=3.0, min_seasons_no_shrink=4. Prior = mean weighted_ppg of above-replacement players at each position. Only applied to above-replacement players (below-replacement are not shrunk toward starter mean). Rookies (seasons_included=0 or is_rookie=True) are always skipped. Key corrections: McCaffrey 21.4->16.3 (1 qualifying season / injury history), Daniels 20.9->19.0, Nabers 14.6->11.4, Bowers 15.0->13.0, Gibbs 18.5->14.9. 81 above-replacement players affected out of 825. Mean delta: +0.32 ppg. raw_weighted_ppg column preserved for transparency.
Rookie development ramp in LTV calculation. Adds per-year additional multipliers applied ON TOP OF the age curve factor for is_rookie=True rows in dynasty_lifetime_value(). Corrects survivor bias: age-21 WR factor (0.72) was derived from all 21-year-olds including experienced ones; true year-1 rookies produce ~55-65% of career peak (implied factor ~0.61), not 72%. Ramps: WR [0.85, 0.92, 1.0], TE [0.80, 0.88, 0.95, 1.0], RB [0.92, 0.97, 1.0], QB [0.88, 0.94, 1.0]. Config-driven in rookie_development_ramp block. Implementation: _load_development_ramps() in age_curves.py; dynasty_lifetime_value() accepts development_ramp parameter; _ltv_row() in enrich_age_curves() detects is_rookie and passes position ramp. No effect on non-rookie players.
Injury type granularity: structural ceiling discount applied to weighted_ppg BEFORE LTV calculation. Distinct from existing injury_discount (post-LTV availability factor). Structural injuries (seasons with <5 games = ACL proxy) now permanently bend the peak ceiling, propagating through the full 7-season LTV projection. New InjuryProfile fields: structural_ceiling_factor, n_structural_injuries, healthy_seasons_since_structural. Formula: base_ceiling[pos] - age_penalty(age_at_injury - cliff_age) + recovery_credit(healthy_seasons_since) × compound_multiplier^(n_structural-1). Config block: injury_model.structural_injury in engine_config.json. Pipeline: rankings_engine._apply_structural_ceiling() runs after rookie injection, before enrich_age_curves(). Availability discount (injury_model.alpha/floor) unchanged and still applied post-LTV.
Superflex QB dynasty corrections. (1) young_qb_ceiling.min_weighted_ppg: 10.0 → 4.0 — was blocking ascending QBs with limited samples (Maye 7.49, Williams 8.24, Stroud 8.11). (2) young_qb_ceiling.bonus_per_year_to_peak: 0.0362 → 0.18 — previous value produced ~11% max bonus for a 23yo QB; new value produces ~30-55% for ascending QBs 3 years from peak. (3) young_qb_ceiling.cap: 1.1 → 1.55 — allows meaningful upside for franchise QBs with thin sample. (4) superflex_qb_dynasty_premium multiplier: 1.12 — applies to all QB dynasty_ppg in superflex dynasty leagues to reflect structural scarcity (20 starter slots vs ~16 reliable starters). Known remaining issue: 30-32yo QBs (Mayfield, Goff, Murray) still overranked relative to PP due to flat QB age curve 31-34 in engine_config; will partially self-correct with 2025 data load. (5) young_qb_ceiling.archetypes_eligible: added dual_threat — Maye and Williams were miscategorized as dual_threat based on rookie scrambling stats, blocking their ceiling bonus. Trajectory filter (ascending/early_peak) already prevents veteran dual-threat QBs like Jackson from qualifying. (6) young_qb_ceiling.trajectories_eligible: [ascending, early_peak] → [ascending] only. early_peak included established starters (Daniels 25.7, Purdy 26.7, Nix 26.5) who were being over-boosted. Developmental ceiling bonus should only apply to true ascending QBs with thin samples.
Reference season advanced to 2025 after full 2025 season data loaded via Sleeper fetch. Recency weight index 0 now maps to 2025. Former-starter discount now correctly identifies QBs who backed up in 2025. ACTIVE_THRESHOLD unchanged at 2024 (broad pool). Backtest re-run and Monte Carlo optimization pending.
Backtest-driven corrections using 2019 reference year (2015–2019 training, 2020–2024 evaluation, n=320 players). Four targeted fixes: (1) RB age curve recalibrated at ages 26–33 — original cliff was too steep, causing systematic -1.9 to -3.8 ppg underranking of 26–30 year old RBs; ages 26/27/28/29 factors raised from 0.94/0.83/0.70/0.55 to 0.97/0.92/0.84/0.72. (2) RB peak window extended from [22,26] to [22,28] — data shows productive prime runs through age 28. (3) QB pocket_passer post_peak_modifier raised from 1.10 to 1.15 — veteran pocket QBs (Brady, Brees, Rodgers, Roethlisberger) were underranked because age cliff was too aggressive after 35; base QB curve also flattened at 36–39. (4) confidence_flags.age_uncertainty_rb_over lowered from 28 to 27 — flagging uncertainty one year earlier to match actual production cliff. Overall backtest: r=0.769 overall, RB bias improved by recalibration. WR model was most accurate (r=0.754, bias +0.18). QB had highest variance (stdev 4.92) driven by backup QBs with inflated starter-season wPPG.
Initial parameters. Baseline from POC build. All values theory-grounded, not yet back-tested against actual outcomes. Includes: young_qb_ceiling block (ascending pocket QB bonus), trade_model block (trade analyzer constants).
This page is server-rendered from engine_config.json at request time.
The version history is append-only and version-controlled. To audit historical changes:
git log -- engine_config.json.