LegDay
KickersDSTQBRBWRTEAccuracyDraft Kit
 Go Premium
EVERY HIT. EVERY MISS. PUBLIC.

Track Record

Every model is evaluated with rolling multi-season backtests: train on every season before year X, predict year X, repeat — never grading a model on data it has seen. The numbers below are the held-out 2025 season, scored exactly once after every model decision was frozen and written down in advance (models selected on 2020–2024 folds; 2025 untouched until then). No in-season results exist yet: grading starts Week 1 of 2026 and will be posted here as it lands, hits and misses alike. For the full story of how the models are built and tested — what we tried, what worked, and what didn't — read our Methodology.

How the backtest works

Fold 120182019202020212022202320242025
Fold 220182019202020212022202320242025
Fold 320182019202020212022202320242025
Fold 420182019202020212022202320242025
Fold 520182019202020212022202320242025
Final test20182019202020212022202320242025
Training seasonsPredicted blindFinal test — scored once, after all model decisions

2025 final test, by position

POSMODELBOOM F1RECALLPRECISIONTOP-5 HIT (CHANCE)RANK CORR.
KRanks: ridge regression, 30 features. Tiers: kick-by-kick simulator blended with a calibrated logistic model, 15 features0.3844%33%43% (~29%)0.18
DSTRanks and tiers: logistic model on the betting lines (implied total, spread, game total)0.4265%31%34% (~20%)0.34
QBRanks: ridge regression, 10 features. Tiers: calibrated logistic model, 10 features0.2940%23%22% (~17%)0.27
RBRanks: ridge regression, 10 features. Tiers: calibrated logistic model, 15 features0.3952%31%40% (~20%)0.42
WRRanks: 20% ridge model / 80% expert-consensus rank blend. Tiers: logistic model, 5 features0.3653%28%43% (~20%)0.37
TERanks: ridge regression, 10 features. Tiers: logistic model, 5 features0.4160%31%43% (~21%)0.44
Boom-tier calls (F1, recall, precision) are pooled over every 2025 regular-season week; top-5 hit and rank correlation grade the published ordering week by week. RB/WR/TE scoring is half-PPR with 4-pt passing TDs (QB is identical in every format). QB is our weakest product and we say so — 2025 was a cold year for quarterback booms (17% base rate) and our QB calls came in below the range we expected.

What to expect live: backtest vs honest estimate

POSBOOM F1 BACKTESTHONEST EST.2025 ACTUALNDCG@5 BACKTESTHONEST EST.2025 ACTUAL
K0.3650.3650.3790.5880.5770.619
DST0.3760.3620.4200.5410.5340.550
QB0.4020.3670.2940.7660.7620.716
RB0.4120.4120.3870.6430.6180.659
WR0.4030.4000.3640.5960.5910.618
TE0.4210.4120.4120.6240.6240.638
Backtest is the 2020–2024 selection average for the model we shipped. Honest estimate subtracts the optimism that comes from picking a winner on the same seasons it is scored on (measured by re-running the selection with each season held out). Live 2026 numbers should be expected to run at the honest estimate, not the backtest number, and to swing around it season to season — on 2025, 9 of the 12 position-by-track numbers landed within or above the range we had written down; QB missed on both ranking and tier calls, and WR tier calls ran slightly under. NDCG@5 grades how well the top of our weekly ordering matches who actually scored (1 = perfect, higher is better).

Do we beat the free expert consensus?

On 2025, no — and we don't lose to it either. Scored against the FantasyPros expert consensus (ECR) on identical 2025 player-weeks (16 matched weeks), paired week by week, every position is a statistical tie: no difference comes close to clearing noise. Ordering quality, NDCG@5 (x100), higher is better:
KLEG DAY63.3EXPERT ECR61.9+1.4 pts · tie (p 0.57)
DSTLEG DAY55.2EXPERT ECR54.9+0.3 pts · tie (p 0.89)
QBLEG DAY70.5EXPERT ECR72.2-1.7 pts · tie (p 0.47)
RBLEG DAY67.5EXPERT ECR68.1-0.6 pts · tie (p 0.83)
WRLEG DAY62.6EXPERT ECR63.0-0.4 pts · tie (p 0.50)
TELEG DAY63.7EXPERT ECR65.0-1.3 pts · tie (p 0.58)
Same rows, top-5 boom hit rate (x100) — start each side's five highest-ranked players every week, the share that boomed:
KLEG DAY35.0EXPERT ECR35.00.0 pts · tie (p 1.00)
DSTLEG DAY30.0EXPERT ECR32.5-2.5 pts · tie (p 0.43)
QBLEG DAY21.2EXPERT ECR21.20.0 pts · tie (p 1.00)
RBLEG DAY42.5EXPERT ECR43.8-1.3 pts · tie (p 0.72)
WRLEG DAY43.8EXPERT ECR43.80.0 pts · tie (p --)
TELEG DAY46.3EXPERT ECR48.8-2.5 pts · tie (p 0.58)
Published anyway, because that's the deal. One season is about 16 week-pairs, which is not enough to separate two rankers this close; parity with the expert consensus on raw weekly ranking is the honest description. The product is the tier calls — boom probabilities and avoid calls — graded on their own terms above and below.

Calling busts — the avoid side

POSMODELBUST F1RECALLPRECISION
KKick-by-kick simulator blended with a calibrated logistic model, 15 features0.4751%43%
DSTLogistic model on the betting lines0.5762%53%
QBCalibrated logistic model, 10 features0.4959%42%
RBCalibrated logistic model, 15 features0.4567%33%
WRLogistic model, 5 features0.4958%42%
TELogistic model, 5 features0.5371%43%
2025 final test, pooled over the season. Precision = when we flag an avoid, how often they actually busted. Judge it against the 2025 bust base rates, which ran from about 23% (RB) to 39% (K) by position: precision of 33–53% is real edge, not a guarantee. Dodging landmines is half the value of a streaming model — often the bigger half.

How live boards differ from the backtest

  • Tuesday boards are previews. They run on Tuesday-morning inputs — betting lines, weather forecasts and the expert consensus as they stood on Tuesday — so they are the actionable-for-waivers version, not the graded one.
  • Saturday boards are the record. Every in-season number on this page will come from the Saturday board, re-run on Saturday-morning lines, forecasts and consensus. Rows for teams playing on Thursday stay frozen at their Tuesday values on the Saturday board (their game already happened), are marked as such, and are graded at those Tuesday values.
  • Route-participation data is not available in season. For running backs, receivers and tight ends, the share of snaps a player ran a route on is published only after the season ends. The 2025 numbers above had it; live 2026 boards will not from Week 2 on, and will fill those inputs with typical values instead. Expect live results nearer the honest estimate than the backtest number.
  • Only pooled players are graded. A player is on the board only after showing startable recent production, so some booms happen off-board (breakout rookies, role changes). Watch-list names, previews and Sunday notes are never counted.

What the numbers mean

Recall
Of the games that actually happened — ceiling games on the boom side, disasters on the bust side — the share we flagged in advance. High recall means we rarely miss one.
Precision
When we make a call, how often it comes true. Judge it against the base rate: in 2025 booms happened 17-29% of the time by position and busts 23-39%, so precision above those numbers is real edge — not a guarantee.
F1 score
Recall and precision combined into one 0-to-1 number (their harmonic mean). It punishes cheating in either direction: flag everyone and precision collapses; flag almost no one and recall collapses. Best single number for comparing models.
Top-5 hit (chance)
Start our five highest-ranked players every week: the share that boomed. The number in parentheses is what you'd get picking five startable players at random — the gap between the two is the whole product.
Rank correlation
Line up our ordering against how players actually finished that week. 0 means our order told you nothing; 1 means it was perfect. Weekly fantasy is extremely noisy — nobody lives near 1, and anything above ~0.25 is meaningful signal.
What MID actually means
Not a prediction of an average game — it's the model declining to make a call. Neither the boom nor the bust probability clears our confidence threshold, so we say so instead of faking a lean.
⚖️
Also published honestly: our preseason "which kicker will lead the season" prediction is close to signal-free (year-over-year rank correlation ≈ 0). That's exactly why we built a weekly streaming product instead of a draft-a-kicker product.
Predictions are probabilistic, not guarantees. Kicker scoring has inherent randomness — we publish our misses. Not affiliated with the NFL. Not gambling advice.
RankingsAccuracyMethodologyDraft KitPremiumPrivacyTerms© 2026 Leg Day