LegDay
KickersDSTQBRBWRTEAccuracyDraft Kit
 Go Premium
HOW WE WORK — WHAT WE TRIED, WHAT WORKED, WHAT DIDN'T

How the sausage is made

This page explains how Leg Day builds and tests its models, in plain language, including the parts that didn't work. We won't hand over the exact recipe — the specific signals and settings are the product — but everything about the process, the testing, and the honest performance numbers is public. The numbers themselves live on the Track Record page; this is the story of where they come from.

The problem we actually solve

Not point projections
Most sites publish point projections — "14.3 points" — a number that looks precise and almost never is. Weekly NFL fantasy scoring is violently random: the same player in the same role can score 4 or 24 depending on one broken tackle or one red-zone call. Pretending otherwise is false precision, so we don't do it.
Boom / bust probabilities instead
We publish the probability that a player has a boom week (a ceiling game that wins you a matchup) or a bust week (a game that sinks one), and we stamp each player BOOM, MID, or BUST when the probability clears our confidence bar. Years of our own testing say this is where the real, measurable skill is — tier probabilities and avoid-calls, not decimal-point projections.
POSBOOM = PTSBUST = PTSSCORING FORMATWHO MAKES THE BOARD
QB25+under 13PPR (no receptions — identical in every format)6-week scoring form ≥ 15/gm
RB18+under 6half-PPR6-week scoring form ≥ 8/gm
WR16+under 5half-PPR6-week scoring form ≥ 7/gm
TE11+under 3half-PPR6-week scoring form ≥ 4/gm
K11+under 7standard kicker scoring3+ career games
DST11+under 4standard DST scoringall 32 units
These boundaries are the product definition, printed on every board, and they were stress-tested: shifting any of them by a few points doesn't change which models win. The "who makes the board" column is our startable-pool rule — we only grade calls on players with startable recent production, and we disclose what that costs (see "What we don't do," below).

What goes into the models

Recent usage and role
How a player is actually being used: snaps, touches, routes, red-zone work, share of the offense, and where they sit on the depth chart — measured over recent weeks, not preseason narratives.
Betting-market context
Point spreads, game totals, and what they imply about each team's expected scoring. Las Vegas prices in a lot of information, and our testing says ignoring it is malpractice — several of our models lean on it heavily.
Opponent tendencies
How the opposing defense has treated this position lately: points allowed to the position, pass-rush and takeaway profiles, pace. Matchups matter, and they're measurable.
Venue, weather, and schedule
Indoor or outdoor, forecast conditions, home or road, divisional games, time of season. These matter most for kickers — a gusty open-air stadium is a different job than a dome.
Expert consensus
The weekly aggregated expert consensus ranking is one input among many. We tested it honestly: it's genuinely hard to beat at raw weekly ranking (experts read injury news; our features don't), so where it earns a seat, it gets one.
What survives the screening
Every candidate signal must prove itself on seasons the model never trained on, then survive a redundancy screen that throws out signals that just restate each other. Most candidates don't make it. The exact lists that do are the product — that part stays home.

How we test — the part that actually matters

Rolling multi-season backtests
Every model is evaluated the same way: train on every season up to year X, predict year X blind, then roll forward and repeat — five separate test seasons (2020–2024), always training only on the past. A model never gets graded on a season it has seen. This is the boring, honest way to test, and it's slower and more humbling than the alternative, which is why most people don't do it. There's a visual walkthrough on the Track Record page.
A season kept in a vault: the 2025 holdout
While all of that model selection was happening, the 2025 season sat untouched — never used for training, never peeked at for tuning. Only after every choice was locked did we score 2025, exactly once. No re-runs, no "just checking one more variant." Whatever came out is what we published, and it's the table below.
Pre-registered, in writing, before the test
Here's the part we're proudest of: every modeling choice for the 2026 season — models, inputs, tier boundaries, publish rules, cadence — was committed to a dated pre-registration document in our git history before the 2025 final test was ever scored. We also wrote down, in advance, the range each number was expected to land in — including an honest correction for the optimism that creeps in whenever you pick winners from a leaderboard. The rule was written down too: a bad final number ships with commentary, not with a quiet model change. One position came in below its expected range (QB — details below), and it shipped exactly as scored.

What we tried — and what we learned

Before freezing the 2026 lineup we raced well over a hundred candidate configurations across every position: regularized linear models, gradient-boosted trees under several objectives, Monte-Carlo game simulations that play out a kicker's day leg by leg, calibrated probability classifiers, ordinal models, multi-stage classifier decompositions, and blends of our models with the expert consensus. Every candidate ran through the identical rolling backtest on identical player-weeks. The shipped lineup is drawn from those families — and we simply chose the configurations we ship.
WHAT WE LEARNED THE EASY WAY
  • Simple, heavily-regularized models are stubbornly hard to beat on NFL-weekly data. Boring wins.
  • Betting-market context carries real signal at every position — and our usage-and-form features add real information beyond the lines at the flex positions.
  • Game simulations earned their keep for kickers, where scoring decomposes into parts you can actually model.
  • Blending with the expert consensus earned a seat where the consensus is genuinely strong.
WHAT WE LEARNED THE HARD WAY
  • Most "improvements" were noise. With five test seasons, small leaderboard gaps are ties — and we treat ties as ties instead of declaring victories.
  • Fancier models mostly bought overfitting, not accuracy. Boosted trees rarely justified their flexibility at this data size.
  • Splitting the three-way call into separate yes/no classifiers looked clever and collapsed in testing.
  • Adding injury-report signals moved nothing and made one position worse. Not adopted.
  • Predicting the season-long best kicker before Week 1 is close to signal-free. That failure is why this is a weekly product.
The honest conclusion after all of it: the edge is not one magic model. The edge is the discipline — testing everything the same brutal way, refusing to promote noise, writing choices down before the answer key arrives, and publishing whatever comes out.

The metrics, in plain English

Precision — when we make a call, does it hit?
Boom precision answers: of all the players we stamp BOOM, what share actually boom? Judge it against guessing. Roughly 1 in 5 startable players booms in a given week, so blind guessing runs about 20%. Our BOOM stamps landed 28–33% at most positions in 2025 — closer to one in three. Bust precision works the same way against a higher base rate (roughly 30–45% of startable players bust).
Recall — how many did we catch?
Recall answers the opposite question: of all the booms that actually happened, what share did we flag in advance? At TE in 2025, 60% of actual boom games were sitting in our BOOM band before kickoff; on the avoid side, 71% of actual TE busts were flagged. Precision and recall pull against each other — stamp everyone BOOM and you catch every boom while your precision collapses to the base rate.
F1 — the two combined, cheat-proof
F1 squeezes precision and recall into one 0-to-1 number (their harmonic mean). It's designed so you can't game it: flag everyone and precision drags it down; flag almost no one and recall drags it down. There's no fixed "good" F1 — it depends on how rare the event is. For events with a ~20% base rate, the 0.36–0.42 our boom callers posted in 2025 represents real, usable edge; 0.29 (our QB number) is thin, and we say so.
One more honesty note: the MID stamp is not a prediction of an average game. It's the model declining to make a call — neither probability cleared the bar, and we'd rather say that than fake a lean.

The 2025 final test — scored once, published as-is

BOOM CALLSBUST CALLS (THE AVOID SIDE)
POSF1PRECISIONRECALLF1PRECISIONRECALL
QB0.2923%40%0.4942%59%
RB0.3931%52%0.4533%67%
WR0.3628%53%0.4942%58%
TE0.4132%60%0.5343%71%
K0.3833%44%0.4743%51%
DST0.4231%65%0.5753%62%
Held-out 2025 season, every regular-season week, scored one time under the pre-registered configuration. Read precision against the base rates: booms run roughly 20% among startable players, busts roughly 30–45%. Nine of our twelve pre-registered position-tracks landed within or above the range we wrote down in advance. QB is the miss: it came in below its pre-registered range on both calling and ranking, in a season where QB booms were scarcer league-wide than in any of our training years. It's our weakest product, the number shipped as scored, and it stays published. Full record, weekly updates, and consensus comparisons live on the Track Record page.

What we don't do

No injury-news wizardry claims
Our models don't read injury reports or scrape beat-writer tweets. We tested injury-report signals; they didn't help, so they're not in. Where news matters — raw weekly ranking, where human experts shine — we say the consensus is hard to beat, because our tests say so.
No decimal-point projections
We won't tell you a player is 'projected for 17.3.' Weekly outcomes don't support that resolution, and we round our published probabilities to whole percentage points for the same reason: anything finer would be noise dressed up as knowledge.
No memory-holed misses
Every graded call stays on the record — the QB miss above, the near-useless preseason kicker prediction, all of it. If a number is bad, it ships with commentary, never with a quiet re-run. The 2025 test was scored once and can never be scored again.
No pretend coverage
The startable-pool rule means some booms happen off our board — breakout rookies and sudden role-changers before they have the recent production to qualify. Historically that's meaningful (around one in six RB booms in a recent season). We publish the rule and a watch list instead of pretending to coverage we never tested.

The weekly rhythm

TUESDAY · 5:00 AM PTNot graded
Preview board
Full boards for the coming week, built for waiver-wire and free-agency decisions, plus the watch list of off-board risers. Same models, Tuesday-vintage inputs. Labeled a preview because that's what it is.
SATURDAY · 5:00 AM PTGraded — this is the record
Final graded board
The models re-run on Saturday-vintage inputs: fresh betting lines, updated weather forecasts, the week's expert-consensus pull. Every call on this board is graded against what actually happened and rolls into the track record.
THURSDAY RULEDisclosed on every board
The freeze
Teams that play Thursday night are frozen at their Tuesday values on the Saturday board — their game already happened, and re-scoring them with information from after kickoff would be grading a different product. Frozen rows are visibly marked, and it's the Tuesday-vintage call that gets graded.
SUNDAY · optionalCommentary only
Late-news notes
If late scratches or weather move a slate, we may publish notes. They never edit a tier, a probability, or a rank on the graded Saturday board. Once graded, graded.
🔒
Why we don't publish the recipe: the exact signals, settings, and tuning behind the models are the product, so they stay private. Everything else — the process, the testing protocol, the pre-registered expectations, and every graded result — is public, because a prediction product you can't audit is just a guy with a newsletter. See the full track record →
Predictions are probabilistic, not guarantees. Kicker scoring has inherent randomness — we publish our misses. Not affiliated with the NFL. Not gambling advice.
RankingsAccuracyMethodologyDraft KitPremiumPrivacyTerms© 2026 Leg Day