Skip to main content

FFT Methodology

How this works, and where it is weak

FFT projects fantasy production from measured on-field opportunity and refuses to look at where players are being drafted. This page is the full account: the inputs, the out-of-sample test, the exact command to re-run it, and every position and cut where the model loses to a baseline that a spreadsheet could produce.

The losses are not an appendix. Publishing them is the entire reason to trust the wins.

FFT model (beta)fft-projection-2.3.0-betabacktest run 2026-08-21

Which model version is which, and why the site quotes more than one

The model shipping right now is fft-projection-2.3.0-beta (shipped 2026-08-21). Every projection, every rank and the live board at /api/edges come from it, and the accuracy figures on this page were measured against fft-projection-2.3.0-beta — the same one, because the holdout was re-run rather than the stamp moved. Where the prose below writes “v2.2.0-beta” it means the tail of that full identifier.

The call record will name older identifiers, and that is deliberate rather than a mismatch. A pre-registered call is stamped with the model that made it and is never restamped. If retiring a version silently re-attributed its calls to the successor, a new model would inherit an old model's record — so every identifier below stays resolvable permanently, and a call registered under fft-projection-1.0.0-beta keeps saying so long after that model stopped shipping.

IdentifierShippedSupersededWhat changed
fft-projection-1.0.0-beta2026-08-072026-08-08First market-independent opportunity model. Per-game volume shrunk toward a single flat positional mean; rookies from players.csv draft capital only.
fft-projection-2.0.0-beta2026-08-082026-08-08v2: per-game volume regresses toward a SNAP-SHARE-CONDITIONAL role prior instead of a flat positional mean (fixes the RB error, which was over-projecting low-snap backs and under-projecting high-snap ones); draft capital may arrive from a second source and 'undrafted' is now distinguished from 'draft capital unknown'.
fft-projection-2.1.0-beta2026-08-082026-08-11Scoring moved onto the FantasyPros standard (-1 per interception, lost fumbles and two-point conversions now scored) so the board is comparable with other evaluated rankers; and the rookie games-played priors were put on the same CONDITIONAL footing as the rookie per-game rates, which were already fitted on rookies who played. Multiplying a conditional rate by an unconditional games count described no coherent population and under-projected rookies who go on to play by a mean 9.73 PPR; that is now 4.21 and no longer statistically distinguishable from zero.
fft-projection-2.2.0-beta2026-08-112026-08-21Games-played projection no longer reorders QB and TE. Last season's games count is ~random year over year, but the persistence term let it dominate those two tightly-clustered positions, burying durable starters who missed time. Expected games for QB/TE now default to a healthy-starter baseline with that persistence zeroed; an active-injury designation still applies its own separate multiplier. QB rush-volume also leans harder on prior-year carries, a sticky role trait, though a two-season history window bounds how far that reaches.
fft-projection-2.3.0-beta2026-08-21still shippingInterceptions now cost 2 points, not 1, in every scoring format. The 1-point charge was adopted in 2.1.0-beta to keep the board comparable with one outside ranker's published rules; that is not what real leagues play, and a cross-check against an authoritative scoring-rules feed found it was the ONLY rule out of step — every other stat already matched to the point. Comparability with an outside ranker was the wrong thing to optimise a user-facing board for, so the board now scores what the user's league scores. The backtest's grading rule moved in the same commit, because a projection built at one interception price and graded at another measures the scoring rule rather than the model; the published /methodology numbers were re-run under the new rule and are not comparable with the ones this page carried before today. Only quarterbacks move.

The edge engine at /api/edges carries a separate fft-edges-* version for the assembly around the model. It moves on its own clock, so it will not match the model version and is not meant to.

01

Why this model exists

FFT previously shipped a projected-points column and a rank column, and built its headline feature — comparing our rank against draft-market ADP to surface value — on top of them. That premise was tested directly against the shipped dataset of 335 players. It failed.

Pairwise concordance in the original shipped dataset (335 players)
ComparisonComparable pairsConcordantInversions
rank vs projected points55,945100.00%0
rank vs average ADP55,88097.26%1,533
rank vs ADP55,90693.38%3,703

Read the first row carefully. Across every one of the 55,945 possible player pairs, sorting by projected points reproduced the rank order exactly — not approximately, exactly. A projection that never once disagrees with a ranking is not a projection; it is the ranking, rescaled.

And the rank was not independent either: it tracked the draft market at 93 to 97 percent concordance, because it was in substance the row order of the ADP table the dataset had been parsed from. The consensus engine found the same thing from the other direction — across the 216 players every ADP source ranked, 23,191 pairs, the model ordered the pool exactly as the market priced it, with zero inversions.

What that meant for the product

Every “our model versus the market” feature was structurally incapable of producing a signal. You cannot disagree with the market using a number copied from the market. The divergence column returning 0.0 / ALIGNED for every comparable player was not a bug — it was an honest instrument reporting that there was nothing to report.

This model exists to fix that at the root: to build a projection out of measured on-field inputs, so that when it disagrees with ADP the disagreement carries actual information.

02

What the model refuses to read

The model must not consume ADP, or anything derived from it, anywhere. The forbidden inputs, by exact field name:

adp   rank   avgAdp   yahooAdp   sleeperAdp   rtSportsAdp   realTimeAdp   proj_fpts   pbt_val

The prohibition is total, and it covers every stage that could smuggle the market back in. Not as a feature. Not as a tie-break, because ordering two equal projections by ADP re-imports the market ranking through the back door. Not as a filter, because projecting only “the top 200 by ADP” makes the player pool itself a market artifact. Not as a calibration target, because tuning the output until it looks right against ADP is precisely the failure being fixed. And not as a blend — a 70/30 model-market blend is 30 percent market, and the result cannot honestly be called independent.

Partial independence is not a weaker form of independence. It is a projection that agrees with the market by construction and therefore cannot tell you when the market is wrong. The engine module has no value imports at all; its only import is a type-only one, erased at compile time.

Where ADP is still used, legitimately

ADP has not been deleted from the app. It is a real measurement of where players are being drafted and the correct answer to “when will he be gone?” The change is that ADP now sits on exactly one side of the comparison instead of both.

03

What the model consumes

Every input is a measurement of something that physically happened on a football field, or a currently-true roster fact.

  • nflverse season statistics

    Games, pass attempts and completions, passing yards and touchdowns, interceptions, carries, rushing yards and touchdowns, targets, receptions, receiving yards and touchdowns. This is the volume spine. Fantasy points appear only as backtest ground truth, never as an input.

  • nflverse snap counts

    Offensive snap share. It does not change a projected rate — it changes how far that rate is regressed, so a player at 90 percent of snaps is treated as a slightly larger sample than one at 25 percent.

  • nflverse player bio

    Birth date for age at the projected season, and draft round and pick as the market-independent proxy for expected rookie opportunity. Draft position is a team's own valuation, made before any fantasy ADP existed.

  • nflverse weekly rosters

    A second, earlier source for the same draft capital. players.csv fills a class in late — for 2026 it carried a pick for 1 rookie of 233 — while the weekly roster file carries an overall pick for 80 of them, and the round is inferred from the measured round boundaries of the 2015-2026 classes. A rookie no source can place is labelled unknown and flagged as a positional floor, not quietly called undrafted.

  • Sleeper live overlay

    Current team, depth-chart position and order, injury designation. Last season's stats describe last season's situation; this is how the model learns the situation changed.

  • Red-zone opportunity

    Red-zone targets and carries, derived by FFT from openly-licensed play-by-play. Kept because it is correctly joined and gives a claim its context — not because it was shown to help. The measurement is below.

The method itself is the standard defensible one: fantasy production is mostly opportunity, and volume is far more stable year over year than efficiency. So the model projects availability, then per-game opportunity, then applies an efficiency rate pulled toward the positional mean in proportion to how little evidence supports it. As of v2.2.0-beta, quarterback and tight-end availability defaults to a healthy-starter games baseline rather than carrying last season’s games count forward: games played is close to unpredictable year over year, and at those two positions letting it carry forward buried durable starters who happened to miss time. A current injury designation still discounts a player separately. Touchdowns are projected from a per-opportunity conversion rate rather than carried forward as a count, and that rate is regressed harder than anything else in the model, because finishing rate is the least persistent thing a player does.

Deliberately not consumed

Target share, air-yards share and WOPR are all present in the source files and are all genuine role signals. The model does not read them. Wiring them in is a real next step, and until it is done this page will not imply otherwise. There is also no experience curve, no offensive-line quality, no coaching-change adjustment, and no kicker or team-defense projection at all.

04

The backtest

Train on data available through 2024. Predict 2025. Score against what actually happened in 2025. The evaluation season appears on exactly one side of the comparison — as ground truth — and never as an input.

Leakage control is explicit. Nothing from 2025 enters the projection path; the positional priors are recomputed from 2024 inside the model itself. The shrinkage constants were fitted on 2023 to 2024 and the age curve and rookie priors on 2021 to 2024, with 2025 held out. The Sleeper live overlay is deliberately excluded, because it describes today's depth charts and injuries rather than the 2025 preseason, and feeding it in would be time travel.

The model is scored against two baselines, and the first one is the hard one:

  • Last season repeated — a player’s 2024 PPR total as his 2025 projection. This is what most casual ranking effectively is, and it is genuinely hard to beat.
  • Positional average — every player at a position gets that position’s mean. The no-information floor.

Beating these baselines is not beating the market

The harness never loads ADP, because the model is forbidden from touching it. Every result on this page measures FFT against naive statistical baselines. No measurement here supports any claim about beating the draft market, and nobody should read one into it. The place where that claim gets tested is the public call record, which can come back against us.

The scored population is every player with at least six games in 2025: 442 of them. The model produced a projection for 422, which is 95.5 percent coverage. The 20 it missed entirely are scored as 0.0 rather than excluded, because a coverage failure is a failure and hiding it would flatter the average.

05

Re-run it yourself

The backtest is one command with no arguments. It needs no ML library, no new dependencies, and no paid data — it is arithmetic and documented heuristics, which is what makes it auditable.

$ node scripts/backtest-projections.mjs

Model version
fft-projection-2.3.0-beta
Figures on this page measured against
fft-projection-2.3.0-beta
Training window
nflverse seasons through 2024
Held-out season
2025
Run transcribed here
2026-08-21
Scoring
PPR season totals, plus a separate per-game cut

It prints the full report to stdout and writes a copy to node_modules/.cache/fft-backtest/backtest-latest.txt, which is the file the model API serves as its backtest summary. Output is deterministic: two runs on the same inputs produce identical numbers, and if they do not, that is a defect. The figures on this page moved from the previous transcription because the model changed, not because the harness did — the same harness was run against both model versions on the same day and the same inputs, which is what makes the before-and-after comparisons below legitimate.

The full engineering write-up, including the parts too detailed for this page, lives in the repository at docs/methodology.md alongside docs/projection-model.md. The harness itself is scripts/backtest-projections.mjs and the engine is src/utils/projectionModel.ts.

06

Results

Spearman is rank correlation against actual finish, which is usually what a drafter cares about. MAE is mean absolute error in PPR points. Higher Spearman is better; lower MAE is better.

All players with at least 6 games in 2025 (n = 442)
PredictorSpearmanMAERMSEBias
FFT model (beta)0.80234.6748.110.72
Baseline: last season repeated0.61044.6165.67-8.95
Baseline: positional average0.28857.6572.71-0.00

This is the headline, and on its own it is misleading. Read the next table before quoting it.

Most of that margin is coverage, not skill

85 of the 442 scored players — 19.2 percent — had no 2024 stat line, and “last season repeated” is structurally forced to answer 0.0 for every one of them. Being able to project a rookie at all is a coverage advantage, not evidence of a better projection. The ablation confirms it from the other side: removing rookies costs 0.172 Spearman, which is nearly the whole headline margin.

Veterans only — players with a prior-season stat line (n = 357). The honest comparison.
PredictorSpearmanMAERMSE
FFT model (beta)0.84035.2149.48
Baseline: last season repeated0.79339.9759.51
Baseline: positional average0.32359.0375.20

+0.047 Spearman and -4.76 MAE. Real, and still modest against a baseline a spreadsheet could produce. The report's own stated noise band at this sample is ±0.02, so the rank-order win clears it. Scored under this app's own half-PPR rules — half a point per reception, two points per interception, lost fumbles and two-point conversions counted. The interception price CHANGED on 2026-08-21 (it was one point, matching FantasyPros); every figure on this page was re-run under the new rule, and figures published here before that date are not comparable with these.

Veterans only, by position — model versus last season repeated
PosnSpear modelSpear baseMAE modelMAE baseWinner
QB360.4290.49568.8978.96mixed
RB870.8560.80437.6539.85model
WR1480.8200.77232.2836.44model
TE860.7760.69023.6929.84model

Quarterback rank order LOSES to the naive baseline on this cut — 0.429 against 0.495 — and it has lost since v2.2.0-beta, which scored 0.432 against 0.486; v2.1.0 scored 0.491 against 0.495, inside noise. That is a deliberate, disclosed trade-off rather than a regression that slipped through: v2.2 stopped the games-played projection from burying durable quarterbacks who missed time, which fixed the QB board's ordering against the market (rank correlation versus market consensus moved 0.658 to 0.831) but necessarily costs accuracy against actual outcomes, because those actuals include the very injuries the projection now declines to forecast. Two rulers, opposite directions, both reported. QB remains the one position this model does not improve on rank, and the model still wins QB on error (MAE 68.89 against 78.96). Running-back season-total error used to go to the baseline — 47.52 against 43.30 — and no longer does, but be careful with that comparison: those two numbers were computed in full PPR, and this table is half-PPR. Across eight rolling-origin folds the RB loss does not appear at all under these rules. The scoring rule was doing more of that work than the model was.

Per-game, with availability stripped out (n = 422)
PredictorSpearmanMAE (points per game)
FFT model (beta)0.8282.04
Baseline: last season repeated0.6462.80

The model is meaningfully better at production rate than at season totals. Both are equally blind to injuries, and that difference is where most of the season-total gap goes.

Touchdowns, scored on touchdowns directly (n = 422)
PredictorSpearmanMAERMSE
Model with red-zone opportunity0.7442.373.74
Model without red-zone (open data only)0.7412.363.72
Baseline: last season's touchdowns repeated0.5483.005.12

The clearest methodological win in the model, and it needs no licensed data — the open-data row is the one to quote. Projecting touchdowns from opportunity rather than from last season's count is what does the work.

The red-zone layer is not the edge, and we measured that rather than assuming it

Adding red-zone opportunity moves touchdown rank correlation by +0.002 and touchdown MAE by -0.01 — nothing, at this sample. Across the only two season-pairs available, red-zone targets per game as a predictor of next-season receiving touchdowns per game points in opposite directions, and in both windows plain target volume, already in the model, captures most of what is there. The input is retained because it is correctly joined and does no harm. No part of this product may describe it as the model’s edge.

Ablation — remove one input at a time, same population
VariantSpearmanMAEChange
Full model0.80234.67
No second season of history0.79434.90-0.008
No snap share0.78237.54-0.019
No age curve0.79236.02-0.010
No rookies0.63040.77-0.172
No red-zone data0.80234.65+0.000
No live roster overlay0.62342.12-0.179

A variant that scores BETTER than the full model is an input that is not earning its place. None does here. Snap share used to be one — removing it improved the previous model version's Spearman by 0.004 — and this page said so; it is now the largest single-input contributor after rookie coverage, because the model was changed to regress volume toward what a player's snap share implies rather than toward the flat positional mean. Anything within ±0.02 at this sample is still noise.

Confidence calibration — does the model's own confidence tier mean anything?
TiernMAEMean actualMAE / mean
High11848.59164.9329.46%
Medium19229.9558.2551.41%
Low11228.5859.4948.04%

MAE / mean is the comparable column; raw MAE is larger in the high tier simply because those players score more. High separates from medium. Low and medium do NOT separate — 48.04% against 51.41%, and the gap is the wrong way round: 'low' now reads marginally BETTER than 'medium' rather than marginally worse. That is the same non-separation this page has reported since the tiers were introduced, with the sign flipped, and a flipped sign inside a gap this size is noise rather than news. Read 'low' as 'little evidence behind this number', not as a prediction about error.

07

Where it loses

Collected in one place, at the same prominence as the wins, because this is the section that makes the rest of the page worth reading.

Quarterback rank order loses to the naive baseline

0.429 Spearman against 0.495 among veterans. This loss WIDENED in v2.2.0-beta — v2.1.0 scored 0.491 against 0.495, which was inside noise — and it widened on purpose. v2.2 stopped projecting last season's missed games forward, which is what had been burying durable quarterbacks and scrambling the QB board against the market. That fix is measured on a different instrument (rank versus market consensus, 0.658 to 0.831) and it costs accuracy against actuals, which include the injuries the model now declines to forecast. It is a real cost and it is not being presented as anything else. QB is still the one position this model does not improve on rank, though it does win QB on error (MAE 68.89 against 78.96).

The tight receiver hit-rate cut goes to the baseline, and got worse

Finding the actual top 12 inside the projection's own top 24: the model goes 7 of 12 at WR against the baseline's 9 of 12. The previous model version went 8 of 12, so this cut regressed. Running backs on the same cut went from 11 of 12 to 10 of 12 — still ahead of the baseline's 8 of 12, but down. Season-total error at WR improved substantially over the same run (42.58 to 38.90 on veterans), so this is a genuine trade, not a contradiction: the change compresses the very top of the receiver distribution slightly while pulling the middle much closer.

Most of the headline margin is coverage

On the full population the model beats the baseline by +0.186 Spearman. On veterans, where both predictors can actually see the player, it is +0.047. The difference is the fifth of the population the baseline cannot project at all.

Availability is unmodelled and is the largest error source

Mean absolute error on games played is 2.90 games. Year-over-year correlation on availability is roughly 0.20 to 0.25 — close to unpredictable. The eight worst over-projections are dominated by quarterbacks whose seasons ended early.

Rookies are a class expectation, and some are still only a floor

Of the 65 rookies actually scored in the backtest, 56 had draft capital and 9 fell to the undrafted-free-agent prior; as a class they were under-projected, bias -10.68 points. On the live 2026 inputs the model resolves draft capital for 80 of 233 skill rookies and confirms 131 more as undrafted, but 22 remain unknown — no source the model can read carries a row for them — and those 22 are flagged as positional floors rather than dressed up as valuations. A rookie number is a statement about a role, never a player-specific forecast.

The failure mode moved rather than vanishing

Regressing volume toward what a player's snap share implies fixed the running-back error, and it raises the projection of a player who held a large snap share and then lost the role. Three such players sit in this run's biggest over-projections: Brian Robinson (168.4 projected, 62.5 actual), Jerome Ford (146.2 / 43.6) and Ray-Ray McCloud (114.7 / 13.9). Net error is much lower than the previous version's; the shape of what is left is different, and losing a job is a thing this model still cannot see coming.

The low and medium confidence tiers do not separate

48.39 percent relative error for low against 48.13 percent for medium — essentially identical, and still marginally the wrong way round. Every tier's relative error improved in the current version; this weakness did not.

20 real players got no projection at all

No prior stat line and not flagged as a rookie. They include J.J. McCarthy at 125.4 PPR, Bam Knight at 92.9 and Darren Waller at 88.7. They are scored as zeroes in the report rather than quietly excluded, and they are absent from the API rather than filled in with a guess.

One season is one season

This is a single 2024-to-2025 holdout, not a multi-year study. It indicates; it does not establish. Only players with a scored season can be evaluated at all, which quietly excludes those injured out of the year — so the reported error is optimistic relative to what a drafter actually experiences.

What changed since the previous model version, including two weaknesses this page used to publish

Two entries that used to sit in the list above are gone, and they are named here rather than deleted quietly — a failures list you can edit without saying so is not a failures list. On the same harness, same inputs, same day: running-back season-total error stopped losing to the baseline — and a caution belongs with that claim. The 47.52-against-43.30 loss this page used to publish was computed in full PPR, a full point per reception. Under this app’s half-PPR rules the same fold reads 37.65 against 39.85, and across eight rolling-origin folds the running-back loss does not appear in a single one. Some of the improvement is the model; some of it is that the old number was measured with the wrong ruler. We cannot cleanly separate the two, so we are not claiming the whole gap. Separately, snap share went from not earning its place to the largest single-input contributor after rookie coverage — removing it moved Spearman by +0.004 before and by -0.019 now. One change caused both: volume is regressed toward what a player’s snap share implies rather than toward the flat positional mean, which stops the model pushing every rotational back up and every workhorse down. Rookie draft capital was the other fix, and it changed no backtest number at all, because the backtest season already had full capital coverage and the live season did not.

08

The edge signals

The projection is one half. The other half is the board of disagreements: what the model says about a player, set against what the draft market is charging for him, with the measured football reason for the difference attached. Four signals feed it, each weighted by how well it held up rather than by how interesting it sounds.

Touchdown regression

Weighted highest

Compares a player's actual prior touchdowns against the model's opportunity-based expectation, rescaled to the games he actually played, and fires past a two-touchdown gap. It is the one market inefficiency measured directly here: the top quintile of prior touchdown rate moved -0.094 per game the following season and the bottom quintile +0.094.

Efficiency

Next Gen Stats

Catch rate adjusted for target depth, plus yards after catch over expectation. Deliberately not raw separation — separation is a style metric, and scoring it directly produced a model that floored every contested-catch receiver in the league.

Usage quality

Weakest of the four

The volume-weighted EPA per target of a player's inferred route-concept mix. The concepts are inferred from where the ball was thrown, not observed from the route the receiver ran. Only about 0.12 EPA per target of the spread between players is real; the rest is sampling noise at 1.53. Even after gating on sample size, roughly a third of a surviving player's measured deviation is signal.

Opportunity

Weighted lowest

Plain target volume. Included for completeness and weighted least, because the market already prices it efficiently — everyone can see targets.

A verdict requires both halves. The signals must sit in the top or bottom quarter of the scored population and the market must be mispricing the player in the same direction, by more than eight picks. Strong signals with the market already agreeing come back ALIGNED — which is most of the board, and is the honest answer. The draft market is largely efficient, and a tool claiming an edge everywhere is not being straight with you.

A signal with no data does not fire

Missing inputs produce fewer signals, never weaker guesses, and never a zero standing in for a measurement. The committed artifact that ships to production carries red-zone opportunity for 487 of its 490 players, route concepts for 289 and Next Gen efficiency for 120 — the last of those because the upstream feed only publishes players past a volume qualifier. A player with no efficiency signal did not clear that qualifier; it is not a statement that he is inefficient.

Known bias on the board

Quarterbacks are structurally tilted on this comparison, and the cause is arithmetic rather than insight: the model ranks on raw points while a draft market prices positional replacement value, and everyone starts exactly one quarterback. Each position is centred on its own mean to compensate, which means a call reads as “the market is late on him relative to how it prices his position” — the board can no longer claim a whole position is mispriced, because nothing here has earned that claim. Rookies are absent from the board entirely: no measured prior season means no signal, and an entry with no evidence behind it is worse than no entry.

09

Known limitations

Things the model cannot see, stated here rather than buried, because someone deciding whether to trust a number needs these more than they need the method.

  • Coaching and scheme changes. A new coordinator can reallocate targets in ways no prior-season statistic anticipates. Not modelled at all.
  • Offensive line quality, and its effect on rushing efficiency. Not modelled.
  • Quarterback changes. A receiver's projection does not move when his quarterback does. A real and known gap.
  • In-season injuries. The model projects a healthy-ish games total from history and cannot know about a Week 3 hamstring.
  • Holdouts, suspensions and late roster moves beyond what the live overlay reflects at build time.
  • Team changes are handled crudely — by widening the uncertainty, which captures 'we are less sure' but not 'this specific offense throws 90 more times a year'.
  • Age curves are positional averages applied uniformly. Individual aging varies far more than the curve implies.
  • Point estimates only. The model outputs an expected value, not a distribution — no ceiling, no floor, no bust probability. Twelve points a game from a volatile player is not the same asset as twelve from a stable one.
  • Name matching across sources is fuzzy. Suffixes, apostrophes and legal-name changes cause mismatches, and a mismatch usually drops a player's history and makes him look like a rookie.
  • No kicker or team-defense projections. Neither is meaningfully opportunity-modellable with these inputs, and producing a confident-looking number anyway is exactly what this work removed.

Standing commitments

  • Everything is labelled FFT model (beta). It is a projection, not a prophecy.
  • Confidence tiers are shown next to numbers, not hidden in a tooltip. A number without its confidence is a stronger claim than the model is making.
  • Backtest results are published including the parts where the model is weak, and the table does not get edited until it looks good.
  • If the model stops beating the naive baseline, this page says so.
  • Missing data shows as missing. Never a zero, never a guess.
  • Calls are pre-registered with timestamps and graded in public, misses included.