How the MVP Meter works
The MVP Meter does not tell you who the MVP is. It ranks quarterbacks by your definition of one: you set how much six factors matter, and the board re-sorts. This page is the whole method — what each factor measures, how the numbers become a score, and where the model is weak.
The six factors
Everything is computed from play-by-play data once a week, after the week's games are final. Nothing is modelled here that is not already modelled elsewhere on the site: EPA, expected pass rate and win probability are consumed exactly as the rest of the platform stores them.
Production
EPA per play across dropbacks (attempts, sacks and scrambles) plus designed QB runs.
His legs count. Scrambles sit with dropbacks because that is where the play started.
Weak support
The strength of the run game, defense and special teams around him — then inverted.
The one factor that runs backwards on purpose: weaker help scores higher, because he is carrying more of it.
Team Success
70% win percentage, 30% standings (division rank and conference seed).
The only factor that rewards winning directly. Division and conference games can be re-weighted inside the win-percentage half.
High Leverage
Third-down EPA per dropback and fourth-quarter/overtime one-score EPA per dropback, averaged.
Measured per DOWN. This is the snap-level version of playing well when it counts.
Clutch
Clutch Drive success rate — tie, take-the-lead and close-out drives inside the final five minutes of one-score games.
Measured per DRIVE, using the Clutch Drive Analyzer's default settings. Season totals are small, often under 15 drives.
Focal Point
Early-down pass rate over expected for his offense, across the games he led in dropbacks.
A usage measure, not a quality one. Trailing teams throw more, so this can flatter a quarterback on a losing team.
How a stat becomes a score
Each factor is measured against the other qualified quarterbacks that week, not against a fixed standard. A quarterback's raw number is converted to how far it sits from the group average, that distance is capped at 2.5 standard deviations in either direction, and the capped range is stretched onto 0–100.
So 50 is the average of the pool, not a passing grade. A 0 or a 100 means 2.5 standard deviations out or beyond — the cap exists so that one outlier week cannot swamp a factor.
Your weights are shares. They are scaled to total 100 before scoring, so setting all six to 100 is identical to setting all six to 17 — only the ratiosmatter. The composite is each factor's score multiplied by its share, added up. That means a composite near 100 is effectively unreachable: it would require being 2.5 standard deviations clear of the field in all six factors at once.
Small samples
Three factors run on thin data early in the year — third-down snaps, late-and-close snaps, and clutch drives. Through week 8, those are pulled toward the pool average in proportion to how little data a quarterback has, with the pull fading to nothing by week 8. A quarterback with four clutch drives in week 3 is therefore rated closer to average than his four drives alone would suggest.
After week 8 there is no correction at all. Clutch in particular stays thin all season — a full year is often fewer than 15 qualifying drives — which is why the board flags it. Treat a Clutch score built on single digits as noise.
Who makes the board
A quarterback qualifies with at least 20 dropbacks per team game played, and at least one dropback in the last two weeks. The second rule is what makes injured and benched quarterbacks fall off rather than sit frozen at an old ranking.
This is our cut-off, not an NFL rule — there is no eligibility criterion for the real award.
The historical comparison
Past seasons are re-scored with whatever weights you currently have set, and the panel shows who your criteria would have crowned against who actually won. Two things it is worth being precise about:
It is not an accuracy score. Agreeing with the voters is not the goal, and a low match count does not mean the model is broken — it means your criteria differ from theirs, which is the entire point of the tool. The more useful number is where the real winners land on your board on average.
It is not out-of-sample.Each season is scored against that season's own pool, and you are free to tune weights until the past looks right. That is fitting, not forecasting. The panel is there to show you what your criteria imply, not to validate them.
Known weaknesses
Stated plainly, because a method page that only lists strengths is marketing:
- Focal Point rewards trailing. Passing over expected on early downs goes up when a team is behind, so a quarterback on a bad team can score well on it. Weight it heavily and you will pull losing teams up the board.
- Weak support cannot tell “carrying a bad roster” from “being on one”. It measures the roster, not his response to it.
- Clutch is small-sample by nature. No amount of shrinkage fixes a season that contains nine qualifying drives.
- Every score is relative to that week's pool.A weak year at the position lifts everyone's numbers; scores are not comparable across seasons in absolute terms.
- Team Success rewards the team. It is included because most voters weigh it, not because we think wins are a quarterback statistic. Set it to zero if you disagree — that is a supported position here.
Where the data comes from
Play-by-play data comes from nflverse, which is open and free to use with attribution. The derived metrics — EPA, expected pass rate, win probability — are covered on our data sources page. The Clutch factor uses the same drive definitions as the Clutch Drive Analyzer, so the two tools agree by construction.