How The Shallows scores models

The Shallows ranks what is worth running on the hardware people actually own. Every score is a weighted average of evidence, and vendor numbers are never the sole input. The full scoring spec lives in the repo and changes only by a human edit.

The trust ladder

RungEvidenceWeight
1Cappy-verified runs30
2Derivative activity22
3Usage rankings (OpenRouter)16
4Contamination-resistant evals13
5Crowd Elo (LMArena)11
6Downloads & likes8
7Vendor launch benchmarks0

A model with only vendor-published benchmarks gets no score at all. Scoring requires a verified run or at least two independent kinds of evidence. Contamination-resistant evals mean LiveBench, LiveCodeBench, and SWE-bench Pro. Rolling benchmarks whose questions postdate training data.

Hardware tiers

  • 8 GB GPU8 GB usable of 8 GB
  • 12 GB GPU12 GB usable of 12 GB
  • 16 GB GPU16 GB usable of 16 GB
  • 24 GB GPU24 GB usable of 24 GB
  • Apple Silicon 16 GB12 GB usable of 16 GB
  • Apple Silicon 32 GB24 GB usable of 32 GB
  • Apple Silicon 64 GB48 GB usable of 64 GB

Apple unified memory is shared with the OS, so tiers assume 75% is usable for the model. The same assumption the rest of the site makes.

When a board stays closed

A tier's board opens only with at least 5 scored models and 1 verified run behind them. Until then the board stays closed rather than ranking on thin evidence. Those numbers are minimums, not targets.

Back to The Shallows