How The Shallows scores models
The Shallows ranks what is worth running on the hardware people actually own. Every score is a weighted average of evidence, and vendor numbers are never the sole input. The full scoring spec lives in the repo and changes only by a human edit.
The trust ladder
| Rung | Evidence | Weight |
|---|---|---|
| 1 | Cappy-verified runs | 30 |
| 2 | Derivative activity | 22 |
| 3 | Usage rankings (OpenRouter) | 16 |
| 4 | Contamination-resistant evals | 13 |
| 5 | Crowd Elo (LMArena) | 11 |
| 6 | Downloads & likes | 8 |
| 7 | Vendor launch benchmarks | 0 |
A model with only vendor-published benchmarks gets no score at all. Scoring requires a verified run or at least two independent kinds of evidence. Contamination-resistant evals mean LiveBench, LiveCodeBench, and SWE-bench Pro. Rolling benchmarks whose questions postdate training data.
Hardware tiers
- 8 GB GPU8 GB usable of 8 GB
- 12 GB GPU12 GB usable of 12 GB
- 16 GB GPU16 GB usable of 16 GB
- 24 GB GPU24 GB usable of 24 GB
- Apple Silicon 16 GB12 GB usable of 16 GB
- Apple Silicon 32 GB24 GB usable of 32 GB
- Apple Silicon 64 GB48 GB usable of 64 GB
Apple unified memory is shared with the OS, so tiers assume 75% is usable for the model. The same assumption the rest of the site makes.
When a board stays closed
A tier's board opens only with at least 5 scored models and 1 verified run behind them. Until then the board stays closed rather than ranking on thin evidence. Those numbers are minimums, not targets.