Local vs Frontier Gap Meter

Each dated snapshot measures the distance between the best model that fits on a home GPU and the frontier APIs, on the same benchmark. A falling line means the gap is closing.

Coding

No snapshots for this use case yet. Cappy is still wading out with the measuring stick.

General reasoning

No snapshots for this use case yet. Cappy is still wading out with the measuring stick.

Agents

No snapshots for this use case yet. Cappy is still wading out with the measuring stick.

Methodology

Only contamination-resistant evals count here: LiveBench, LiveCodeBench, and SWE-bench Pro. The gap is the frontier baseline's score minus the best catalog model's score on the same benchmark, in that benchmark's own points. Higher line, wider gap. A model qualifies for a VRAM tier when its recommended (or minimum) VRAM fits the tier and it is local-capable. Snapshots are computed once a month and archived below. Past rows are never rewritten except by an explicit recompute.

Frontier baselines are reference data points, not catalog entries. They are entered by hand from the linked leaderboards. They are not scraped at page load. Each shows its source and observation date.

Frontier baselines (non-catalog reference points)

No frontier baselines entered yet.

Monthly snapshot archive

No snapshots yet. When a measurement is taken, it lands here with its date.

← Back to catalog