Model fit coverage
Runs fully on the GPU across 0 catalog entries
- 1RTX 40600 fit
- 2RTX 30600 fit
- 3RTX 4060 Ti0 fit
See what fits, how fast it may run, and which software stack each card supports. Every comparison uses the same catalog and workload assumptions.
Choose up to twelve cards.
Quick read
Among the GPUs you selected
Runs fully on the GPU across 0 catalog entries
Memory available before runtime headroom
7B Q4_K_M reference
GPU test bench
Fit across 0 published catalog entries. Every row uses the same grading rules.
0 / 0
0 / 0
0 / 0
0 / 0
0 / 0
0 / 0
0 / 0
0 / 0
0 / 0
0 / 0
0 / 0
0 / 0
Runs well and runs tight stay fully on the GPU. Partial offload uses system memory and is slower.
Capacity and output
Each point is one selected GPU. Farther right means more usable memory. Higher means a faster estimate for the shared reference workload.
Capacity
Discrete cards use their listed VRAM. Apple unified memory reserves operating system headroom before model fit is graded.
Estimated output
Estimated tokens per second for a 7B Q4_K_M model at 8,192 tokens. This is a comparison baseline, not a measured benchmark.
Native stacks
These are platform specific compute stacks, not an overall compatibility score. Other runtimes may still support a card.
| GPU | Vendor | CUDA | ROCm | Metal |
|---|---|---|---|---|
| RTX 4060 | NVIDIA | No | No | |
| RTX 3060 | NVIDIA | No | No | |
| RTX 4060 Ti | NVIDIA | No | No | |
| RTX 4090 | NVIDIA | No | No | |
| RTX 3090 | NVIDIA | No | No | |
| RX 7900 XTX | AMD | No | No | |
| RTX 5090 | NVIDIA | No | No | |
| RTX A6000 | NVIDIA | No | No | |
| M4 16 GB | Apple | No | No | |
| M4 32 GB | Apple | No | No | |
| M4 Max 64 GB | Apple | No | No | |
| M4 Max 128 GB | Apple | No | No |
Evidence boards
Model rankings open only when a tier has enough scored models and at least one verified run. Closed boards show their progress.
Discrete GPU
8 GB usable for models
Coding
General reasoning
Agents
Boards open when the evidence gate is met
Discrete GPU
12 GB usable for models
Coding
General reasoning
Agents
Boards open when the evidence gate is met
Discrete GPU
16 GB usable for models
Coding
General reasoning
Agents
Boards open when the evidence gate is met
Discrete GPU
24 GB usable for models
Coding
General reasoning
Agents
Boards open when the evidence gate is met
Unified memory
12 GB usable for models
Coding
General reasoning
Agents
Boards open when the evidence gate is met
Unified memory
24 GB usable for models
Coding
General reasoning
Agents
Boards open when the evidence gate is met
Unified memory
48 GB usable for models
Coding
General reasoning
Agents
Boards open when the evidence gate is met