Compare GPUs for local AI.

See what fits, how fast it may run, and which software stack each card supports. Every comparison uses the same catalog and workload assumptions.

GPUs in this comparison

Choose up to twelve cards.

  • GeForce RTX 4060
  • GeForce RTX 3060 12GB
  • GeForce RTX 4060 Ti 16GB
  • GeForce RTX 4090
  • GeForce RTX 3090
  • Radeon RX 7900 XTX
  • GeForce RTX 5090
  • RTX A6000
  • Apple M4 (16GB unified memory)
  • Apple M4 (32GB unified memory)
  • Apple M4 Max (64GB unified memory)
  • Apple M4 Max (128GB unified memory)

Quick read

Highlights

GPU test bench

What fits

Fit across 0 published catalog entries. Every row uses the same grading rules.

Runs wellRuns tightPartial offloadWill not fitCloud onlyNot graded
  1. 0 / 0

  2. 0 / 0

  3. 0 / 0

  4. 0 / 0

  5. 0 / 0

  6. 0 / 0

  7. 0 / 0

  8. 0 / 0

  9. 0 / 0

  10. 0 / 0

  11. 0 / 0

  12. 0 / 0

Runs well and runs tight stay fully on the GPU. Partial offload uses system memory and is slower.

Capacity and output

Memory versus speed

Each point is one selected GPU. Farther right means more usable memory. Higher means a faster estimate for the shared reference workload.

  1. M4 Max 128 GB96 GB, 64 tok/s est.
  2. RTX A600048 GB, 91 tok/s est.
  3. M4 Max 64 GB48 GB, 64 tok/s est.
  4. RTX 509032 GB, 211 tok/s est.
  5. RTX 409024 GB, 119 tok/s est.
  6. RTX 309024 GB, 110 tok/s est.
  7. RX 7900 XTX24 GB, 113 tok/s est.
  8. M4 32 GB24 GB, 14 tok/s est.
  9. RTX 4060 Ti16 GB, 34 tok/s est.
  10. RTX 306012 GB, 42 tok/s est.
  11. M4 16 GB12 GB, 14 tok/s est.
  12. RTX 40608 GB, 32 tok/s est.

Capacity

Usable model memory

Discrete cards use their listed VRAM. Apple unified memory reserves operating system headroom before model fit is graded.

  1. 1M4 Max 128 GB
    96 GB
  2. 2RTX A6000
    48 GB
  3. 3M4 Max 64 GB
    48 GB
  4. 4RTX 5090
    32 GB
  5. 5RTX 4090
    24 GB
  6. 6RTX 3090
    24 GB
  7. 7RX 7900 XTX
    24 GB
  8. 8M4 32 GB
    24 GB
  9. 9RTX 4060 Ti
    16 GB
  10. 10RTX 3060
    12 GB
  11. 11M4 16 GB
    12 GB
  12. 12RTX 4060
    8 GB

Estimated output

Generation speed

Estimated tokens per second for a 7B Q4_K_M model at 8,192 tokens. This is a comparison baseline, not a measured benchmark.

  1. 1RTX 5090
    211 tok/s est.
  2. 2RTX 4090
    119 tok/s est.
  3. 3RX 7900 XTX
    113 tok/s est.
  4. 4RTX 3090
    110 tok/s est.
  5. 5RTX A6000
    91 tok/s est.
  6. 6M4 Max 64 GB
    64 tok/s est.
  7. 7M4 Max 128 GB
    64 tok/s est.
  8. 8RTX 3060
    42 tok/s est.
  9. 9RTX 4060 Ti
    34 tok/s est.
  10. 10RTX 4060
    32 tok/s est.
  11. 11M4 16 GB
    14 tok/s est.
  12. 12M4 32 GB
    14 tok/s est.

Native stacks

Software support

These are platform specific compute stacks, not an overall compatibility score. Other runtimes may still support a card.

GPUVendorCUDAROCmMetal
RTX 4060NVIDIANoNo
RTX 3060NVIDIANoNo
RTX 4060 TiNVIDIANoNo
RTX 4090NVIDIANoNo
RTX 3090NVIDIANoNo
RX 7900 XTXAMDNoNo
RTX 5090NVIDIANoNo
RTX A6000NVIDIANoNo
M4 16 GBAppleNoNo
M4 32 GBAppleNoNo
M4 Max 64 GBAppleNoNo
M4 Max 128 GBAppleNoNo

Evidence boards

Models by memory tier

Model rankings open only when a tier has enough scored models and at least one verified run. Closed boards show their progress.

  • Discrete GPU

    8 GB GPU

    8 GB usable for models

    0 open

    Coding

    0 models5 needed
    0 verified1 needed

    General reasoning

    0 models5 needed
    0 verified1 needed

    Agents

    0 models5 needed
    0 verified1 needed

    Boards open when the evidence gate is met

  • Discrete GPU

    12 GB GPU

    12 GB usable for models

    0 open

    Coding

    0 models5 needed
    0 verified1 needed

    General reasoning

    0 models5 needed
    0 verified1 needed

    Agents

    0 models5 needed
    0 verified1 needed

    Boards open when the evidence gate is met

  • Discrete GPU

    16 GB GPU

    16 GB usable for models

    0 open

    Coding

    0 models5 needed
    0 verified1 needed

    General reasoning

    0 models5 needed
    0 verified1 needed

    Agents

    0 models5 needed
    0 verified1 needed

    Boards open when the evidence gate is met

  • Discrete GPU

    24 GB GPU

    24 GB usable for models

    0 open

    Coding

    0 models5 needed
    0 verified1 needed

    General reasoning

    0 models5 needed
    0 verified1 needed

    Agents

    0 models5 needed
    0 verified1 needed

    Boards open when the evidence gate is met

  • Unified memory

    Apple Silicon 16 GB

    12 GB usable for models

    0 open

    Coding

    0 models5 needed
    0 verified1 needed

    General reasoning

    0 models5 needed
    0 verified1 needed

    Agents

    0 models5 needed
    0 verified1 needed

    Boards open when the evidence gate is met

  • Unified memory

    Apple Silicon 32 GB

    24 GB usable for models

    0 open

    Coding

    0 models5 needed
    0 verified1 needed

    General reasoning

    0 models5 needed
    0 verified1 needed

    Agents

    0 models5 needed
    0 verified1 needed

    Boards open when the evidence gate is met

  • Unified memory

    Apple Silicon 64 GB

    48 GB usable for models

    0 open

    Coding

    0 models5 needed
    0 verified1 needed

    General reasoning

    0 models5 needed
    0 verified1 needed

    Agents

    0 models5 needed
    0 verified1 needed

    Boards open when the evidence gate is met

How to read this

Fit and speed are estimates derived from published model facts and GPU specifications. Measured setup counts come from accepted community reports. Check a GPU page for the full model list and calculation details.