● CALIBRATED 2026-10-03 · REC 000
Local AI Frontier

How Much VRAM Do You Actually Need? The VRAM-to-Model Lookup

VRAM is the single number that decides what you can run locally, and most lookup tables online are somebody’s guess. Ours is anchored to measurements from our four-machine lab: an 8B model at Q4_K_M used 4.6-6.8GB and ran anywhere from 27 to 269 tok/s depending on the card; a 30B MoE coder needed 18.1GB; and a Q8_0 quant saturated a 24GB card with zero context headroom. Below is the tier-by-tier lookup, the quant multiplier, and the rules of thumb the numbers actually support. When a combination matters to you, run it through Model Fit and then check whether we’ve benchmarked it.

How to read this page

“Fits” means weights plus KV cache at a practical context length, with a little headroom — not “the file is smaller than the card.” Speed classes come from our measured runs (llama.cpp-family backends, Q4_K_M unless stated). Where we haven’t measured a model class, we say so and give a rule-of-thumb estimate, labelled as such. Q4_K_M is the default throughout because our quantization testing found it’s no longer a compromise tier — see our 2026 quantization piece.

The lookup: what fits at each VRAM tier

All cells assume Q4_K_M. Footprints in quotes are measured on our fleet; estimates use our validated rule of ~0.6GB per billion parameters.

VRAM 7-9B 12-14B 27-35B 70B dense 100B+ MoE
8GB Fits — 4.6-6.8GB measured; 27-269 tok/s No (est. ~8-9GB + context) No (17-23GB measured) No No
12GB Fits easily, context headroom Fits at Q4 (est. ~8-9GB); tight at long context* No No No
16GB Fits trivially Fits comfortably* No — 27B measured 16.8-21.2GB; 30B MoE 18.1GB No No
20GB Fits Fits Tight — 27B Q4 fits (16.8GB, 25.2 tok/s measured); 30B MoE (18.1GB) leaves ~2GB; 35B no No No
24GB Fits Fits Fits — the sweet spot: 27B comfortable; 30B at Q4 (18.1GB) and Q6_K (22.6GB) fit; Q8_0 saturates; 35B Q4 (23.2GB) leaves no context room Split only (2×24GB) No
32GB Fits Fits Fits — 35B Q4 measured 23.2GB at 78.3 tok/s; 30B at Q8_0 with headroom Tight — Q3 or partial offload; Q4 wants 48GB No
64-128GB unified Fits Fits Fits Fits (est. ~42GB at Q4) — bandwidth-bound, slow Fits (est. ~72GB at Q4 for a 120B) — fits the 96GB-usable pool of a 128GB box

*12-14B: we have not measured a model in this class on our fleet. The footprint is a rule-of-thumb estimate; expect a speed class between our 8B and 27B results.

The quant multiplier

Same model, same card, three quants — our sweep of Qwen3-Coder-30B-A3B on the Arc Pro B60 at 8K context (bm-003):

Quant VRAM used Generation Prompt processing Notes
Q4_K_M 18.1GB 38.6 tok/s 412.3 tok/s baseline
Q6_K 22.6GB 27.4 tok/s 398.1 tok/s ~1.4GB headroom left on a 24GB card
Q8_0 24.0GB (+4.1GB in RAM) 19.8 tok/s 380.5 tok/s saturates 24GB — no context headroom

Q6_K bought a +25% footprint for a 29% speed penalty. Q8_0 was worse than that: it filled the card completely and still spilled 4.1GB to system RAM. If you’re shopping by VRAM tier, assume Q4_K_M and treat anything above it as a different tier entirely.

Speed classes we actually measured

Model class (Q4_K_M) Measured footprint Generation speed Hardware
8B dense (Llama-3.1-8B) 4.58GB 27.0 tok/s Arc B60 (Vulkan) — bm-008
8B dense 4.58GB 105.2 tok/s RX 7900 XT — bm-008
8B MoE (LFM2.5-8B-A1B) 4.79-6.81GB 68.8-269.0 tok/s Arc B60 / RX 7900 XT — bm-007
8B MoE on unified iGPU 4.79GB 150.2 tok/s Ryzen AI Max+ 395 — bm-011
27B dense (Qwen3.6-27B) 16.8-21.2GB 24.3-26.7 tok/s 7900 XT / dual R9700 / dual 5070 — bm-013
30B MoE (Qwen3-Coder-30B-A3B) 17.9-18.1GB 38.6-52.1 tok/s Arc B60 / RTX 5070 CUDA — bm-001, bm-002
30B MoE, dual-GPU split 19.0GB 172.1 tok/s dual RTX 5070 laptop — bm-009
35B MoE (Ornith-1.0-35B) 23.2GB 78.3 tok/s R9700 (Ollama) — bm-005
35B MoE, overflow case 34.2GB (spilled to RAM) 27.3 tok/s R9700 (llama.cpp Vulkan) — bm-005

One caveat from the 27B row: the same model ran at just 8.4 tok/s on Arc when the backend fell back to Vulkan — a backend problem, not a hardware ceiling. Check the benchmark page for your exact combination before trusting a tier.

Rules of thumb

  • Q4_K_M ≈ 0.6GB per billion parameters. Validated across our fleet: 8B → 4.6-4.8GB, 27B → 16.8GB, 30B → 18.1GB, all measured.
  • Context is a separate line item. KV cache grows with context length; a model that fits at 2K may not at 32K. Long context also costs speed: the same 27B setup generated at 64.5 tok/s at 8K context and 55.7 tok/s at 80K — 13.6% slower (bm-006).
  • MoE: total parameters set memory, active parameters set speed. The 30B-A3B (3.3B active) outpaces the dense 27B despite a similar footprint, and the 35B hit 78.3 tok/s — credit optimised expert routing (bm-005).
  • The backend can matter more than the tier. Same 35B, same card: 78.3 tok/s on Ollama versus 27.3 with a Vulkan overflow. On the same Arc B60, OpenVINO Model Server generated at 68.0 tok/s where Vulkan managed 38.6 (bm-012 vs bm-001 — caveat: different model format and context length).
  • Dual GPUs buy capacity first, speed second. For the 27B that already fit one card, dual 12GB ran 24.3 tok/s against 25.2 on a single 20GB card (bm-013). But the dual-5070 split hit 172.1 tok/s on the 30B MoE (bm-009). Don’t buy a second card assuming a 2× speedup — see single vs dual GPU.

Which card buys each tier

Tier Representative cards Price band (moves weekly)
8GB — We don’t recommend buying this tier; our budget-build guide calls it a false economy. Step up to 12GB.
12GB Arc B580, used RTX 3060 12GB B580 ~$300-330; 3060 used pricing moves weekly
16GB RX 9070 XT / RTX 5070 Ti / RTX 5080 ~$600-650 / ~$750-800 / ~$1,200
20GB RX 7900 XT pricing moves weekly; the 24GB XTX (~$900) is often the better buy
24GB Arc Pro B60 / RX 7900 XTX / used RTX 3090 / used RTX 4090 ~$450-500 / ~$900 / ~$600-900 / ~$1,400-1,600
32GB RTX 5090 ~$2,000+
64-128GB unified Mac mini M6 (16-32GB) / EVO-X2 128GB / DGX Spark 64GB / Mac Studio M5 Max 128GB $899-$1,299 / ~$3,499 and rising / ~$4,999 / $5,099 as spec’d

All prices are bands, not quotes — they move weekly, and the late-2026 memory shortage is pushing memory-heavy hardware up, not down. For the unified tier, note the trade-off we published in our mini-PC guide: a 70B fits the pool but generates at single-digit tok/s, because unified memory is bandwidth-bound. External pricing references: EVO-X2 per datahardware.ai, 96GB-usable figure per sunkcost.ai, and gpt-oss 120B decoding in the low-to-mid 30s tok/s on Spark-class hardware per runaihome.com. For card-by-card measured picks, see best GPUs for local AI, and for the underlying arithmetic see how much VRAM do I need plus our interactive will-it-fit guide.

Where this data comes from

Every footprint and speed figure above comes from a published benchmark run on our lab machines: the 30B quant sweep, the 35B backend showdown, the 27B cross-machine baseline, the 30B Arc vs CUDA runs and its dual-5070 rerun, the 8B general-work benchmark, the LFM2.5-8B real-app test, the unified-memory APU run, and the long-context decay study. Anything labelled “est.” in the tables is rule-of-thumb arithmetic, not a measurement — we mark it that way on purpose.