This is the late-2026 refresh of our measured GPU guide — it updates that piece, it doesn’t replace it. The framework hasn’t changed: VRAM decides what you can run, and everything else decides how fast. What has changed is the market around it. A DRAM shortage is pushing memory prices up across new and used cards alike, and the current model cycle — MoE models like the Qwen3-Coder-30B class we benchmark constantly — has shifted the value sweet spot squarely onto the 16–24GB mid tier. Here’s every tier, with our own lab numbers where we have them and explicit spec-sheet analysis where we don’t.
What changed in late 2026
Two things, and both affect the buying decision right now.
The DRAM shortage is real and it’s pushing prices up. The late-2026 memory crunch — some corners of the community have taken to calling it the “rampocalypse” — has raised DRAM costs across the board, and that leaks into GPU and mini-PC pricing. Unified-memory machines illustrate it best: the GMKtec EVO-X2 128GB SKU sat around $3,499 in August 2026 and rising (per datahardware.ai), NVIDIA’s DGX Spark 64GB config launched at roughly $4,999 in early October while the 128GB config was pushed to around $6,950, and even Apple’s tier moved (the Mac mini M6 starts at $899 for 16GB — low-end for serious local AI, with 9B-class models estimated around 26 tok/s per llmcheck.net). The used market isn’t immune: 24GB cards are drifting up, not down. Our honest read: if you’ve done the VRAM math and know the tier you need, buying sooner rather than later is the defensible position this quarter. Prices below are bands and they move weekly.
MoE models moved the sweet spot to 16–24GB. The models people actually run in late 2026 — Qwen3-Coder-30B-A3B, Qwen3.6-27B, gpt-oss 120B in MoE form — are sparse: only a few billion parameters active per token, so they generate fast if the weights fit. Our measurements put Qwen3-Coder-30B-A3B at Q4_K_M at 18.1GB of VRAM (bm-001) — which a 16GB card cannot hold with usable context, but a 24GB card handles comfortably. That single number explains why the mid tier is where the value lives now.
The tier table, with our measured status
| Tier | Cards | Price band (moves weekly) | Our measured status |
|---|---|---|---|
| 12GB | RTX 3060 12GB, Arc B580 | ~$300–$330 | Spec-sheet analysis — we have not measured this hardware |
| 16GB | RTX 5070 Ti, RX 9070 XT, RX 7600 XT 16GB, used RTX 4080 | ~$330–$800 | Spec-sheet analysis — we have not measured this hardware |
| 20–24GB | RX 7900 XT, RX 7900 XTX, Arc Pro B60, used RTX 3090, used RTX 4090 | ~$450–$1,600 | Our measured backyard — see below |
| 32GB | Radeon AI PRO R9700, RTX 5090 | ~$2,000+ (5090) | R9700 measured; 5090 spec-sheet only |
| Multi-GPU | Dual RTX 5070, dual R9700 | varies | Measured — see below |
The 20–24GB tier: our measured backyard
This is the tier our lab actually runs, so these are first-party numbers, not community folklore.
- Arc Pro B60 (24GB, ~$450–$500): Qwen3-Coder-30B-A3B at Q4_K_M generated at 38.6 tok/s with 740ms time-to-first-token, 18.1GB VRAM, 245W (bm-001). The quantization sweep (bm-003) shows the trade: Q6_K drops generation to 27.4 tok/s at 22.6GB; Q8_0 saturates the card at 24.0GB and 19.8 tok/s. On small models the Arc is slower — Llama-3.1-8B managed 27.01 tok/s vs the RX 7900 XT’s 105.22 tok/s (bm-008) — but OpenVINO changes the picture: the same 30B-A3B class hit 67.95 tok/s with 363ms TTFT via OVMS (bm-012). Cheap VRAM, real caveats.
- RX 7900 XT (20GB): the same small-model cross-test had it at 269.01 tok/s generation on LFM2.5-8B (bm-007), and in the cross-machine Qwen3.6-27B baseline it ranked second overall at 25.21 tok/s (bm-013) — within 10% of machines costing far more. The RX 7900 XTX (24GB, ~$900) is the same family with more VRAM — spec-sheet analysis, we have not measured it.
- Used RTX 3090 (24GB, ~$600–$900): we don’t run one on the fleet, so its performance figures are community-reported — but the 24GB tier itself we’ve measured extensively on the B60 and 7900 XT, and the VRAM math transfers. Full buying checklist in our 3090 piece.
The 32GB tier
Radeon AI PRO R9700 (32GB) — measured. This is our fastest single card. Qwen3.6-27B with MTP speculative decoding hit 66.0 tok/s vs 31.9 tok/s with it off (bm-004), and Ornith-1.0-35B at 96K context reached 78.3 tok/s (bm-005) — though note that llama.cpp Vulkan run spilled to 34.2GB and collapsed to 27.3 tok/s, a backend/memory lesson in itself. At long context, MTP’s advantage decays: 64.46 tok/s at 8K down to 55.7 tok/s at 80K (bm-006). It’s a pro card at a pro price; the RTX 5090 (32GB, ~$2,000+) is the consumer alternative — spec-sheet analysis, we have not measured one.
Multi-GPU: what the second card actually buys
Our dual-GPU data says: buy the second card for capacity, not speed. On Victor, a dual-RTX-5070 tensor split re-measured Qwen3-Coder-30B-A3B at 172.05 tok/s — a 3.3x jump over the earlier 52.1 tok/s CUDA reference (bm-009). But in bm-013, dual R9700s managed only 26.74 tok/s vs the single 7900 XT’s 25.21 — the second card earns nothing once the model fits on the first. Victor’s dual-5070 prompt-eval (344.85 tok/s, roughly 2x the field) is the honest exception: splitting helps prompt processing even when generation doesn’t move. See single vs dual GPU.
The vendor software tax, honestly
CUDA remains the path of least resistance — every tool just works. ROCm 7.x on RDNA4 (gfx1201) is genuinely good now; Vulkan beat HIP by ~1.18x on the R9700 in bm-004. Arc is the fiddliest: Vulkan-only on Battlemage, with one honest caveat from our own data — our Arc entry in bm-013 ran at 8.41 tok/s under a Vulkan fallback, not the card’s real capability. The full breakdown is in our CUDA/ROCm/Vulkan/OpenVINO comparison.
What to buy at each budget (late 2026)
- ~$350: RX 7600 XT 16GB (
$330) — the cheapest 16GB entry; 7B–13B class with headroom. Arc B580 12GB ($300–$330) if budget is absolute. Spec-sheet analysis, not measured. - ~$500: Arc Pro B60 24GB (~$450–$500) — measured, and the cheapest new path to the 24GB sweet spot, if you accept the software tax.
- ~$800: Used RTX 3090 (
$600–$900) for 24GB + CUDA, or RTX 5070 Ti 16GB ($750–$800) / RX 9070 XT 16GB (~$600–$650) new with warranty. All spec-sheet on performance; the tier’s fit is measured. - ~$1,200: RTX 5080 16GB (
$1,200) — fast but still 16GB; or stretch to a used RTX 4090 24GB ($1,400–$1,600). Spec-sheet analysis. - $2,000+: RTX 5090 32GB (~$2,000+) — the consumer ceiling. Spec-sheet analysis; our measured 32GB reference is the R9700.
Where this data comes from
All measured numbers come from our own benchmark reports: bm-001, bm-002, bm-003, bm-004, bm-005, bm-006, bm-007, bm-008, bm-009, bm-012, and bm-013. Market pricing reflects the US market as of October 2026; bands move weekly. Check your target model against Model Fit and the VRAM lookup before buying.