CALIBRATION 001 · EVIDENCE-FIRST
What local AI actually runs on hardware you own.
You've read the "best GPU for local AI" lists. You still don't know what will actually run on the card in your cart — because nobody tested it.
We benchmark local AI on real consumer GPUs — Intel Arc, AMD RDNA, NVIDIA CUDA — and publish the configs, the raw numbers, and the runs that failed. So you buy the right card once, watch your model hit a usable speed the day it arrives, and stop wondering whether that last recommendation was bought.
BM-003
Qwen3-Coder-30B-A3B quantization sweep: Q4_K_M vs. Q6_K vs. Q8_0 on Arc B60 Pro
0.0
GEN SPEED
0
TIME TO FIRST TOK
LATEST FINDING
From the bench
OpenVINO beats Vulkan on the Arc B60 — and we were wrong about SYCL
We benchmarked OpenVINO Model Server 2026.2.1 against llama.cpp Vulkan on the same Intel Arc Pro B60. OpenVINO is ~1.76x faster on a 30B-A3B model — and it works cleanly, contradicting our own earlier note that Battlemage was Vulkan-only. Here's the data and the correction.
AMD Radeon for local AI: the ROCm reality check
Is AMD a real option for local LLMs in 2026, or still the 'it works but...' alternative? The honest answer, from running RDNA3 and RDNA4 cards in the lab: usable, genuinely good value at the 24GB tier, with one persistent caveat you need to know before you buy.
Apple Silicon and MLX for local AI
A Mac is the only machine where 'unified memory' means a 70B model fits without a discrete GPU. Is MLX on Apple Silicon a real local-AI path in 2026, or a niche? The honest answer for Mac owners — and the one thing that decides whether it's worth it.
Best local TTS models in 2026
The open-weight text-to-speech landscape finally has a real default: Kokoro-82M for almost everything, XTTS-v2 for zero-shot voice cloning, F5-TTS for maximum quality, and Piper for edge. A practical pick-by-use-case guide — with the licensing catches that decide which you can actually ship.
All dispatches →WHY THIS EXISTS
The benchmark problem
Before you trust another list, ask why you still don't have an answer. You've done the research, compared the specs, watched the videos — and you still can't tell whether that 24GB card will run your model at a usable speed. That's not a gap in your knowledge. The evidence was never published.
01
Specs lie
Hardware spec sheets don't tell you what real models do. A "24GB" card isn't equal to another 24GB card once you account for backend, driver, and memory bandwidth.
02
Benchmarks hide their work
Most "best GPU for local AI" lists are AI-summarized, undated, and never tested. Even good ones omit quantization, backend, and run-to-run variance.
03
Recommendations are bought
Affiliate and sponsorship interests are rarely disclosed. You're reading ads shaped like advice.
EVIDENCE
Latest benchmark reports
Qwen3-Coder-30B-A3B quantization sweep: Q4_K_M vs. Q6_K vs. Q8_0 on Arc B60 Pro
On 24GB Arc, Q6_K is the sweet spot: quality recovers to within 2 points of Q8_0 on HumanEval while staying 38% faster and fitting comfortably. Q8_0 barely fits and leaves no room for context. Q4_K_M is the budget pick when VRAM is tight.
Qwen3-Coder-30B-A3B: NVIDIA CUDA vs. Arc Vulkan at Q4_K_M
SUPERSEDED 2026-07-18 — see bm-009. Provisional Victor numbers in this record were ~3x too slow (real: 172 tok/s gen, 3403 tok/s prompt) and based on an incorrect single-24GB-GPU hardware spec. Retained for history only; do not cite. The original summary read: 'CUDA on the NVIDIA reference produces 52.1 tok/s vs. 38.6 tok/s on Arc Vulkan — a 35% generation-speed lead at the same 24GB VRAM tier.'
Qwen3-Coder-30B-A3B on Arc B60 Pro vs. dual Radeon AI PRO R9700
Vulkan on Intel Battlemage lands within 12% of dual-Radeon AI PRO R9700 for code generation at Q4_K_M, at roughly half the system cost. Arc is the price-performance leader for sub-$2K builds; multi-GPU AMD wins raw throughput once you accept the complexity.
THE LAB
Documented test machines
Every result is traceable to a specific machine with locked configuration. No anonymized "test rig" — these are the actual systems.
Codex Librarian
Jitori PC
- GPU
- Intel Arc B60 Pro
- VRAM
- 24 GB
- Backend
- vulkan
Marshal Alpha + Marshal Beta
Ray
- GPU
- 2× AMD Radeon AI PRO R9700
- VRAM
- 64 GB
- Backend
- vulkan
Architect + Router
EvoX2
- GPU
- AMD Radeon RX 7900 XT (gfx1100, 20GB) + Ryzen AI MAX+ 395 iGPU (gfx1151, unified memory)
- VRAM
- 20 GB
- Backend
- vulkan
NVIDIA CUDA reference (mobile)
Victor (msi-command)
- GPU
- NVIDIA GeForce RTX 5070 + RTX 5070 Ti Laptop (Blackwell sm_120)
- VRAM
- 24 GB
- Backend
- cuda
DECIDE
Pick your next move
TOOL
Model Fit
Will that 30B model actually run on your 16GB card? A transparent estimator with the math shown.
RECIPES
Tested Builds
Complete rigs by budget and workload. Parts, settings, and total cost — all dated.
PATH
Start Here
New to local AI? A 5-step path from "what do I want to do" to "here's what to buy."
FRONTIER DISPATCH
Get new benchmarks first.
One email when a significant result drops. No spam, no roundups, no "top 10" filler. Unsubscribe in one click.