Most local-AI disappointment is an expectations problem before it is a hardware problem. Someone reads a launch post, downloads a model, and expects cloud-assistant quality at cloud-assistant speed from a $500 card — then concludes local AI is overhyped when it isn’t. The reality is measurable, and we have measured it: our four-machine fleet (Intel Arc, AMD Radeon, NVIDIA Blackwell, and a unified-memory desktop APU) has run the same model classes under controlled conditions. An 8B-class model runs anywhere from 27 to 269 tok/s on our hardware depending on card and architecture; a 27B lands between 24 and 44 tok/s on every mid-tier machine we own. Here is what to actually expect in 2026.
Speed expectations by hardware class
| Hardware class | Measured workload | Generation speed | Record |
|---|---|---|---|
| Desktop iGPU, unified memory (Ryzen AI Max+ 395) | 8B MoE (LFM2.5-8B-A1B), Q4_K_M | 150.16 tok/s | bm-011 |
| Budget discrete 24GB (Arc B60) | 8B dense (Llama-3.1-8B), Q4_K_M | 27.01 tok/s | bm-008 |
| Budget discrete 24GB (Arc B60) | 8B MoE (LFM2.5-8B), Q4_K_M | 68.79 tok/s | bm-007 |
| Mid discrete 20-32GB (RX 7900 XT, dual R9700, dual RTX 5070) | 27B (Qwen3.6-27B), Q4_K_M | 24.28-26.74 tok/s | bm-013 |
| Mid discrete 24-32GB (Arc B60, R9700) | 30B MoE coder (Qwen3-Coder-30B), Q4_K_M | 38.6-43.8 tok/s | bm-001 |
| High-end single GPU (R9700 + MTP) | 27B MoE with speculative decoding | 66.0 tok/s | bm-004 |
| High-end single GPU (R9700, Ollama) | 35B MoE (Ornith-1.0-35B) | 78.3 tok/s | bm-005 |
| High-end dual GPU (dual RTX 5070, CUDA) | 30B MoE coder | 172.05 tok/s | bm-009 |
Three things this table should reset. First, architecture matters as much as the card: the same 8B class runs 27.01 tok/s as a dense model on the Arc B60, 68.79 tok/s as an MoE on the same card, and 269.01 tok/s as an MoE on the RX 7900 XT. Since the MoE shift, total parameters set the quality ceiling while active parameters set the speed — a 30B-A3B is not a 30B dense. Second, desktop iGPUs are real now: 150 tok/s on an 8B MoE from a mini-PC APU, faster than a budget discrete card runs dense 8B, with unified memory that holds models no discrete card of that price can (see unified memory; Apple silicon such as the Mac mini M6 is listed at an estimated ~26 tok/s for 9B-class per llmcheck.net — an estimate, not our measurement). Third, at the 27B tier, hardware class matters less than you’d hope: three different machine classes clustered within 10% of each other (24.28-26.74 tok/s). The cards behind these numbers span roughly $450-$500 (Arc B60) to $1,000-$1,200 (R9700) — bands move weekly; see best GPUs for local AI, measured and model fit.
Two caveats: these are single-stream numbers (multi-user concurrency is a different question we haven’t fully measured), and the records use different protocols (workload packs, llama-bench, single-prompt probes) — treat them as class expectations, not precisely cross-comparable figures.
What quantization actually costs you
Same model, same card, three quants (bm-003):
| Quant | Generation speed | VRAM | HumanEval pass@1 |
|---|---|---|---|
| Q4_K_M | 38.6 tok/s | 18.1 GB | 0.81 |
| Q6_K | 27.4 tok/s | 22.6 GB | 0.84 |
| Q8_0 | 19.8 tok/s | 24.0 GB (saturated) | 0.86 |
Going from Q4_K_M to Q8_0 costs just under half your speed (about 49%) and buys five HumanEval points. For chat and drafting, Q4_K_M is usually the right trade; for code generation the delta is real but small. The bigger trap is the Q8_0 row’s last column: at 24.0 GB the card is saturated, leaving no headroom for KV cache as context grows — “fits” is not “runs well” (how much VRAM do I need, VRAM lookup). And if your model feels dumb, check the quant first: below Q4_K_M the degradation stops being subtle (quantization without compromise).
Why your local LLM feels dumber (and when it isn’t the config)
The models behind major cloud assistants are 100B+ class — even the open-weights frontier, gpt-oss 120B, is 120B parameters. A local 27B is a different class of machine, and no amount of tuning changes that. Local wins on privacy, latency, and cost; it does not win on frontier capability (local vs cloud, fully compared).
But a measurable chunk of “dumb” is configuration, not model class. In why your local LLM feels dumber we catalog six traps, all measured: quant below Q4, context negotiated down by the runtime, backend quirks, and tier overreach — asking a 27B to run agentic workflows built for frontier models. Before you blame the hardware or conclude local AI can’t do your task, work through that list. And if you’re choosing a model rather than debugging one, start with best local LLMs 2026.
Long context: the speed you bought is not the speed you get
Context length is a hidden performance dial. On the R9700 with MTP (bm-006):
| Target context | Generation | Prompt processing | MTP acceptance |
|---|---|---|---|
| 8K | 64.46 tok/s | 762.78 tok/s | 0.660 |
| 32K | 61.44 tok/s | 836.07 tok/s | 0.694 |
| 56K | 59.26 tok/s | 685.10 tok/s | 0.738 |
| 80K | 55.70 tok/s | 579.46 tok/s | 0.752 |
Generation decayed 13.6% from 8K to 80K, and prompt processing fell about a quarter. One honesty note: the larger context targets tokenized to roughly 22.5K actual prompt tokens, so treat the labels as tiers rather than exact token counts. MTP acceptance actually rose with context (0.660 → 0.752), partially offsetting the decay — the mechanics are in speculative decoding, measured. Practically: set the context you need, not the maximum your VRAM allows, and remember that long sessions also degrade instruction-following (model fit).
Backends matter less than you think — except when they don’t
Without speculative decoding, every backend on the R9700 landed in a narrow band: 26.7 tok/s (Ollama), 27.1 (HIP), 31.9 (Vulkan, no MTP). The 2× jump to 66.0 came from MTP, not backend magic (bm-004). Don’t spend weeks chasing backends for single-digit gains (CUDA vs ROCm vs Vulkan vs OpenVINO, Ollama vs llama.cpp).
Except when the backend is broken. Stock Ollama’s Vulkan fallback on the Arc B60 ran the 30B coder at 8.41 tok/s; OpenVINO on the identical card ran a sibling model at 67.95 tok/s (bm-012) — roughly an 8× software swing on the same silicon. Against llama.cpp Vulkan’s 38.6 on the same card, OpenVINO’s 67.95 is a directional 1.76× (different model variants, so treat it as directional, not exact). And before crediting a backend, check memory fit: Ornith-35B ran 78.3 tok/s via Ollama but 27.3 via llama.cpp Vulkan — because the Vulkan run spilled 34.2 GB past the 32 GB card into system RAM (bm-005). That was a memory-fit failure, not a backend verdict.
The rule: before blaming hardware, test a second backend; before crediting a backend, confirm the model actually fit in VRAM.
The expectations, compressed
| Expectation | Measured reality |
|---|---|
| “Local feels as smart as ChatGPT” | Different model class — cloud is 100B+; local 27B-30B is the practical tier |
| “Q8 is twice as smart as Q4” | Five HumanEval points (0.81 → 0.86) for about 49% of your speed |
| “My card’s tok/s is fixed” | 64.5 → 55.7 tok/s from 8K to 80K context; backends swing up to ~8× |
| “Bigger GPU means proportionally faster” | Three machine classes within 10% on the 27B baseline (24.28-26.74 tok/s) |
Where this data comes from
Every number above traces to a measured record in our benchmark lab: bm-001, bm-003, bm-004, bm-005, bm-006, bm-007, bm-008, bm-009, bm-011, bm-012, and bm-013. Cross-machine context lives in the dual-vs-single-GPU baseline and Arc B60 vs RX 7900 XT.
Two disclosures: we have not benchmarked the RTX 3090 in our lab — our used-3090 analysis is spec-sheet and market reasoning, not measurement — and we have not measured Apple silicon either. Set the expectation first, then buy the hardware; the full archive is in benchmarks.