Last checked: 2026-10-03. This page updates monthly — the next full refresh is scheduled for early November 2026.
Most “best local LLM” lists are written from other people’s Reddit threads. This one is written from our lab. We run four machines — Intel Arc B60 Pro, dual Radeon AI PRO R9700, RX 7900 XT, and a dual-RTX-5070 laptop — and every speed figure below marked measured comes from our own benchmark runs, published in full on the benchmarks index. Where we haven’t measured a model ourselves, we say so explicitly and label it spec-sheet or community-reported rather than dressing it up as data. That distinction is the whole point of this page.
The roster at a glance
All speeds are generation tok/s at Q4_K_M unless noted. “VRAM” is the model footprint we recorded at the stated context length — your real requirement is higher once KV cache is added (see how much VRAM you need).
| Model | Class | VRAM @ Q4_K_M | Measured speed (tok/s) | Status |
|---|---|---|---|---|
| Qwen3-Coder-30B-A3B | 30B MoE, coding | 18.1GB @ 8K ctx | 38.6–52.1 single GPU; 172.0 dual-laptop | measured |
| Qwen3.6-27B-A3B (+MTP) | 27B MoE, general | 21.9GB @ 64K ctx | 31.9 without MTP; 66.0 with MTP (Vulkan) | measured |
| LFM2.5-8B-A1B | 8B MoE, fast daily driver | 4.8GB @ 2K ctx | 68.8–269.0 across our fleet | measured |
| Ornith-1.0-35B | 35B MoE, long-context | 23.2GB @ 96K ctx | 78.3 (Ollama, single R9700) | measured |
| Llama-3.1-8B-Instruct | 8B dense, general | 4.6GB @ 2K ctx | 27.0–105.2 across 2 GPUs | measured |
| gpt-oss 120B | 120B MoE, large | needs ~64GB+ unified | ~30 (community-reported, unverified) | community-reported |
| Qwen3.8 27B | 27B, current cycle | not yet measured by us | — | spec-sheet |
| Z-Image / FLUX.2-klein | image/media models | varies | — | community-reported |
Qwen3-Coder-30B-A3B — the coding daily driver
Who it’s for: anyone whose main workload is code generation on a single 24GB-class GPU.
This is the most-measured model on our site. On the Arc B60 Pro it generates at 38.6 tok/s and processes prompts at 412.3 tok/s (bm-001); on NVIDIA CUDA the same model hits 52.1 tok/s (bm-002). On Victor’s dual-RTX-5070 laptop with a tensor split, it reaches 172.05 tok/s (bm-009) — the fastest single result we’ve recorded for this model. It fits in 18.1GB at 8K context, so a single 24GB card runs it with headroom, and our quantization sweep shows Q6_K still fits at 22.6GB if you want the extra quality.
Qwen3.6-27B-A3B with MTP — the general-purpose pick
Who it’s for: general chat, writing, and agent work where you want one model that does everything reasonably well.
The interesting result here is speculative decoding. With the MTP draft head loaded on a Radeon AI PRO R9700 via Vulkan, generation hits 66.0 tok/s; the same backend without MTP manages 31.9 tok/s, and the HIP and Ollama routes trail at 26.7–27.1 tok/s (bm-004) — roughly a 2.1x gain from a software flag. The catch, which we measured in bm-006, is that the gain decays with context: 64.5 tok/s at 8K, 61.4 at 32K, 55.7 at 80K. Still worth having, but not free. In our cross-machine baseline, a single RX 7900 XT at 25.21 tok/s was within 10% of the dual-GPU rigs — for a model that fits on one card, the second GPU adds little.
LFM2.5-8B-A1B — the small-model speed demon
Who it’s for: always-on assistants, classification, summarisation — anything latency-sensitive that doesn’t need 30B-class reasoning.
The spread across our fleet tells the story: 68.8 tok/s on the Arc B60 Pro, 150.2 tok/s on the Ryzen AI MAX+ 395 iGPU (bm-011), 269.0 tok/s on the RX 7900 XT (bm-007). At 4.8GB it fits on almost anything, including 12GB starter cards. The MoE architecture means only ~1B parameters are active per token, which is why an 8B-class model keeps up with much larger dense models on speed.
Ornith-1.0-35B — the long-context specialist
Who it’s for: document analysis and retrieval work where you need 90K+ context in VRAM.
At 78.3 tok/s on a single R9700 via Ollama (bm-005), this was a notable result for a 35B model — and a cautionary one: the same model via llama.cpp Vulkan overflowed the 32GB card (34.2GB reported) and spilled to system RAM, dropping to 27.3 tok/s. Backend and memory-fit matter as much as raw hardware here.
Llama-3.1-8B-Instruct — the compatibility workhorse
Who it’s for: tooling and integrations where universal support matters more than peak speed.
The oldest model on this roster and still everywhere. 27.0 tok/s on the Arc B60 Pro, 105.2 tok/s on the RX 7900 XT (bm-008). Nothing about it is exciting; everything about it works.
gpt-oss 120B — the unified-memory tier (community-reported)
Who it’s for: owners of 128GB unified-memory boxes — Mac Studio, Strix Halo mini PCs — who want the largest open model that class can hold.
We have not measured this model ourselves; our fleet tops out at 32GB per card. The ~30 tok/s figure circulating for gpt-oss 120B on 128GB unified-memory machines (per runaihome.com’s comparison of the ASUS Ascent GX10) is community-reported and unverified by us. Treat it as directional, not data. If you’re shopping this tier, see our mini PC review and the unified memory explainer.
Qwen3.8 27B — the current cycle (spec-sheet)
Who it’s for: early adopters who want the newest general model and accept unmeasured territory.
Qwen3.8 27B is the current model cycle as of October 2026. We have not benchmarked it yet — this entry is spec-sheet analysis based on published model cards, and we won’t quote speeds we haven’t run. It’s on our bench queue for the next refresh; until then, Qwen3.6-27B with MTP remains the measured recommendation in this size class.
Z-Image and FLUX.2-klein — the media tier (community-reported)
Who it’s for: image generation alongside your LLM stack.
Both are current-cycle image models with active community discussion. We have no measured numbers for either — this entry is community-reported and included for roster completeness, not as a recommendation. Our benchmarking focus is text inference; treat any speed claims you see elsewhere with the same skepticism we’d apply.
How we verify
Every measured figure on this page comes from a published benchmark run with the full methodology attached: hardware, backend, quantization, context length, and variance. The short version: identical GGUF weights, llama-bench-style prompt-processing and generation tests, three runs each, with VRAM and power captured where the runtime reports them. The full protocol is on the methodology page, and the raw runs are linked from each benchmark page. Anything we couldn’t measure is labelled spec-sheet or community-reported in the table above — those labels are the contract.
What changed this month
- Added: Qwen3.8 27B (spec-sheet) and the media tier (Z-Image, FLUX.2-klein) as current-cycle entries.
- Context: the late-2026 DRAM shortage is pushing unified-memory machine prices up — if you’re shopping a 128GB box, the buy-now-or-wait calculus has shifted (see the mini PC landscape).
- Hardware note: GPU prices move weekly; current bands put the Arc Pro B60 24GB around $450–500 and the RX 7900 XT 24GB around $900 — both measured hosts on this page.
Where this data comes from
All measured figures link to their full benchmark reports: bm-001, bm-002, bm-003, bm-004, bm-005, bm-006, bm-007, bm-008, bm-009, bm-011, bm-012, and bm-013. Community-reported figures are attributed inline and marked unverified. This page updates monthly; last full refresh 3 October 2026.