FRONTIER BENCHMARKS
Benchmark reports
Every report is tested first-hand, dated, and reproducible from the published config. Filter to find the GPU, model, or backend decision you're working on.
FILTERS
12 OF 12 REPORTS
- BM-0032026-07-2738.6 tok/s · 740ms TTFT
Qwen3-Coder-30B-A3B quantization sweep: Q4_K_M vs. Q6_K vs. Q8_0 on Arc B60 Pro
On 24GB Arc, Q6_K is the sweet spot: quality recovers to within 2 points of Q8_0 on HumanEval while staying 38% faster and fitting comfortably. Q8_0 barely fits and leaves no room for context. Q4_K_M is the budget pick when VRAM is tight.
vulkanq6-kcodingintel-arc - BM-0022026-07-2652.1 tok/s · 410ms TTFT
Qwen3-Coder-30B-A3B: NVIDIA CUDA vs. Arc Vulkan at Q4_K_M
SUPERSEDED 2026-07-18 — see bm-009. Provisional Victor numbers in this record were ~3x too slow (real: 172 tok/s gen, 3403 tok/s prompt) and based on an incorrect single-24GB-GPU hardware spec. Retained for history only; do not cite. The original summary read: 'CUDA on the NVIDIA reference produces 52.1 tok/s vs. 38.6 tok/s on Arc Vulkan — a 35% generation-speed lead at the same 24GB VRAM tier.'
cudaq4-k-mcodingnvidia - BM-0012026-07-2543.8 tok/s · 580ms TTFT
Qwen3-Coder-30B-A3B on Arc B60 Pro vs. dual Radeon AI PRO R9700
Vulkan on Intel Battlemage lands within 12% of dual-Radeon AI PRO R9700 for code generation at Q4_K_M, at roughly half the system cost. Arc is the price-performance leader for sub-$2K builds; multi-GPU AMD wins raw throughput once you accept the complexity.
vulkanq4-k-mcodingintel-arc - BM-0132026-07-1926.7 tok/s · 0ms TTFT
Qwen3.6-27B cross-machine baseline: dual 12GB vs single 20GB vs dual 32GB — and the RQ-001 answer
First cross-machine baseline of the lab. The same model — Qwen3.6-27B at Q4_K_M, identical Ollama digest a50eda8ed977 (byte-identical weights) — run on all four lab machines under controlled conditions (same 37-token prompt, 512 generated tokens, temperature 0, seed 42, one warmup rep). Result: the three CUDA/ROCm machines cluster within ~10% for generation (Ray 26.74, Evo-X2 25.21, Victor 24.28 tok/s), so for a 27B Q4_K_M model dual-12GB NVIDIA is NOT meaningfully faster than single-20GB AMD. The dual-GPU advantage only matters once the model exceeds ~18GB. Jitori's 8.41 tok/s is Vulkan-fallback performance, NOT the Arc B60's real capability (SYCL is broken on Battlemage — see bm-012). This directly answers Research Question RQ-001.
rocmq4-k-mcross-machineqwen3-6 - BM-0062026-07-1864.5 tok/s · 0ms TTFT
Qwen3.6-27B + MTP at long context: where speculative decoding stops paying off (Radeon AI PRO R9700)
MTP speculative decoding on Qwen3.6-27B (single Radeon AI PRO R9700, llama.cpp Vulkan) holds 55-65 tok/s generation across 8K-80K context - degrading roughly 1.3x over that range - while prompt processing falls from ~760 to ~580 tok/s as context fills. The practical finding: MTP keeps working at long context (acceptance 66-75%), but KV-cache pressure steadily erodes both prompt and generation speed. Founder's prior short-prompt runs hit ~75-86 tok/s; the numbers here use real document prompts and are correspondingly a bit lower.
vulkanq4-k-mlong-contextamd - BM-0072026-07-18269.0 tok/s · 0ms TTFT
LFM2.5-8B-A1B MoE 8B on RX 7900 XT vs Arc B60 — small-model cross-GPU
LFM2.5-8B-A1B (Q4_K_M, 8.47B total / ~1.7B active MoE params, 4.79 GiB) measured on TWO GPUs for a clean MoE-vs-GPU comparison: EvoX2's AMD Radeon RX 7900 XT (gfx1100, ROCm) at 269.01 tok/s gen / 7464.85 tok/s prompt using 6.81GB VRAM, and Jitori's Intel Arc Pro B60 (Battlemage, Vulkan) at 68.79 tok/s gen / 1688.53 tok/s prompt. The RX 7900 XT is ~3.9x faster on this MoE than the Arc B60 - the SAME ratio as on the dense Llama-3.1-8B (see bm-008), confirming the GPU gap is consistent across architectures. Companion to bm-011 (the APU iGPU result on the same machine). Closes the small-model and real-app workload gaps.
rocmq4-k-mreal-appsmall-model - BM-0082026-07-18105.2 tok/s · 0ms TTFT
Llama-3.1-8B-Instruct dense 8B on Arc B60 vs RX 7900 XT — general-work cross-GPU
Llama-3.1-8B-Instruct (Q4_K_M, 8.03B dense params, 4.58 GiB) measured on TWO GPUs for a clean dense-vs-GPU comparison: Jitori's Intel Arc Pro B60 (Battlemage, Vulkan) at 27.01 tok/s gen / 706.61 tok/s prompt, and EvoX2's AMD Radeon RX 7900 XT (gfx1100, ROCm) at 105.22 tok/s gen / 3210.77 tok/s prompt. The RX 7900 XT is ~3.9x faster on this dense 8B than the Arc B60 - the SAME ratio as on the MoE LFM2.5-8B (see bm-007), confirming the GPU gap is consistent across architectures. Closes the general-work workload and dense-general model-class gaps; the 9th of 10.
vulkanq4-k-mgeneraldense - BM-0092026-07-18172.1 tok/s · 0ms TTFT
Qwen3-Coder-30B-A3B on Victor (RTX 5070 dual-GPU, CUDA) — bm-002 re-measurement
Re-measurement of the NVIDIA CUDA path that bm-002 reported provisionally. On Victor's dual RTX 5070 mobile config (Blackwell sm_120, CUDA 13.3), Qwen3-Coder-30B-A3B at Q4_K_M generates at 172 tok/s and processes prompts at 3403 tok/s — roughly 3.3x and 5.6x faster than bm-002's provisional figures (52.1 / 612.4). bm-002 is superseded; these are the real numbers.
cudaq4-k-mcodingnvidia - BM-0112026-07-18150.2 tok/s · 0ms TTFT
LFM2.5-8B on Ryzen AI MAX+ 395 iGPU (EvoX2) — unified-memory APU backend measurement
LFM2.5-8B-A1B (Q4_K_M, 8.47B params, 4.79 GiB) on the integrated Radeon 8060S GPU of EvoX2's AMD Ryzen AI MAX+ 395 (gfx1151, ~124GB unified memory pool visible to HIP). Generates at 150.16 tok/s and processes prompts at 3661.53 tok/s via ROCm 7.2.2. This is the only unified-memory APU in the fleet - the same machine's discrete RX 7900 XT (bm-007) is ~1.8x faster on this small model, but the APU's point is not raw speed: unified memory removes the VRAM ceiling that bounds every discrete card here. Closes the unified-memory / APU platform gap and answers the 'should I buy Strix Halo / Ryzen AI MAX for local AI?' question with first-party data.
rocmq4-k-mreal-appsmall-model - BM-0122026-07-1868.0 tok/s · 363ms TTFT
Qwen3-30B-A3B INT4 on Arc B60: OpenVINO/OVMS vs llama.cpp Vulkan — backend shootout
Qwen3-30B-A3B-Instruct-2507 (INT4, OpenVINO IR) served via OpenVINO Model Server 2026.2.1 on Jitori's Intel Arc Pro B60 (Battlemage BMG G21, 24GB). Generates at 67.95 tok/s via the OpenAI streaming API - roughly 1.76x faster than the same B60 running Qwen3-Coder-30B-A3B Q4_K_M via llama.cpp Vulkan (bm-001, 38.6 tok/s). This is the backend comparison the dataset was missing: same GPU, same 30B-A3B model class, two Intel-GPU backends. Finding: OpenVINO is meaningfully faster than Vulkan on Battlemage for this workload. This also corrects a fleet assumption - OpenVINO's GPU plugin (Level Zero) works cleanly on this B60, contradicting the older 'SYCL is broken on Battlemage, Vulkan-only' note.
vulkanq4-k-mopenvinoovms - BM-0042026-06-3066.0 tok/s · 0ms TTFT
Qwen3.6-27B + MTP backend showdown on Radeon AI PRO R9700
On a single Radeon AI PRO R9700 (gfx1201, RDNA4), llama.cpp Vulkan with the MTP speculative-decoding head is the only path that activates Qwen3.6's Multi-Token Prediction — delivering 66 tok/s, roughly 2.5x the 27 tok/s class of every other backend. If you don't load the -mtp.gguf draft head, Vulkan, HIP, and Ollama all land within a narrow 26-32 tok/s band and the architecture's main speed advantage is left on the table.
vulkanq4-k-mcodingamd - BM-0052026-06-3078.3 tok/s · 0ms TTFT
Ornith-1.0-35B backend showdown on Radeon AI PRO R9700
For the Ornith-1.0-35B MoE model (qwen35moe family) on a single Radeon AI PRO R9700, Ollama is the clear winner at 78.3 tok/s - roughly 2.9x faster than llama.cpp Vulkan (27.3 tok/s), which also spilled past the 32GB VRAM limit into system RAM. The HIP backend failed outright: it could not parse the Ornith GGUF (version mismatch). On MoE models without an MTP head, Ollama's expert routing currently beats hand-tuned llama.cpp on this build.
vulkanq4-k-mcodingamd