CALIBRATED 2026-07-27 · REC 012
Local AI Frontier
BM-002TESTED 2026-07-26REVISED 2026-07-18

Qwen3-Coder-30B-A3B: NVIDIA CUDA vs. Arc Vulkan at Q4_K_M

SUPERSEDED 2026-07-18 — see bm-009. Provisional Victor numbers in this record were ~3x too slow (real: 172 tok/s gen, 3403 tok/s prompt) and based on an incorrect single-24GB-GPU hardware spec. Retained for history only; do not cite. The original summary read: 'CUDA on the NVIDIA reference produces 52.1 tok/s vs. 38.6 tok/s on Arc Vulkan — a 35% generation-speed lead at the same 24GB VRAM tier.'

cudaq4-k-mllama.cpp b3500 (CUDA 12.5)codingnvidiaintel-arccudavulkan30b-classsuperseded

Configuration

Model
Qwen3-Coder-30B-A3B-Instruct
Artifact
Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf
Checksum
sha256:pending-real-checksum
Quantization
q4-k-m
Context length
8,192 tokens
Backend
cuda
Runtime
llama.cpp b3500 (CUDA 12.5)
Settings
-ngl 99 -c 8192 -t 8 -b 512
Workload pack
workload-pack-v1
Author
Edgar
Reviewer
Edgar

Results

Same workload pack as bm-001 to enable direct comparison.

52.1tok/s

GEN · Victor (msi-command)

612tok/s

PROMPT

410ms

TTFT

18.0GB

VRAM USED

MACHINEPROMPT tok/sGEN tok/sTTFT msVRAM GBRAM GBPOWER W
Victor (msi-command)
NVIDIA CUDA reference. cuBLAS, standard memory pool.
612.452.141018.02.8285
Jitori PC
Arc B60 Pro, Vulkan. Reference from bm-001.
412.338.674018.13.2245

Quality scores

CODING

0.81/ 0-1

HumanEval pass@1

Limitations

  • Victor configuration is provisional pending identity confirmation (see blueprint Appendix A1).
  • CUDA version pinned to 12.5; cuBLAS behavior varies across minor versions.
  • Identical model artifact and quant to bm-001 to isolate the backend variable.

Corrections

  • 2026-07-18

    SUPERSEDED by bm-009. The Victor numbers in this record (gen 52.1 tok/s, prompt 612.4 tok/s, vram 18.0GB on a claimed single 24GB card at CUDA 12.5) were provisional placeholders and are incorrect. Real measurements on the actual hardware (MSI Vector 16 HX laptop, 2x RTX 5070-class mobile GPUs, CUDA 13.3) are ~3.3x faster for generation and ~5.6x faster for prompt processing — see bm-009. The original hardware spec ('24GB NVIDIA reference') was also wrong. This record is retained for editorial history; do not cite its numbers.

NO COMMERCIAL INTEREST IN THIS RESULT