TAG
#nvidia
3 items — 0 dispatches3 benchmark reports.
Benchmark reports
BM-013Qwen3.6-27B cross-machine baseline: dual 12GB vs single 20GB vs dual 32GB — and the RQ-001 answer
2026-07-19First cross-machine baseline of the lab. The same model — Qwen3.6-27B at Q4_K_M, identical Ollama digest a50eda8ed977 (byte-identical weights) — run on all four lab machines under controlled conditions (same 37-token prompt, 512 generated tokens, temperature 0, seed 42, one warmup rep). Result: the three CUDA/ROCm machines cluster within ~10% for generation (Ray 26.74, Evo-X2 25.21, Victor 24.28 tok/s), so for a 27B Q4_K_M model dual-12GB NVIDIA is NOT meaningfully faster than single-20GB AMD. The dual-GPU advantage only matters once the model exceeds ~18GB. Jitori's 8.41 tok/s is Vulkan-fallback performance, NOT the Arc B60's real capability (SYCL is broken on Battlemage — see bm-012). This directly answers Research Question RQ-001.
rocmq4-k-mcross-machineqwen3-627bdual-gpuBM-002Qwen3-Coder-30B-A3B: NVIDIA CUDA vs. Arc Vulkan at Q4_K_M
2026-07-18SUPERSEDED 2026-07-18 — see bm-009. Provisional Victor numbers in this record were ~3x too slow (real: 172 tok/s gen, 3403 tok/s prompt) and based on an incorrect single-24GB-GPU hardware spec. Retained for history only; do not cite. The original summary read: 'CUDA on the NVIDIA reference produces 52.1 tok/s vs. 38.6 tok/s on Arc Vulkan — a 35% generation-speed lead at the same 24GB VRAM tier.'
BM-009Qwen3-Coder-30B-A3B on Victor (RTX 5070 dual-GPU, CUDA) — bm-002 re-measurement
2026-07-18Re-measurement of the NVIDIA CUDA path that bm-002 reported provisionally. On Victor's dual RTX 5070 mobile config (Blackwell sm_120, CUDA 13.3), Qwen3-Coder-30B-A3B at Q4_K_M generates at 172 tok/s and processes prompts at 3403 tok/s — roughly 3.3x and 5.6x faster than bm-002's provisional figures (52.1 / 612.4). bm-002 is superseded; these are the real numbers.