CORRECTIONS LOG
Corrections
GUARDRAIL: FEWER THAN 2% OF RECORDS REQUIRING A PREVENTABLE DATA-ERROR CORRECTION
We don't silently edit published numbers. When something is wrong or materially changed, we log it here, date it, and link to the revised record.
Active corrections
2026-07-19
Initial publication. First cross-machine baseline of the lab. Answers RQ-001 (dual-12GB vs single-24GB) for the 27B Q4_K_M model class: at this size, dual-12GB NVIDIA is not meaningfully faster than single-20GB AMD — the three CUDA/ROCm machines cluster within ~10%. Jitori's Vulkan-fallback number is explicitly flagged as not representative of the Arc B60 (SYCL regression, see bm-012 for the working OpenVINO path).
2026-07-18
BM-002 — Qwen3-Coder-30B-A3B: NVIDIA CUDA vs. Arc Vulkan at Q4_K_M
SUPERSEDED by bm-009. The Victor numbers in this record (gen 52.1 tok/s, prompt 612.4 tok/s, vram 18.0GB on a claimed single 24GB card at CUDA 12.5) were provisional placeholders and are incorrect. Real measurements on the actual hardware (MSI Vector 16 HX laptop, 2x RTX 5070-class mobile GPUs, CUDA 13.3) are ~3.3x faster for generation and ~5.6x faster for prompt processing — see bm-009. The original hardware spec ('24GB NVIDIA reference') was also wrong. This record is retained for editorial history; do not cite its numbers.
2026-07-18
BM-001 — Qwen3-Coder-30B-A3B on Arc B60 Pro vs. dual Radeon AI PRO R9700
Ray GPU identity corrected from 'RX 9700 (RDNA3)' to 'Radeon AI PRO R9700 (gfx1201, RDNA4)'. Measurement values unchanged — this is a labeling correction only. [D-007]
2026-07-18
METHODOLOGY CORRECTION (same day as initial publication). The first version of bm-006 used -ngl 99 (and -c 262144), which under-offloaded layers to the GPU and produced numbers ~2.5-3x too slow (28.7 tok/s @ 8K, 20.4 @ 65K). Re-run with the founder's documented winning recipe (-ngl 999 full offload, -c 98304, otherwise identical) yields 64.5 tok/s @ 8K and 55.7 @ 80K - matching the founder's independently-recovered memory of ~60-86 tok/s on this exact hardware. The earlier wrong numbers have been replaced in this record; the lesson (-ngl must be high enough to fully offload all layers) is captured for future runs.
2026-07-18
BM-007 — LFM2.5-8B-A1B MoE 8B on RX 7900 XT vs Arc B60 — small-model cross-GPU
Extended from single-GPU (RX 7900 XT only) to a 2-GPU comparison by adding the Arc B60 row. Same artifact (sha256-matched), same flags. Also captured the model_checksum (was 'pending' in v1). v1 RX 7900 XT numbers unchanged.
2026-07-18
BM-008 — Llama-3.1-8B-Instruct dense 8B on Arc B60 vs RX 7900 XT — general-work cross-GPU
Extended from single-GPU (Arc B60 only) to a 2-GPU comparison by adding the RX 7900 XT row. Same artifact (sha256-matched), same flags. Corrects the original publication's framing: the 'dense-vs-MoE ~10x speed gap' insight in the v1 summary was confounded (it compared Llama-on-Arc to LFM-on-RX, mixing architecture with GPU). The clean architecture effect at matched GPU is ~2.5x (see bm-007 for the companion LFM rows), and the GPU effect is ~3.9x consistently across both architectures. v1 numbers unchanged; framing corrected.
2026-07-18
BM-009 — Qwen3-Coder-30B-A3B on Victor (RTX 5070 dual-GPU, CUDA) — bm-002 re-measurement
Initial publication. Supersedes bm-002, whose provisional Victor numbers (52.1 gen / 612.4 prompt tok/s, claimed single-24GB-GPU at CUDA 12.5) were ~3x too slow and based on an incorrect hardware spec. bm-002 is marked superseded; its record is retained for history.
2026-07-18
BM-012 — Qwen3-30B-A3B INT4 on Arc B60: OpenVINO/OVMS vs llama.cpp Vulkan — backend shootout
Initial publication. Establishes that OpenVINO Model Server works on the Arc B60 (contradicting the older fleet note that 'SYCL is broken on Battlemage, Vulkan-only' - that note referred to an ad-hoc SYCL build, not Intel's curated OVMS container) and that it delivers ~1.76x generation speedup over llama.cpp Vulkan on the same GPU for the 30B-A3B model class. Caveats: model variant, quant method, and measurement protocol all differ from bm-001, so the 1.76x is directional, not precise. A same-model GGUF-vs-IR comparison is the documented follow-up.
Report an error
If you spot something wrong — a miscalculation, a missing setting, a stale price, a contradiction with our disclosure page — tell us. Email lab@localaifrontier.com with the benchmark ID and the issue. We investigate every report and publish a correction within 72 hours if it's valid.