TAG
#hardware
18 items — 18 dispatches.
Dispatches
AMD Radeon for local AI: the ROCm reality check
2026-07-31Is AMD a real option for local LLMs in 2026, or still the 'it works but...' alternative? The honest answer, from running RDNA3 and RDNA4 cards in the lab: usable, genuinely good value at the 24GB tier, with one persistent caveat you need to know before you buy.
Apple Silicon and MLX for local AI
2026-07-31A Mac is the only machine where 'unified memory' means a 70B model fits without a discrete GPU. Is MLX on Apple Silicon a real local-AI path in 2026, or a niche? The honest answer for Mac owners — and the one thing that decides whether it's worth it.
The best mini PC for local AI (Ryzen AI Max+ 395)
2026-07-31A mini PC that runs a 70B model in a box the size of a hardback book — marketing claim or reality? We run the same Ryzen AI Max+ 395 chip in our lab's EvoX2 machine. Here's what it actually does, what it doesn't, and whether the unified-memory mini PC is the real deal.
A home AI lab build under $1,000
2026-07-31A complete, tested parts list for a real local-AI machine at the four-figure tier — what to spend on, what to cheap out on, and the one upgrade that's actually worth the money. Built around the 12GB sweet spot for 7B–13B daily and 30B MoE when you need it.
How much VRAM do I need for local AI?
2026-07-31The one number that decides your entire local-AI experience. A practical VRAM-to-model map with real measured numbers from the lab — what fits at each tier (6GB → 48GB+), what 'fits' actually means once you account for context, and the fastest way to answer it for your exact card.
Used RTX 3090 — still the value king for local AI?
2026-07-3124GB of VRAM for ~$700 used, when the cheapest new 24GB card costs double that. Is the RTX 3090 still the best local-AI buy in 2026? The honest answer: yes for most people, with a clear checklist for what to verify before you buy.
Battlemage one year in: Arc B580/B60 as a local-AI card
2026-07-25A year ago Intel Arc was still a punchline for AI workloads. The Battlemage refresh (B580/B60) is the thing that changed that — and a year of driver, backend, and ecosystem work has made it genuinely good, with caveats. Here's where Arc actually stands for local AI in mid-2026.
How to run AI models locally in 2026: the complete setup
2026-07-24A start-to-finish guide to running LLMs on your own hardware in 2026 — pick the right GPU, choose a backend, pick a quantization, and pick a runtime. Every recommendation is backed by measured benchmarks, not guesses. For beginners and builders.
Best GPUs for local AI in 2026: measured cross-vendor
2026-07-19Every GPU roundup ranks the same NVIDIA cards. Almost none run the same model file on Arc, AMD, and NVIDIA and report the actual tok/s gap. We did. The cross-vendor gap is real. So is the price gap — and that flips the recommendation for budget buyers.
Dual 12GB vs single 24GB: we ran the same 27B on four machines and the GPU count barely mattered
2026-07-19We took Qwen3.6-27B at Q4_K_M — byte-identical weights, verified by Ollama digest — and ran it on all four lab machines: dual 12GB NVIDIA, single 20GB AMD, dual 32GB AMD, and a 24GB Intel Arc. The three CUDA/ROCm machines landed within 10% of each other. The honest answer to 'do I need two GPUs?' is: only once the model stops fitting in one.
Home AI server: a working four-node fleet
2026-07-19Every 'home AI lab' guide describes a single workstation. None describe running models across the machines you already own. Here's the four-node cross-vendor fleet we run — Intel Arc, dual AMD, AMD APU, dual NVIDIA — the routing pattern that ties it together, and what a fleet gets you that a single big box doesn't.
Local video generation in 2026: Wan, Hunyuan, and CogVideoX on your own GPU
2026-07-19Running Wan, Hunyuan Video, and CogVideoX locally — what's actually runnable on consumer GPUs, what isn't, and the optimizations that help. The honest version, because video is where local hardware hurts most.
What a quantization tier costs in watts: measured
2026-07-19Every 'how much power does local AI use' piece multiplies GPU TDP specs by hours and calls it a day. Almost nobody measures actual wall-power during inference. We measured a narrow slice — and the headline finding is that quantization barely moves power draw. The honest version of what we have, what we don't, and the rough cost math.
Running Flux and Stable Diffusion locally: the honest cost and hardware guide
2026-07-19What it actually takes to run Flux and Stable Diffusion on your own GPU — VRAM floors by model, GPU picks by budget tier, and the five optimizations that make marginal cards viable. No 'best GPU' list — the math.
Single vs dual GPU for local LLMs: when the second card stops idling
2026-07-19Forum threads answer 'is dual GPU worth it' with 'depends.' Our four-machine baseline (bm-013) gives the actual crossover: dual-12GB-NVIDIA tied single-32GB-AMD on a 27B model. Below ~18GB of model weight the second card idles; above, dual becomes mandatory or faster. The measured rule, with the trade-offs single-card buyers forget.
Arc B60 vs RX 7900 XT: a real comparison, not a chart fight
2026-07-18We ran the same two 8B models on both cards with byte-identical weights and identical settings. The RX 7900 XT is ~3.9x faster — consistently across both architectures. Here's the full 2x2, what it actually tells you, and where we got the framing wrong the first time.
OpenVINO beats Vulkan on the Arc B60 — and we were wrong about SYCL
2026-07-18We benchmarked OpenVINO Model Server 2026.2.1 against llama.cpp Vulkan on the same Intel Arc Pro B60. OpenVINO is ~1.76x faster on a 30B-A3B model — and it works cleanly, contradicting our own earlier note that Battlemage was Vulkan-only. Here's the data and the correction.
Run a 30B model on a $300 GPU
2026-07-18Yes, really. A Q4_K_M 30B coder model fits in 12GB of VRAM and runs at usable speed on a single Intel Arc B580. Here's the recipe — with the numbers to back it.