FRONTIER DISPATCH
Blog
News on local AI power and the practical things you can do with it — model releases, hardware shifts, and real workflows running on your own hardware. Written from the bench, not the press release.
AMD Radeon for local AI: the ROCm reality check
2026-07-31Is AMD a real option for local LLMs in 2026, or still the 'it works but...' alternative? The honest answer, from running RDNA3 and RDNA4 cards in the lab: usable, genuinely good value at the 24GB tier, with one persistent caveat you need to know before you buy.
hardwareamd10 minApple Silicon and MLX for local AI
2026-07-31A Mac is the only machine where 'unified memory' means a 70B model fits without a discrete GPU. Is MLX on Apple Silicon a real local-AI path in 2026, or a niche? The honest answer for Mac owners — and the one thing that decides whether it's worth it.
hardwareapplemlx10 minBest local TTS models in 2026
2026-07-31The open-weight text-to-speech landscape finally has a real default: Kokoro-82M for almost everything, XTTS-v2 for zero-shot voice cloning, F5-TTS for maximum quality, and Piper for edge. A practical pick-by-use-case guide — with the licensing catches that decide which you can actually ship.
fundamentalsaudiovoice9 minThe best mini PC for local AI (Ryzen AI Max+ 395)
2026-07-31A mini PC that runs a 70B model in a box the size of a hardback book — marketing claim or reality? We run the same Ryzen AI Max+ 395 chip in our lab's EvoX2 machine. Here's what it actually does, what it doesn't, and whether the unified-memory mini PC is the real deal.
hardwaremini-pc10 minA home AI lab build under $1,000
2026-07-31A complete, tested parts list for a real local-AI machine at the four-figure tier — what to spend on, what to cheap out on, and the one upgrade that's actually worth the money. Built around the 12GB sweet spot for 7B–13B daily and 30B MoE when you need it.
hardwarebuild-guide11 minComfyUI: a beginner's guide
2026-07-31ComfyUI is the most powerful local image-generation tool — and the one with the steepest learning curve. Here's the honest on-ramp: what the node graph actually is, the minimum workflow to get an image, where beginners get stuck, and when to use ComfyUI versus a simpler alternative.
how-toimage-generationcomfyui11 minHow much VRAM do I need for local AI?
2026-07-31The one number that decides your entire local-AI experience. A practical VRAM-to-model map with real measured numbers from the lab — what fits at each tier (6GB → 48GB+), what 'fits' actually means once you account for context, and the fastest way to answer it for your exact card.
fundamentalshardwaredecision-hub9 minIs local AI actually private? An honest threat model
2026-07-31Local isn't a privacy magic spell — it shifts which adversaries you're defending against. A concrete threat-model breakdown: what local protects you from, what it doesn't, and the five questions that decide whether your setup earns the word 'private.'
privacyfundamentals8 minLearn local AI: the complete roadmap
2026-07-31A five-stage path from 'what is local AI' to running a multi-GPU home lab — each stage a clear outcome with the exact articles and steps to get there. No prerequisites beyond a willingness to use a terminal once or twice.
fundamentalslearncurriculum9 minHow much does a local AI setup cost per month?
2026-07-31The honest total-cost number everyone asks for and nobody publishes. Upfront hardware by tier, the real monthly electricity cost (measured, not guessed), the hidden costs nobody mentions, and the break-even point against a cloud API.
fundamentalscost11 minLocal vs cloud AI: the full comparison
2026-07-31The honest version of a question everyone asks and most articles answer with a slogan. Four axes — cost, privacy, latency, capability — with where each side wins, where it's a wash, and the decision rule for your specific case.
fundamentalscomparison10 minHow to run DeepSeek locally
2026-07-31DeepSeek's reasoning models are among the most capable open weights you can run — and the V4 generation is enormous. Here's what actually fits on consumer hardware, what doesn't, and the honest setup path for the DeepSeek models that matter for local inference.
modelshow-to9 minUsed RTX 3090 — still the value king for local AI?
2026-07-3124GB of VRAM for ~$700 used, when the cheapest new 24GB card costs double that. Is the RTX 3090 still the best local-AI buy in 2026? The honest answer: yes for most people, with a clear checklist for what to verify before you buy.
hardwarebuyer-guide10 minvLLM vs SGLang: the serving deep-dive
2026-07-31Ollama is the right default for single-user desktop chat — but when does it stop being enough? This is the answer: when you're serving concurrent users, running agents with shared prefixes, or need structured output at throughput. Here's when to move to vLLM or SGLang, and which one.
softwareserving11 minWhy local AI is better for privacy
2026-07-31Privacy is not a feature you can toggle on a cloud API — it's a property of where the computation happens. A threat-model breakdown of what 'private' actually requires, where cloud 'private' modes fall short, and what a genuinely private local setup looks like.
fundamentalsprivacy8 minBatch-process 10,000 documents locally for the cost of electricity
2026-07-30The deep-dive on use-case #5. Classify, extract, summarize, or transform a pile of documents headlessly — running the model against a queue, not a chat. Where local AI's cost-at-scale argument actually bites, the throughput math, and the honest 'one stream at a time' limit.
use-caseshow-toautomation11 minOffline voice assistant with Whisper + a local LLM
2026-07-29The deep-dive on use-case #3. Wire a local speech-to-text model, a local LLM, and a local TTS engine into a voice assistant that works offline, has no wake-word corporation listening, and runs on hardware you own. Architecture, latency budget, and the honest '1–2 seconds' reality.
use-caseshow-tovoice11 minRun a private RAG knowledge base on a single GPU
2026-07-28The deep-dive on use-case #2. Drop contracts, research dumps, or internal docs into a local RAG stack and the model reads, retrieves, and answers — with nothing leaving your disk. Architecture, model choices, the honest limits of local retrieval, and a working starting config.
use-caseshow-torag12 minBuild a private coding agent that never phones home
2026-07-27The deep-dive on use-case #1 from '5 things you can do with a local LLM.' A working coding agent — refactors, tests, explains codebases — running on a local model, with your source never leaving your machine. Setup, model choice, where it's good, and where it falls down.
use-caseshow-tocoding11 minSelf-host your own 'ChatGPT' for $300
2026-07-26A working private chat assistant — model, web UI, and OpenAI-compatible API — running entirely on hardware you own, for the price of one card. The shortest path from zero to 'it works,' with the honest limits and where to go next.
how-touse-casesgetting-started10 minBattlemage one year in: Arc B580/B60 as a local-AI card
2026-07-25A year ago Intel Arc was still a punchline for AI workloads. The Battlemage refresh (B580/B60) is the thing that changed that — and a year of driver, backend, and ecosystem work has made it genuinely good, with caveats. Here's where Arc actually stands for local AI in mid-2026.
hardwareintel-arcbenchmarks11 minBest local LLMs in 2026: measured, not summarized
2026-07-24Every 'best local LLM' list ranks models by reputation and parameter count. None of them actually ran the models on consumer hardware and reported measured tok/s, VRAM, and quality scores. We did. Here are the models that genuinely run well on hardware you own — with the receipts.
comparisonbenchmarksfundamentalsmodels14 minHow to run AI models locally in 2026: the complete setup
2026-07-24A start-to-finish guide to running LLMs on your own hardware in 2026 — pick the right GPU, choose a backend, pick a quantization, and pick a runtime. Every recommendation is backed by measured benchmarks, not guesses. For beginners and builders.
how-tofundamentalshardwaremodels16 minOllama vs llama.cpp in 2026: which to actually pick
2026-07-24The framing most posts use — 'Ollama is easier, llama.cpp is faster' — is from 2024. In 2026 the real question is whether you're running one model for yourself or serving many. Here's the honest decision rule, including where vLLM and SGLang have eaten Ollama's edge.
toolsruntimescomparison10 minQuantization in 2026: Q4_K_M is no longer the compromise
2026-07-23In 2024, Q4_K_M meant 'noticeably broken.' In 2026 it's the practical default. We ran the Q4 vs Q6 vs Q8 sweep on Qwen3-Coder-30B — here's the measured quality delta, the VRAM math, and when you actually should step up.
fundamentalsquantizationbenchmarks9 minThe MoE shift: why every new local model is a Mixture-of-Experts
2026-07-22Every model worth running locally in 2026 — Qwen3-30B-A3B, LFM2.5-8B-A1B, the GLM/Kimi frontier — is MoE. Here's what 'active parameters' actually means, why it's the reason these models fit on your card, and the measured 2.5× speed data behind it.
fundamentalsarchitecturebenchmarks9 minThe open-source LLM landscape in July 2026
2026-07-21GLM-5.2, DeepSeek V4, Kimi K2.7, Qwen3.6, MiniMax M3 — every roundup lists them. Almost none tell you which actually run on a card you own. Here's the landscape, sorted by what fits on consumer hardware, with VRAM math and the honest 'API-only' calls.
newsmodel-selectionroundup11 minBest GPUs for local AI in 2026: measured cross-vendor
2026-07-19Every GPU roundup ranks the same NVIDIA cards. Almost none run the same model file on Arc, AMD, and NVIDIA and report the actual tok/s gap. We did. The cross-vendor gap is real. So is the price gap — and that flips the recommendation for budget buyers.
hardwarecomparisonbenchmarksfundamentals12 minBest local LLM for coding in 2026: measured
2026-07-19Every coding-LLM roundup ranks models by leaderboard score. Almost none tell you what the score costs in tok/s when you step down a quant tier — which is the only trade-off that matters once the model has to fit on a card you own. Here's the measured table for the coder we run, across three quant tiers, on the same hardware.
use-caseshow-tobenchmarkscoding11 minCloud image generation APIs compared: Replicate, fal.ai, Midjourney, and OpenAI in 2026
2026-07-19Replicate, fal.ai, Midjourney, and OpenAI compared for cloud image generation. The billing models differ more than the models do — per-second, per-image, subscription, and token-bundled each win for different usage patterns. Here's the decision tree.
comparisonuse-cases9 minCloud video generation APIs compared: Runway, fal.ai, Pika, and Sora for cost and control
2026-07-19Cloud video generation has the widest price range in AI media — $0.13 to $3.20 per clip. Runway, fal.ai, Pika, and Sora compared on pricing models, quality, and content control, with a decision tree for which provider wins for which use case.
comparisonuse-cases9 minComfyUI optimization in 2026: Dynamic VRAM, quantization, and attention backends explained
2026-07-19How to make ComfyUI actually fast on a card you can afford. Dynamic VRAM is the big 2026 leap, but the quantization and attention-backend choices matter as much. Here's what each optimization does, when it helps, and the honest trade-offs.
how-totoolsuse-cases13 minDual 12GB vs single 24GB: we ran the same 27B on four machines and the GPU count barely mattered
2026-07-19We took Qwen3.6-27B at Q4_K_M — byte-identical weights, verified by Ollama digest — and ran it on all four lab machines: dual 12GB NVIDIA, single 20GB AMD, dual 32GB AMD, and a 24GB Intel Arc. The three CUDA/ROCm machines landed within 10% of each other. The honest answer to 'do I need two GPUs?' is: only once the model stops fitting in one.
hardwarebenchmarksvram-wars8 minCUDA vs ROCm vs Vulkan vs OpenVINO: measured
2026-07-19Every backend explainer describes CUDA, ROCm, Vulkan, and SYCL correctly and then tells you to 'pick what works on your card.' Almost none measure them on the same hardware. We did, across nine benchmark records. The backend gaps are large, vendor lock-in is real, and one documented failure is worth a thousand abstract paragraphs.
how-tocomparisonbenchmarksfundamentals13 minHome AI server: a working four-node fleet
2026-07-19Every 'home AI lab' guide describes a single workstation. None describe running models across the machines you already own. Here's the four-node cross-vendor fleet we run — Intel Arc, dual AMD, AMD APU, dual NVIDIA — the routing pattern that ties it together, and what a fleet gets you that a single big box doesn't.
use-caseshow-tohardwarefundamentals12 minLocal video generation in 2026: Wan, Hunyuan, and CogVideoX on your own GPU
2026-07-19Running Wan, Hunyuan Video, and CogVideoX locally — what's actually runnable on consumer GPUs, what isn't, and the optimizations that help. The honest version, because video is where local hardware hurts most.
how-tohardwareuse-cases12 minLocal vs cloud AI generation: the honest decision for image and video workloads
2026-07-19The local-vs-cloud question for AI image and video generation isn't 'which is better.' It's a cost, privacy, and volume calculation — and the answer is different from the one for text. Here's the framework, the break-even math, and where each side actually wins.
fundamentalsuse-casescomparison9 minOllama vs llama.cpp on the same hardware: measured
2026-07-19The 'Ollama is easier, llama.cpp is faster' framing is from 2024 and it was correct then. What it never came with was the actual measured gap on the same card with the same weights. Here's that gap — across four backend configurations on a single R9700, plus the one result that looks like an Ollama win and isn't.
how-tocomparisonbenchmarksfundamentals11 minWhat a quantization tier costs in watts: measured
2026-07-19Every 'how much power does local AI use' piece multiplies GPU TDP specs by hours and calls it a day. Almost nobody measures actual wall-power during inference. We measured a narrow slice — and the headline finding is that quantization barely moves power draw. The honest version of what we have, what we don't, and the rough cost math.
benchmarkshardwarehow-tofundamentals10 minRunning Flux and Stable Diffusion locally: the honest cost and hardware guide
2026-07-19What it actually takes to run Flux and Stable Diffusion on your own GPU — VRAM floors by model, GPU picks by budget tier, and the five optimizations that make marginal cards viable. No 'best GPU' list — the math.
how-tohardwareuse-cases11 minSingle vs dual GPU for local LLMs: when the second card stops idling
2026-07-19Forum threads answer 'is dual GPU worth it' with 'depends.' Our four-machine baseline (bm-013) gives the actual crossover: dual-12GB-NVIDIA tied single-32GB-AMD on a 27B model. Below ~18GB of model weight the second card idles; above, dual becomes mandatory or faster. The measured rule, with the trade-offs single-card buyers forget.
hardwarecomparisonbenchmarksfundamentals11 minSpeculative decoding on consumer GPUs: MTP measured 2.07×
2026-07-19Every speculative decoding explainer calls it a 'free speedup.' Almost none measure the actual multiplier on consumer hardware or show the accept-rate curve across context lengths. Here's both — Qwen3.6-27B on the same R9700 with and without MTP, plus the long-context sweep that explains when the speedup holds and when it shrinks.
how-tobenchmarksfundamentalscomparison12 minArc B60 vs RX 7900 XT: a real comparison, not a chart fight
2026-07-18We ran the same two 8B models on both cards with byte-identical weights and identical settings. The RX 7900 XT is ~3.9x faster — consistently across both architectures. Here's the full 2x2, what it actually tells you, and where we got the framing wrong the first time.
hardwarebenchmarkscomparisonintel-arcamd9 min5 things you can do with a local LLM today
2026-07-18Not future use-cases — things you can set up this week on hardware you already own. Practical workflows, what they're good at, and where they fall down.
use-casesworkflows7 minOpenVINO beats Vulkan on the Arc B60 — and we were wrong about SYCL
2026-07-18We benchmarked OpenVINO Model Server 2026.2.1 against llama.cpp Vulkan on the same Intel Arc Pro B60. OpenVINO is ~1.76x faster on a 30B-A3B model — and it works cleanly, contradicting our own earlier note that Battlemage was Vulkan-only. Here's the data and the correction.
hardwarebenchmarksintel-arcopenvino7 minRun a 30B model on a $300 GPU
2026-07-18Yes, really. A Q4_K_M 30B coder model fits in 12GB of VRAM and runs at usable speed on a single Intel Arc B580. Here's the recipe — with the numbers to back it.
how-tohardwarebenchmarks8 minWhy local AI matters in 2026
2026-07-18Cloud LLMs got good. So why are people still buying GPUs to run models at home? Four reasons — and where the cloud still wins.
fundamentalsopinion6 min