TAG
#fundamentals
4 items — 4 dispatches.
Dispatches
Quantization in 2026: Q4_K_M is no longer the compromise
2026-07-23In 2024, Q4_K_M meant 'noticeably broken.' In 2026 it's the practical default. We ran the Q4 vs Q6 vs Q8 sweep on Qwen3-Coder-30B — here's the measured quality delta, the VRAM math, and when you actually should step up.
The MoE shift: why every new local model is a Mixture-of-Experts
2026-07-22Every model worth running locally in 2026 — Qwen3-30B-A3B, LFM2.5-8B-A1B, the GLM/Kimi frontier — is MoE. Here's what 'active parameters' actually means, why it's the reason these models fit on your card, and the measured 2.5× speed data behind it.
Local vs cloud AI generation: the honest decision for image and video workloads
2026-07-19The local-vs-cloud question for AI image and video generation isn't 'which is better.' It's a cost, privacy, and volume calculation — and the answer is different from the one for text. Here's the framework, the break-even math, and where each side actually wins.
Why local AI matters in 2026
2026-07-18Cloud LLMs got good. So why are people still buying GPUs to run models at home? Four reasons — and where the cloud still wins.