DISPATCH
The best mini PC for local AI (Ryzen AI Max+ 395)
A mini PC that runs a 70B model in a box the size of a hardback book — marketing claim or reality? We run the same Ryzen AI Max+ 395 chip in our lab's EvoX2 machine. Here's what it actually does, what it doesn't, and whether the unified-memory mini PC is the real deal.

The unified-memory mini PC is the surprise hardware story of 2026 for local AI: a machine the size of a hardback book, with 64–128GB of shared CPU/GPU memory, marketed as running 70B models locally. The GMKtec EVO-X2 and its peers, built on AMD's Ryzen AI Max+ 395, are genuinely interesting — and unlike most reviewers, we run this exact chip in our lab. Our EvoX2 machine uses a Ryzen AI Max+ 395 (with its Radeon 8060S integrated GPU) alongside a discrete RX 7900 XT, so we can speak to the unified-memory side from direct experience, not just a spec sheet.
Here's the honest take: the unified-memory mini PC is real, it's a legitimate category, and for a specific reader it's the best single-box local-AI option that isn't a Mac. It also has real limits the marketing glosses over. This is what it actually does.
What the Ryzen AI Max+ 395 actually is
The chip at the center of these mini PCs — AMD's Ryzen AI Max+ 395 — is the PC-world answer to Apple Silicon's unified memory. It pairs a fast CPU (up to ~5.1 GHz, 16 cores) with a capable integrated GPU (the Radeon 8060S, gfx1151) that can address most of the system's LPDDR5X memory as GPU memory. Configure it with 64GB, 96GB, or 128GB of RAM and you get a single tiny box with a large shared memory pool — the same architectural trick that makes Apple Silicon interesting for large models, in a PC form factor.
The flagship product is the GMKtec EVO-X2, priced around $1,800–$2,000 depending on configuration (64GB or 128GB). It's the machine that turned "unified-memory mini PC for local AI" from a curiosity into a real buying category in 2026.
One honest architecture note that matters: the Ryzen AI Max+ 395's memory is shared between CPU and GPU, but it's not identical to Apple's unified memory architecture (UMA). The split between CPU and GPU memory is configurable but not as seamlessly unified as Apple's. In practice it behaves similarly for inference — you get a large addressable pool for the model — but the comparison to Apple isn't 1:1, and anyone telling you it's "exactly like a Mac Studio for less" is overselling.
What we measured on the EvoX2
Our lab's EvoX2 machine runs the Ryzen AI Max+ 395 with 124GB of system RAM, paired with a discrete RX 7900 XT (20GB). In bm-013, we ran the same Qwen3.6-27B model — byte-identical weights — across all four lab machines. On the EvoX2's discrete RX 7900 XT (not the iGPU), it generated at 25.21 tok/s — within 10% of the dual-NVIDIA and dual-AMD machines. That tells you the platform is a serious inference host.
The iGPU (Radeon 8060S, gfx1151) side — the part that uses the unified memory — is what the mini-PC pitch is really about. Our EvoX2 runs it as a second endpoint (the "Architect," serving Qwen3.6-35B) on the integrated GPU while the discrete card serves a different model. That dual-endpoint topology is something a unified-memory mini PC enables that a single-GPU desktop can't easily match: two models served simultaneously from one box, one on the iGPU's unified-memory pool, one on a discrete card.
The honest takeaway from running this hardware: the unified-memory iGPU genuinely can address a large memory pool and run models that wouldn't fit on a small discrete GPU. The 70B claim isn't fiction — with 96GB+ configured, the memory pool is large enough. What the marketing underplays is the speed: unified memory on this platform has lower bandwidth than a discrete GPU, so a 70B fits but generates at single-digit tok/s. It works; it's not fast.
Where the unified-memory mini PC wins
- Large-model capacity in a tiny, quiet box. This is the real pitch and it's genuine. A 128GB EVO-X2 can hold a 70B model in its memory pool — no multi-GPU build, no server rack, near-silent, drawing a fraction of a multi-GPU PC's power. For "I want to occasionally run a 70B without building a fleet," it's the simplest PC path.
- Multiple models simultaneously. The unified-memory iGPU plus (optionally) a discrete card lets you serve more than one model from one box. Our EvoX2 does exactly this. A single-GPU desktop serves one model well; a unified-memory mini PC can juggle.
- Power, noise, footprint. If you want local AI in a living room, on a desk, or in a small apartment — not a server closet — this form factor is dramatically more livable than a multi-GPU rig.
- Price vs. a Mac Studio. A 128GB EVO-X2 at ~$2,000 undercuts a comparably-memory-configured Mac Studio meaningfully, for readers who want the unified-memory large-model story without leaving the PC ecosystem.
Where it loses
- Generation speed on large models. Same caveat as Apple Silicon: unified memory has less bandwidth than a discrete GPU, so a 70B model fits but runs at single-digit tok/s. A multi-GPU PC or a high-end discrete card is meaningfully faster on the same model. If you want 70B fast, this isn't the tier.
- It's not a discrete GPU replacement for small/mid models. For 7B–30B work, a $300 discrete GPU on a cheap host is faster and cheaper. The unified-memory advantage only pays off at the large-model tier where the memory pool matters. If your workload fits in 24GB, buy 24GB of discrete VRAM instead.
- Memory is non-upgradable. Like a Mac, you pick the memory tier at purchase and it's soldered. A 128GB decision is permanent.
- Integrated GPU, not CUDA. The iGPU is AMD — the same ROCm/Vulkan software caveats apply, slightly sharpened because the gfx1151 is a newer target with less mature tooling than gfx1100. It works; expect the occasional fiddle.
- Pricey for what many readers need. At ~$2,000, it's a serious investment. Most readers' workloads don't need unified memory — they need a 24GB discrete card in a cheaper host. The mini PC only earns its price when you specifically need the large-memory-pool-in-a-tiny-box combination.
Who should buy one
- You want a single tiny, quiet box that can occasionally run 70B. This is the genuine fit. No other PC form factor gives you large-model capacity in this footprint at this price.
- You want multiple models served from one machine. The iGPU + optional discrete card topology is real and useful for that.
- You're priced out of a Mac Studio but want the unified-memory story. The EVO-X2 is the PC-world alternative at a lower price point.
Who shouldn't
- Your workload fits in 24GB. Then you don't need unified memory — you need a 24GB discrete GPU in a cheaper host. Faster, cheaper, simpler.
- You want the fastest 70B inference. Unified memory fits the model but doesn't run it fast. A multi-GPU PC wins on speed.
- You're budget-constrained. At ~$2,000, this is well above the $1,000 build tier that covers most readers' actual needs.
The takeaway
The unified-memory mini PC is a legitimate and exciting category in 2026, not marketing fluff — we run the chip in our lab and it does what it claims within honest limits. The Ryzen AI Max+ 395 in the GMKtec EVO-X2 (and peers) genuinely delivers large-model capacity in a tiny, quiet box, and for the reader who wants "70B-capable without a server rack," it's arguably the best single-box PC option. The honest qualifiers: it fits large models but runs them slowly (bandwidth-bound), it's not a discrete-GPU replacement for smaller workloads, and at ~$2,000 it's a serious investment that only pays off if you specifically need the unified-memory large-pool capability.
For most readers, the math still points to a 24GB discrete GPU in a cheaper host — faster, cheaper, and covers the daily-driver tier that's the actual workload. The mini PC is the right answer for a specific reader, not the default one.
Next: Model Fit to check whether your target model actually needs unified memory or fits in 24GB; the VRAM guide for the tier decision; Apple Silicon and MLX for the direct unified-memory comparison.
The EvoX2 is one of our four lab machines — see the fleet for its role and specs. Mini PC pricing reflects the US market as of July 2026. Sources: GMKtec EVO-X2, r/LocalLLaMA EVO-X2 thread.
KEEP READING
AMD Radeon for local AI: the ROCm reality check
Is AMD a real option for local LLMs in 2026, or still the 'it works but...' alternative? The honest answer, from running RDNA3 and RDNA4 cards in the lab: usable, genuinely good value at the 24GB tier, with one persistent caveat you need to know before you buy.
2026-07-3110 minApple Silicon and MLX for local AI
A Mac is the only machine where 'unified memory' means a 70B model fits without a discrete GPU. Is MLX on Apple Silicon a real local-AI path in 2026, or a niche? The honest answer for Mac owners — and the one thing that decides whether it's worth it.
2026-07-3110 minA home AI lab build under $1,000
A complete, tested parts list for a real local-AI machine at the four-figure tier — what to spend on, what to cheap out on, and the one upgrade that's actually worth the money. Built around the 12GB sweet spot for 7B–13B daily and 30B MoE when you need it.
2026-07-3111 min