● CALIBRATED 2026-10-03 · REC 000
Local AI Frontier

Unified Memory for Local AI: Mac Studio vs DGX Spark vs Strix Halo

Unified memory is the reason a book-sized box can hold a 70B model at all. In 2026 the category has exploded: Apple’s M6 and M5 Macs, NVIDIA’s DGX Spark, and AMD Strix Halo mini PCs all sell the same architectural idea — one pool of fast memory addressed by both CPU and GPU — at prices from $899 to roughly $6,950. But the boxes are not interchangeable. They differ by up to 3x on memory bandwidth, which is the spec that decides how fast models generate, and only one of them has been measured in our lab. This page is the decision framework: what unified memory changes, what each contender offers, what our data says, and how the late-2026 RAM shortage should bend your timing.

What unified memory actually changes

On a conventional PC, system RAM and VRAM are separate pools, and the VRAM ceiling decides what fits — that’s the whole logic of our VRAM guide. Unified memory removes the ceiling: one pool the GPU can address nearly in full. We’ve seen this literally on our own hardware — in bm-011, the Ryzen AI MAX+ 395’s iGPU reported its device memory as 126,976 MiB: the entire ~124GB system pool, exposed as VRAM.

The trade is bandwidth. A discrete GPU reads its GDDR memory at a 900+ GB/s class rate; unified platforms run from 170 GB/s (Mac mini M6) to ~819 GB/s (Mac Studio M5 Ultra, per runaihome.com). Capacity is the easy part — bandwidth is the tax, and bandwidth sets generation speed on dense models.

One big exception: MoE models. A mixture-of-experts model reads only its active parameters per token, not the whole file. That’s why our measured unified-memory numbers look better than the bandwidth arithmetic suggests, and why the current model cycle (Qwen3-Coder-30B-A3B, gpt-oss 120B, Qwen3.8 27B) is MoE-heavy. 2026 is the first year unified-memory boxes make sense for more than bragging rights.

The 2026 contenders

Platform Memory Bandwidth Price band Buy where Our data
Mac mini M6 16–32GB 170 GB/s $899–$1,299 Amazon spec-sheet only
Mac mini M5 Pro up to 64GB 307 GB/s ~$1,600–$2,200 Amazon spec-sheet only
Mac Studio M5 Max up to 128GB 546 GB/s (as listed) $2,500–$5,100 (128GB spec’d at $5,099) Amazon spec-sheet only
Mac Studio M5 Ultra up to 192GB+ ~819 GB/s (per runaihome.com) $5,499+ Amazon spec-sheet only
NVIDIA DGX Spark 64 / 128GB 273 GB/s ~$4,999 / ~$6,950 OEM direct — not on Amazon spec-sheet only
GMKtec EVO-X2 (Strix Halo) 64–128GB (96GB usable at 128) 256 GB/s ~$1,700–$3,500, 128GB rising Amazon measured (bm-011, bm-013)
Beelink GTR9 Pro (Strix Halo) 64–128GB 306 GB/s (as listed) ~$1,800 earlier 2026 Amazon spec-sheet only
Reference: used RTX 3090 24GB VRAM 900+ GB/s class $600–$900 used Amazon 24GB tier measured (Arc B60, R9700)

Every price is a band and moves weekly — the RAM shortage (below) is actively pushing several of these up. “Spec-sheet only” means we have not measured that machine; treat its speed claims accordingly.

What our measured Strix Halo data says

We run one unified-memory platform in our lab: the GMKtec EVO-X2 (Ryzen AI Max+ 395, Radeon 8060S iGPU). In bm-011 we ran LFM2.5-8B-A1B (Q4_K_M, a 4.79GB MoE with ~1B active parameters) on the iGPU: 150.16 tok/s generation and 3,661.53 tok/s prompt processing under ROCm 7.2.2, with tight variance across three runs.

Two honest footnotes. First, the same model on the same machine’s discrete RX 7900 XT (bm-007) hit 269.01 tok/s — the iGPU is ~1.8x slower on this small model. The iGPU’s point was never speed; it’s the pool. Second, gfx1151 is a newer ROCm target than the mature gfx1100: our short runs were stable, but sustained multi-hour loads are unverified, and the report notes the VRAM figure is the model footprint rather than a clean runtime allocation — an inherent reporting ambiguity on APU hardware we document rather than hide.

The platform itself is real. In bm-013 we ran byte-identical Qwen3.6-27B weights across all four lab machines, and the EvoX2’s discrete card generated at 25.21 tok/s — within 10% of the dual-R9700 (26.74) and dual-RTX-5070 (24.28) machines. Strix Halo is a serious inference host; the unified-memory iGPU is the part that removes the capacity ceiling.

Capacity vs bandwidth: what fits vs what runs fast

What fits (capacity math, anchored to our measured footprints):

  • 96GB usable (the EVO-X2 128GB SKU exposes 96GB to the GPU, per sunkcost.ai): a 70B-class model at Q4 — roughly 40GB, scaled from our measured 18.1GB for a 30B at Q4_K_M — fits with real context headroom.
  • 128GB (DGX Spark 128GB, Mac Studio M5 Max): 120B-class models fit. ASUS’s Spark-class Ascent GX10 decodes gpt-oss 120B in the low-to-mid 30s tok/s, per runaihome.com — a community figure, not ours.
  • 192GB+ (Mac Studio M5 Ultra): 120B-class plus large context, or several large models resident at once.

What runs fast (bandwidth arithmetic — a ceiling, not a measurement): a dense 70B at Q4 reads ~40GB per generated token. At 256 GB/s that’s a theoretical ceiling around ~6 tok/s; at ~819 GB/s, around ~20 tok/s; at 900+ GB/s, ~22+ tok/s. Real-world lands below the ceiling. That’s the whole trade in three numbers: unified memory fits the big dense model; only the top-end Mac pays bandwidth close to discrete speed, at more than double a Spark’s price.

The MoE escape hatch: our measured Qwen3-Coder-30B-A3B (3B active) ran at 38.6 tok/s on a 24GB Arc B60 (bm-001), and the 1B-active LFM2.5 hit 150.16 tok/s on Strix Halo. If your large models are MoE — and the 2026 lineup mostly is — unified memory’s bandwidth tax shrinks dramatically.

The honest Mac comparison

We have not measured a Mac. Our fleet is x86 and discrete-GPU-centric — Arc B60 Pro, dual Radeon AI PRO R9700, dual RTX 5070 laptop, and the EvoX2 — so every Mac number here is spec-sheet or community-sourced, and labeled as such. What the specs imply:

  • The M5 Ultra’s ~819 GB/s (per runaihome.com) is the only unified-memory bandwidth in consumer reach that approaches discrete-GPU territory. If you want big dense models at respectable speed in one quiet box, it’s the only machine on this page architected for it — at $5,499 and up.
  • The M6 mini’s 170 GB/s is an entry tier: llmcheck.net estimates 9B-class models around ~26 tok/s there (community estimate). Fine for small models; the 32GB ceiling caps the unified-memory pitch.
  • The software story differs too: MLX is Apple-native and strong, llama.cpp is the portable fallback — we cover that split in our Apple Silicon piece. Memory is soldered and non-upgradable on all of this, and there’s no CUDA.

The RAM-shortage pricing reality

Late 2026’s DRAM shortage is documented, and it lands hardest on exactly the SKUs this page is about. The EVO-X2 128GB sat around ~$3,499 in August 2026 and was rising (per datahardware.ai); the DGX Spark 128GB config was pushed to ~$6,950 amid the shortage; the Bosman M5 128GB mini was $1,699 earlier in the year (per hardware-corner.net). Memory-heavy configs are the ones moving.

We don’t forecast prices — our timing framework exists because guessing markets isn’t our job. But two things are knowable. First, the documented direction of travel on big-RAM SKUs is up. Second, memory is the one spec you can never add later on any machine in this table — it’s soldered at purchase. Put together: if you’ve already decided you need the 96–192GB tier, the trend argues against waiting on the big-RAM config. If you haven’t decided the tier yet, wait on the decision — Model Fit — not on the price.

Who should buy what

Profile 1 — your workload fits in 24GB (most readers). Don’t buy unified memory. A used RTX 3090 ($600–$900, moves weekly) or an Arc Pro B60 ($450–$500) is faster per dollar than anything in the table above, and the 30B-class MoE models that dominate daily work run brilliantly there. Run your target model through Model Fit first; most people discover they never needed the big pool.

Profile 2 — big models, small box, real budget. A Strix Halo mini PC (GMKtec EVO-X2 or Beelink GTR9 Pro) is the measured, cheapest entry into genuine unified memory: 96GB usable at the top SKU, near-silent, and the only platform here with first-party numbers on this site. Dense 70B fits but crawls; MoE models run well. See our mini PC guide and the /mini-pcs/ hub for the model-by-model picture.

Profile 3 — capacity, bandwidth, and silence, budget secondary. Mac Studio M5 Max or M5 Ultra — or the DGX Spark if you specifically want NVIDIA’s software stack, bought direct from OEMs since it isn’t on Amazon. The Ultra is the only unified-memory machine in this tier that also has near-discrete bandwidth. Read our DGX Spark reality check before committing to the NVIDIA path; its 273 GB/s is Spark-class across both configs, and the price bands moved up this quarter.

Where this data comes from

Mac, DGX Spark, and pricing figures are spec-sheet and community-sourced (named above: runaihome.com, llmcheck.net, datahardware.ai, hardware-corner.net), not our measurements. Bandwidth ceilings are labeled arithmetic. Prices are bands and move weekly.