Sooner or later every local-AI builder hits the same wall: the model you want doesn’t fit in the box you can afford. And the internet’s favourite answer is “just cluster it” — two cheap boxes, network them together, pool the memory. The pitch is seductive: two Strix Halo mini PCs at roughly $1,700 each pool 256GB of unified memory for less than the price of one big multi-GPU box, and each box runs quietly on wall power. We run a four-machine lab with multi-GPU nodes in it, so we can answer the clustering question with measurements rather than forum optimism.
Here’s the short version: clustering adds memory, not speed. Combining memory across GPUs and boxes works, and for giant MoE models it’s sometimes the only affordable path. But if the model fits on one device, a second one adds almost nothing — and once you cross a network, latency and complexity eat the rest. For most buyers, one bigger box beats two cheap ones. There’s one important exception, and we’ll get to it.
The pitch: two cheap boxes vs one big box
The 2026 hardware market makes the cluster pitch unusually tempting. On the cheap side: a GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB unified memory, 256GB/s bandwidth) was around $3,499 in August 2026 and rising per datahardware.ai — the 128GB tier has been reported around $1,700–$3,500 depending on SKU and timing (prices move weekly in the current DRAM shortage). On the big-box side: a Mac Studio M5 Ultra starts at $5,499 with up to 192GB+ of unified memory at ~819GB/s, and NVIDIA DGX Spark-class boxes (64GB, 273GB/s) run about $4,999 via OEMs, with the 128GB config pushed to roughly $6,950 amid the same shortage.
So the arithmetic looks like this:
| Path | Approx. cost | Pooled memory | Bandwidth per pool |
|---|---|---|---|
| 2× Strix Halo mini PC (128GB each) | ~$3,400–$7,000 | 256GB | 256GB/s per box |
| 1× Mac Studio M5 Ultra (192GB+) | from $5,499 | 192GB+ | ~819GB/s |
| 1× DGX Spark-class box (64GB) | ~$4,999 | 64GB | 273GB/s |
| 2× Radeon AI PRO R9700 (32GB each) | ~$2,000–$2,600 | 64GB | high, per-card |
The cluster path pools the most memory per dollar. The big-box path buys far more bandwidth. Both claims are true at once, and which one matters depends entirely on whether your bottleneck is capacity or speed.
What our lab data actually shows about combining GPUs
We’ve measured this directly, twice, on real workloads.
Memory pooling works; speed doesn’t double. In bm-013, we ran the same Qwen3.6-27B model — byte-identical weights — across four machines. The dual-R9700 node (2× 32GB) generated at 26.74 tok/s; the single RX 7900 XT (20GB) hit 25.21 tok/s. Two 32GB cards beat one 20GB card by about 6%. The model fits on either, so the second card contributed almost nothing to speed — it just added 44GB of headroom. The dual-NVIDIA laptop node (2× RTX 5070 12GB, tensor-split) was actually slower at 24.28 tok/s despite being Rank 1 on prompt processing by roughly 2×. That’s the pattern: combining memory works, speed doesn’t magically double.
MoE models are the exception that proves the rule. In bm-009, the same dual-RTX-5070 laptop node generated at 172.05 tok/s on Qwen3-Coder-30B-A3B — a Mixture-of-Experts model that activates only a fraction of its parameters per token. The dual-GPU tensor split (~19.0GB across the two cards) works brilliantly for MoE because each token only needs a few experts, so the inter-GPU traffic stays low. That’s the one workload class where a multi-GPU split genuinely pays.
Backend choice matters as much as topology. In bm-004, llama.cpp Vulkan with MTP speculative decoding hit 66.0 tok/s on a single R9700 — 2.1× the HIP baseline’s 31.9 tok/s on the same card. Software tuning can matter more than adding a second GPU.
Why LLM clustering is harder than game clustering
Gamers cluster GPUs (SLI/CFX) for frames; LLM builders cluster for memory. The mechanics are completely different, and the failure modes are too.
- Pipeline parallelism is serial. Most multi-box LLM setups split model layers across machines (pipeline parallelism). While layer 1 computes on box A, box B waits. You never get 2× the speed; you get roughly the speed of the slowest link, plus network hops.
- Tensor-split overhead is real. Splitting one model across GPUs (as our dual-5070 laptop does with
-ngl 999and a tensor-split config) requires constant inter-GPU communication every layer. For dense models, that communication is overhead, not acceleration — which is exactly what bm-013 showed. - Network latency compounds. Two boxes on 2.5GbE add real per-token latency once activations cross the wire. Even Thunderbolt 5 RDMA setups — the fastest consumer interconnect — trade latency for capacity.
- Software is the hard part. Multi-node llama.cpp, EXO, and similar tools work, but they’re fiddly, version-sensitive, and break in ways single-box setups don’t.
What the community reports (not our measurements)
There’s genuine community work on clustering consumer boxes. We haven’t reproduced any of it, so treat these as reported results, not our data:
- Strix Halo RDMA clusters: GitHub user kyuz0 has published guides on clustering Strix Halo boxes over RDMA (reported ~232 upvotes on r/LocalLLaMA). Reported as working, with the usual multi-node caveats.
- Mac Studio TB5 RDMA: Jeff Geerling documented clustering Mac Studios over Thunderbolt 5 RDMA, pooling up to 1.5TB across machines (228 comments on his write-up). Impressive capacity pooling; not a speed solution.
- EXO on Spark + Studio: ExoLabs’ blog claims 4× scaling running models across DGX Spark and Mac Studio nodes. Vendor-adjacent claim; we haven’t verified it.
The pattern across all of them: clustering demonstrably works for pooling memory across boxes. None of them claim it makes inference faster than a single well-configured box — because it doesn’t.
When clustering wins
- Giant MoE capacity. If your target is gpt-oss 120B or larger MoE models that don’t fit any single affordable box, pooling is sometimes the only path. MoE’s sparse activation makes cross-box traffic tolerable — our dual-5070 MoE result (172 tok/s) shows the pattern.
- Availability. Two cheap boxes mean one can fail without taking your lab down. A single big box is a single point of failure.
- Spreading cost over time. Buy one box now, add the second when budget allows. Each box is independently useful from day one.
When clustering loses
- Latency-sensitive speed. If you want fast interactive generation on a model that fits on one device, clustering adds latency and complexity for near-zero speed gain — the bm-013 result in miniature.
- Complexity budget. Two OS installs, two runtimes, network config, and a distributed runtime that breaks differently than single-box tools. Your time has a cost.
- Power and space. Two boxes, two PSUs, more wall power, more heat, more noise sources.
- Resale. Two mid-tier boxes are easier to sell than one exotic one, but neither holds value like a single high-demand card.
The honest verdict
For most buyers, one bigger box beats two cheap ones. If your target model fits in 24GB, buy a 24GB card and stop (how much VRAM do you need). If it needs 64–96GB, a single unified-memory box — Strix Halo mini PC or Mac — delivers it with far less friction than any cluster. Clustering only wins when capacity is the binding constraint and the model is MoE, where sparse activation keeps cross-node traffic low.
The exception: if you’re targeting giant MoE models (gpt-oss 120B-class and up) that no single affordable box can hold, two Strix Halo boxes pooling 256GB is a legitimate, community-demonstrated path — just go in expecting capacity, not speed.
Next: Single vs dual GPU for the in-box version of this decision; the four-node fleet for how our lab machines divide this labour; Model Fit to check whether your target model fits on one box.
Community results (kyuz0’s Strix Halo RDMA guides, Geerling’s Mac Studio TB5 clustering, ExoLabs’ EXO claims) are reported, not reproduced in our lab. Pricing reflects the US market as of October 2026; prices move weekly amid the ongoing DRAM shortage. Sources: bm-013, bm-009, bm-004.
The EvoX2 is one of our four lab machines — see the fleet for its role and specs.