If your team runs on ChatGPT and you handle client data, contracts, or anything a regulator would call personal information, you have a compliance exposure no enterprise plan fully closes: the data still leaves the building. The air-gapped alternative — a private AI stack serving ChatGPT-style tools from a box you own, with no external network path — has moved from paranoid novelty to a procurement line item, and late 2026 is a good moment to buy: the DRAM shortage is pushing unified-memory prices up, which strengthens the buy-now case rather than weakening it.
We run a four-machine lab and have measured every tier of this decision on real hardware. This is the sober version: the compliance case, three reference stacks by budget, the software layer, the math against SaaS, and the maintenance reality nobody puts in the pitch deck.
The compliance case, stated plainly
The pitch is simple: no data leaves the building. A properly air-gapped stack has no external network path at all — the model, the documents, and the queries all live on hardware you control. That is a categorically different privacy posture from any cloud plan, enterprise or otherwise.
We have written the detailed versions elsewhere: why local AI is better for privacy makes the core case; is local AI actually private? A threat model covers the honest caveats — local inference protects the prompt, not your whole security posture; and local vs. cloud AI maps the tradeoffs. One caveat before the hardware: air-gapping the model box is the easy part. If your workflow involves RAG over company documents, the retrieval layer is where most privacy designs quietly leak — see our private RAG guide.
Three reference stacks by budget
Stack 1: $1,700–$3,500 — one Strix Halo box
The single-box option is a unified-memory mini PC on AMD’s Ryzen AI Max+ 395 — the GMKtec EVO-X2 class of machine, with the Beelink GTR9 Pro as the value alternative. The EVO-X2’s 128GB tier was around $3,499 in August 2026 and rising (per datahardware.ai); the GTR9 Pro was roughly $1,800 earlier in 2026 (per hardware-corner.net, which also lists the comparable Bosman M5 128GB at $1,699). Prices move weekly right now — treat every figure as a band, not a quote.
This is the tier where we have the most direct evidence, because the EVO-X2 is one of our lab machines:
- Small models are genuinely fast. In bm-011, the LFM2.5-8B MoE (Q4_K_M) generated at 150.16 tok/s on the integrated Radeon 8060S, with prompts processed at 3,661.53 tok/s. Comfortably interactive for a small team on an 8B-class assistant.
- 27B-class models are usable. In bm-013, the same machine’s discrete RX 7900 XT generated Qwen3.6-27B at 25.21 tok/s — within 10% of the dual-GPU machines in that test. That is the daily-driver tier for coding and document work, and it is real.
The honest limits: unified memory has less bandwidth than discrete VRAM, so 70B-class models fit in the pool but run slowly — the value here is capacity and simplicity, not speed at the top end. And the memory is soldered: the 128GB decision is permanent.
Stack 2: $5,500+ — Mac Studio M5 Ultra
The Apple path is a Mac Studio M5 Ultra: base at $5,499, configurable to 192GB+ of unified memory at ~819GB/s bandwidth (per runaihome.com). At that memory tier it runs 120B-class models like gpt-oss 120B locally — the spec-sheet case is strong.
Spec-sheet analysis — we have not measured this hardware. Our lab is AMD, Intel Arc, and NVIDIA Blackwell; there is no Mac Studio in the fleet and no first-party numbers. The bandwidth figure and the 120B-class capability claims are from published comparisons, not our bench. Treat performance expectations as directional until you can test your own workload.
The advantages are architectural: no VRAM ceiling, a mature host for local tooling, a silent box. The costs: soldered memory that is expensive to configure up, and a lock to Apple’s ecosystem for the life of the machine.
Stack 3: $10,000+ — dual-box or DGX Spark cluster
At the top tier you are buying either redundancy or capacity, and you should know which.
The dual-box path is two discrete-GPU machines, each serving a different workload — one on coding, one on documents/RAG. Our lab runs this topology. The Radeon AI PRO R9700 (32GB, RDNA4) is the card we would build around: in bm-004, a single R9700 generated Qwen3.6-27B with MTP speculative decoding at 66.0 tok/s — 2.07x the no-MTP baseline of 31.9 tok/s on the same card. In bm-005, the same card ran the Ornith-1.0-35B MoE at 78.3 tok/s at 96K context. Those are our measurements, on our hardware.
The DGX Spark path is NVIDIA’s integrated answer: the new 64GB config launched October 2, 2026 at ~$4,999 via OEMs (Acer, ASUS, Dell, Gigabyte, HP, MSI), with the 128GB tier pushed to ~$6,950 amid the DRAM shortage. Two spec-sheet caveats: bandwidth is 273GB/s LPDDR5X — well below the M5 Ultra’s 819GB/s — and it is not sold on Amazon; you buy through OEM/direct channels. Community comparisons (per runaihome.com) put Spark-class 128GB boxes decoding gpt-oss 120B in the low-to-mid 30s tok/s. Spec-sheet and community figures — not our measurements.
Also price the boring stuff: a second machine means a second PSU, a second maintenance surface, and a second thing to patch. Dual-box earns its cost when two teams or two workloads genuinely need isolation; otherwise one well-configured box is the better buy.
The software stack
Hardware is the cheap half of this decision. The software layer is where private stacks succeed or fail, and it is simpler than the vendor pitch suggests:
- Serving: Ollama for team-friendly model management, or llama.cpp directly for control. For multi-user serving at real concurrency, vLLM is the serious option.
- Interface: Open WebUI on top, which gives your team a ChatGPT-style chat interface against your local models. This is the piece that makes the stack feel like a product instead of a terminal.
- Documents: if the point is querying company knowledge, the private RAG layer is a separate build with its own failure modes — budget time for it.
None of this requires a specialized platform vendor. The stack is open-source and identical whether you buy the $1,700 box or the $10k cluster.
The math vs. SaaS
A ChatGPT-class plan runs about $20/user/month. A ten-person team is $2,400/year — roughly the price of one Strix Halo box, before the DRAM shortage finishes pushing that number up. Break-even on the $1,700–$3,500 stack is roughly one to two years for ten people, faster at twenty.
Three honest qualifiers:
- The box does not replace SaaS at parity. Local models are weaker at frontier tasks, and the workflow change is real. Budget for a hybrid period where both run.
- Electricity and maintenance are real costs. A mini PC draws tens of watts idle and a few hundred under load; a multi-GPU box draws more. It is free of per-prompt billing, not free to run.
- Your time is the hidden line item. Someone on the team owns the box. If nobody owns it, the stack decays into a shelf ornament.
For most regulated teams the compliance case, not the cost case, is the stronger argument. The math is favorable; the privacy posture is the reason to actually do it.
What nobody tells you
- Maintenance is ongoing. Model updates, driver updates, host-OS security patches. A private stack is infrastructure, not an appliance you set and forget.
- Model churn is a decision you now own. New models land monthly — Qwen3.8 27B, gpt-oss 120B, the current coder models. Someone evaluates them and re-validates the stack each cycle.
- There is no magic RAG. “Just point it at our documents” is the most common failure mode we see. Retrieval quality determines answer quality, and building that layer well is a project — see the private RAG guide.
- Air-gapped does not mean secure. It means no external network path for model traffic. The host OS, the update process, and the people using it are still attack surface — the threat model post covers what local inference does and does not protect.
Decision checklist
- What data triggers the requirement? Client contracts or regulated personal data make the air-gapped case strong; general productivity work may be cheaper to solve with a cloud enterprise plan.
- What model tier does the work? Run your target workload through Model Fit first. If 8B–30B models cover it, the $1,700–$3,500 stack is sufficient and the bigger boxes are overkill.
- How many concurrent users? A single Strix Halo box serves a small team interactively on small models; heavier concurrency or sustained 27B-class serving pushes you toward the dual-box tier.
- Who owns it? Name the person. An unowned stack is a dead stack.
- Buy now or wait? The DRAM shortage is pushing unified-memory prices up, not down. If the compliance case is real, buying sooner costs less than buying later.
Where this data comes from
The measured numbers come from our own lab benchmarks: bm-011 (LFM2.5-8B on the Ryzen AI Max+ 395 iGPU), bm-013 (Qwen3.6-27B across all four lab machines), and bm-004 / bm-005 (Radeon AI PRO R9700 backend showdowns). Mac Studio, DGX Spark, and mini-PC pricing are spec-sheet and market figures from published sources (runaihome.com, datahardware.ai, hardware-corner.net), labeled as such — we have not measured that hardware. Prices are bands as of October 2026 and move weekly.