No box in local AI carries more badge value than the NVIDIA DGX Spark: the DGX name runs datacenter training clusters, and this is a desk-sized version of it. Launch coverage treated it as the definitive local-AI machine. Six months, one new configuration, and one OEM variant later, the receipts from people who actually own one are more complicated than the keynote. We have not benched a Spark in our lab and will not pretend otherwise — but we can put its spec sheet next to hardware we have measured, and next to owners’ first-hand reports, and ask whether the price buys speed or just buys NVIDIA.
What the DGX Spark actually is
The Spark is a desk machine built on NVIDIA’s GB10 superchip, with 128GB of LPDDR5X unified memory shared between CPU and GPU and 273GB/s of memory bandwidth (as listed by NVIDIA). One big shared pool, no discrete VRAM — architecturally a cousin of the unified-memory platforms in our unified-memory guide, not of the multi-GPU towers we benchmark.
Two configurations matter as of October 2026:
- A new 64GB config launched October 2, 2026 at around $4,999, sold through OEM partners (Acer, ASUS, Dell, Gigabyte, HP, MSI) and direct — pricing in this segment moves weekly.
- The 128GB config has been pushed to roughly $6,950 amid the late-2026 DRAM shortage — shortage pricing, and it moves weekly too.
One purchase fact, stated plainly: the DGX Spark is not on Amazon. It is OEM/direct-only; NVIDIA’s DGX Spark page is the pointer for current configurations.
And one honesty fact: this is spec-sheet analysis — we have not measured this hardware. Every performance claim below is a named external source or a measurement from our own fleet on different silicon, labeled as such.
The hype versus the receipts
The Spark’s first months produced an unusually readable paper trail, and the pattern across it is consistent: well-made hardware, with software and marketing ahead of the delivered performance.
- Simon Willison, after hands-on time, called it “great hardware, early days” — a verdict that has aged well.
- The lmsys team’s in-depth review remains the best public technical walkthrough of what the Spark does out of the box and what it doesn’t.
- John Carmack — hardly a hostile witness — publicly assessed that the Spark delivers roughly “half advertised performance” in real workloads. His claim, not our measurement.
- A Reddit thread titled “NVFP4 still missing after 6 months” captures the software gap: flagship quantization support buyers expected at launch still had not landed half a year in. Community-reported, but consistent with the theme.
- Jeff Geerling’s review of the Dell variant found it fixes several of the original Spark’s pain points — encouraging, and equally telling that the first revision needed to.
None of this says the Spark is bad. It says it is premium-priced developer hardware whose real-world throughput has repeatedly come in below the headline, per the people who measured it. For a machine whose pitch is the DGX name, that gap is the story.
What 273GB/s actually buys
Bandwidth is the honest way to price a unified-memory box, because token generation is bandwidth-bound: every token streams the active weights through memory once. So what does the Spark’s 273GB/s buy?
Our lab’s closest measured analog is the GMKtec EVO-X2 (Ryzen AI Max+ 395, Strix Halo), whose unified pool runs at 256GB/s — about 6% less paper bandwidth than the Spark. In bm-011, that iGPU generated 150.16 tok/s on LFM2.5-8B-A1B at Q4_K_M, with prompt processing at 3,661.53 tok/s. Same bandwidth class, same architecture class — the honest expectation for the Spark on small models is this speed class, not discrete-GPU speed.
Bigger models scale sublinearly but unforgivingly: generation speed tracks the parameters each token activates, not the total on-box count. Per runaihome.com’s measurement of the Spark-class ASUS Ascent GX10, gpt-oss 120B decodes in the low-to-mid 30s tok/s — a community measurement, not ours, but it lands where the bandwidth math says it should.
- Software matters as much as bandwidth. On the same EVO-X2 machine, the discrete RX 7900 XT ran the identical model at 269.01 tok/s — 1.8x the iGPU (bm-007). Paper bandwidth is not delivered tok/s.
- The Mac comparison is brutal on paper. The Mac Studio M5 Ultra’s 819GB/s (as listed) is 3x the Spark’s bandwidth, with up to 192GB+ of unified memory from a $5,499 base — spec-sheet analysis; we have not measured a Mac.
| Platform | Memory | Bandwidth | Price band | Evidence basis |
|---|---|---|---|---|
| NVIDIA DGX Spark (GB10) | 64GB / 128GB LPDDR5X | 273GB/s | ~$4,999 / ~$6,950, moves weekly | Spec sheet + community reports; not measured by us |
| ASUS Ascent GX10 (GB10) | 128GB | 273GB/s | $3,000-$4,000 (verify — moves weekly) | Spec sheet; gpt-oss 120B low-to-mid 30s tok/s per runaihome.com |
| GMKtec EVO-X2 (Strix Halo) | 128GB (96GB usable) | 256GB/s | ~$3,499 and rising | Measured by us: bm-011, bm-013 |
| Mac Studio M5 Ultra | up to 192GB+ | 819GB/s | $5,499+ | Spec sheet; not measured by us |
Who the Spark is actually for
Strip the hype and two real buyers remain:
- Developers who need CUDA-everywhere. The GB10 runs the same CUDA stack as NVIDIA’s datacenter hardware, so code validated on the desk behaves the same in production. If your job is building against NVIDIA’s toolchain — and you accept the “early days” caveats above — the Spark is the only desk box that guarantees that parity. An ecosystem purchase, and a legitimate one.
- Cluster tinkerers. The community networks multiple Sparks with EXO into larger effective pools — as reported practice it works, as a hobby rather than a productivity plan.
If neither description is you, the Spark’s price is buying you a badge.
Who should skip it
Most buyers, honestly. Three concrete exits:
- Your models fit in 24GB. The 7B-30B class — where the models we actually recommend live — is covered by the lab rule from our platform-decision framework: a ~$300 discrete GPU is faster and cheaper than any unified-memory box. Run your target model through Model Fit and the VRAM guide before spending five figures’ worth of consideration on a 64GB pool.
- You want unified-memory capacity at a sane price. Strix Halo delivers the same architecture for roughly half the money: the EVO-X2’s 128GB SKU sat around $3,499 in August 2026 and has been rising (per datahardware.ai), while earlier Strix Halo boxes came in lower still — the Beelink GTR9 Pro at about $1,800 earlier in 2026, and the Bosman M5 128GB at $1,699 (per hardware-corner.net). We measured this platform class — see our mini-PC guide.
- You want the fastest large-model inference per dollar. At the 128GB Spark’s ~$6,950, the Mac Studio M5 Ultra’s $5,499 base buys triple the paper bandwidth — spec-sheet reasoning on our side, but the direction is hard to argue with.
The shortage warning deserves its own line: the late-2026 DRAM shortage is pushing unified-memory prices up across the board, and the Spark’s 128GB config at ~$6,950 is shortage pricing stacked on an already-premium product. Assume the price moves weekly, and check our buy-now-or-wait framework first.
The takeaway
Is the DGX Spark worth it? For most local-AI buyers: no. It is well-built hardware with a genuine niche — CUDA parity on the desk, cluster experiments — at a price that buys NVIDIA’s ecosystem rather than throughput. The receipts say “half advertised performance” (per Carmack), “early days” (per Willison), and missing flagship quantization six months in (per the community). Our measured analog says 273GB/s is the same speed class as a Strix Halo mini PC at roughly half the price, and a third of the bandwidth of a similarly-priced Mac Studio M5 Ultra.
Buy it if you are the CUDA-everywhere developer or the cluster tinkerer and accept the premium as the price of admission. Skip it if you wanted the fastest local inference for the money — that money goes further on discrete GPUs, Strix Halo, or Apple silicon, depending which constraint binds you.
Where this data comes from
- bm-011 — LFM2.5-8B on the Ryzen AI Max+ 395 iGPU: the 150.16 tok/s unified-memory anchor at 256GB/s.
- bm-007 — LFM2.5-8B on the same machine’s RX 7900 XT: the 269.01 tok/s discrete-GPU contrast.
- bm-013 — Qwen3.6-27B cross-machine baseline: how the EVO-X2 platform ranks against our dual-GPU machines.
We have no DGX Spark, ASUS Ascent GX10, or Mac in our benchmark fleet. All Spark, GX10, and Mac figures here are spec-sheet values or named community measurements (runaihome.com, Simon Willison, lmsys, John Carmack, Jeff Geerling, r/LocalLLaMA), labeled as such throughout. Pricing reflects the US market as of October 2026 and moves weekly.