● CALIBRATED 2026-10-03 · REC 000
Local AI Frontier

CPU vs GPU for Local LLMs: When 128GB of RAM Beats a Graphics Card

Every local-AI buying guide starts the same way: buy a GPU. That advice is right about speed and wrong about everything else, because it skips the question that actually decides most builds — where do the model weights live? A GPU generates tokens fast but can only hold what fits in its VRAM; system RAM is slow but cheap and enormous. This page is the honest map of that trade: when the graphics card wins, when 128GB of RAM beats it, and where the new generation of MoE models has quietly rewritten the rules.

Why GPUs dominate: bandwidth and parallel math

LLM generation is memory-bandwidth-bound. Every token requires streaming the active weights from memory through the compute units, so the speed of inference tracks the speed of memory far more than the speed of the processor. Discrete GPUs win on both axes: GDDR6X/HBM memory moves hundreds of gigabytes per second, and thousands of parallel cores do the matrix multiplication that makes up nearly all of the work. A CPU has a handful of fat cores and dual-channel DDR5 — it computes the same math, just with a fraction of the memory throughput and parallelism.

Our own cross-GPU data shows the gap on identical workloads. In bm-008, the same Llama-3.1-8B-Instruct at Q4_K_M generated at 105.22 tok/s on our EvoX2 machine’s RX 7900 XT versus 27.01 tok/s on our Arc B60 Pro — a ~3.9x difference between two discrete cards. Prompt processing was even starker: 3,210.77 tok/s versus 706.61 tok/s. The GPU’s parallel-matmul advantage is not a marketing story; it shows up in every measurement we run.

One honesty note on “CPU” numbers: our lab measures GPU-class hardware, so we don’t have a pure-CPU-only figure in the digest. What we do have is an APU data point — in bm-011, the integrated Radeon 8060S in our Ryzen AI Max+ 395 (which draws from the same DDR5-class memory pool a CPU uses) generated the same LFM2.5-8B model at 150.16 tok/s. That’s iGPU-with-fast-unified-memory performance, not a desktop CPU in a socket — treat it as the ceiling of what “CPU-adjacent” hardware does, not a typical CPU result.

The CPU path anyway: capacity you can’t buy as VRAM

Here’s the counterweight: 96GB of DDR5 costs roughly $400–$500 as a kit (moves weekly — the late-2026 DRAM shortage is pushing memory prices up, so treat that band as directional). No GPU at that price holds anything close. A $400–$500 Arc Pro B60 gives you 24GB of VRAM; a $300–$330 Arc B580 gives you 12GB. For the price of one mid-range GPU you can populate a board with 96–128GB of system RAM — and RAM holds models no consumer GPU can.

A 70B model at Q4_K_M needs roughly 40GB+ just for weights, before context. That fits in 96GB of RAM with room to spare; it does not fit in any single consumer GPU under $900. So the CPU path is real: load the big model into system RAM, run inference on the CPU, accept the speed penalty. Expect roughly 1–5 tok/s on a pure CPU for a 70B-class model — this is spec-sheet reasoning based on memory-bandwidth math, not something we have measured; we have not benchmarked pure-CPU inference in our lab. Community reports land in that same band, but verify against your own hardware before you plan around it.

The trade, stated plainly: capacity versus speed. RAM buys you the ability to run a model class the GPU simply cannot hold. The GPU buys you the speed to make a smaller model pleasant to use. Most people overestimate how much speed they need and underestimate how much capacity their target model wants — run your candidate through Model Fit before assuming either.

Hybrid offload: the middle path that mostly disappoints

The obvious compromise is --n-gpu-layers in llama.cpp: keep most layers in RAM, offload some to the GPU, split the work. In practice this is the most misunderstood option in local AI, because generation speed on a bandwidth-bound model is set by whichever memory pool the active weights stream from. Offload enough layers to matter and you’re VRAM-bound again; offload too few and the CPU’s slow memory path dominates the whole loop.

Our bm-013 cross-machine run shows what partial-offload realities look like at the top of the consumer range: Qwen3.6-27B at Q4_K_M used 21.24GB across dual Radeon AI PRO R9700 32GB cards and generated at 26.74 tok/s — with the model fully resident in VRAM. The same model on a single RX 7900 XT (16.82GB used) hit 25.21 tok/s, and dual RTX 5070 12GB cards managed 24.28 tok/s despite the fastest prompt processing (344.85 tok/s, roughly 2x the AMD machines). The lesson: once a model fits in VRAM, adding memory or cards barely moves generation speed — and if it doesn’t fit, partial offload drags you back toward CPU-class speeds. The useful middle ground is narrow.

The NPU reality check

Modern CPUs ship with NPUs marketed for “local AI.” For LLM inference, ignore them for now — NPUs are built for low-power small-model work (and vendor demo workloads), not for serving a 27B model. We cover what they actually do in our NPU reality check. If you’re buying a CPU for local AI, you’re buying cores, memory capacity, and memory bandwidth — not the NPU block.

MoE changes the math

The rule “CPU-class hardware is too slow” was written for dense models, and 2026’s model cycle has quietly invalidated it. Mixture-of-Experts models activate only a fraction of their parameters per token — Qwen3-Coder-30B-A3B activates ~3B parameters per token, not all 30B. Less active math per token means memory bandwidth matters less, and suddenly CPU-class and APU-class hardware produces usable speeds.

Our measurements make this concrete. In bm-009, Qwen3-Coder-30B-A3B at Q4_K_M generated at 172.05 tok/s on Victor’s dual RTX 5070 laptop GPUs — fast, but note the model only needed ~19GB across two cards. More striking is bm-011: the LFM2.5-8B-A1B MoE generated at 150.16 tok/s on an integrated GPU drawing from shared DDR5-class memory — within touching distance of the discrete RX 7900 XT’s 269.01 tok/s on the same model and workload (bm-007). An iGPU at 150 tok/s would have been absurd two years ago. MoE is why CPU and APU inference suddenly works, and it’s why a unified-memory box like the GMKtec EVO-X2 — 128GB of shared memory, Ryzen AI Max+ 395 — has become a serious local-AI platform rather than a curiosity.

Decision table: pick by workload

Your situation Best fit Why
7B–8B models, chat/coding daily driver Discrete GPU (12–16GB) Fastest per dollar; bm-008 shows 105 tok/s on a 7900 XT
27B–30B MoE, full context 24GB discrete GPU Fits entirely in VRAM; 26–44 tok/s measured across our fleet
70B dense at usable quant 96–128GB RAM, CPU or APU No affordable GPU holds it; expect single-digit tok/s (spec-sheet reasoning, unmeasured)
Many models, one box, low noise Unified-memory mini PC EVO-X2-class hardware holds big models in shared memory
Occasional big model, budget build CPU + 96GB DDR5 kit Capacity for ~$400–$500; slow but it runs
Laptops with NPUs Ignore the NPU for LLMs See NPU reality check

The takeaway

The GPU-versus-RAM question is a capacity-versus-speed trade, and the right answer depends entirely on your target model. For the 7B–30B tier that covers most daily work, a discrete GPU is faster and the VRAM math is settled — check how much VRAM you need. For 70B-class models, 128GB of RAM beats any graphics card you can afford, at single-digit speeds. And for the new MoE generation, the gap is narrower than the old advice assumes — an iGPU hitting 150 tok/s on an 8B MoE is our lab-measured proof. Buy capacity for the models you want to run; buy speed for the ones you run all day.

Next: Model Fit to check what your target model actually needs; the VRAM lookup for per-model numbers; unified memory explained for the Apple-Silicon-style alternative; and what to expect for realistic speed targets at every tier.

Every measured figure on this page comes from our own benchmark reports — see the benchmark index for methodology, machine specs, and raw evidence. Prices are US-market bands as of October 2026 and move weekly, particularly memory kits amid the current DRAM shortage.