DISPATCH
What a quantization tier costs in watts: measured
Every 'how much power does local AI use' piece multiplies GPU TDP specs by hours and calls it a day. Almost nobody measures actual wall-power during inference. We measured a narrow slice — and the headline finding is that quantization barely moves power draw. The honest version of what we have, what we don't, and the rough cost math.
Every "how much electricity does local AI use" piece follows the same shape: take the GPU's TDP from the spec sheet, multiply by hours of use, multiply by your kWh rate, done. That is a reasonable starting point. It is also almost never grounded in measured wall-power, which is the only number that captures what the card actually pulls under inference load — not peak synthetic load, not idle, not the manufacturer's worst-case TDP. The TDP-multiplication school is fine for an order-of-magnitude estimate and misleading for anything more precise than that.
This is the narrow slice of measured wall-power we actually have. The headline finding is uncomfortable for anyone who assumed quantization saves power: on the Arc B60, quantization barely moves power draw at all. Q4_K_M, Q6_K, and Q8_0 all measure within 3% of each other under sustained load. The card's wattage is set by the silicon running, not by the precision of the math.
For the broader quantization thesis — why Q4_K_M is the practical default for quality and speed — see quantization in 2026: no compromise. This article is specifically about the watts.
The two data points we have
Honesty first. Of the thirteen benchmark records in our library, only two captured real wall-power: bm-001 and bm-003. Every other record reports power_watts = 0, which is our sentinel for "not measured." We're backfilling. This article is what we have so far, not the comprehensive piece we'd like to write once the rest of the fleet is characterized.
The two records we do have both measured the same workload — Qwen3-Coder-30B-A3B at Q4_K_M (and Q6_K, Q8_0 for bm-003) running our 100-prompt HumanEval suite — on two different cards.
| Card | Workload | Quant | Wall-power |
|---|---|---|---|
| Intel Arc B60 Pro (24GB) | Qwen3-Coder-30B-A3B, sustained | Q4_K_M | 245W |
| Intel Arc B60 Pro (24GB) | Qwen3-Coder-30B-A3B, sustained | Q6_K | 250W |
| Intel Arc B60 Pro (24GB) | Qwen3-Coder-30B-A3B, sustained | Q8_0 | 252W |
| Dual AMD R9700 (single card active) | Qwen3-Coder-30B-A3B, sustained | Q4_K_M | 410W |
That is the entire measured wall-power dataset. Four readings, two cards, one model. The rest of this article interprets those four readings honestly and tells you what we don't know.
Quantization barely moves power draw
The bm-003 finding, made explicit: across Q4_K_M, Q6_K, and Q8_0 on the same Arc B60 under the same workload, wall-power measured 245W, 250W, and 252W. That's a 7W spread — under 3% — across three meaningfully different quant tiers.
This contradicts a folk assumption you'll see in forum threads: "lower quants save power because the math is simpler." The math is simpler, but the silicon doing it is the same. The GPU's power draw is set by how hard its compute units, memory controllers, and VRAM are running, and at sustained inference load those are all running regardless of whether each MAC operation is 4-bit or 8-bit. The card is power-bound by utilization, not precision.
The practical implication: if you were choosing Q4_K_M over Q8_0 to save on your power bill, that's not a real reason. Choose Q4_K_M for VRAM headroom and tok/s — see the quantization piece and the coding-model comparison — not for watts.
The Arc B60's power envelope
245W under sustained 30B-coder load, across three quant tiers, on the Arc B60 Pro. That's a real measured number, not a TDP spec. The B60 Pro's nominal TDP is in the same neighborhood, which is reassuring but not the same thing — TDP is a thermal design target, not a measurement of what the card pulls under any specific workload.
For the cost math (below), 245W is the number to use. If you're running a 30B coder on a B60 Pro for several hours a day, expect the card to sit at or near 245W the whole time. Idle and light-load behavior will be lower; we haven't measured those curves.
The dual R9700's power envelope
410W on the dual-R9700 Ray machine with one card active on the same 30B coder (bm-001). The second card idles at near-zero — recall from Single vs dual GPU that the second card sits idle below the ~18GB crossover, and a 30B at Q4_K_M is right at that line.
The dual-rig premium is real but not 2×. Combined-mode power (both cards tensor-splitting on a model big enough to need them) is not in our library; based on TDP specs you'd expect roughly 1.6–2× the single-card figure, but treat that as an estimate, not our data. The honest read: a dual-GPU rig under single-card-utilization workloads pulls notably more than one card alone, because the rest of the system (CPU, motherboard, the second card's idle draw) doesn't zero out.
The rough electricity-cost math
For the Arc B60 at 245W, the math:
- 245W × 8 hours/day = 1.96 kWh/day
- 1.96 kWh × 30 days = 58.8 kWh/month
- 58.8 kWh × $0.16/kWh (US average residential) = $9.40/month
So running a 30B coder for 8 hours a day on the Arc B60 costs roughly $9–$10 a month in electricity at average US rates. Your number will differ by GPU (the dual R9700 at 410W roughly triples this to ~$16/month), by utilization pattern (most home users run far less than 8 hours of sustained inference), and by local electricity rate (EU rates roughly double the US figure; cheap-rate regions halve it).
The order of magnitude is "tens of dollars a month for active daily use," not "hundreds." For the break-even-vs-cloud calculation, see Why local AI matters in 2026 — the cost-inverts-at-scale argument is the load-bearing one, and electricity is a small slice of it.
What we did NOT measure
The full list, because honesty is the point of this piece:
- Idle power. Every reading above is sustained load. Idle draw (the card sitting in a warm machine doing nothing) is unmeasured.
- Peak vs sustained curves. A reading at the start of a workload may differ from one after thermal soak; we have single snapshots, not curves.
- NVIDIA power. Victor (dual RTX 5070) has zero power data. Every chart that ranks GPUs by efficiency is using TDP specs for the NVIDIA cards, not our measurements.
- Multi-GPU combined-mode power. Ray's 410W is one card active. Both cards actively tensor-splitting is unmeasured.
- APU power. EvoX2's Ryzen AI MAX+ 395 has no wall-power data. The unified-memory APU story is interesting precisely because its power envelope is different from a discrete GPU's, but we cannot speak to it numerically.
- Sleep and wake behavior. Whether a card drops to a low-power state between requests is unmeasured. Most do, in our anecdotal experience; we don't have the data to claim it.
- Quantization on other cards. The "quantization barely moves power" finding is one card (Arc B60) on one model (Qwen3-Coder-30B). It should generalize — the silicon-level argument is vendor-agnostic — but the only measured support we have is that one sweep.
This is roughly 90% of the fleet uncovered. The next benchmark protocol revision captures wall-power across all four machines; until then, treat this article as "here is what we have," not "here is the comprehensive power picture."
How to measure your own
If you want real numbers for your own hardware, the cheap path:
- Kill-A-Watt meter (~$25) plugged between the wall and your machine. Captures whole-system wall-power, which is what shows up on your bill. The GPU-only number is lower; the whole-system number is what you pay for.
nvidia-smi dmonorrocm-smifor the GPU's self-reported power draw during a run. Easier than a Kill-A-Watt, slightly less accurate, and gives you per-card rather than whole-system numbers.- Run a sustained inference workload (a coding benchmark, a long generation, anything that keeps the card busy for several minutes), record the wattage at the meter or in the SMI tool after thermal soak.
The whole exercise takes twenty minutes and gives you a number grounded in your hardware and your electricity rate, which is better than any blog post's estimate.
The takeaway
Power draw under sustained local-LLM load is roughly the GPU's TDP, and quantization barely changes it. The Arc B60 sits at 245W across Q4/Q6/Q8 on a 30B coder — 7W of variation, under 3%. The dual R9700 sits at 410W with one card active. At US average electricity rates, that's roughly $9–$16/month for 8 hours of daily use. The broader picture — most of the fleet, idle curves, NVIDIA measurements, combined-mode dual-GPU — is a gap in our coverage that we're closing. Until then, the TDP-multiplication school is a reasonable fallback for cards we haven't measured, and a Kill-A-Watt is a twenty-dollar way to get the real number for your own hardware.
The quantization thesis behind why Q4_K_M is the default — for quality and tok/s, not for watts — is in quantization in 2026: no compromise. The full per-card performance data is in Best GPUs for local AI, measured. The raw records behind every number above are in the benchmark library.
Last verified: July 2026. The wall-power coverage gap is closing — re-check the benchmark library for updated records before relying on the "we didn't measure X" caveats above.