Local image generation is the most common “second workload” after chat — and it breaks a lot of LLM buying advice. A card that’s brilliant for 30B text models can be mediocre for diffusion, and the video models multiply the problem again. This hub aggregates what our lab has actually measured on image and video generation, separates it from community-reported claims, and ends with a decision table by budget tier. If you’ve already read how much VRAM you need for local AI, treat this as the media-workload companion.
Image generation VRAM tiers
The honest tiering, from our ComfyUI coverage:
| Tier | VRAM | What runs |
|---|---|---|
| SDXL-class | 8–12GB | SDXL, SD 1.5, most community fine-tunes |
| FLUX-class | 12–24GB | FLUX.1 dev/schnell, FLUX.2-klein, larger DiT models |
Two things worth knowing that the tier table hides:
- Quantization changes everything. A FLUX-class model at Q4 fits in roughly half the VRAM of its full-precision form. Our ComfyUI optimization guide covers the settings that make a 12GB card behave like a 16GB one.
- The current model cycle keeps raising the bar. FLUX.2-klein and Z-Image are the 2026 names to watch; the FLUX run-through covers what they demand.
One community claim we can’t verify: FLUX.2-klein has been reported running on a $150 BC250 card — a striking datapoint if true, but it’s community-reported and unverified; we have not measured it and treat it as anecdote until we see the setup.
What our lab’s cards actually deliver
We run FLUX-class workloads on two of our four lab machines, and both handle it comfortably:
- Intel Arc Pro B60 (24GB) — runs FLUX-class comfortably per our FLUX run-through and ComfyUI beginner guide. The 24GB buffer means quantized FLUX fits with headroom for larger resolutions and ControlNet-style extras.
- AMD RX 7900 XT (20GB) — also runs FLUX-class comfortably per the same coverage. Note this is the 20GB card; the 24GB RX 7900 XTX is the stronger media option in that family.
The honest caveat: our measured LLM numbers on these cards don’t transfer directly to diffusion throughput. Diffusion workloads stress different parts of the GPU, and we haven’t published a standardized image-gen benchmark pack yet — so treat “runs comfortably” as our operational experience from the posts above, not a tok/s figure.
Video generation: the reality check
Local video generation is where marketing and reality diverge most. The current open models — Wan, Hunyuan, CogVideoX — need 16–24GB of VRAM for short clips, and they are slow on consumer cards. Our local video generation deep-dive is blunt about this: a short clip is a multi-minute render even on strong hardware, and anything beyond short clips pushes past what a single consumer card holds.
If you want video generation as a practical daily tool rather than a weekend experiment, our cloud video API comparison is the honest counterpoint. Local video is real and improving; it is not yet the frictionless experience the demos imply.
Audio and TTS: the cheap tier
Text-to-speech is the outlier — it’s the one media workload that’s CPU-friendly. Our best local TTS models guide covers models that run acceptably without any GPU at all. If media generation is your goal but your budget is tiny, TTS is where you can start today on hardware you already own.
AMD vs NVIDIA for media generation
The CUDA ecosystem is more mature for diffusion tooling — ComfyUI, custom nodes, and most community workflows assume NVIDIA first. That’s a real friction difference, not marketing. But our measured results complicate the “AMD can’t do media” story:
- OpenVINO is Arc’s ace. In bm-012, OpenVINO Model Server on the Arc B60 hit 67.95 tok/s generation with a 363ms TTFT — beating llama.cpp Vulkan on the same card. Intel’s OpenVINO stack is the strongest non-CUDA path for diffusion-adjacent workloads on Arc.
- AMD ROCm works but needs patience. Our ROCm reality check covers the setup friction honestly; the cross-vendor comparison puts all three stacks side by side.
The practical summary: NVIDIA is the path of least resistance, AMD and Arc work with more setup, and Arc’s OpenVINO support is genuinely strong.
Mac as a media box
We have not measured any Mac for image or video generation — our fleet is Arc, AMD, and NVIDIA, and we won’t pretend otherwise. What we can say from the spec sheet: the Mac Studio M5 Max with up to 128GB unified memory is architecturally attractive for large models, and Apple’s MLX ecosystem is maturing. But for diffusion specifically, community tooling still centers on CUDA, and we can’t verify Mac media performance from our own data. Treat Mac media-gen claims as spec-sheet analysis — we have not measured this hardware.
Decision table by budget tier
| Budget | Card | Media verdict |
|---|---|---|
| ~$300 | Arc B580 12GB | SDXL-class comfortably; FLUX-class at quantized settings. Prices move weekly. |
| ~$450–500 | Arc Pro B60 24GB | FLUX-class with headroom; OpenVINO is the strong software path. Prices move weekly. |
| ~$900 | RX 7900 XTX 24GB | FLUX-class with headroom; ROCm setup friction is the tradeoff. Prices move weekly. |
| $2,000+ | RTX 5090 32GB | The no-compromise media card: 32GB handles video-class models and FLUX at full precision. Prices move weekly. |
| $5,099+ | Mac Studio M5 Max | Unmeasured by us; unified memory up to 128GB, but CUDA-centric tooling is the friction. |
All prices are bands and move weekly — the late-2026 DRAM shortage is pushing memory-heavy hardware upward, so treat these as directional, not quotes. For the buy-now-or-wait angle, see our hardware timing analysis.
The takeaway
Image generation is the easy media add-on: 12GB gets you started, 24GB makes it comfortable, and both our Arc B60 and RX 7900 XT run FLUX-class workloads without drama. Video generation is the honest hard one: 16–24GB for short clips and slow renders on consumer cards — real, but not yet frictionless. TTS is the free entry point. And the software stack matters as much as the silicon: CUDA is the path of least resistance, but OpenVINO on Arc and ROCm on AMD are viable with more setup.
Next: Model Fit to check what your target media workload actually needs; the VRAM guide for the tier math; our ComfyUI beginner guide to get image generation running today.
Where this data comes from
The only hardware-throughput number in this piece from our own lab is the bm-012 OpenVINO vs Vulkan shootout on the Arc B60; the image and video VRAM tiers come from our ComfyUI, FLUX, and video generation coverage. The BC250 claim is community-reported and unverified, and all Mac claims are spec-sheet analysis — we have not measured that hardware. Full methodology lives in our benchmark reports.
Pricing reflects the US market as of October 2026 and moves weekly.