← Notes

The Toolkit Learns to See: Local Image Generation on the Same 12 GB Card

Standing up ComfyUI on the RTX 5070 and benchmarking SDXL against Z-Image-Turbo — measured seconds per image, measured VRAM, and a fox-off the newer model wins.

The Toolkit Learns to See: Local Image Generation on the Same 12 GB Card

Same card, new modality

The local-llm toolkit so far has been about text: serving 30B-class language models on a 12 GB RTX 5070, measuring what the card actually does rather than what leaderboards claim, and running three graded experiments on the results.

This post extends the toolkit to images, with the same rules: everything measured on this card, everything scriptable, samples shown rather than described.

The landscape, briefly

Local image generation in 2026 has consolidated into three families. FLUX (Black Forest Labs) is the quality and prompt-adherence reference — the 12B dev model at FP8 fits a 12 GB card, and the 4B klein runs far below that. Qwen-Image owns text-in-image — labels, UI mockups, typography. SDXL remains the ecosystem king: three years old, ~3.5B parameters, thousands of LoRAs. The speed tier belongs to distilled models like Alibaba’s Z-Image-Turbo (6B), built to generate in single-digit steps.

We started with two that need no license gate: SDXL (the baseline everyone has) and Z-Image-Turbo (the newest architecture that fits the card).

Setup: boringly smooth, which is news

The one landmine we expected never went off. The RTX 5070 is a Blackwell card (compute capability 12.0), which through much of 2025 meant hand-rolling PyTorch nightlies. The current ComfyUI portable build (v0.37) ships torch 2.13 + CUDA 13.0, which recognises the card out of the box:

torch 2.13.0+cu130 | cuda 13.0
gpu: NVIDIA GeForce RTX 5070
capability: (12, 0)

The whole installation was: download the 1.9 GB portable archive, extract, drop a model file into models/checkpoints/, launch. The server was answering its API 54 seconds after start. After weeks of debugging agent harnesses where five different failures returned exit code 0, it is worth recording that mature inference infrastructure simply works: the entire image stack came up in under half an hour including model downloads, and nothing failed silently.

One habit from the text-side work transferred directly: don’t drive the UI, drive the API. ComfyUI exposes its full node graph over HTTP, so the toolkit gained bench-image.py — queue a workflow, poll history, record wall-clock and peak VRAM from nvidia-smi. Same discipline as bench.ps1 for tokens.

The numbers

Both models, 1024×1024, same prompt, three warm runs after a cold load:

SDXL 1.0 (3.5B)Z-Image-Turbo (6B int8)
Steps (per model design)258
Cold first image11.9 s7.4 s
Warm, per image7.3 s4.5 s
Peak VRAM9.2 GB11.0 GB
Headroom on 12 GBcomfortable~0.9 GB

Both fit. The distilled model is 1.6× faster despite being larger, because it needs a third of the denoising steps — that is the entire trade the distillation makes. Its VRAM peak of 11.0 GB is close to the ceiling, though: Z-Image on this card runs alone or not at all, the same one-resident-model-at-a-time rule the text side lives under.

The fox-off

The benchmark prompt asks for specific, checkable things: “a detailed photograph of a red fox standing in a misty pine forest at dawn, shafts of golden light, dew on grass, sharp focus.” Four concrete elements — mist, dawn shafts, dew, pine forest — which makes adherence countable rather than vibes.

SDXL (7.3 s):

SDXL fox — handsome, but no mist and no dew

Z-Image-Turbo (4.5 s):

Z-Image-Turbo fox — mist, dawn shafts, and dew all present

Count the elements. SDXL produced a handsome fox with golden light — and no mist, no visible dew: 2 of 4. Z-Image-Turbo delivered mist hanging between the pines, sun shafts, and dew beaded on every blade of grass: 4 of 4, in 40% less time. Three years of architecture and training progress, visible in one image pair: the newer model is simultaneously faster and more literal about what you asked for.

One pair of images is one draw from each model — the same n=1 caveat every paper in this series carries. But the direction matches what the current comparisons report about post-SDXL architectures, and unlike blog claims, these two foxes came off this machine and the script that made them is in the repo.

What this sets up

  • FLUX next. dev at FP8 reportedly fits 12 GB exactly; klein (4B) and Qwen-Image GGUF are the other candidates. Same benchmark script, growing table.
  • The 6 GB tier. Unlike a language model, an image model is bursty rather than resident — it loads, generates, and can unload — which makes small-VRAM cards far more useful for images than for text. SDXL with a Lightning LoRA and FLUX at Q4 reportedly fit in 6-8 GB. To be measured, not believed.
  • A graded image task? The three working papers each needed an objective score. Prompt-element counting, as in the fox-off, is a start; whether it can anchor a Bakeoff-style trial for image models is an open question we find genuinely interesting.

Hardware: RTX 5070 12 GB, Ollama stack untouched and running alongside. ComfyUI 0.37.0 portable, torch 2.13.0+cu130. Both foxes generated unattended via the API, seeds fixed, images unretouched.