Chasing Sharpness: Two-Pass Rendering, FLUX.2 klein, and a Judge That Needed Glasses
The first local generations were soft. Two fixes measured on the same 12 GB card — a second denoising pass with the model you already have, and a new-generation 4B model — plus a judging bug worth knowing about.
The complaint
The foxes in the first image post were respectable, but not sharp. Fair. This post is the sharpness chase: what actually improves detail on a 12 GB card, measured, with the same fixed prompt and seeds throughout.
Two routes were tested — one costing nothing but time, one costing an 8 GB download — and along the way our image judge turned out to need corrective lenses, which is its own finding.
Route 1: a second pass with the model you already have
The oldest trick in the diffusion book is hires-fix: render normally, upscale the latent 1.5×, then run a second denoising pass at half strength so the model re-draws fine detail at the higher resolution. Same SDXL checkpoint, no new weights.
Route 2: a newer generation of model
FLUX.2 klein — Black Forest Labs’ open 4B model from January 2026 — was the obvious candidate from the landscape survey. A pleasant architectural surprise made it cheap to add: klein uses the same Qwen3-4B text encoder Z-Image-Turbo uses, which was already on disk. The incremental download was the 7.8 GB diffusion model and a 340 MB VAE.
The numbers
Same card, same prompt, warm runs at 1024×1024 (two-pass outputs 1536):
| SDXL 1-pass | SDXL 2-pass | Z-Image-Turbo | FLUX.2 klein | |
|---|---|---|---|---|
| Warm s/image | 7.3 | 26.4 | 4.5 | 8.2 |
| Peak VRAM | 9.2 GB | 11.4 GB | 11.0 GB | 11.5 GB |
| Steps | 25 | 25+25 | 8 | 8 |
Everything fits, and everything except the two-pass stays under ten seconds. The crops tell the quality story — same region of the same-seed fox, native resolution, left to right: SDXL 1-pass, SDXL 2-pass, FLUX.2 klein:

The two-pass render transforms SDXL — fur goes from soft suggestion to individual strands, at 3.6× the time. Klein reaches its detail in a single 8-step pass, and its full frame is the best image this card has produced:

On the strict-critic element checklist from the previous post, klein scores 3/3 — mist, sun shafts, and dew all unmistakably present — matching Z-Image and beating both SDXL variants, with visibly finer texture than any of them.
The judge needed glasses
We ran blind A/B comparisons (sides randomized, three votes each) expecting the judge to confirm the obvious. It didn’t: the plainly sharper two-pass render lost to its own soft single-pass version, and klein lost to Z-Image.
The cause was the instrument, not the images. Our A/B composite resized both full frames to 512px — and downscaling erases precisely the fine detail a sharpness comparison exists to measure. Composition survives a resize; texture does not. The judge was being asked about sharpness while being shown thumbnails.
After switching the comparator to native-resolution center crops, klein vs Z-Image flipped from 1/3 to a decisive 3/3 for klein. The SDXL two-pass pairing did not flip, so we chased it one step further: splitting the question into separate sharpness-only and naturalness-only votes (five each). The judge landed at 2/5 on both axes — a coin flip. Not contrarian; blind to this particular gap. The working characterization: this judge reliably resolves large quality differences (a model generation apart) and is at chance on moderate ones (a rendering trick apart). The flip is still the lesson:
When a judge disagrees with your eyes, check what the judge can actually see before concluding anything. Evaluation instruments fail as silently as the things they evaluate. That is now a three-instrument lesson in this series: pixel metrics reward mush, naive checklists say yes to everything, and downscaled A/B judges can’t see sharpness.
Where this leaves the stack
For quality-per-second on a 12 GB card, FLUX.2 klein is the new default: best detail, full prompt adherence, 8.2 s/image, one resident model. Z-Image keeps the speed slot at 4.5 s. The SDXL two-pass remains worth knowing as the zero-download sharpness upgrade for any checkpoint-based model — 3.6× slower, dramatically sharper.
Next in the queue: video (Wan and LTXV both claim 12 GB viability for short clips), and the graded image-generation trial the last post sketched — for which today’s instrument repair was necessary homework.
RTX 5070 12 GB, ComfyUI 0.37, fixed prompt and seeds; all four workflows in
bench-image.py;
judge changes in
judge-image.py.
Images unretouched.