← Notes

Your Upscaler Is Better Than Its Score: Metrics, Eyes, and a Local Judge

A ground-truth upscaling benchmark where the objective metric picks the wrong winner — and a local vision model, framed as a harsh critic, sides with human eyes.

Your Upscaler Is Better Than Its Score: Metrics, Eyes, and a Local Judge

The setup: a task with an answer key

After standing up local image generation, the natural next step for this series was a task that can be graded without opinions — the structure that worked for the heat-simulation trial: hold back ground truth, let methods reconstruct it, measure the difference.

Upscaling fits perfectly. Take real photographs at 1024px and hide them as truth. Downscale to 256px. Hand every method the same small image and ask for 1024px back. Score with PSNR and SSIM against the hidden original. As the null hypothesis: plain bicubic interpolation — any learned upscaler that cannot beat bicubic is decoration.

We ran Real-ESRGAN (the standard open upscaler, one node in ComfyUI, half a second per image on the RTX 5070) against that baseline over four photographs.

The metrics pick the wrong winner

Mean over 4 photosBicubicReal-ESRGAN
PSNR29.08 dB26.50 dB
SSIM0.7570.722

The learned upscaler loses, on both metrics, on three of four photos. If you stopped here — and automated benchmarks routinely stop here — you would delete Real-ESRGAN and ship bicubic.

Then you look at the images. Center crops, same photo: truth on the left, bicubic in the middle, Real-ESRGAN on the right.

Truth vs bicubic vs Real-ESRGAN — the blurry one has the better PSNR

The middle image — the smear — is the one with the winning score. The right image, which restores branches, bark and foliage you can actually see, loses by three decibels.

This is the perception–distortion tradeoff, well documented in the super-resolution literature: a GAN-trained upscaler invents plausible detail, and invented detail is never pixel-aligned with the original. PSNR punishes every hallucinated branch; a human rewards it. Bicubic hedges by blurring toward the mean, which is exactly what pixel metrics love and eyes despise.

For a series built on “objective scoring beats judgment”, this is the instructive counterexample: an objective metric is only as good as what it measures. The heat-equation trial could use pixel-level truth because a temperature either matches the PDE or it doesn’t. Images are not like that.

Recruiting a judge that sees

If pixel metrics can’t capture perceptual quality, can a local vision-language model? The toolkit already serves gemma3:12b (vision-capable) for text work, so the judge costs nothing new. Two experiments, both with known answers.

First, a calibration check on the fox images from the previous post, where we hand-counted prompt elements. Asked naively — “is mist visible? is dew visible?” — the judge said yes to everything on both images: total yes-bias, zero discrimination. Asked as “a harsh photography critic: answer true only if unmistakably visible, false when in doubt”, it reproduced the manual counts exactly — SDXL’s fox has sun shafts only; Z-Image’s has mist, shafts, and dew.

The same lesson our text-model trials learned about verification applies to vision judges: a default-framed judge agrees with everything; only an adversarially framed judge discriminates.

Second, the real test. Blind A/B: both reconstructions of each photo, side by side, sides randomized so the judge can’t develop a positional habit, one forced question — which is sharper, more detailed, more natural?

PhotoPSNR prefersJudge prefers
forestbicubicReal-ESRGAN
architecturebicubicReal-ESRGAN
coastlinebicubicReal-ESRGAN
portraitbicubicbicubic

The judge sides with Real-ESRGAN 3 of 4 — against the pixel metrics, with human eyes. On four photos this is a demonstration rather than a measurement (the n-caveat this series always carries), but the direction is exactly what the perception-distortion literature predicts, and every step of it ran on one 12 GB card: the upscaler on the GPU, then the judge, with a VRAM handoff in between because they don’t fit together.

What this buys the series

The three working papers each needed an evaluator nobody could argue with: repairs counted, a holdout scored, a PDE solved. Images now have the outline of one — but it has to be a pair of instruments, because this experiment shows a single number lies:

  • fidelity: PSNR/SSIM against ground truth, for tasks where truth exists — catches a broken upscaler instantly;
  • perception: an adversarially framed local VLM in blind, side-randomized A/B — catches the case where fidelity metrics reward mush.

Both scripts (bench-upscale.py, judge-image.py) are in the toolkit. A graded image-generation trial in the style of the working papers — multiple models, hidden prompt suite, element checklists plus blind preference — is now mostly a matter of scale, and it is the obvious candidate for the next paper in the series.

Setup: RTX 5070 12 GB, ComfyUI 0.37 for reconstruction, Ollama + gemma3:12b as judge. Four CC0 test photographs, 1024px truth, 256px inputs. Sides randomized per photo with fixed seeds; judge temperature 0; images unretouched.