← Notes

The Reference Shot: What 12 GB Video Actually Looks Like at Its Best

After a series of honest disappointments, the lever that finally produced a compelling clip: anchor the motion model with the best still the card can make. Image-to-video, measured, embedded.

The Reference Shot: What 12 GB Video Actually Looks Like at Its Best

The fair complaint

Six posts into the video series, a fair reading of the evidence was: nothing generated so far is compelling. Text-to-video at 768×512 on the speed-tier model produces previz — useful, measurable, and visibly below what anyone has seen from modern video models.

The missing piece was never a bigger model. It was the lever behind nearly every impressive local-video demo in circulation: image-to-video. Don’t ask the motion model to invent a beautiful scene and animate it — hand it the beautiful scene, and let it spend its entire capacity on motion.

We already had the beautiful scene: FLUX.2 klein’s fox, the best still this card produces.

The recipe

klein still (1024², 8 steps, 8 s) → Wan 2.2 5B image-to-video, 704×704, 81 frames, 35 steps, motion-explicit prompt. One node different from text-to-video: the still enters the latent as start_image.

The result — 3.4 seconds, 190 seconds of compute:

The reference shot: klein still animated by Wan i2v

Frames at 0/33/66/100%:

Hero shot contact sheet: quality holds while the fox advances

The measurements

Best t2v shot (Acorn film)This shot (i2v)
Motion score7.07.6 — series high
Resolution768×512704×704
Source image qualitymodel’s own inventionklein, the card’s best
Quality across framessoftensholds at every sample point
Compute~130 s / 4 s190 s / 3.4 s

The fox walks toward camera through dew bokeh and light shafts — real locomotion, klein-grade texture, no drift, no softening. This is the clip to judge the hardware by, and the answer to “what should I use as reference”: i2v from your best still is the ceiling of a 12 GB card today, at about a minute of compute per second of video.

What this recalibrates

  • The pipeline is stills-first. The image stack isn’t a separate hobby from the video stack; it is the video stack’s front end. Better stills are now the highest-leverage video investment.
  • The Acorn recipe upgrades directly: generate a klein still per shot, i2v each one. Character consistency should improve too — the same rendered pup can seed every shot instead of a description string.
  • If this still isn’t enough, the next rung is Wan 2.2 14B in GGUF quantization with offload — slower, reportedly sharper, unverified on this card. That is the remaining untested quality lever.

RTX 5070 12 GB. Recipe reproducible from the graphs in bench-video.py; clip and frames unretouched.