The Reference Shot: What 12 GB Video Actually Looks Like at Its Best
After a series of honest disappointments, the lever that finally produced a compelling clip: anchor the motion model with the best still the card can make. Image-to-video, measured, embedded.
The fair complaint
Six posts into the video series, a fair reading of the evidence was: nothing generated so far is compelling. Text-to-video at 768×512 on the speed-tier model produces previz — useful, measurable, and visibly below what anyone has seen from modern video models.
The missing piece was never a bigger model. It was the lever behind nearly every impressive local-video demo in circulation: image-to-video. Don’t ask the motion model to invent a beautiful scene and animate it — hand it the beautiful scene, and let it spend its entire capacity on motion.
We already had the beautiful scene: FLUX.2 klein’s fox, the best still this card produces.
The recipe
klein still (1024², 8 steps, 8 s) → Wan 2.2 5B image-to-video, 704×704, 81
frames, 35 steps, motion-explicit prompt. One node different from
text-to-video: the still enters the latent as start_image.
The result — 3.4 seconds, 190 seconds of compute:

Frames at 0/33/66/100%:

The measurements
| Best t2v shot (Acorn film) | This shot (i2v) | |
|---|---|---|
| Motion score | 7.0 | 7.6 — series high |
| Resolution | 768×512 | 704×704 |
| Source image quality | model’s own invention | klein, the card’s best |
| Quality across frames | softens | holds at every sample point |
| Compute | ~130 s / 4 s | 190 s / 3.4 s |
The fox walks toward camera through dew bokeh and light shafts — real locomotion, klein-grade texture, no drift, no softening. This is the clip to judge the hardware by, and the answer to “what should I use as reference”: i2v from your best still is the ceiling of a 12 GB card today, at about a minute of compute per second of video.
What this recalibrates
- The pipeline is stills-first. The image stack isn’t a separate hobby from the video stack; it is the video stack’s front end. Better stills are now the highest-leverage video investment.
- The Acorn recipe upgrades directly: generate a klein still per shot, i2v each one. Character consistency should improve too — the same rendered pup can seed every shot instead of a description string.
- If this still isn’t enough, the next rung is Wan 2.2 14B in GGUF quantization with offload — slower, reportedly sharper, unverified on this card. That is the remaining untested quality lever.
RTX 5070 12 GB. Recipe reproducible from the graphs in
bench-video.py;
clip and frames unretouched.