← Notes

Thirty Seconds, Two Ways: One Shot Collapses, Eight Shots Make a Film

Pushing a single generation to 30 seconds produces literal nothing — an orange shimmer. Composing eight 4-second shots produces a 31-second film with a plot, a consistent character, and measured motion in every cut. The film is embedded.

Thirty Seconds, Two Ways: One Shot Collapses, Eight Shots Make a Film

Raising the bar

The duration post ended with a correction: long single shots don’t crash, they quietly stop moving — coherence purchased with stillness, and a motion score to catch it. A reader’s blunt framing stuck: a drifting camera around a frozen scene isn’t video. So this post holds everything to a functional standard — does it move, does it tell, would you watch it — and takes on the real target: thirty seconds of content. Two ways.

Way one: a single 30-second shot

LTXV, 721 frames, one continuous generation. It rendered without complaint — 6.6 minutes, no OOM, and a motion score of 2.42, technically above our static threshold.

Then you look at it:

Thirty seconds of nothing: the single-shot collapse

Thirty seconds of orange gradient. No fox, no forest, no scene — at more than double the model’s healthy window, generation doesn’t degrade, it evaporates. And the motion score of 2.42 is pure texture shimmer, which teaches the instrument lesson of the day: the motion metric catches stillness, but it cannot tell locomotion from flicker. It needs the contact sheet beside it. (Every instrument in this series has needed a second instrument; this is not an exception.)

So the single-shot verdict is now complete and monotonic on this card: full motion to ~5–7 s, motion collapse to ~15 s, total collapse by 30 s.

Way two: eight shots and a cut

Nothing about a commercial or an animated short is one shot — real content cuts every few seconds. So the composition experiment: a shot list with a plot, every shot inside the motion-healthy window, one locked character description repeated verbatim in all eight prompts, rendered on Wan 2.2 5B and concatenated with ffmpeg.

The plot, in the spirit of a certain acorn-obsessed ice-age squirrel: an acorn bonks a fox pup on the head; curiosity, pounce, downhill chase, a stream mishap, triumphant retrieval, and a nap with the prize.

The result — 31.3 seconds, 17 minutes of compute, every second of it moving:

Mid-frames from all eight shots:

One frame from each of the eight shots

The measurements

ShotActionMotion score
1acorn bonk2.9
2wide-eyed stare7.0
3pounce and tumble4.2
4downhill chase5.0
5skid to the stream7.0
6splashing6.5
7triumphant retrieval4.1
8curl up and sleep4.8

Every shot clears the functional bar — most by a wide margin. For scale: a single Wan shot pushed to just 6.7 seconds scored 1.5. Thirty-one seconds of composed video carries three times the motion of seven seconds of continuous video. Cutting is not a workaround; on this hardware it is the mechanism.

Character consistency — the known weakness of the approach — held better than expected: the white-tipped tail and oversized ears survive all eight cuts, and the pup reads as one character across the film, carried by nothing more than a repeated description string. Honest flaws, visible above: the “acorn” of shot 2 is frankly a pinecone, lighting temperature drifts between shots, and there is no audio. This is a storyboard-quality short, not a finished commercial — but it is watchable, which no single generation at this length can claim.

The recipe

The pipeline is one script and one JSON: make-film.py takes a shot list (character string, style, per-shot actions and durations), renders each shot inside the healthy window with the right frame quantization per model, and assembles the mp4. The Acorn’s shot list ships as a worked example — a 30-second commercial is the same file with different lines.

The 12 GB verdict on longform: duration is not a rendering problem, it is an editing problem — and editing is cheap. Shots plus cuts gets you to any length; the per-shot motion window is the only wall that matters.

RTX 5070 12 GB, Wan 2.2 5B for shots, LTXV for the single-shot probe, 768×512, 25 steps. Motion scores via motion-score.py. Film and frames unretouched.