← Notes

How Long Can a Clip Get? Duration, Health, and Why Cartoons Don't Help

Pushing local video generation to 15 seconds on 12 GB: no memory wall in sight, cost scales linearly — and the limit that actually bites is temporal health, measured with contact sheets.

How Long Can a Clip Get? Duration, Health, and Why Cartoons Don't Help

The question

The previous post benched 2–4 second clips. The obvious follow-ups: how long can a clip get on 12 GB before something breaks — and does content matter? Can a cartoon, with less fine detail to maintain, go longer than a photoreal scene?

Long render times are acceptable; the real question is whether long clips stay healthy. So the sweep below measures cost, and the contact sheets measure health — frames pulled at 0%, 33%, 66% and 100% of each clip, because the first frame always looks fine and decay lives in the later ones.

The sweep: no wall in reach

768×512, 25 steps, both models pushed upward until something broke. Nothing broke:

ModelFramesVideo lengthWall clockPer second of videoPeak VRAM
LTXV 2B1616.7 s58 s8.6 s11.2 GB
LTXV 2B25710.7 s95 s8.8 s11.3 GB
LTXV 2B36115.0 s146 s9.7 s10.0 GB
Wan 2.2 5B813.4 s113 s33.6 s11.2 GB
Wan 2.2 5B1215.0 s167 s33.1 s9.8 GB
Wan 2.2 5B1616.7 s239 s35.6 s10.2 GB

Two findings. Cost per second of video is flat — duration scales linearly, no super-linear blowup in the tested range. And VRAM never climbed — the same offload behaviour that made 64K text context a dial rather than a cliff absorbs the growing video latent too. Fifteen seconds of video on a 12 GB card, two and a half minutes of compute, no out-of-memory in sight.

The health evidence

LTXV at 15 seconds — the actual clip, then its frames at 0/33/66/100%:

LTXV 361 frames, the full clip

LTXV 361 frames: scene and fox stable across all fifteen seconds

Identity holds completely: same fox, same forest, no morphing. But watch the clip itself and the problem is obvious — almost nothing moves. What the contact sheet politely called “stability through stillness” is, bluntly, a failure mode, and it deserves a number rather than a euphemism. We added one: mean per-pixel difference between consecutive frames, where a held frame scores ~0.

ClipMotion score
LTXV 4 s3.5–4.2
LTXV 6.7 s4.4
LTXV 10.7 s1.4
LTXV 15 s1.6
Wan 2 s2.4–3.3
Wan 6.7 s1.5

The pattern is unambiguous: past each model’s comfort zone, motion collapses before identity does. The long clips pass every frame-coherence check precisely because the model quietly stops animating. That is the real single-shot duration wall, and it arrives far earlier than the memory wall — around 7 seconds for LTXV, under 5 for Wan, on this card.

Wan at 6.7 seconds (3.3× its benched length) — clip, then sheet:

Wan 161 frames, the full clip

Wan 161 frames: coherent throughout, but softer than its short clips

Coherent, but doubly degraded: the texture has gone painterly versus the crisp 2-second clip, and its motion score has halved. Long clips don’t fail loudly; they quietly stop being video.

Do cartoons go longer? No — and it cost them

The hypothesis was reasonable: flat-shaded content has no fur to shimmer, so it should drift less. Two halves to the answer.

The hard limit can’t move, and now that’s measured: a cartoon prompt at 361 frames cost 146.1 s against the photoreal clip’s 146.0 s, same VRAM. Memory and compute are set by width × height × frames before the model knows what it’s drawing. Style cannot buy you a longer ceiling.

And the health limit moved the wrong way:

Cartoon at 361 frames: the fox dissolves by the final third

The cartoon fox starts charming, loses its face by two-thirds, and dissolves into smears by the end — while the photoreal clip at the identical length held every frame. Our guess at the mechanism: these models are trained overwhelmingly on real footage, so the temporal prior that keeps a scene coherent is strongest exactly there. A cartoon isn’t easier — it’s off-distribution, and off-distribution is where long generations lose their grip first. (One pair of clips; a hypothesis with the usual n=1 caveat, not a law.)

Verdict

  • The memory wall is out of reach, but that is the wrong wall to watch. 15 s renders fine in 2.5 minutes — and scores as functionally static.
  • The real single-shot limit is motion collapse, measurable and early: roughly 7 s for LTXV, under 5 s for Wan on this card. Past it, models keep identity by giving up movement — clips that pass every coherence check and fail at being video.
  • Getting to genuinely moving 30-second content therefore means composition — many short shots inside the motion-healthy window — which is the next experiment.
  • Style doesn’t extend duration — costs are geometry, not content — and stylized content may age worse, not better.
  • Contact sheets are the cheap instrument that makes all of this visible; they’re four lines of PIL and they caught everything the single-frame checks missed.

RTX 5070 12 GB, ComfyUI 0.37, 768×512, 25 steps throughout. Sweep and style runs via bench-video.py (--frames, --prompt). Sheets unretouched.