How Long Can a Clip Get? Duration, Health, and Why Cartoons Don't Help
Pushing local video generation to 15 seconds on 12 GB: no memory wall in sight, cost scales linearly — and the limit that actually bites is temporal health, measured with contact sheets.
The question
The previous post benched 2–4 second clips. The obvious follow-ups: how long can a clip get on 12 GB before something breaks — and does content matter? Can a cartoon, with less fine detail to maintain, go longer than a photoreal scene?
Long render times are acceptable; the real question is whether long clips stay healthy. So the sweep below measures cost, and the contact sheets measure health — frames pulled at 0%, 33%, 66% and 100% of each clip, because the first frame always looks fine and decay lives in the later ones.
The sweep: no wall in reach
768×512, 25 steps, both models pushed upward until something broke. Nothing broke:
| Model | Frames | Video length | Wall clock | Per second of video | Peak VRAM |
|---|---|---|---|---|---|
| LTXV 2B | 161 | 6.7 s | 58 s | 8.6 s | 11.2 GB |
| LTXV 2B | 257 | 10.7 s | 95 s | 8.8 s | 11.3 GB |
| LTXV 2B | 361 | 15.0 s | 146 s | 9.7 s | 10.0 GB |
| Wan 2.2 5B | 81 | 3.4 s | 113 s | 33.6 s | 11.2 GB |
| Wan 2.2 5B | 121 | 5.0 s | 167 s | 33.1 s | 9.8 GB |
| Wan 2.2 5B | 161 | 6.7 s | 239 s | 35.6 s | 10.2 GB |
Two findings. Cost per second of video is flat — duration scales linearly, no super-linear blowup in the tested range. And VRAM never climbed — the same offload behaviour that made 64K text context a dial rather than a cliff absorbs the growing video latent too. Fifteen seconds of video on a 12 GB card, two and a half minutes of compute, no out-of-memory in sight.
The health evidence
LTXV at 15 seconds — the actual clip, then its frames at 0/33/66/100%:


Identity holds completely: same fox, same forest, no morphing. But watch the clip itself and the problem is obvious — almost nothing moves. What the contact sheet politely called “stability through stillness” is, bluntly, a failure mode, and it deserves a number rather than a euphemism. We added one: mean per-pixel difference between consecutive frames, where a held frame scores ~0.
| Clip | Motion score |
|---|---|
| LTXV 4 s | 3.5–4.2 |
| LTXV 6.7 s | 4.4 |
| LTXV 10.7 s | 1.4 |
| LTXV 15 s | 1.6 |
| Wan 2 s | 2.4–3.3 |
| Wan 6.7 s | 1.5 |
The pattern is unambiguous: past each model’s comfort zone, motion collapses before identity does. The long clips pass every frame-coherence check precisely because the model quietly stops animating. That is the real single-shot duration wall, and it arrives far earlier than the memory wall — around 7 seconds for LTXV, under 5 for Wan, on this card.
Wan at 6.7 seconds (3.3× its benched length) — clip, then sheet:


Coherent, but doubly degraded: the texture has gone painterly versus the crisp 2-second clip, and its motion score has halved. Long clips don’t fail loudly; they quietly stop being video.
Do cartoons go longer? No — and it cost them
The hypothesis was reasonable: flat-shaded content has no fur to shimmer, so it should drift less. Two halves to the answer.
The hard limit can’t move, and now that’s measured: a cartoon prompt at 361 frames cost 146.1 s against the photoreal clip’s 146.0 s, same VRAM. Memory and compute are set by width × height × frames before the model knows what it’s drawing. Style cannot buy you a longer ceiling.
And the health limit moved the wrong way:

The cartoon fox starts charming, loses its face by two-thirds, and dissolves into smears by the end — while the photoreal clip at the identical length held every frame. Our guess at the mechanism: these models are trained overwhelmingly on real footage, so the temporal prior that keeps a scene coherent is strongest exactly there. A cartoon isn’t easier — it’s off-distribution, and off-distribution is where long generations lose their grip first. (One pair of clips; a hypothesis with the usual n=1 caveat, not a law.)
Verdict
- The memory wall is out of reach, but that is the wrong wall to watch. 15 s renders fine in 2.5 minutes — and scores as functionally static.
- The real single-shot limit is motion collapse, measurable and early: roughly 7 s for LTXV, under 5 s for Wan on this card. Past it, models keep identity by giving up movement — clips that pass every coherence check and fail at being video.
- Getting to genuinely moving 30-second content therefore means composition — many short shots inside the motion-healthy window — which is the next experiment.
- Style doesn’t extend duration — costs are geometry, not content — and stylized content may age worse, not better.
- Contact sheets are the cheap instrument that makes all of this visible; they’re four lines of PIL and they caught everything the single-frame checks missed.
RTX 5070 12 GB, ComfyUI 0.37, 768×512, 25 steps throughout. Sweep and style
runs via
bench-video.py
(--frames, --prompt). Sheets unretouched.