The Hardest Style: A Line-Art Short, and the Best Video the 5070 Can Make
A 12-second black-and-white line-art film about a father, a toddler, and a flower — the deliberately worst-case style for these models — used to find the real ceiling of local video on a 12 GB card.
Picking the hardest possible target
Every earlier video test used photoreal or painterly scenes — squarely inside what these models were trained on. This one deliberately goes the other way: minimalist black-and-white line art, a father holding his toddler daughter’s hand as she meanders, stops to pick a flower, and walks on. Flat ink on white is about as far off-distribution as local video gets — the cartoon dissolved in the duration test, and line art is starker still.
The goal was two-fold: make the nicest 12-second clip this card can produce, and, along the way, settle which model is actually best — with measurements rather than a hunch.
Every model, one seed, one shot
The whole image stack earns its place here: FLUX.2 klein drew a clean line-art keyframe, and each video model then animated that identical still. Same input, three animators:
| Model | Time / shot | Motion | Held the line-art style? |
|---|---|---|---|
| LTXV 2B | 33 s | 4.1 | child dissolves into smudge |
| Wan 2.2 5B | ~9 min | 5.0 | yes — clean lines, real steps |
| LTXV 13B distilled | ~60 s | 8.2 | lively, faint late smudging |

Wan is the quality champion — it kept the ink clean and produced an actual walk cycle. But the surprise is the middle column of times.
The 13B model runs almost as fast as the 2B
LTXV’s 13-billion-parameter model is 15.7 GB of weights on a 12 GB card. By every earlier lesson it should crawl — it has to stream half its weight across PCIe. It renders a shot in about a minute, roughly 17× faster than Wan and within striking distance of the 2B model. The reason is the same shape as the sparse-MoE surprise from the text benchmarks: it is distilled to 8 denoising steps, so the offload penalty is paid so few times it nearly disappears. On a 12 GB card, a distilled big model is a genuinely different thing from a dense one — the “won’t fit in VRAM” rule keeps being less of a wall than it looks.
That makes a real two-tier workflow: 13B to draft at a minute a shot, Wan to finish the keeper.
The prompt is the director
The first attempt at the flower shot failed instructively. Wan turned the figures into solid black silhouettes partway through:

Nothing was wrong with the model — the prompt underspecified. “Line art” alone lets a footage-trained model drift toward filled shapes. The fix was directorial, not technical: a negative prompt banning solid black, filled shapes, silhouette, and positive language insisting on uniform thin line weight, no filled shapes. The lesson generalizes to every shot in the series that misbehaved — the model animates forward from what you actually pin down, and off-distribution styles need pinning down hard.
The 13B model, run as a rapid burst of three seeds on that same cursed shot, shows why drafting cheaply matters: same prompt, wildly different takes — one gives the flower, one shakes a giant fringe of hair, one barely moves. At a minute each, you simply roll again and pick.

The film
Three shots — walk together, stop for the flower, walk on with it — each seeded from a klein keyframe, best take chosen per shot, assembled with ffmpeg into twelve seconds:
It is not flawless — the father’s proportions wander between shots, and the “consistent character across cuts” problem is only half-solved by reusing a description. But it holds one clean visual style across three shots, tells a tiny complete story with real motion in each, and renders end-to-end on a desktop with 12 GB of VRAM. For the hardest style we could pick, that is a genuine result.
The verdict on “best local video, 12 GB”
- Quality king: Wan 2.2 5B, at image-to-video, native 1280×704, seeded by the best still the card can draw. That combination is the ceiling today.
- Speed/draft king: LTXV 13B distilled — big-model quality at near-small- model speed, because distillation makes offload nearly free.
- The real workflow is both: draft cheaply, finish carefully, and treat the prompt as shot direction.
- The one rung left unclimbed is Wan 2.2 14B in GGUF quantization — the last model this card can physically hold, and the subject of whether the dense-offload tax finally bites. That is the next measurement.
RTX 5070 12 GB, ComfyUI 0.37. Keyframes: FLUX.2 klein. Video: Wan 2.2 5B and
LTXV (2B and 13B distilled). Scripts:
bench-video.py,
motion-score.py.
Film and frames unretouched.