← Notes

The Card Learns to Move: Local Text-to-Video on 12 GB

LTXV 2B and Wan 2.2 5B measured on the RTX 5070 — seconds of compute per second of video, peak VRAM, and the actual clips, embedded and unretouched.

The Card Learns to Move: Local Text-to-Video on 12 GB

The question

Images on this card settled at 4–8 seconds each. Video is the obvious next modality, and the blog numbers around it are the least trustworthy of any — “runs on 12 GB” claims rarely say what resolution, how many frames, or how long you wait. So, same method as always: install, measure on this machine, show the artifact.

Two models, both with native ComfyUI support and ungated weights, filling the same two slots the image side settled into — a speed pick and a quality pick:

  • LTXV 2B (Lightricks) — the small fast one, ~11 GB of downloads with its T5 encoder
  • Wan 2.2 5B (Alibaba) — the quality one, ~18 GB with encoder and VAE

Setup was another quiet install: every node needed for both models ships in ComfyUI 0.37 — no custom nodes, no plugin manager, nothing. One download URL was wrong and produced a 200-byte error page in place of a 5 GB text encoder; a size check caught it. (Verify artifacts, not exit codes — the rule holds at every scale.)

The measurements

Same prompt as the whole fox saga, 768×512, 25 steps, measured warm:

LTXV 2BWan 2.2 5B
Clip benched97 frames (~4 s)49 frames (~2 s)
Wall clock, warm31.0 s72.3 s
Compute per second of video7.7 s35.4 s
Peak VRAM11.6 GB11.5 GB

Both fit — barely, in the now-familiar 11.5 GB envelope that means one resident model at a time. And both are far faster than expected: a four-second clip in half a minute from LTXV, and Wan at roughly half a minute per second of video, not the multi-minute waits the 12 GB folklore suggests.

The clips

These are the actual outputs, animated, unretouched (they play in any modern browser):

LTXV 2B — 4 seconds, 31 seconds of compute:

LTXV: fox in a misty forest, 4 seconds

Wan 2.2 5B — 2 seconds, 72 seconds of compute:

Wan 2.2: fox in a misty forest, 2 seconds

The image-side pattern repeats exactly. LTXV delivers atmosphere — the light shafts are lovely — but its fox is a soft suggestion at distance. Wan puts a sharp, backlit, anatomically respectable fox mid-stride against detailed bark, and looks like footage rather than a moving painting. Speed pick, quality pick: 4.6× the compute buys a generation of visible quality, the same trade Z-Image-Turbo and FLUX.2 klein settled into for stills.

Prompt-adherence, judged by eye against the usual checklist: both deliver forest, dawn light, and a walking fox; LTXV loses the dew, Wan loses the mist. Motion quality has no instrument yet — our image judge can score frames but nothing scores movement, so this post reports timings and shows the artifacts rather than claiming a graded result.

What this means for the stack

A 12 GB card in 2026 is a genuine short-clip video machine: storyboard-frame quality at half a minute per second of video, drafts at a quarter of that. The practical ceiling is length and resolution — a few seconds at 768×512 — not whether it runs.

The pipeline this obviously wants: FLUX.2 klein still → Wan image-to-video (its latent node takes an optional start image), which would combine the best image this card produces with the best motion. That, and an instrument for judging motion, are the queue.

RTX 5070 12 GB, ComfyUI 0.37, fixed prompt and seeds, warm timings after one cold load. Script: bench-video.py.