Gus Moves: Your Own Photos Are the Best Video Seeds
The highest motion score of the entire series came from animating a real photo of a real dog — locally, in eight minutes, without the picture ever leaving the machine.
The best seed was in a photo folder all along
The reference-shot post established the lever: image-to-video, anchored by the best still available. We had been generating those stills. The obvious next question — what happens when the seed is a real photograph — turned out to have the best answer of the series.
Meet Gus, animated from a single snapshot:

3.4 seconds at 704×1280, 8¼ minutes of compute on the same 12 GB card, from one JPG. Motion score: 12.8 — nearly double the previous series high, and it is real motion: the frames catch a head-shake with genuine blur in the ears, a raised paw, weight shifting between haunches:

Everything that should stay still, stays still — the window light, the chair, the collar tag, the exact pattern of his coat. Everything that should move, moves.
Why a photograph beats a generated still
A generated seed, however good, carries its model’s fingerprint — and the motion model spends some of its capacity reconciling that. A photograph is perfectly in-distribution for a model trained on footage of the real world. The motion model recognises everything in the frame and spends its entire budget on what it’s actually for: making it move. Series motion scores tell the story in one line:
| Seed | Motion score |
|---|---|
| No seed (text-to-video, best shot) | 7.0 |
| Generated still (FLUX.2 klein) | 7.6 |
| A real photo | 12.8 |
The recipe, including the traps
- Fix the EXIF rotation first. Cameras store orientation as metadata; the pipeline sees raw pixels. Our first attempt fed the model a sideways dog.
- Match the video to the subject. Portrait photo, portrait video — 704×1280 is as native to Wan 2.2 as widescreen is.
- Center-crop to the exact target ratio yourself rather than letting any node stretch it.
- Describe motion the photo supports. “Wags its tail, tilts its head, shifts its front paws” — things this dog, in this pose, could do next. The model animates forward from what it sees; ask for plausible futures.
One honest caveat: i2v can get creative with limbs mid-sequence — inspect before you trust a clip, same as everything else in this series.
The local-first point, which writes itself
This is the first result in the series that is not a benchmark artifact but something a person might actually want: a living photo of their own dog. And it was produced entirely on a desktop GPU — the photograph of Gus never touched a cloud API, never entered anyone’s training pipeline, never left the house he sits in. For personal media, that is not a technicality; it is the whole argument for local generation, made by a bernedoodle.
RTX 5070 12 GB, Wan 2.2 5B image-to-video, 81 frames, 35 steps, native portrait resolution. Clip and frames unretouched. Gus was not compensated for this appearance.