Upload a photo, pick something being argued about online, add your own director notes — and get a short video of your pet weighing in. This page documents how it is built.
720×1280 · 40s · h264 + AAC · 7.2 MB
Everyone stays on screen the whole time, like a real call. The audio runs in sequence so only one speaks at a time, and the live tile carries a lit border. Each dog was generated alone, then composited — asking one model for four characters in a single shot produced the same dog four times.
720×1280 · 8s · 3.3 MB
The dog, the speech and the mouth movement all come out of a single text prompt — no reference photo and no separate lip-sync pass. This is the unit the grid is built from.
| Model | google/veo-3.1-lite |
|---|---|
| Route | OpenRouter video API |
| Mode | Text-to-video, native audio |
| Generation time | 62 seconds |
| List price | $0.05 / second at 720p |
Breed, coat colour and markings are described in words. Naming the breed explicitly carries the likeness further than image resolution does.
A real discussion — a thread, a podcast clip, an argument. The claim being made is preserved; only the stakes change.
The owner supplies the attitude. Loud and self-important, or flat and unimpressed. This is what makes it sound like their pet.
Each character is generated alone, then the four are composited into one grid with captions and branding.
Honest first pass, not a finished result. Moose’s line picked up a fragment of Butcher’s and needs regenerating in a clean context. The laptops still visible in each room are leftovers from an earlier approach and come out of the prompt next. Butcher’s coat also drifted tan-heavy on the single-shot version where the character is meant to be predominantly white.