What the model still gets to decide
An i2v run has one job: produce frames that continue plausibly from the one you gave it. The subject, palette, lens character, and framing are already settled. What remains open is motion, and that includes three separate things the model will decide whether or not you mention them: camera movement, subject movement, and ambient movement like cloth, water, hair, and smoke.
If you name none of them, most models default to a slow push in with mild ambient drift. That default is why so much i2v output looks alike. Naming one specific movement is the fastest quality upgrade available on this task.
Choosing a first frame that animates well
This is the part that separates good i2v output from mush, and it happens before you ever write a prompt.
- Match the target aspect ratio and hit the model's native resolution. Anything else gets cropped or padded on the way in.
- Leave room for the move. If you want the camera to drift right, do not put the subject flush against the right edge.
- Keep the subject whole. Limbs cropped at the frame edge get invented when they re-enter, and invented limbs are where warping shows.
- Avoid an already-blurred source. The model reads existing motion blur as texture and smears it forward.
- Prefer a frame caught mid-action. A figure leaning into a step animates; a figure standing symmetrically at rest tends to produce a slow zoom and nothing else.
- Watch the background complexity. Dense crowds, foliage, and fine patterns are where temporal artifacts appear first.
Prompting motion instead of appearance
With the look already fixed, spend every word on movement.
[camera move], [subject action], [ambient motion], [what stays still]
One camera move and one subject action per clip. Two competing moves is the fastest way to get a wobble that reads as neither. Naming what should stay still is underused and works well: it gives the model a stability target instead of letting it drift everything at once.
Restating appearance is the common waste. Writing "a woman in a red coat in an alley" when that is exactly what the input frame shows adds nothing and can pull the output away from the frame it was handed.
Which video task to reach for
| You have | Task |
|---|---|
| Only an idea | Text to video |
| A frame you want to move | Image to video |
| A start and an end you both control | First and last frame |
| Footage you want restyled | Video to video |
The middle two are worth distinguishing. Supplying a last frame as well turns generation into interpolation between two fixed points, which is tighter control but a different creative act: you are specifying a transition rather than a shot.
Failure modes and their real fixes
Morphing subjects, a camera that drifts when it should be locked, backgrounds that pump and breathe, and hands that fold into themselves are the recurring four. None of them are prompt-wording problems.
Morphing usually means the requested motion is beyond what the model can solve from one frame, so ask for less. Unwanted camera movement responds to explicitly requesting a locked camera. Background pumping is usually a busy source frame, so simplify the still. Hands are a known weak point, and the practical answer is to frame them out or keep them still rather than to describe them more carefully.