AI Motion Control

Give it a reference clip and your character, and the motion transfers over intact — dance, gesture, expression. The movement is exact; the character is still yours.

Key features

Motion follows the reference

A reference clip specifies the movement far more precisely than words — "do a dance" gets you the model's improvisation; a reference gets you that dance.

The character stays yours

Motion comes from the reference, appearance from your character image. Two separate controls.

Make a photo move

One still portrait plus a reference motion gives you video of that person performing it.

One motion, many characters

Apply the same reference to different characters for a set of clips with matching motion and different people.

Use cases

Make a photo move

Make a photo move

A still portrait performs the reference motion — no shoot required.

Dance and action video

Dance and action video

Reference a dance and have your character perform it, hitting the same beats.

Consistent motion across characters

Consistent motion across characters

Apply one motion to several different characters to produce a matched set.

Virtual presenters and character content

Virtual presenters and character content

A fixed character plus recorded motion, for producing content at volume with a consistent person.

How to use

01

Open AI Motion Control and upload your character image or clip.

02

Upload a reference motion clip to specify the movement to transfer.

03

Download the result, or send it to the canvas to keep cutting.

Deep dive

Specify the motion with a clip, not a description

Write "do a dance" in a video generator and you get the model's idea of dancing — possibly nice, not the dance you meant. Motion is too high-bandwidth for text: tempo, amplitude, the details of a pose don't fit in a prompt.

Motion control changes the input: hand it a reference clip. Motion is extracted from the video, appearance comes from your character image, and the two are controlled separately. So your character performs that specific movement.

The reference needs readable motion

The reference's image quality doesn't matter — it only supplies motion, and appearance comes entirely from your character. What matters is that the motion is readable:

  • Complete subject — full-body movement needs the full body; gestures need visible hands
  • No occlusion — a hidden limb is motion that can't be extracted
  • Not too dark — no visible contour, no readable pose
  • Preferably one person — several people moving at once and the model doesn't know who to follow

One more thing worth matching is framing. A full-body dance reference against a waist-up character image forces the model to invent the missing lower half, which may not match your character. Closest results come when both framings are similar.

When to reach for a different tool

  • The exact motion doesn't matter — the video generator is faster; one sentence and let it improvise.
  • What you're aligning is the mouth — lip-sync works against audio and governs speech.
  • You need exact start and end frames — first-last-frame.
  • You want to change the look — video-edit.

A reasonable way to think about it

This page's value is motion certainty. So it fits two situations: you have a movement that must be reproduced (a specific dance, a specific gesture), or you need a batch of clips with different characters and identical motion.

Conversely, if the motion is arbitrary and "moving" is all you need, the video generator is less work — no reference to prepare, one sentence is enough.

Frequently asked questions

How is this different from image-to-video?+

In image-to-video the model decides the motion and you only steer it loosely with words — "dancing" gets you the model's idea of dancing. Motion control uses a real clip as the motion template, so what happens is that specific movement. Use this page when the motion must be exact; use image-to-video when you want the model's idea.

What does the reference clip need?+

Clear motion and a complete subject matter most. Full-body movement needs the full body visible; gestures need visible hands. Occlusion, very low light, or several people moving at once all degrade extraction. The reference's image quality is irrelevant — it only supplies motion, while appearance comes from your character.

Will the reference affect how my character looks?+

By design, no — motion from the reference, appearance from your character. But when the two differ a lot in framing (a full-body dance reference against a waist-up character image), the model has to invent the missing part, and that part may not match your character exactly.

How does this relate to lip sync?+

They govern different parts: lip sync aligns the mouth to audio; motion control transfers body movement and pose. For a character that both moves and speaks, the usual order is motion first, then lip sync.

More models