AI First & Last Frame Video

Give it a first frame and a last frame, and AI fills the motion between — you set both ends, so the shot can't wander off.

Key features

You define both ends

Upload two images and the model fills the motion. Far more controllable than a text description — where the shot lands is decided before you generate.

Transitions become controllable

Use the outgoing shot's last frame and the incoming shot's first frame, and the bridge between them generates naturally.

Thirty-plus models to choose from

Most video models support first/last frame, and they differ in motion style and amplitude. Re-run the same pair on another.

One image also works

Supply only a first frame and it continues forward from it; supply both when you want precise interpolation.

Use cases

Bring a still to life

Bring a still to life

Start from an image you like, specify what it should become, and the motion in between fills itself in.

Keyframe the motion

Keyframe the motion

Lock the trajectory at both ends so the model can't improvise somewhere you didn't want.

Product reveals

Product reveals

Closed to open, wide to tight — both the start and end states are product shots you chose.

Transitions between shots

Transitions between shots

Generate the bridge from the boundary frames of two shots, so the seam reads better than a hard cut.

How to use

01

Open AI First & Last Frame Video and upload the first frame.

02

Upload the last frame, then add a sentence about what should happen in between.

03

Pick a duration and model, generate, then download or send to the canvas.

Deep dive

Pin both ends, let the model fill the middle

There's a familiar frustration in animating a single image: the motion looks great and goes somewhere you didn't want. You supplied a start, and the model decided the rest.

First/last frame inverts that: you supply both ends and the model only fills the middle. Where the picture lands is settled before you generate.

That constraint is stronger than any prompt. To take a product from closed to open, a camera from wide to tight, a scene from day to night — showing both ends beats describing the process.

The two frames need a plausible relationship

Quality mostly comes down to the ends you pick:

Two states of one scene are ideal — closed and open, wide and tight, day and night. The model reads it as one thing changing, so the bridge is smooth.

Unrelated images cause trouble — the model can only force a transition, and the middle tends to deform or jump. If you genuinely need to go from A to B with no relationship, that was always a hard cut, not an interpolation.

One frame is allowed — it then degrades to continuing forward from that image, close to image-to-video. The last frame is what you add when the landing point matters.

When to reach for a different tool

  • You want the model's own idea — the video generator, from one start frame or pure text, lets it improvise.
  • You want an existing clip to run longer — extend-video continues from the footage's final frames, with no end frame to prepare.
  • You want a specific action transferred — motion-control maps a reference clip's movement onto your subject.
  • You need mouth shapes to match speech — lip-sync.

A reasonable way to think about it

First/last frame suits the moment when you already know the answer. The start and end are in your head and all you need is the bridge — showing both ends is more precise than writing a description.

Conversely it's awkward when the direction is still open, since you'd have to produce two images first. At that stage, running a few takes in the video generator is faster.

One practical combination: make two key frames in the image generator, then bring them here to interpolate. That way both ends of the shot are pictures you chose, rather than whatever the model happened to produce.

Frequently asked questions

How is this different from ordinary image-to-video?+

Image-to-video gives the model only a starting point and lets it decide where to go — sometimes great, sometimes not what you pictured. First/last frame adds an endpoint, which fences in the trajectory. Use it when you want a determined result; use image-to-video when you want to see the model's idea.

Do the two images need to look similar?+

They need a plausible relationship. Two states of the same scene — closed/open, wide/tight, day/night — work best. If the two are unrelated, the model can only force a bridge, and the middle often deforms or jumps.

Can I control the motion in between?+

You can add a sentence — say whether it's a push-in, a rotation, or a morph. But the two frames themselves are the strongest constraint; in most cases choosing the right ends beats iterating on the prompt.

Can I use just one image?+

Yes — fill only the first frame and it continues forward from there, close to image-to-video. The last frame is what you add when the landing point matters.

More models