
One-take 30-second stories
A single generation covers a complete narrative beat — a continuous camera move, a full ad read, a walkthrough — instead of three clips that have to be matched in an editor.
Alibaba's all-in-one video renderer — 30-second single-pass clips, native audio, and one reference set that takes images, video, and voice together.
Betal en gang for kreditter - brug dem på tværs af hver model på ZOOOP. · Fyld op, når du har brug for det, ingen månedlig forbrænding.
Powered by Wan AI's API on ZOOOP
One request produces up to 30 seconds of continuous video — long enough for a full narrative beat, a one-take walkthrough, or a complete ad read, without stitching clips together afterwards.
Up to ten reference images, five reference clips and five reference audio tracks go into one request together. Faces, products and voices stay aligned across the whole shot instead of drifting between takes.
Lock the opening frame, optionally lock the closing frame, and Wan generates the motion that bridges them — the most reliable way to land an exact ending for a cut.
Dialogue, ambient sound and music are generated with the picture in the same pass and arrive as a stereo track on the finished file. Toggling audio off costs nothing extra.

A single generation covers a complete narrative beat — a continuous camera move, a full ad read, a walkthrough — instead of three clips that have to be matched in an editor.

Feed up to ten reference images of the same person, costume or product and the identity holds through the whole clip — the same face at second 2 and second 28.

Reference audio goes into the same request as the visual references, so a spoken scene lands with the voice you supplied rather than a generic synthetic read.

Start on the packshot, end on the hero frame. First/last frame control fixes both ends so the reveal cuts cleanly into whatever follows it.

The top tier renders at 2560×1440 for landing-page heroes and large-screen playback, where a 720p master visibly softens.

Wan 3.0 Prime is the same model on an express lane — identical controls, prioritized generation — for the pass where you are still deciding the shot.
Wan 3.0 is the pick when a shot has to run long and stay consistent with real reference material. Switch when the job wants something else.
Open Wan 3.0 from this page or pick it in the Video Generator.
Write the prompt, and add reference images, clips or audio if the shot needs them.
Pick aspect ratio, resolution up to 1440p, and a duration between 2 and 30 seconds.
Generate — the finished clip arrives with its audio track already synced.
Wan 3.0 is Alibaba's answer to the question every video model has been dodging: what happens when the shot needs to last longer than a few seconds? Most flagships cap a single generation somewhere between five and fifteen seconds, which means anything resembling a scene gets assembled in an editor from clips that never quite match — the face drifts, the light shifts, the room changes shape at the cut. Wan 3.0 renders up to thirty seconds in one pass, and that single change reorganizes how you plan a shot. A continuous camera move through a space, a complete thirty-second ad read, a one-take product walkthrough: these stop being edit problems and become prompt problems.
The second thing that sets it apart is All-in-One Reference. Reference systems usually take one kind of input — some models take images, a few take a clip, a rare one takes a voice. Wan 3.0 takes all three in the same request: up to ten reference images, five reference clips and five reference audio tracks, capped at fifteen seconds of reference video and fifteen seconds of reference audio. In practice that means a whole brand kit or a whole character sheet goes in at once, and the identity holds for the length of the clip rather than for the first few seconds. On ZOOOP the routing is automatic — attach any reference and the request goes to the reference pipeline; leave them empty and the same model runs plain text-to-video.
Two more capabilities matter day to day. First and last frame control lets you pin both ends of a clip and have Wan generate the motion between them, which is how you land an exact ending that cuts cleanly into the next shot. And native synchronized audio — dialogue, ambience, music — is generated with the picture in the same pass and arrives as a stereo track on the file, with the audio toggle costing nothing either way.
Where it's weaker: it is not open-weight. Wan 2.7 shipped under Apache 2.0 with published weights, and Wan 3.0 did not — no weights, API only. For a pipeline that needs self-hosting or open provenance, that is a hard stop and 2.7 is still the answer. Its 1440p tier is an enhanced upscale, not a native render — a real 2560×1440 file, good for large-screen delivery, but a model with a native high-resolution pass will hold detail better. There is no negative prompt, so exclusions have to be handled by describing what you do want. And on image-to-video the output follows the aspect ratio of your input frame — there is no ratio control on that path, so crop the source before you generate.
A reasonable mental model: reach for Wan 3.0 when the shot is long, or when consistency against real reference material is the thing that decides whether the take is usable. For open weights and instruction-based edits, Wan 2.7. For top-end single-shot fidelity, Veo 3.1. For quick throwaway drafts on the same family, Wan 2.6 Flash. And when you are still deciding what the shot is, run the Prime tier and switch back for the final render.
Between 2 and 30 seconds in a single generation. That is the headline difference from the previous generation, which topped out well short of that — a 30-second beat that used to need three clips and an edit is now one request.
Up to ten reference images, five reference clips and five reference audio tracks, all in the same request. Reference clips and reference audio are each capped at 15 seconds in total. Adding any reference automatically routes the request to the reference pipeline; a text-only prompt runs text-to-video.
No. Wan 2.7 shipped under Apache 2.0 with open weights; Wan 3.0 is API-only, with no published weights. If open-weight provenance or self-hosting matters for your pipeline, Wan 2.7 remains the model to use.
The 480p, 720p and 1080p tiers are native renders. The 1440p tier is produced through an enhanced upscale of a native render rather than a native 1440p pass — it delivers a genuine 2560×1440 file, and it is the right pick for large-screen delivery, but it is not the same thing as a native 1440p model.
Same model, same parameters, same limits — Prime runs on an express lane for quicker turnaround at a higher rate. Use the standard tier for final renders and Prime when you are iterating and waiting is the expensive part.
No. Wan 3.0 does not take a negative prompt — describe what you want in the positive prompt instead. Being specific about camera angle, lighting and motion does more here than trying to exclude things.
Første og sidste rammevideoTekst og henvisning til video
Første og sidste rammevideoTekst og henvisning til video
Første og sidste rammevideoTekst og henvisning til video
Første og sidste rammevideoTekst og henvisning til video
Første og sidste rammevideoTekst og henvisning til video
Første og sidste rammevideoTekst og henvisning til video
Wan 3.0
Prompt*
Billeder
Videoer
Lyd
Aspektforhold*
Opløsning*
Varighed*