MiniMax

MiniMax H3

MiniMax's omni-modal video model — mix images, video, and audio as one reference, and get 2K clips with native stereo sound in a single call.

No subscription
Credits never expire
Learn more

Pay once for credits — use them across every model on ZOOOP. · Top up when you need to, no monthly burn.

Powered by MiniMax's API on ZOOOP

Key features

Omni-modal reference in one call

Mix up to 9 images, 3 videos, and 3 audio clips into a single reference context — H3 reads them together to hold a character, camera move, or voice across shots instead of treating each input as a separate slot.

Native stereo audio, generated in the same pass

Every clip ships with sound. Dialogue, sound effects, and music are generated alongside the picture and timed to the action, so footsteps, lines, and cuts land together — no separate scoring or dubbing step.

Prompt-level editing

Swap a subject, replace a background, relight a scene, or rewrite a line of dialogue in plain language on footage you supply — H3 keeps the untouched parts stable while applying the change.

First & last frame control

Animate a first-frame image and optionally pin the last frame for a precise start-to-end transition — the same model that drives reference work also handles clean image-to-video.

Use cases

Character-consistent series

Character-consistent series

Feed one image per subject and H3 keeps the same face, product, or mascot recognizable across every shot — the practical way to build multi-clip stories around a recurring character.

Music videos with native sound

Music videos with native sound

Supply a beat or track as an audio reference and H3 syncs the motion to it, delivering the clip with built-in stereo audio — ready to publish without a separate audio pass.

Voice-matched dialogue

Voice-matched dialogue

Send a voice sample alongside your visuals and the generated character speaks in that timbre — carry one recognizable voice across an entire series.

Product & e-commerce spots

Product & e-commerce spots

Turn a product still into a 360-style showcase or feature loop with first/last-frame control and native sound — one reference image becomes a paid-social-ready clip.

Prompt-level footage edits

Prompt-level footage edits

Replace a subject, change the background, adjust the lighting, or rewrite a line on an already-approved clip — iterate on the edit without reshooting anything.

Vertical short drama

Vertical short drama

Scripted 9:16 shorts — costume mystery, family drama — with the dialogue voiced in the same generation, so a 15-second hook is ready straight out of the model.

Pick the right model

Pick the right model for the scene, not the brand. Your credits work everywhere on ZOOOP.

Omni-modal reference + native audio in one callMiniMax H3
Multi-reference + beat-aware audioSeedance 2.0
Anime / micro-expressions / cost-effectiveHailuo 2.3
Cinematic fidelity + 4K upscaleVeo 3.1
Multi-shot storyboard sequencesKling V3
Open-weight + instruction editsWan 2.7

How to use

01

Open MiniMax H3 from this page or the Video Generator, then pick the mode — text / reference, or first-last-frame.

02

Write the prompt; for reference mode, add up to 9 images, 3 videos, and 3 audio clips as anchors for subject, motion, and voice.

03

Choose aspect ratio and duration (up to 15s) — output is 2K with native stereo audio.

04

Generate.

Deep dive

What MiniMax H3 is good at — and what it's not

MiniMax H3 is the model to reach for when a shot has to carry a specific character, voice, and sound all at once. Where most pipelines chain a video model, a text-to-speech model, and an editor together, H3 reads text, images, video, and audio as one shared context and returns a finished clip with native stereo sound in a single call. That makes it the practical choice for series work and reference-driven production — the pieces that used to need three tools and a stitching pass now come out of one generation.

The capability that separates H3 from the field is its omni-modal reference. In one request you can hand it up to nine images, three video clips, and three audio tracks, and it treats them as a single context rather than isolated input slots — pulling a face from a photo, a camera move from a clip, and a mood or beat from a track into the same scene. That is what keeps a character, product, or mascot recognizable across a multi-shot story, and what lets a supplied voice sample carry one identity across an entire series. Because references are expressed in natural language rather than a fixed task menu, the same model also does prompt-level editing — swap a subject, replace a background, relight a scene, or rewrite a line of dialogue on footage you already have, with the untouched regions held stable.

The second flagship capability is native audio. Every clip is delivered with sound generated in the same pass as the picture — dialogue, sound effects, and music timed to the action, so a line, a footstep, and a cut all land together. For short ads and social-native clips that means a result is publish-ready straight out of the model, without a separate scoring or dubbing step. Output is 2K across aspect ratios from cinematic 21:9 to vertical 9:16, at 5–15 second clip lengths (first-last-frame image-to-video runs up to 10 seconds), which covers trailers, product loops, and vertical short drama without switching models.

Where it's weaker: for the absolute top tier of cinematic resolution and 4K delivery, Veo 3.1 still leads. For multi-shot storyboarding described in a single prompt, that's Kling V3's lane. And for the strongest anime and micro-expression work at the best price-per-quality, Hailuo 2.3 is the pick. H3's sweet spot is reference-driven consistency, voice-matched dialogue, and getting picture-plus-sound out of one generation.

A reasonable mental model: default to MiniMax H3 when a shot needs a consistent character or voice, mixed references, or built-in audio. For cinematic 4K, switch to Veo 3.1; for multi-shot sequences, Kling V3; for anime and value drafting, Hailuo 2.3. Whichever you pick, the same credits work across all of them.

Frequently asked questions

What makes MiniMax H3 different from a standard video model?+

Most video models take a prompt and one image and return a silent clip. H3 reads a mixed set of text, images, video, and audio as one shared context and returns a finished clip with native sound — folding generate, edit, and reference into a single model instead of chaining a video model, a voice model, and an editor.

Does every H3 video have audio?+

Yes — every result is delivered with native stereo sound. Dialogue, sound effects, and ambient audio are generated in the same pass as the picture. Provide a voice reference and the character speaks in that timbre, which removes a separate text-to-speech step from the pipeline.

What can I use as a reference?+

In reference mode, up to 9 images, 3 videos, and 3 audio clips in one request (12 files total). At least one image or video is required; audio must accompany a visual reference and each audio clip must be 2–15 seconds. H3 reads them as one context — pulling a character from a photo, a camera move from a clip, and a mood from a track into the same scene.

Can I edit existing footage with H3?+

Yes — editing is one of its core modes. Send a clip and an instruction to swap a character or object, replace a background, adjust lighting, or rewrite a line of dialogue, and H3 keeps the unmentioned parts close to the original. Multiple edits can be stacked into one request.

What resolution, length, and aspect ratios does H3 support?+

Output is 2K with native audio. Reference and text modes produce 5–15 second clips; first-last-frame runs 5–10 seconds. Aspect ratios span 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, so one model covers cinematic widescreen and vertical social.

Is MiniMax H3 cost-effective?+

Yes — for a 2K model with native audio it sits among the lower-cost options on ZOOOP, and reference inputs are not billed separately (you pay by the second of video generated). Iterate freely, and your credits never expire.

More models

ByteDance
Seedance V2.0
ByteDance
Hailuo AI
Hailuo 2.3
Hailuo AI
Google
Veo 3.1
Google
Kling AI
Kling V3
Kling AI
OpenAI
GPT Image 2.0
OpenAI
MiniMax
Minimax Music V2.6
MiniMax
MiniMax
Minimax Music V2
MiniMax
MiniMax
Speech-2.8-HD
MiniMax
MiniMax
Speech-2.8-Turbo
MiniMax
ByteDance
Seedance V2.0 Fast
ByteDance
ByteDance
Seedance V1.5 Pro
ByteDance
ByteDance
Seedance V1.0 Pro
ByteDance
ByteDance
Seedance V1.0 Pro Fast
ByteDance
ByteDance
Seedance V1.0 Lite
ByteDance
ByteDance
Seedream 5.0 Pro
ByteDance
ByteDance
Seedream 5.0 Lite
ByteDance
ByteDance
Seedream 4.5
ByteDance
ByteDance
Seedream 4
ByteDance
ByteDance
Dreamactor V2
ByteDance
Kling AI
Kling O3
Kling AI
Kling AI
Kling V3 Pro
Kling AI
Kling AI
Kling V2.6 Pro
Kling AI
Kling AI
Kling V2.6
Kling AI
Kling AI
Kling Lipsync
Kling AI
Kling AI
Kling Avatar V2
Kling AI
Kling AI
Kling O1
Kling AI
Midjourney
Midjourney V8.2
Midjourney
Midjourney
Midjourney V8.1
Midjourney
Midjourney
Midjourney
Midjourney
Alibaba
Happy Horse V1.1
Alibaba
Alibaba
Happy Horse
Alibaba
xAI
Grok Imagine V1.5
xAI
xAI
Grok Imagine
xAI
xAI
xAI TTS
xAI
Google
Veo 3.1 Fast
Google
Google
Veo 3
Google
Google
Nano Banana Pro
Google
Google
Nano Banana 2
Google
Google
Nano Banana
Google
Google
Lyria 3 Pro
Google
Google
Lyria 3
Google
Google
Lyria2
Google
Google
Gemini 3.1 Flash TTS
Google
Wan AI
Wan V2.2
Wan AI
Wan AI
Wan V2.2 Turbo
Wan AI
Wan AI
Wan V2.5
Wan AI
Wan AI
Wan V2.6
Wan AI
Wan AI
Wan V2.6 Flash
Wan AI
Wan AI
Wan V2.7
Wan AI
Pixverse AI
Pixverse V6
Pixverse AI
Pixverse AI
Pixverse V5.5
Pixverse AI
Pixverse AI
Pixverse V5
Pixverse AI
Pixverse AI
Pixverse Lipsync
Pixverse AI
Vidu AI
Vidu Q3 Pro
Vidu AI
Vidu AI
Vidu Q3
Vidu AI
Vidu AI
Vidu Q3 Turbo
Vidu AI
Vidu AI
Vidu Q2 Pro
Vidu AI
Vidu AI
Vidu Q2 Turbo
Vidu AI
Luma AI
Luma Ray 2
Luma AI
Luma AI
Luma Ray 2 Flash
Luma AI
Flux AI
Flux 2 Pro
Flux AI
Flux AI
Flux 2
Flux AI
Flux AI
Flux 2 Flash
Flux AI
ElevenLabs
Multilingual V3
ElevenLabs
ElevenLabs
Multilingual V2
ElevenLabs
ElevenLabs
Sound Effects V2
ElevenLabs
Hailuo AI
Hailuo 2.3 Fast
Hailuo AI
Hailuo AI
Hailuo 02
Hailuo AI
Lightricks
LTX-2.3 Pro
Lightricks
Lightricks
LTX-2.3 Fast
Lightricks
Lightricks
LTX-2.3
Lightricks
Pika AI
Pika V2.2
Pika AI
Qwen
Qwen3-TTS
Qwen
Inworld
Inworld TTS
Inworld
Bilibili Index
Index TTS 2
Bilibili Index
Resemble AI
Chatterbox TTS Multilingual
Resemble AI
Open Source
ACE-Step
Open Source
Open Source
LUX TTS
Open Source
CassetteAI
Music Generator
CassetteAI
Recraft
Text to Vector V4.1
Recraft
Recraft
Text to Vector V4.1 Pro
Recraft
Recraft
Image to Vector
Recraft