New · 5–15 seconds · Up to 4K · Native stereo audio

MiniMax H3

Create 5–15-second videos from text, start and end frames, or image, video and audio references, with output up to 4K and native stereo audio.

Use one MiniMax H3 engine for text-to-video, image-to-video and reference-to-video. Choose the duration and resolution, then direct character continuity, motion, camera and sound in one workflow.

MiniMax H3 lighthouse keeper rescue sequence with storm ambience and native stereo dialogue
15s
2K16:9

Real MiniMax H3 render

Elara races through a storm to relight a lighthouse and guide a lifeboat home

View full render

Three creation workflows

Move between text-to-video, image-to-video and reference-to-video inside one H3 engine.

5–15 seconds at up to 4K

Choose any whole-second duration and render at 768P, 2K or 4K at 24 FPS.

Multimodal reference direction

Combine image, video and audio references to guide identity, performance, motion and sound.

Native stereo audio

Generate ambience, effects, music or speech with the picture in the same render.

MiniMax H3 pricing for your render

Choose workflow, duration, resolution and references, then review the exact price before rendering.

Generate with MiniMax H3

entryText

$0.52

audio incl.

imageProduction

$1.69

Most popular

audio incl.

referenceFinal

$3.12

audio incl.

fullImageReferenceSet

$2.11

audio incl.

The first five reference images are included; each additional reference image adds its displayed surcharge.

Real MiniMax H3 outputs

Review accepted MaxVideoAI renders that test character continuity, camera direction, reference control and synchronized sound.

Character continuity

Inspect face, wardrobe and performance across the full sequence.

Camera direction

Follow framing, camera movement and transitions from beat to beat.

Reference control

Compare the intended identity, movement and sound with the final video.

Native stereo sound

Evaluate dialogue timing, ambience and synchronized effects.

Real production renders

Every example on this page is generated through MaxVideoAI.

Starting from a written scene?

Use text-to-video and choose Auto or a fixed aspect ratio for a complete shot from one prompt.

Generate from text

Need a controlled opening and ending?

Animate one start image and add an optional end image; H3 follows the source framing automatically.

Generate from frames

Directing a recurring character?

Use reference mode to assign identity, wardrobe, motion and voice cues across the scene.

Open Prompt Lab

Prompt Lab — MiniMax H3

Build the prompt in three passes: define the performance, direct the shot sequence, then lock continuity and sound.

Give each shot one visible action, one camera instruction and one sound event.

Source: MiniMax H3 guide

A character-first workflow for H3 scenes

Text-to-video

Describe the character, action, setting, camera and sound, then choose a fixed ratio or Auto.

Image-to-video

Upload a start image and optionally an end image; the output follows the source aspect ratio.

Reference-to-video

Combine up to 9 images, 3 videos and 3 audio clips while staying within 12 unique references.

Native audio direction

Write dialogue, ambience and synchronized sound cues directly into the scene prompt.

Define the character and action

Establish identity, expression, body language and the visible change in the scene.

Character:
[Appearance, wardrobe, expression and voice]

Setting:
[Location, time, weather and visual anchors]

Action:
[One clear physical action and reaction]

Sound:
[Dialogue, ambience and synchronized effects]
EXAMPLE

Sound: [Dialogue, ambience and synchronized effects]

Global principles

  • Put character and action before style adjectives.
  • Link every camera movement to a visible story beat.
  • Describe sounds at the moment their visible cause occurs.
  • Keep identity, wardrobe, lighting and screen direction explicit across shots.

Engine quirks / what to watch for

  • H3 combines text, frames and multimodal references in one production model.
  • Multi-shot direction supports a short sequence instead of one isolated movement.
  • Native stereo audio lets dialogue, ambience and effects develop with the picture.

Lighthouse keeper rescue sequence

Text-to-video

Subject: Elara, a lighthouse keeper in a mustard raincoat  •  Action: Races up the tower, powers the lens and calls a lifeboat home
Camera: Handheld exterior tracking, stairwell close follow, lantern-room orbit and aerial pullback  •  Style: Storm-lit cinematic naturalism at blue hour
2K render: 15-second 2K multi-shot scene with native stereo dialogue

View full prompt
Create a 15-second cinematic 16:9 film with native multi-shot continuity and native stereo sound. Original fictional character: ELARA, a 34-year-old lighthouse keeper with wind-burned olive skin, dark wavy hair tied back, a weathered mustard raincoat, charcoal sweater, and practical sea boots. Keep her face, clothing, and proportions consistent in every shot. 0.0–3.5 seconds: wide exterior at blue hour, a stone lighthouse on a violent Atlantic cliff, rain sweeping sideways, Elara runs toward the tower carrying a small weathered signal case; handheld tracking camera, realistic wet fabric and foot contact. 3.5–7.5 seconds: interior spiral stairwell, medium close tracking shot as she climbs fast, breath visible, one hand gripping the rail, practical amber lamps, no cuts that break identity. 7.5–11.5 seconds: lantern room, she reaches the controls and throws the emergency lever; orbit from her determined face to the huge Fresnel lens as it powers on. At 10.5 seconds Elara says clearly, with natural lip sync: “Harbor light is alive. Bring them home.” 11.5–15.0 seconds: exterior aerial pullback as the rescue beam sweeps across black waves and reveals a distant lifeboat turning toward safety. Audio: native stereo only—storm wind outside, rain on glass, boots on iron stairs, strained breathing, lever clunk, electrical rise, deep lighthouse mechanism, distant surf, and Elara's exact spoken line. No music. No subtitles, captions, on-screen text, logos, watermarks, beauty-ad framing, or product pack shot. Photorealistic skin and motion, coherent hands, restrained cinematic contrast.
15s16:9Native stereo audio
MiniMax H3 lighthouse keeper rescue sequence
15s16:9

Tips and boundaries

Best practices, common fixes, and important limitations to help you get the strongest results with H3.

What works best

  • Name the character, action, setting, camera and sound in that order.
  • Use one visible action and one camera move for each shot beat.
  • Assign references explicit roles such as identity, motion, wardrobe or voice.
  • Repeat only the continuity anchors that must survive every shot.

Common problems → fast fixes

  • Feels random / inconsistent → simplify to: subject + action + camera + lighting. Re-run 2–3 takes.
  • Motion looks weird → reduce movement: one camera move, slower action, fewer props.
  • Subject drifts off-brand → start from a reference image and lock palette + lighting.
  • Text looks wrong → avoid readable signage, tiny UI, micro labels. Keep text off-screen.
  • Dialogue drifts → keep lines short and punchy; avoid long monologues.

Hard limits to keep in mind

  • Output is short-form (5-15s). For longer edits, stitch multiple clips.
  • Resolution tops out at 768p / 2K / 4K for this tier.
  • No fixed seeds — iteration = re-run + refine.

Compare H3 vs other AI video models

These side-by-side comparisons break down price, resolution, audio, speed, and motion style so you can pick the right engine fast.

Each page includes real outputs and practical best-use cases.

H3 vs Kling 3.0 Omni Pro

Use Kling 3.0 Omni Pro for reference-guided Kling video, storyboard images, @Image prompt anchors, optional start/end frames, video-to-video, and native audio.

Compare H3 vs Kling 3.0 Omni Pro →

H3 vs Google Veo 3.1

Generate cinematic Veo 3.1 videos with text prompts, start-image animation, multi-reference guidance, optional last-frame control, and extend workflows in one unified MaxVideoAI model page.

Compare H3 vs Google Veo 3.1 →

MiniMax H3 overview

The limits that shape your renders.

Price / second

768P $0.10/s2K $0.17/s4K $0.21/s

Text-to-Video

Supported

Image-to-Video

Supported

Video-to-Video

Supported as a reference-to-video input

First/Last frame

Supported (start image + optional end image)

Start / reference image

Up to 9 image references

Reference video

Up to 3 video references; 2-15s each and 15s combined

Max resolution

768p / 2K / 4K

Max duration

5-15s

Aspect ratios

Auto / 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16

FPS options

24 fps

Output format

MP4

Audio output

Native stereo audio

Native audio generation

Supported

Lip sync

Supported

Camera / motion controls

Prompt-based camera and multi-shot control

Watermark

No (MaxVideoAI)

Release date

Jul 2026

MiniMax H3 overview

Details
  • Duration: Every integer second from 5 to 15 seconds
  • Output: 768P, 2K or 4K · Auto, 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16 · 24 FPS
  • Text and frames: Text-to-video, or a start image with an optional end image
  • Reference workflow: Up to 9 images, 3 videos and 3 audio clips, with 12 unique references in total
  • Audio: Native stereo audio generated with every video

Safety & people / likeness

Built-in safeguards and best practices for responsible creation with H3.

  • Use original characters and owned references.
  • Avoid real people, celebrities and protected characters.
  • Do not use someone's likeness without consent.
  • Avoid copyrighted franchises, logos and protected IP.

FAQ

What inputs does MiniMax H3 accept?

Use a text prompt, a start image with an optional end image, or a reference set containing images, videos and audio. Audio references must accompany at least one image or video reference.

What duration and resolution can I choose?

Choose any whole-second duration from 5 through 15 seconds and render at 768P, 2K or 4K at 24 FPS.

How many references can H3 use?

Reference mode accepts up to 9 images, 3 videos and 3 audio clips, with no more than 12 unique references in one request.

Does MiniMax H3 generate audio?

Yes. H3 produces native stereo audio with the video; there is no separate audio switch.

Can I generate with MiniMax H3 from this page?

Yes. MiniMax H3 is available in MaxVideoAI, and the exact price is shown before every render.