AI Video Sensei
Just in

๐ŸŽฌ MiniMax H3 Guide: 2K Video, Native Audio & Open Weights

MiniMax H3 landed July 31 with native 2K, built-in dialogue and open weights. The real API limits, the five input modes, and where it still breaks.

Jordan Reyes ยท AI Video Producer

ยท 6 min read

โœ“ Fact-checked & production-testedBased on our own paid generations and published videos. Last reviewed 2026-08-01.How we test โ†’
MiniMax H3 Guide: 2K Video, Native Audio & Open Weights

MiniMax and ByteDance both shipped on July 31 โ€” Bloomberg covered the two launches as a single story. MiniMax's own release notes date H3 to July 31, 2026, though some trackers list July 29 from when listings first appeared. The headline number is native 2K. The number that actually changed our workflow is the audio track.

By the numbers

SpecMiniMax H3
Model IDMiniMax-H3
ResolutionNative 2K, no upscale pass
Frame rate24 fps, fixed
Clip length4โ€“15 seconds, integers only
AudioDialogue, SFX and ambience, same pass
Aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, 9:16
Prompt budget7,000 characters
Reference inputs9 images + 3 videos + 3 audio, 12 files max
Concurrent jobs2 free tier, 15 paid
Listed price~$0.13/sec (third-party, unofficial)
WeightsOpen, rolling out

The audio is the feature

Every other model in our rotation hands back a silent file. The pipeline that follows is always the same: generate the picture, write the line, run it through a voice model, then spend twenty minutes in an editor nudging the waveform until the lips stop lying.

H3 collapses that. You write the line inside the prompt, in quotes, and it comes back spoken over the matching mouth movement with room tone underneath. MiniMax's own reference examples pair a spoken line with an action-triggered effect and a room-tone bed in one request โ€” three audio layers, one call, no edit pass.

Two caveats worth knowing before you budget around it. Paraphrased intent ("she greets him") produces generic mumbling; you have to write the exact words. And there is no audio toggle in the API, so a silent plate means asking for silence in the prompt and stripping the track later.

Budget about two words of dialogue per second of clip. A six-second shot holds one twelve-word line without rushing.

Five input modes, not four

Most coverage lists four. The API documents five, and the one everybody misses is useful:

ModeWhat you send
Text-to-videoPrompt only โ€” aspect ratio required, cannot be adaptive
First framePrompt + opening image
Last frame onlyPrompt + closing image
First + lastPrompt + both, interpolated between
ReferencePrompt + up to 12 reference files

Last-frame-only is the interesting one. You hand H3 the frame you are cutting on and it invents the approach to it. When you have a hero composition from a still model and no idea how the camera gets there, that is a faster route than guessing at a first frame.

Frames and references are mutually exclusive โ€” the API rejects a request containing both.

Reference mode is where consistency lives

Nine images, three video clips, three audio clips, twelve files total. Each reference should carry exactly one job: this image is the face, this one is the wardrobe, this one is the set. Two references disagreeing about a hairstyle blend into a third wrong hairstyle.

Reference video does motion transfer โ€” it copies a camera path and pacing without copying appearance, if you say so explicitly. Reference audio carries a voice, and it will not travel alone: the API rejects audio-only reference sets unless at least one image or video comes with it.

The billing detail nobody mentions: reference video duration is billed as input seconds on top of your output seconds. A ten-second reference on a five-second render bills fifteen. The finished task reports the real figure in usage.total_seconds.

Prompt structure that works

H3 reads a prompt as a brief, not a caption. The order MiniMax's documented examples follow, and the one we are building our shot templates against:

SUBJECT โ†’ SCENE โ†’ ACTION (in order) โ†’ CAMERA โ†’ TIMING โ†’ STYLE โ†’ AUDIO

Keep audio in its own block or it gets rendered as an on-screen element. State camera movement in plain language โ€” "slow clockwise orbit, then settle" survives; "cinematic" destabilises the frame. Name what must stay constant ("her red scarf is unchanged throughout") and drift drops noticeably over longer takes.

One thing to unlearn from the older Hailuo models: the bracket camera syntax, [Push in] and [Truck left], belongs to the v1 Director models covered in our Hailuo/MiniMax guide. H3 does not use it. Write the move out.

Where it still breaks

Character drift over the full fifteen seconds is the most consistently reported weakness in launch-week testing โ€” wardrobe and facial features shift on long takes, which argues for keeping shots in the five-to-eight-second range until it settles. Small text, logos and packaging detail remain unreliable, so verify any on-screen type at 100% before it ships. Lip-sync accuracy and speaker identity hold up in short takes but have not been stress-tested at scale by anyone yet.

On third-party benchmarks H3 leads on video editing, trails Gemini Omni Flash on text-to-video, and sits behind both Seedance 2.0 and Gemini Omni Flash on image-to-video. It is not a straight upgrade over everything, and it does not displace the current best-in-class list outright.

Open weights change the calculation

MiniMax released H3 with open weights while ByteDance kept Seedance 2.5 in a closed API. If the download lands as promised, H3 becomes the first frontier-grade video model you can point at your own hardware โ€” no footage crossing someone else's servers, which matters for client work under NDA. Where it would slot against the models you can already run today is covered in open-weight AI video models.

Nobody has credible consumer-GPU numbers yet. Treat any VRAM figure you see this week as a guess.

How we assessed this

This is a documentation-first breakdown, not a render review โ€” H3 is days old and we have not put production footage through it yet. Every limit here is read off MiniMax's own v2 API reference and cross-checked against the create and query endpoint schemas, which is where the fifth input mode and the reference-billing rule surfaced. Quality judgements are attributed to third-party launch-week testing, not to us. Launch timing and the open-weights decision come from Bloomberg, SCMP and TechTimes coverage of the July 31 releases. Benchmark placements are third-party and pre-stabilisation; H3 is days old and still in early access.

What didn't make the cut

The "instruction-based editing" in MiniMax's launch material is a Hailuo app feature. There is no edit parameter on the v2 API, so through the API your only levers are re-rolling with the same references or feeding a previous output back as a reference video. Budget for full re-rolls.

We also left out the 30-second Extend workflow โ€” it is app-side, and the API caps at 15.

Pros and cons

Worth it for: dialogue shots, character work that needs voice and face locked together, anything where a silent clip means an hour of audio post.

Skip it for: bulk silent b-roll, on-screen text or logos, and single continuous scenes longer than fifteen seconds. Seedance 2.5 still owns that last one, and cheaper models own the first.

Frequently asked questions

โ–ธHow long can a MiniMax H3 clip be?

Four to fifteen seconds, integers only. The API rejects 5.5 or anything above 15. Fifteen seconds is where launch-week testers consistently report character drift, so treat it as a ceiling rather than a default and keep shots nearer 5 to 8 seconds.

โ–ธDoes MiniMax H3 generate audio automatically?

Yes, and you cannot turn it off. Dialogue, sound effects and room tone are produced in the same pass as the picture, and there is no parameter to disable them. If you need a silent plate, ask for silence in the prompt and strip the track afterwards.

โ–ธWhat does MiniMax H3 cost?

OpenRouter lists it from $0.13 per second, which puts a 15-second clip near $1.95. MiniMax has not published an official H3 rate card, so read usage.total_seconds off each finished task for the real number โ€” reference video seconds get billed as input on top of your output seconds.

โ–ธCan I run MiniMax H3 on my own hardware?

That is the plan. MiniMax released H3 as an open-weight model, in contrast to ByteDance keeping Seedance 2.5 behind a closed API. As of August 1 the downloadable weights are still landing, so nobody has credible consumer-GPU VRAM figures yet. Treat any you see as a guess.

The 5 best AI video finds, every week

New models, tested prompts, and what actually worked in our production โ€” one short email a week. No spam, unsubscribe anytime.

Written by Jordan Reyes

AI Video Producer

Runs multiple faceless YouTube channels and tests every major AI video model against the same prompts before recommending one. Tracks render time and credit cost like other people track calories.

#minimax h3#hailuo 3#minimax h3 api#minimax h3 pricing#hailuo 3.0 review

Keep learning