๐ฌ MiniMax H3 Guide: 2K Video, Native Audio & Open Weights
MiniMax H3 landed July 31 with native 2K, built-in dialogue and open weights. The real API limits, the five input modes, and where it still breaks.
Jordan Reyes ยท AI Video Producer
ยท 6 min read

MiniMax and ByteDance both shipped on July 31 โ Bloomberg covered the two launches as a single story. MiniMax's own release notes date H3 to July 31, 2026, though some trackers list July 29 from when listings first appeared. The headline number is native 2K. The number that actually changed our workflow is the audio track.
By the numbers
| Spec | MiniMax H3 |
|---|---|
| Model ID | MiniMax-H3 |
| Resolution | Native 2K, no upscale pass |
| Frame rate | 24 fps, fixed |
| Clip length | 4โ15 seconds, integers only |
| Audio | Dialogue, SFX and ambience, same pass |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| Prompt budget | 7,000 characters |
| Reference inputs | 9 images + 3 videos + 3 audio, 12 files max |
| Concurrent jobs | 2 free tier, 15 paid |
| Listed price | ~$0.13/sec (third-party, unofficial) |
| Weights | Open, rolling out |
The audio is the feature
Every other model in our rotation hands back a silent file. The pipeline that follows is always the same: generate the picture, write the line, run it through a voice model, then spend twenty minutes in an editor nudging the waveform until the lips stop lying.
H3 collapses that. You write the line inside the prompt, in quotes, and it comes back spoken over the matching mouth movement with room tone underneath. MiniMax's own reference examples pair a spoken line with an action-triggered effect and a room-tone bed in one request โ three audio layers, one call, no edit pass.
Two caveats worth knowing before you budget around it. Paraphrased intent ("she greets him") produces generic mumbling; you have to write the exact words. And there is no audio toggle in the API, so a silent plate means asking for silence in the prompt and stripping the track later.
Budget about two words of dialogue per second of clip. A six-second shot holds one twelve-word line without rushing.
Five input modes, not four
Most coverage lists four. The API documents five, and the one everybody misses is useful:
| Mode | What you send |
|---|---|
| Text-to-video | Prompt only โ aspect ratio required, cannot be adaptive |
| First frame | Prompt + opening image |
| Last frame only | Prompt + closing image |
| First + last | Prompt + both, interpolated between |
| Reference | Prompt + up to 12 reference files |
Last-frame-only is the interesting one. You hand H3 the frame you are cutting on and it invents the approach to it. When you have a hero composition from a still model and no idea how the camera gets there, that is a faster route than guessing at a first frame.
Frames and references are mutually exclusive โ the API rejects a request containing both.
Reference mode is where consistency lives
Nine images, three video clips, three audio clips, twelve files total. Each reference should carry exactly one job: this image is the face, this one is the wardrobe, this one is the set. Two references disagreeing about a hairstyle blend into a third wrong hairstyle.
Reference video does motion transfer โ it copies a camera path and pacing without copying appearance, if you say so explicitly. Reference audio carries a voice, and it will not travel alone: the API rejects audio-only reference sets unless at least one image or video comes with it.
The billing detail nobody mentions: reference video duration is billed as input seconds on top of your output seconds. A ten-second reference on a five-second render bills fifteen. The finished task reports the real figure in usage.total_seconds.
Prompt structure that works
H3 reads a prompt as a brief, not a caption. The order MiniMax's documented examples follow, and the one we are building our shot templates against:
SUBJECT โ SCENE โ ACTION (in order) โ CAMERA โ TIMING โ STYLE โ AUDIO
Keep audio in its own block or it gets rendered as an on-screen element. State camera movement in plain language โ "slow clockwise orbit, then settle" survives; "cinematic" destabilises the frame. Name what must stay constant ("her red scarf is unchanged throughout") and drift drops noticeably over longer takes.
One thing to unlearn from the older Hailuo models: the bracket camera syntax, [Push in] and [Truck left], belongs to the v1 Director models covered in our Hailuo/MiniMax guide. H3 does not use it. Write the move out.
Where it still breaks
Character drift over the full fifteen seconds is the most consistently reported weakness in launch-week testing โ wardrobe and facial features shift on long takes, which argues for keeping shots in the five-to-eight-second range until it settles. Small text, logos and packaging detail remain unreliable, so verify any on-screen type at 100% before it ships. Lip-sync accuracy and speaker identity hold up in short takes but have not been stress-tested at scale by anyone yet.
On third-party benchmarks H3 leads on video editing, trails Gemini Omni Flash on text-to-video, and sits behind both Seedance 2.0 and Gemini Omni Flash on image-to-video. It is not a straight upgrade over everything, and it does not displace the current best-in-class list outright.
Open weights change the calculation
MiniMax released H3 with open weights while ByteDance kept Seedance 2.5 in a closed API. If the download lands as promised, H3 becomes the first frontier-grade video model you can point at your own hardware โ no footage crossing someone else's servers, which matters for client work under NDA. Where it would slot against the models you can already run today is covered in open-weight AI video models.
Nobody has credible consumer-GPU numbers yet. Treat any VRAM figure you see this week as a guess.
How we assessed this
This is a documentation-first breakdown, not a render review โ H3 is days old and we have not put production footage through it yet. Every limit here is read off MiniMax's own v2 API reference and cross-checked against the create and query endpoint schemas, which is where the fifth input mode and the reference-billing rule surfaced. Quality judgements are attributed to third-party launch-week testing, not to us. Launch timing and the open-weights decision come from Bloomberg, SCMP and TechTimes coverage of the July 31 releases. Benchmark placements are third-party and pre-stabilisation; H3 is days old and still in early access.
What didn't make the cut
The "instruction-based editing" in MiniMax's launch material is a Hailuo app feature. There is no edit parameter on the v2 API, so through the API your only levers are re-rolling with the same references or feeding a previous output back as a reference video. Budget for full re-rolls.
We also left out the 30-second Extend workflow โ it is app-side, and the API caps at 15.
Pros and cons
Worth it for: dialogue shots, character work that needs voice and face locked together, anything where a silent clip means an hour of audio post.
Skip it for: bulk silent b-roll, on-screen text or logos, and single continuous scenes longer than fifteen seconds. Seedance 2.5 still owns that last one, and cheaper models own the first.
Frequently asked questions
โธHow long can a MiniMax H3 clip be?
Four to fifteen seconds, integers only. The API rejects 5.5 or anything above 15. Fifteen seconds is where launch-week testers consistently report character drift, so treat it as a ceiling rather than a default and keep shots nearer 5 to 8 seconds.
โธDoes MiniMax H3 generate audio automatically?
Yes, and you cannot turn it off. Dialogue, sound effects and room tone are produced in the same pass as the picture, and there is no parameter to disable them. If you need a silent plate, ask for silence in the prompt and strip the track afterwards.
โธWhat does MiniMax H3 cost?
OpenRouter lists it from $0.13 per second, which puts a 15-second clip near $1.95. MiniMax has not published an official H3 rate card, so read usage.total_seconds off each finished task for the real number โ reference video seconds get billed as input on top of your output seconds.
โธCan I run MiniMax H3 on my own hardware?
That is the plan. MiniMax released H3 as an open-weight model, in contrast to ByteDance keeping Seedance 2.5 behind a closed API. As of August 1 the downloadable weights are still landing, so nobody has credible consumer-GPU VRAM figures yet. Treat any you see as a guess.
The 5 best AI video finds, every week
New models, tested prompts, and what actually worked in our production โ one short email a week. No spam, unsubscribe anytime.
Written by Jordan Reyes
AI Video Producer
Runs multiple faceless YouTube channels and tests every major AI video model against the same prompts before recommending one. Tracks render time and credit cost like other people track calories.
Keep learning
ComparisonsMiniMax H3 vs Seedance 2.5: Open Weights vs Native 4K
Two frontier video models shipped within hours on July 31. One went open-weight at 2K, the other stayed closed at 4K. Which one fits your work.
2026-08-01
GuidesHailuo AI (MiniMax) Guide: The Free-Credits Video Model (2026)
Hailuo 2.3 tested: what 200 free credits actually buy, the 768p-vs-1080p credit math, real pricing tiers, and where MiniMax's model beats the big names.
2026-07-23
GuidesOpen-Weight AI Video Models You Can Actually Run
MiniMax H3 going open-weight restarts this argument. What runs on 8GB, what needs 24GB, and what the quantization actually costs you in quality.
2026-08-01
GuidesSeedance 2.5: Native 4K, 30s Shots & What Changes
ByteDance's Seedance 2.5 launched with native 4K, 30-second outputs and 3D pre-visualization. What's confirmed, what it means for 2.0 workflows, and our migration plan.
2026-07-16
GuidesKling 3.0 Omni Guide: 4K Editing, 15s Clips & 6 Cuts in One Take
Kling 3.0 Omni is the editing-capable tier, not just a bigger renderer. What the June upgrade changed, when to pick it over Turbo, and the upscaler it replaces.
2026-07-31