๐ธ MiniMax H3 vs Veo 3.1: The Real Price Gap Explained
Veo 3.1 Standard costs 3x MiniMax H3 per second, but Veo 3.1 Fast undercuts it. Both ship native audio bundled. Which one your shot actually needs.
Jordan Reyes ยท AI Video Producer
ยท 5 min read

Google publishes a real rate card for Veo 3.1. MiniMax does not publish one for H3. That asymmetry makes this the rare AI video comparison where the cost side can actually be calculated on one half and has to be hedged on the other.
By the numbers
| MiniMax H3 | Veo 3.1 | |
|---|---|---|
| Launched | July 31, 2026 | Current Google line |
| Resolution | 2K only | 720p / 1080p / 4K |
| Standard price | ~$0.13/sec at 2K | $0.40/sec (720p & 1080p), $0.60/sec 4K |
| Fast tier | โ | $0.10/sec 720p, $0.12/sec 1080p, $0.30/sec 4K |
| Lite tier | โ | $0.05/sec 720p, $0.08/sec 1080p |
| Single-call length | 4โ15 sec | 4, 6 or 8 sec |
| Max finished length | 15 sec | ~148 sec via extend |
| Native audio | Yes, no toggle | Yes, bundled in price |
| Frame rate | 24 fps fixed | Not tier-split publicly |
| Weights | Announced, not shipped | Closed |
| Official rate card | None | Yes |
Google's prices are from the Gemini API pricing page. H3's are third-party listings โ OpenRouter quotes from $0.13/sec, ModelsLab around $0.156/sec, and other resellers say $0.14. MiniMax itself has published nothing.
The price comparison depends entirely on which Veo you mean
This is where most comparisons go wrong. "Veo 3.1" is three products.
Against Veo 3.1 Standard, H3 wins on cost outright: $0.40 versus $0.13 per second is a 3.08x gap, and H3 is delivering 2K while Standard's $0.40 tier tops out at 1080p. On a 10-second shot that is $4.00 against $1.30.
Against Veo 3.1 Fast, the gap closes and then inverts. Fast at 1080p is $0.12/sec โ about 8% cheaper per second than H3's 2K. You are trading resolution for a marginal saving, which mostly argues for H3.
Against Veo 3.1 Lite at $0.05/sec for 720p, nothing MiniMax offers competes on price โ and if cost per clip is the whole decision, our AI video cost breakdown compares the field on the same basis. If your output is a 720p social clip and the budget is the constraint, Lite is the cheapest native-audio video in this comparison by a factor of two and a half.
MiniMax's own marketing claims 2K generation costs less than one-third of mainstream rivals. Against Standard that checks out. Against Fast and Lite it does not.
Length is the structural difference
H3 does 4 to 15 seconds in a single generation. Veo does 4, 6 or 8 โ shorter per call โ but its extend feature chains roughly 7 seconds at a time to about 148 seconds total.
So the honest framing is: H3 for the longest single unbroken take, Veo for the longest finished piece. If you need a 40-second continuous scene, neither does it in one pass, but only Veo gets there at all.
Extend is not free continuity. Every hop is a seam where the model can drift, which is the same tax you pay stitching clips manually โ just automated. Seedance 2.5 is the only current model that reaches 30 seconds in a genuinely single pass, covered in its multi-shot guide. But it exists, and H3's v2 API has no equivalent.
Audio: both do it, one lets you turn it off
Both models generate synchronized dialogue and ambience in the same pass as the picture, and in both cases the price includes it. Google's listed Veo figures are explicitly the video-with-audio default, not a surcharge.
The difference is control. H3 has no audio parameter at all โ no toggle, no mute, nothing on the v2 API. Every clip comes back with a stereo track. If you need a silent plate you ask for silence in the prompt and strip the track in post.
For dialogue work that is fine. For an editor assembling silent b-roll against a scored timeline, it is an extra step on every single asset.
Quality, as measured rather than asserted
Artificial Analysis places H3 first in Video Editing and top-three in both text-to-video and image-to-video. Their text-to-video-with-audio board has Gemini Omni Flash at Elo 1246 and H3 second at 1242 โ a four-point gap. On image-to-video with audio, Seedance 2.0 leads at 1196, Gemini Omni Flash takes 1195, and H3 sits third at 1185.
One caveat that matters: Gemini Omni Flash is Google's newer video model, announced at I/O 2026, and it is a different thing from Veo 3.1 โ our Veo tool hub tracks which model sits where in Google's line. These arena standings tell you where H3 sits in the field. They are not a controlled H3-versus-Veo-3.1 test, and nobody has published one.
Decision table
| Your constraint | Take |
|---|---|
| Cheapest native-audio clip, 720p is fine | Veo 3.1 Lite ($0.05/sec) |
| Best cost per second at high resolution | H3 (2K at ~$0.13) |
| Finished video longer than 15 seconds | Veo 3.1 (extend to ~148s) |
| Longest single unbroken generation | H3 (15s in one pass) |
| 4K master | Veo 3.1 ($0.30โ0.60/sec) |
| Voice carried from a reference clip | H3 (reference audio) |
| Predictable, published pricing | Veo 3.1 |
| Silent b-roll at volume | Veo โ H3 cannot be muted |
How we picked
Every Veo 3.1 figure is from Google's official Gemini API pricing and Veo documentation. Every H3 spec is from MiniMax's published v2 API reference. The benchmark placements are Artificial Analysis's public leaderboards, cited as arena standings.
The H3 prices are the soft spot and we have flagged each one as third-party. We have not run matched shots through both models โ this is a specification and pricing comparison, not a shootout, and anyone publishing a quality verdict on a model this new is describing a handful of clips.
One further thing worth knowing before you build on H3: TechTimes reported its launch alongside an active copyright lawsuit touching the model. That is a live commercial risk that has no equivalent on Google's side.
What didn't make the cut
Cost per finished usable shot, which is the number that actually matters. It requires a known re-roll rate on both models, and two-day-old models do not have one. When we have burned real budget on both, that becomes its own post.
We also skipped the "which looks better" verdict. See the H3 guide for its documented weak spots โ character drift past ten seconds and unreliable on-screen text โ and judge against your own shot list.
Frequently asked questions
โธIs MiniMax H3 cheaper than Veo 3.1?
Against Veo 3.1 Standard, yes and by a lot โ $0.13/sec at 2K versus $0.40/sec at 720p or 1080p, roughly a 3x gap. Against Veo 3.1 Fast at 1080p ($0.12/sec) H3 is actually about 8% more expensive per second, though you are getting 2K for it. Veo 3.1 Lite at $0.05/sec undercuts everything.
โธDo both generate audio, or is that an add-on?
Both generate audio in the same pass, and in both cases it is bundled rather than surcharged. Google's listed Veo 3.1 prices are the video-with-audio default. H3 has no audio toggle at all โ it always produces a stereo track, and there is no parameter to request silence.
โธWhich makes longer videos?
Veo 3.1, by a wide margin, but not in one pass. A single Veo call gives you 4, 6 or 8 seconds; its extend feature then adds roughly 7 seconds per hop up to about 148 seconds total. H3 does 4-15 seconds in one generation with no extend path on the API. Longest single call goes to H3; longest finished video goes to Veo.
โธWhich is better quality?
Artificial Analysis ranks H3 first in Video Editing and top-three in both text-to-video and image-to-video. On text-to-video with audio, Gemini Omni Flash leads at Elo 1246 with H3 second at 1242. Veo 3.1 is a different generation of Google's line than Omni Flash, so treat these as arena standings rather than a head-to-head verdict.
The 5 best AI video finds, every week
New models, tested prompts, and what actually worked in our production โ one short email a week. No spam, unsubscribe anytime.
Written by Jordan Reyes
AI Video Producer
Runs multiple faceless YouTube channels and tests every major AI video model against the same prompts before recommending one. Tracks render time and credit cost like other people track calories.
Explore these topics
Every guide, comparison and prompt library we have on each.





