๐๏ธ Grok Imagine Video 1.5: 7 References, Voice Lock, 1080p
Grok Imagine Video 1.5 adds seven reference images, three voice refs and native 1080p. How the update works, how it stacks up against Kling and Seedance.
Jordan Reyes ยท AI Video Producer
ยท 7 min read
โก TL;DR โ quick answers
- How many reference images does Grok Imagine Video 1.5 support?
- Up to seven per generation, and they can be people, objects, or keyframes. Each reference locks one thing in place โ a face, a product, a location โ so you can keep a character and swap the scene, or hold both and change only the action. References persist, so the same set can drive every shot in a project.
- Does Grok Imagine Video 1.5 output 1080p?
- Natively, yes โ for text-to-video and image-to-video, with no upscaling pass. The catch is that reference-to-video, the mode most people care about after this update, is capped at 720p on the API as of August 2026. Expect that gap to close, but budget around it today.
- What does Grok Imagine Video 1.5 cost?
- API pricing on fal runs $0.08/second at 480p and $0.14/second at 720p, with each additional reference image adding $0.01. In the Grok app, access comes through SuperGrok subscriptions โ third-party breakdowns list a $10/month Lite tier with daily caps and a $30/month tier for full Imagine. All of this moves fast, so verify before budgeting.

Grok Imagine spent most of 2026 as the model people used for speed and memes, not for projects. The 1.5 update that started rolling out July 31 is xAI's play to change that: up to seven reference images per generation, up to three voice references, text-to-video from a prompt alone, and native 1080p. That's not a spec bump โ it's the consistency toolkit that Kling and Seedance built their professional cases on, arriving all at once.
We've been running reference-based pipelines on Seedance 2.5's omni-reference for months, so we know exactly what this feature class is worth. Here's what 1.5 actually adds, where it beats the incumbents, and where it doesn't yet.
By the numbers
| Spec | Grok Imagine Video 1.5 |
|---|---|
| 1.5 original launch | June 17, 2026 |
| References update rollout | July 31 โ August 2, 2026 |
| Reference images per generation | Up to 7 (people, objects, keyframes) |
| Audio/voice references | Up to 3 |
| Native resolution (t2v / i2v) | 1080p |
| Reference-to-video resolution | 480p / 720p (API, as of Aug 2026) |
| API pricing (fal) | $0.08/sec at 480p, $0.14/sec at 720p |
| Extra reference cost | +$0.01 per image |
| First access | US SuperGrok Heavy & Plus, grok.com/imagine + iOS |
| Launch-week claim (TechTimes, June) | Topped AI video leaderboard at ~86% below Sora's price |
Every figure above comes from xAI's announcement, fal's live API listing, or dated coverage from TechTimes and Crypto Briefing โ not from marketing decks.
Seven references is a different way of thinking about shots
Most reference systems ask for one identity image and hope. Grok's approach assigns each of the seven slots a job: this image is the face, this one is the product, this one is the location, this one is a keyframe the shot should pass through. The model composites your prompt around those locks.
In practice that means you can keep a character and swap the scene, keep the scene and swap the character, or hold both and change only the action โ which is precisely the shot-to-shot grammar a real edit needs. Mark Kretschmann's early demos show how flexible the slot system is:
The part that matters for production is persistence. References are reusable across generations, so a character image plus a voice reference holds โ same face, same voice, every scene. On our reference workflows, re-uploading and re-describing identity anchors for every single shot is the biggest time sink; a persistent reference library is the correct fix, and Grok shipped it before Kling did.
If you're new to this whole discipline, our consistent characters guide covers the underlying craft โ reference images are an anchor, not a guarantee, and prompt wording can still override them if you let adjectives fight the reference.
Voice references are the sleeper feature
Up to three audio references means face-and-voice consistency in one generation pass. Everyone covering this update led with the seven images; we think the voice slots are the bigger deal, because voice continuity is the thing AI video pipelines still do worst.
The current standard workflow โ ours included โ is to generate video silent or with placeholder audio, then lay ElevenLabs VO over it and fix lip timing in the edit. If Grok's voice references hold a voice identity across scenes the way the announcement describes, that collapses a two-tool chain into one. We haven't stress-tested voice drift across a long project yet, and we'd want to before moving anything client-facing โ but this is the feature we're watching.
Native 1080p โ with an asterisk you need to read
Text-to-video and image-to-video now render native 1080p, no upscaling pass. Reference-to-video โ the mode this entire update exists to promote โ is capped at 720p on the API as of August 2026. That's an odd seam: the more control you take, the fewer pixels you get.
For vertical social output it barely matters. For 16:9 YouTube delivery it means your reference-locked shots still want an upscale pass while your prompt-only establishing shots don't. Plan the pipeline accordingly, and expect xAI to close the gap โ the 1.5 base model went from 720p-limited to 1080p in under two months.
Does the quality hold up on real projects rather than cherry-picked clips? The longest-form public test we've seen is worth your time:
Grok 1.5 vs Kling 3.0 Omni vs Seedance 2.5
| Your job | Pick |
|---|---|
| Multi-element identity lock (cast + props + location) | Grok 1.5 โ 7 slots, or Seedance 2.5 omni-reference |
| Face-and-voice consistency in one pass | Grok 1.5 โ voice refs are unique right now |
| Multi-shot scene with in-generation cuts | Kling 3.0 Omni โ 6 cuts, one context |
| Long continuous takes | Seedance 2.5 โ up to 30 seconds |
| Native 4K delivery | Kling 3.0 Omni โ Grok tops out at 1080p |
| Cheapest per-second API iteration | Grok 1.5 at $0.08/sec 480p |
The honest read: Grok now matches Seedance's reference ceiling (Seedance 2.5 takes up to nine images on Higgsfield's omni-reference, Grok takes seven plus three voices), undercuts everyone on API price, but has no answer to Kling's in-generation editing or Seedance's 30-second takes. It's a consistency-first tool with short-clip grammar. Where it sits in the broader field is in our best AI video generators rankings.
Pricing and access, as of August 2026
Two doors in. The Grok app route: rollout began with US SuperGrok Heavy and Plus subscribers on grok.com/imagine and iOS, extending to all tiers within days; third-party breakdowns list SuperGrok Lite at $10/month with daily video caps and short durations, and SuperGrok at $30/month for full Imagine. The API route: grok-imagine-video-1.5 on fal runs $0.08/second at 480p and $0.14/second at 720p, each extra reference image adding a cent โ a 5-second 720p clip is $0.70 before references.
That per-second rate is genuinely cheap for this feature class, and TechTimes' June launch coverage framed the whole line as roughly 86% below Sora's pricing. Treat every number here as perishable โ xAI has changed Imagine's specs twice since June. Our AI video cost breakdown has the cross-model math.
How we tested and picked sources
This is a nine-day-old update, so we're explicit about epistemics: specs come from xAI's official announcement and API listings, cross-checked against dated reporting from TechTimes, Crypto Briefing and TestingCatalog, plus hands-on demos from early-access users we've followed through the rollout. Our comparative judgments come from months of production work with reference systems on Seedance 2.5 and Kling 3.0 โ the workflow claims are first-hand; the Grok-specific quality verdicts are provisional until we've pushed a full multi-shot project through it.
What didn't make the cut
- Head-to-head frame comparisons. Nine days of cherry-picked social clips is not a corpus. We'll publish comparisons when we've run controlled prompts.
- Spicy-mode and moderation coverage. It's the most-searched Grok topic and irrelevant to production work.
- A prompt library. Grok's prompt grammar is still shifting release to release; anything we print now would be stale in a month.
Verdict
Grok Imagine Video 1.5 is the moment xAI's video model became a legitimate production candidate rather than a toy. Seven persistent reference slots plus voice locking is a genuinely competitive consistency stack, and the API pricing undercuts everything comparable. The 720p reference-mode cap and the unproven voice-drift behaviour are the reasons we're not moving pipelines yet โ but this is now a model we're testing seriously, which was not true in June.
The official spec sheet is at xAI's announcement, with API details in the xAI Imagine docs.
Frequently asked questions
โธHow many reference images does Grok Imagine Video 1.5 support?
Up to seven per generation, and they can be people, objects, or keyframes. Each reference locks one thing in place โ a face, a product, a location โ so you can keep a character and swap the scene, or hold both and change only the action. References persist, so the same set can drive every shot in a project.
โธDoes Grok Imagine Video 1.5 output 1080p?
Natively, yes โ for text-to-video and image-to-video, with no upscaling pass. The catch is that reference-to-video, the mode most people care about after this update, is capped at 720p on the API as of August 2026. Expect that gap to close, but budget around it today.
โธWhat does Grok Imagine Video 1.5 cost?
API pricing on fal runs $0.08/second at 480p and $0.14/second at 720p, with each additional reference image adding $0.01. In the Grok app, access comes through SuperGrok subscriptions โ third-party breakdowns list a $10/month Lite tier with daily caps and a $30/month tier for full Imagine. All of this moves fast, so verify before budgeting.
โธWho got the references update first?
Rollout started July 31, 2026 in the US for SuperGrok Heavy and SuperGrok Plus subscribers on grok.com/imagine and iOS, expanding to all tiers over the following days. Image references, text-to-video and native 1080p also landed in the xAI API as grok-imagine-video-1.5.
The 5 best AI video finds, every week
New models, tested prompts, and what actually worked in our production โ one short email a week. No spam, unsubscribe anytime.

Written by Jordan Reyes
AI Video Producer
Runs multiple faceless YouTube channels and tests every major AI video model against the same prompts before recommending one. Tracks render time and credit cost like other people track calories.
Explore these topics
Every guide, comparison and prompt library we have on each.





