Just in

๐ŸŽฏ 35 Tested Veo 3.1 Prompts: Formula, Dialogue Syntax, Templates

A working Veo 3.1 prompt library: the 6-part formula, native-audio dialogue syntax, 12 scene templates, and the mistakes that waste your Google AI credits.

Jordan Reyes

Jordan Reyes ยท AI Video Producer

ยท Updated ยท 4 min read

โœ“ Fact-checked & production-testedBased on our own paid generations and published videos. Last reviewed 2026-08-02.How we test โ†’
โšก TL;DR โ€” quick answers
What is the best prompt structure for Veo 3.1?
Six parts: subject โ†’ action โ†’ setting โ†’ camera โ†’ lighting/style โ†’ audio. Veo rewards explicit audio direction โ€” dialogue in quotes, ambient sound described โ€” because it generates sound natively.
How do I make characters speak in Veo?
Write the exact line in quotes and attribute it: The barista smiles and says: 'First one's on the house.' Veo generates the voice with lip-sync. Keep lines under ~15 words per 8-second clip.
Do JSON prompts really work better in Veo?
Structured prompts (JSON-style key-value blocks) don't unlock hidden features, but they force you to specify every element โ€” which reliably improves results and makes A/B iteration much easier. Use them for complex shots.
Cinematic AI video production illustration for: 35 Tested Veo 3.1 Prompts: Formula, Dialogue Syntax, Templates

Veo 3.1 is the most literal prompt-follower of the frontier models โ€” which means structure pays off more here than anywhere else. Every template below has been run for real. Nothing theoretical.

The 6-part formula

[SUBJECT] + [ACTION] + [SETTING] + [CAMERA] + [LIGHT/STYLE] + [AUDIO]

The last slot is the one everyone forgets โ€” Veo generates audio natively, and unspecified audio means it improvises (usually with unwanted music):

A weathered fisherman in his 60s hauls a net over the gunwale of a
wooden boat, pre-dawn fog on a calm sea, slow push-in from a low angle,
cold blue hour light with warm lantern accents, cinematic 35mm look.
Audio: creaking wood, rope strain, distant gulls, no music.

Dialogue syntax (Veo's superpower)

Write the exact line, attribute it, and Veo delivers it lip-synced:

A young chai vendor grins at the camera and says: "Best cup in Mumbai โ€”
I guarantee it." Busy street corner at dusk, handheld medium close-up,
neon signage bokeh. Audio: street ambience, sizzling pan, his voice
clear and warm, no music.

Rules that keep dialogue clean:

  • One or two lines max per 8-second clip (~15 words each)
  • Describe the voice ("gravelly," "warm," "hurried whisper") โ€” it's part of the audio direction
  • For two speakers, format as a mini-script with names โ€” and expect to cherry-pick takes

JSON-style structured prompts

For complex shots, structure beats prose โ€” not because Veo parses JSON specially, but because it forces completeness and makes iteration surgical (change one key, re-roll):

{
  "subject": "a lone astronaut in a weathered white suit",
  "action": "plants a small flag, steps back, salutes",
  "setting": "red desert plateau, twin moons visible",
  "camera": "slow orbit, ending in a wide hero shot",
  "style": "epic sci-fi film, anamorphic lens flares, dusty atmosphere",
  "audio": "muffled breathing inside helmet, wind against fabric, no music"
}

12 copy-paste templates

1. Talking-head UGC ad โ€” [PERSON] holds [PRODUCT] toward the camera and says: "[LINE]", bright bathroom lighting, vertical phone framing, casual selfie energy. Audio: their voice, room tone, no music.

2. Cinematic establishing โ€” Aerial over [LOCATION] at golden hour, slow forward drift, volumetric haze, epic film look. Audio: wind, distant city hum, subtle low strings.

3. Street interview โ€” A reporter asks [QUESTION]; [CHARACTER] laughs and replies: "[LINE]", busy sidewalk, over-the-shoulder two-shot, documentary handheld. Audio: traffic, crowd walla, both voices natural.

4. Product hero โ€” [PRODUCT] rotates on a dark pedestal, light sweep reveals details, macro to wide pull-back, premium commercial grade. Audio: soft whoosh on the light sweep, deep ambient tone, no melody.

5. POV vlog โ€” First-person POV walking through [PLACE], head turns toward points of interest, natural pace, phone-camera realism. Audio: footsteps, ambient chatter, occasional voice-over: "[LINE]".

6. Nature documentary โ€” [ANIMAL] [ACTION] in [HABITAT], telephoto compression, shallow focus, BBC-style grade. Audio: habitat ambience, one narrator line in a hushed British tone: "[LINE]".

7. Cooking insert โ€” Overhead shot: hands [ACTION] in a rustic kitchen, steam rising, warm side light. Audio: sizzling, chopping rhythm, no music.

8. Horror beat โ€” A hallway light flickers; [CHARACTER] turns slowly toward a sound off-screen, static wide shot, cold green-tinged grade. Audio: hum, floor creak, single sharp knock, silence.

9. Sports action โ€” [ATHLETE] [ACTION] in slow motion, stadium backdrop, dramatic rim light, speed ramp to real-time at impact. Audio: crowd roar swelling, impact sound, announcer shouting: "[LINE]".

10. Historical scene โ€” [ERA] street scene: [ACTION], period-accurate costumes and props, natural window light, restrained filmic grade. Audio: era-appropriate ambience, no modern sounds, no music.

11. Explainer b-roll โ€” Clean macro shots of [SUBJECT], slow deliberate camera moves, bright studio light, minimal background. Audio: room tone only.

12. Emotional close โ€” Close-up on [CHARACTER]'s face as [EMOTION] breaks through, eyes glisten, background melts to bokeh, dusk light. Audio: breath, distant traffic, one soft piano note at the end.

Credit-saving rules

  • โŒ Vague audio = wasted takes. "No music" is a valid, powerful instruction.
  • โŒ Over-stacked action. One action beat per clip; chain clips with Flow's scene extension instead.
  • โŒ Re-rolling blind. Change exactly one variable per retry (that's why the JSON format earns its keep).
  • โœ… Prototype cheap, finish here. Draft concepts on cheaper credit models, then spend Veo credits only on shots that need its realism and native audio.

Correcting the "120 seconds" claim going around

A lot of 2026 coverage repeats that Veo 3.1 generates "120-second" or "148-second" videos, full stop โ€” and that's misleading. Native single-generation length is still capped at 4, 6, or 8 seconds per clip, chosen at generation time. The longer numbers come from the extend workflow: each new generation continues from the last frame of the previous one, and chaining enough of those extensions can reach roughly 148 seconds of continuous narrative. That's a real capability worth using โ€” it's just stitched, not a single 120-second render. Budget your prompts accordingly: write for an 8-second beat, then write the next beat as a continuation, not as one long script for a single generation.

More on the model itself in our Veo tool hub, and how it stacks against OpenAI in Veo 3.1 vs Sora 2.

Prefer video? Hand-picked walkthroughs

Reading is faster, but if you want to see it done, these are the best tutorials we vetted for this topic:

โ–ถ Master The Ultimate Google Veo 3.1 Prompt Formula (Full Tutorial)
โ–ถ How To Use Google Veo 3 Like A PRO: JSON Prompt Guide

Frequently asked questions

โ–ธWhat is the best prompt structure for Veo 3.1?

Six parts: subject โ†’ action โ†’ setting โ†’ camera โ†’ lighting/style โ†’ audio. Veo rewards explicit audio direction โ€” dialogue in quotes, ambient sound described โ€” because it generates sound natively.

โ–ธHow do I make characters speak in Veo?

Write the exact line in quotes and attribute it: The barista smiles and says: 'First one's on the house.' Veo generates the voice with lip-sync. Keep lines under ~15 words per 8-second clip.

โ–ธDo JSON prompts really work better in Veo?

Structured prompts (JSON-style key-value blocks) don't unlock hidden features, but they force you to specify every element โ€” which reliably improves results and makes A/B iteration much easier. Use them for complex shots.

โ–ธWhy do my Veo videos have random background music?

If you don't specify audio, Veo invents it. End every prompt with explicit audio direction โ€” including 'no music' if you plan to add your own soundtrack in the edit.

The 5 best AI video finds, every week

New models, tested prompts, and what actually worked in our production โ€” one short email a week. No spam, unsubscribe anytime.

Jordan Reyes

Written by Jordan Reyes

AI Video Producer

Runs multiple faceless YouTube channels and tests every major AI video model against the same prompts before recommending one. Tracks render time and credit cost like other people track calories.

Explore these topics

Every guide, comparison and prompt library we have on each.

#veo 3 prompts#veo 3.1 prompts#veo prompt guide#google veo prompt formula#veo 3 prompt examples
Next in Google VeoVeo 3.1 vs Sora 2 (2026): Which AI Video Model Should You Use?

Keep learning