๐๏ธ ElevenLabs v3 Guide: Tags, Settings & Realistic AI Voiceovers
How to get genuinely human AI narration from ElevenLabs v3 โ audio tags, key settings, multi-speaker dialogue, and the workflow we use for every voiceover.
Jordan Reyes ยท AI Video Producer
ยท Updated ยท 4 min read
โก TL;DR โ quick answers
- What are ElevenLabs v3 audio tags?
- Inline bracketed directions โ like [whispers], [laughs], [sighs], [excited] โ that you place in your script to control emotional delivery. v3 interprets them as performance directions rather than reading them aloud.
- Why does my ElevenLabs voice sound flat?
- Usually three causes: stability set too high (lower it for more expressive range), a script with no emotional cues (add audio tags), or generating one giant block of text (generate per paragraph so the model resets its energy).
- Can ElevenLabs v3 do conversations between two speakers?
- Yes โ v3 supports multi-speaker dialogue generation with distinct voices and natural turn-taking, which is what makes AI podcast and character-dialogue content practical.

Voice is 50% of a video's perceived quality โ viewers forgive average visuals but click away from robotic narration in seconds. ElevenLabs v3 is what we use for every voiceover across our channels. This is the complete practical guide.
What makes v3 different
Older TTS read text. v3 performs it. The model understands emotional context and โ critically โ accepts audio tags: inline stage directions that shape delivery without being spoken.
[thoughtful] The strange thing is... nobody noticed for years.
[whispers] And when they finally did โ [beat] it was too late.
[excited] But here's where the story turns!
Tags we use constantly:
- Emotion:
[excited],[thoughtful],[sad],[nervous],[confident] - Delivery:
[whispers],[shouts],[slowly],[rushed] - Non-verbal:
[laughs],[sighs],[gasps],[clears throat] - Pacing:
[beat],[pause], ellipses...for natural hesitation
Full walkthrough of the tag system.
Settings that actually matter
- Stability โ the big one. High stability = consistent but flat; low stability = expressive but occasionally wild. For narration we run low-to-mid stability and regenerate the rare weird take. For character dialogue, go lower still.
- Similarity: keep high to stay true to the chosen voice.
- Model: use the v3 model for anything with emotional range. Older models are fine for utility audio only.
Rule of thumb: fix delivery in the script (tags), not in the sliders. Settings set the range; tags direct the performance.
The workflow we ship with
- Write for the ear โ short sentences, contractions, concrete words.
- Tag the script while writing, not after: mark the emotional beats as you feel them.
- Generate per paragraph, not one giant block. Energy stays fresh, retakes are cheap, and one bad take doesn't cost the whole render.
- Keep one voice per channel. Voice = brand for faceless content. Save the voice and settings; never drift.
- Master lightly: normalize loudness (target around -14 LUFS for YouTube), add a subtle music bed at -20dB or lower under narration.
Prompt-engineering deep dive โ how phrasing and punctuation shape delivery.
Multi-speaker dialogue
v3 handles multi-speaker generation with natural turn-taking โ the technology behind AI podcasts and character scenes. Assign distinct voices per speaker, keep lines short (real people interrupt and react), and use non-verbal tags ([laughs], [hmm]) for listener realism.
What v3 costs in mid-2026 (updated)
The pricing picture improved since this guide first ran. After ElevenLabs' $500M raise in February 2026, the company cut consumer plan pricing by roughly half โ the current lineup runs Free ($0), Starter ($6/mo, 30k credits), Creator ($22/mo, 121k credits) and Pro ($99/mo, 600k credits), with Scale and Business tiers above that for teams. On the API side, v3 and Multilingual v2 bill around $0.10 per 1,000 characters, with the Flash/Turbo models at half that for latency-sensitive work.
Practical read: the $6 Starter tier now covers a casual weekly channel's narration, and Creator at $22 comfortably runs a daily short-form pipeline. Prices shift โ verify against the current plans page before subscribing โ but the direction of travel has been firmly downward.
Cost control
Pricing is character-based, so waste = regeneration. Three habits keep our bills sane:
- Generate per-paragraph (retake only what failed)
- Proof the script before generating, aloud
- Draft new voices/styles on short test lines, not full scripts
For where the voiceover fits in the full production stack โ script โ voice โ visuals โ edit โ see our faceless YouTube channel guide.
Prefer video? Hand-picked walkthroughs
Reading is faster, but if you want to see it done, these are the best tutorials we vetted for this topic:
Frequently asked questions
โธWhat are ElevenLabs v3 audio tags?
Inline bracketed directions โ like [whispers], [laughs], [sighs], [excited] โ that you place in your script to control emotional delivery. v3 interprets them as performance directions rather than reading them aloud.
โธWhy does my ElevenLabs voice sound flat?
Usually three causes: stability set too high (lower it for more expressive range), a script with no emotional cues (add audio tags), or generating one giant block of text (generate per paragraph so the model resets its energy).
โธCan ElevenLabs v3 do conversations between two speakers?
Yes โ v3 supports multi-speaker dialogue generation with distinct voices and natural turn-taking, which is what makes AI podcast and character-dialogue content practical.
The 5 best AI video finds, every week
New models, tested prompts, and what actually worked in our production โ one short email a week. No spam, unsubscribe anytime.

Written by Jordan Reyes
AI Video Producer
Runs multiple faceless YouTube channels and tests every major AI video model against the same prompts before recommending one. Tracks render time and credit cost like other people track calories.
Explore these topics
Every guide, comparison and prompt library we have on each.





