SPIRITT logoSPIRITTElevenLabsEleven v4

Eleven v4. Give every line a performance.

Whisper an aside. Land the punchline. Take one voice across 90+ languages. Meet ElevenLabs’ expressive speech model, plus the Turbo variant built for live conversations.

Sound on. Hear the difference.

Official demos, creator experiments and independent results. Some videos pair Eleven v4 audio with separate avatar, video or robot tools; those visuals are not generated by the speech model.

A voice that follows your direction.

For ads, stories, characters and conversations that need more than words read aloud.

One launch. Two ways to give AI a voice.

ElevenLabs launched Eleven v4 and Eleven v4 Turbo on September 28, 2026. v4 focuses on produced speech; Turbo targets real-time conversations.

Direct delivery with inline tags, carry a speaker’s identity across scenes, and create multi-speaker dialogue. ElevenLabs reports stronger emotional control and support for 90+ languages.

Eleven v4 leads Artificial Analysis’s Provider Voice Arena at 1,320 Elo in the October 2 snapshot. It ranks second in the separate Controlled Voice comparison: a useful reminder to test the voice you will actually use.

Turbo’s roughly 100 ms inference figure is not end-to-end response time. ElevenLabs’ separate speech-start test reports 150 ms, excluding network latency.

The direct API launch offer is $22 per million characters for v4 and $11 for Turbo until October 12. These promotional rates are separate from SPIRITT usage and ElevenLabs app subscriptions.

Expressive speech90+ languagesMulti-speaker dialogueInline directionTurbo for real-time useReleased September 28

Provider Voice Arena · Elo

Winner
Eleven v4
1320
Qwen-Audio-3.1-TTS-Plus
1291
Sonic 3.6
1273
Gemini 3.8 Flash TTS
1272
Qwen-Audio-3.0-TTS-Plus
1259
Eleven v3
1171

Artificial Analysis, October 2, 2026. All accents and categories. Eleven v4: 1,803 samples, eight provider voices, 95% interval 1,302–1,338. Blind listener preference for model + voice, not accuracy. Intervals can overlap; the ranking can change.

The voice quality. The speed. The trade-offs.

Compare the voice you hear, the time you wait and the cost of production. Independent evaluations and vendor-run tests are labeled separately.

Independent voice-matched comparison

Controlled Voice Arena · Elo

Qwen-Audio-3.1-TTS-Plus
1184
Eleven v4
1156
Realtime TTS-2
1145
Sonic 3.6
1138
Qwen-Audio-3.0-TTS-Plus
1126
Eleven v3
1071

Artificial Analysis, October 2. Shared voice references, separate from Provider Voice. Eleven v4 is #2: 1,606 samples, 95% interval 1,141–1,171. Qwen leads this comparison. Selected models shown.

Independent pronunciation test
Winner

Pronunciation Robustness

Eleven v4
91.7%
Gemini 3.8 Flash TTS
89.5%
Gemini 3.1 Flash TTS
88.2%
Eleven v3
85.6%

Artificial Analysis’s September 28 launch snapshot. A separate pronunciation test, not an arena preference rating. These dated results are not presented as the live leaderboard.

Vendor-run blind listening test

Eleven v4 expressiveness win rate

v4 vs Cartesia Sonic 3.6
81%
v4 vs Inworld TTS-2
81%
v4 vs Gemini 3.8 Flash-Lite TTS
72%
v4 vs Gemini 3.8 Flash TTS
65%

ElevenLabs, September 2026. Each bar is Eleven v4’s share against the named opponent, not that opponent’s score. Same-line blind pairs; ties count as half. Sample counts and confidence intervals were not published.

Real-time speech · vendor test
Winner

Turbo time to first speech · ms

Eleven v4 Turbo
150
Cartesia Sonic 3.6
262
xAI TTS
362
Gemini 3.8 Flash-Lite TTS
685
GPT-4o mini TTS
814

ElevenLabs, September 2026. Median request-to-audible-speech time, network latency removed; identical scripts and default settings. Turbo uses WebSocket streaming. This is not its ~100 ms inference metric or full application latency. Lower is better.

Independent generation throughput

Generation speed · characters/sec

Eleven v4
73.4
Eleven v3
42.5

Artificial Analysis’s September 28 launch snapshot. Characters processed per second of generation time. Throughput is not speech playback speed, time to first audio, or a guarantee for your workload.

Published language coverage

Language coverage · lower bounds

Eleven v4 family · 90+
90
Eleven v3 · 70+
70
Flash v2.5
32
Multilingual v2
29

ElevenLabs model documentation. Bars for 90+ and 70+ plot the stated lower bounds, not exact counts. Language support does not mean equal pronunciation, accents or quality in every language.

Published input limits

Characters per speech request

Eleven v4
10000
Eleven v3
5000
Flash v2.5
40000
Multilingual v2
10000

ElevenLabs model overview. Request limits are not continuous-duration guarantees. The Text to Dialogue endpoint separately recommends at most 2,000 total text characters for reliable generation. Longer productions need chunking and review.

Temporary direct API offer

Launch API price · USD / 1M chars

Eleven v4
$22.00
Eleven v4 Turbo
$11.00

ElevenLabs’ displayed $0.022 / $0.011 per 1,000 characters, scaled to one million. Offer until October 12, 2026; cutoff timezone unspecified. Taxes excluded. Not app-plan credits or SPIRITT pricing.

Independent listed API prices

AA price snapshot · USD / 1M chars

Eleven v4
$80.00
Sonic 3.6
$49.00
Qwen-Audio-3.1-TTS-Plus
$19.31
Gemini 3.8 Flash TTS
$16.49

Artificial Analysis, October 2. Listed API prices, not ElevenLabs’ temporary launch offer above. ElevenLabs shows $80 as v4’s reference rate; it is not a guaranteed future quote. Check provider terms before budgeting. Lower is better.

Sources checked October 2, 2026: Artificial Analysis Provider Voice and Controlled Voice leaderboards; its September 28 launch post; ElevenLabs’ launch article, model documentation, Text to Dialogue reference and API pricing. Dated snapshots can differ from current rankings. Elo, percentages, characters/sec and milliseconds measure different things. Voice quality and cost depend on the voice, language, settings and workload.

From script to finished audio.

Eleven v4 performs the voice. SPIRITT runs the production around it.

01

Bring your script

Open a SPIRITT Workspace with your script, audience and delivery goal. Start with a product ad, a lesson, a story or a character conversation.

Glass cloud-workspace illustration for organizing scripts and production assets
02

Direct the performance

Ask for Eleven v4 narration. Specify the voice, language, tone and pacing, then review a short take before producing the full script. Use only voices you have permission to use.

Glass microphone, audio waveform and direction controls
03

Make it a production workflow

Have SPIRITT organize takes, prepare language versions and connect your approved publishing destinations. Turn one recording into a repeatable workflow that keeps up when your content changes.

Glass globe surrounded by audio tracks for multilingual production

Give your next idea a voice.

A product story. A lesson. A character. Bring the script and the direction; let SPIRITT turn it into finished audio and a workflow you can keep using.

Questions

Eleven v4, explained.

01What is Eleven 4?+
Eleven v4 is ElevenLabs’ speech-generation model, released September 28, 2026 alongside Eleven v4 Turbo. It turns written material and performance direction into spoken audio.
02What is the difference between v4 and v4 Turbo?+
Choose v4 for produced narration and dialogue. Turbo is designed for low-latency interactive speech. Neither replaces the reasoning, transcription or turn-taking system around a complete voice agent.
03How should I read the benchmark charts?+
Provider Voice compares models using their provider voices; Controlled Voice uses shared reference voices. Elo measures listener preference, not percentage accuracy. The vendor-run charts and dated launch snapshots are labeled separately.
04Is Turbo really a 100 ms voice agent?+
No: roughly 100 ms is the reported median inference time, excluding application and network latency. A complete voice agent also needs listening, reasoning, turn-taking and audio delivery. The separate vendor speech-start result is 150 ms.
05How much does Eleven v4 cost?+
The direct API offer shown on October 2 is $0.022 per 1,000 characters for v4 and $0.011 for Turbo until October 12, 2026, excluding taxes. Displayed reference rates are $0.08 and $0.04. App credits, app promotions and SPIRITT usage are separate; check current terms before budgeting.
06Can I clone a voice?+
ElevenLabs supports voice cloning, but you need the speaker’s verified permission. Do not use a cloned voice to mislead people about who is speaking. Review the provider’s current verification and voice-compatibility requirements before production.
07Can it produce a whole audiobook in one request?+
No single-request promise is implied. The model overview lists 10,000 characters for v4, while the dialogue endpoint recommends at most 2,000 total text characters for reliable generation. Split longer work into reviewable sections and check continuity.
08Does it generate the avatars and videos in these demos?+
No. Eleven v4 generates speech. Creators may combine it with separate video, avatar, animation or robotics tools. Music generation and transcription are also separate capabilities, not the v4 model.
09Where can I check the original sources?+
Use the official launch, API pricing and independent leaderboard links on this page. The native X posts retain creator attribution and link back to each original demo or evaluation.
10How do I use Eleven v4 with SPIRITT?+
Open a workspace and ask for a concrete audio outcome with Eleven v4. SPIRITT can prepare the script, generate narration, organize the files and build the surrounding production workflow. Eleven v4 is an audio tool, not a chat-model-picker choice. Generation consumes usage; the vendor API promotion is not a SPIRITT pricing promise.
Buy from builders who use what they sellBuilt usingSPIRITT