Controlled Voice Arena · Elo
Artificial Analysis, October 2. Shared voice references, separate from Provider Voice. Eleven v4 is #2: 1,606 samples, 95% interval 1,141–1,171. Qwen leads this comparison. Selected models shown.
Eleven v4Whisper an aside. Land the punchline. Take one voice across 90+ languages. Meet ElevenLabs’ expressive speech model, plus the Turbo variant built for live conversations.
Official demos, creator experiments and independent results. Some videos pair Eleven v4 audio with separate avatar, video or robot tools; those visuals are not generated by the speech model.
Introducing Eleven v4 and Eleven v4 Turbo, our fastest and most emotive voice models yet.
— ElevenLabs (@ElevenLabs) September 28, 2026
Eleven v4 is INSANE! 🤯
— Burak Tuyan (@buraktuyan) October 2, 2026
Truly blown away by the results from Eleven v4.
— Justine Moore (@venturetwins) September 28, 2026
ElevenLabs’ Eleven v4 takes #1 on the Artificial Analysis Provider Voice TTS Arena Leaderboard and Pronunciation Robustness benchmark, and #2 on Controlled Voice, surpassing Cartesia’s Sonic 3.6 and Google’s Gemini 3.8 Flash TTS on Provider Voice
— Artificial Analysis (@ArtificialAnlys) September 28, 2026
Direct the performance of Eleven v4 with inline tags.
— ElevenLabs (@ElevenLabs) September 28, 2026
On the back of Eleven v4 being released, we paired @elevenlabs v4 with Cara-4 on @Anam__AI and got an avatar performing a dramatic speech from Shakespeare’s Julius Caesar. Director notes handles the face (you cue anger, sadness etc as the speech turns) while v4 handles the voice, so they build together.
— Ben Carr (@benatanam) September 30, 2026
The most emotive model meets the most expressive body
— Boris Starkov (@ktoya_me) September 29, 2026
النموذج الصوتي الجديد Eleven v4 من ElevenLabs ممتاز في الأداء الشِّعري..
— عبدالرحمن السيد | Easy AI (@abdulrahmanway) October 1, 2026
ElevenLabs Eleven v4の日本語音声を可能な限り自然な日本語に最適化してみました。 コメントで「海外出身の方っぽい声の響き」、「スタジオ録音っぽい」、「舞台役者っぽい」という声をたくさん頂いたのでその原因を1つずつ潰した結果です。
— 中崎工房 | AIで1時間の仕事を5分で終わらせる人 (@nakazakifam) October 2, 2026
Eleven v4 Turbo is now available in ElevenAgents, bringing the #1 rated quality of Eleven v4 to live conversations with a median inference latency of ~100 ms.
— ElevenLabs (@ElevenLabs) September 29, 2026
Eleven v4 turbo is insanely fast - all while generating super high quality speech.
— Thomas Yu (@thommyyu) September 30, 2026
Introducing ElevenReader, powered by @ElevenLabs v4.
— ElevenReader (@elevenreader) October 2, 2026
For ads, stories, characters and conversations that need more than words read aloud.
ElevenLabs launched Eleven v4 and Eleven v4 Turbo on September 28, 2026. v4 focuses on produced speech; Turbo targets real-time conversations.
Direct delivery with inline tags, carry a speaker’s identity across scenes, and create multi-speaker dialogue. ElevenLabs reports stronger emotional control and support for 90+ languages.
Eleven v4 leads Artificial Analysis’s Provider Voice Arena at 1,320 Elo in the October 2 snapshot. It ranks second in the separate Controlled Voice comparison: a useful reminder to test the voice you will actually use.
Turbo’s roughly 100 ms inference figure is not end-to-end response time. ElevenLabs’ separate speech-start test reports 150 ms, excluding network latency.
The direct API launch offer is $22 per million characters for v4 and $11 for Turbo until October 12. These promotional rates are separate from SPIRITT usage and ElevenLabs app subscriptions.
Artificial Analysis, October 2, 2026. All accents and categories. Eleven v4: 1,803 samples, eight provider voices, 95% interval 1,302–1,338. Blind listener preference for model + voice, not accuracy. Intervals can overlap; the ranking can change.
Compare the voice you hear, the time you wait and the cost of production. Independent evaluations and vendor-run tests are labeled separately.
Artificial Analysis, October 2. Shared voice references, separate from Provider Voice. Eleven v4 is #2: 1,606 samples, 95% interval 1,141–1,171. Qwen leads this comparison. Selected models shown.
Artificial Analysis’s September 28 launch snapshot. A separate pronunciation test, not an arena preference rating. These dated results are not presented as the live leaderboard.
ElevenLabs, September 2026. Each bar is Eleven v4’s share against the named opponent, not that opponent’s score. Same-line blind pairs; ties count as half. Sample counts and confidence intervals were not published.
ElevenLabs, September 2026. Median request-to-audible-speech time, network latency removed; identical scripts and default settings. Turbo uses WebSocket streaming. This is not its ~100 ms inference metric or full application latency. Lower is better.
Artificial Analysis’s September 28 launch snapshot. Characters processed per second of generation time. Throughput is not speech playback speed, time to first audio, or a guarantee for your workload.
ElevenLabs model documentation. Bars for 90+ and 70+ plot the stated lower bounds, not exact counts. Language support does not mean equal pronunciation, accents or quality in every language.
ElevenLabs model overview. Request limits are not continuous-duration guarantees. The Text to Dialogue endpoint separately recommends at most 2,000 total text characters for reliable generation. Longer productions need chunking and review.
ElevenLabs’ displayed $0.022 / $0.011 per 1,000 characters, scaled to one million. Offer until October 12, 2026; cutoff timezone unspecified. Taxes excluded. Not app-plan credits or SPIRITT pricing.
Artificial Analysis, October 2. Listed API prices, not ElevenLabs’ temporary launch offer above. ElevenLabs shows $80 as v4’s reference rate; it is not a guaranteed future quote. Check provider terms before budgeting. Lower is better.
Sources checked October 2, 2026: Artificial Analysis Provider Voice and Controlled Voice leaderboards; its September 28 launch post; ElevenLabs’ launch article, model documentation, Text to Dialogue reference and API pricing. Dated snapshots can differ from current rankings. Elo, percentages, characters/sec and milliseconds measure different things. Voice quality and cost depend on the voice, language, settings and workload.
Eleven v4 performs the voice. SPIRITT runs the production around it.
Open a SPIRITT Workspace with your script, audience and delivery goal. Start with a product ad, a lesson, a story or a character conversation.

Ask for Eleven v4 narration. Specify the voice, language, tone and pacing, then review a short take before producing the full script. Use only voices you have permission to use.

Have SPIRITT organize takes, prepare language versions and connect your approved publishing destinations. Turn one recording into a repeatable workflow that keeps up when your content changes.

A product story. A lesson. A character. Bring the script and the direction; let SPIRITT turn it into finished audio and a workflow you can keep using.