Tools/Web engineering
ElevenLabs: speech synthesis as an API
A review of the ElevenLabs audio API: model lineup, latency figures, credit and per-character pricing, tier-gated formats, and where OpenAI and Amazon Polly win.
- Type
- Voice and audio API
- Pricing
- Free tier · from $5 per month
Balázs Csorba··10 min read
- Text to speech
- Voice API
- Speech to text
- Voice cloning

Key takeaways
- The API charges $0.04 per 1,000 characters for Flash and v4 Turbo and $0.08 for v3 and Multilingual v2, roughly three to five times the commodity band where OpenAI, Google and xAI sit.
- Every latency figure in the model table is median inference time, marked in the docs as excluding application and network latency.
- Useful output formats are gated by plan: MP3 at 192 kbps needs Creator and PCM or WAV at 44.1 kHz needs the $99 Pro plan.
- The free plan gives 10,000 credits a month without a commercial licence, and concurrency starts at two simultaneous requests on Multilingual v2.
- Nothing is self-hostable, so the cloned voice and the prosody a product depends on are rented rather than owned.
ElevenLabs is a hosted audio API: text to speech, speech to text, dubbing, sound effects, music and a voice-agent pipeline behind a single key. The output is still the reason teams choose it, and the position taken here is deliberately narrow: buy it for speech quality and language coverage, not for its pricing architecture, which is built so that a per-character rate never quite lines up with a competitor's.
It occupies the slot a cloud text-to-speech service used to occupy, an HTTP call inside a content pipeline or a websocket inside a voice agent, and it competes with OpenAI's audio models, Amazon Polly and Google's Chirp on price. All three are cheaper per minute of finished speech. None of them attaches professional voice cloning to an expressive model line, which is what the product actually is. Teams building a realtime voice loop care about that difference more than about the model card.
What it is
ElevenLabs was founded in 2022 by Piotr Dąbkowski and Mati Staniszewski, runs from London with offices in New York, Warsaw and San Francisco, and the Wikipedia entry puts headcount at about 400 in 2026. The company sells three surfaces: ElevenCreative in the browser, ElevenAPI for metered use, and ElevenAgents for voice agents. Every model behind them is proprietary. There are no open weights and no self-hosted deployment, only regional endpoints.
- Proprietary models behind a REST API: nothing to deploy, no weights to download, no on-premises option, and regional hosts only.
- Text to speech spans eleven_v4, eleven_v4_turbo, eleven_v3, eleven_multilingual_v2 and eleven_flash_v2_5, with 90+ languages on the v4 family and 29 on Multilingual v2.
- Speech to text uses Scribe v2, Scribe v2 Realtime at about 150 ms and Scribe v2 Medical, with diarisation for up to 32 speakers and 65 entity types.
- Free plan gives 10,000 credits a month without a commercial licence; commercial use starts with Starter at $6.
- Six public plans: $0, $6, $22, $99, $299 and $990 a month, plus custom Enterprise terms with SSO, DPA and SLA.
- Regional endpoints for api.elevenlabs.io alongside US, EU, India and Singapore residency hosts.
- Concurrency is per plan: 2 simultaneous Multilingual v2 requests on Free, 5 on Creator, 10 on Pro, 15 on Business.
How it works
A synthesis call is one POST to /v1/text-to-speech/{voice_id} with the output format in the query string and a JSON body of text, model_id and optional voice settings. The response is the audio itself, streamed as bytes rather than wrapped in JSON, and the official SDKs expose the same shape. For realtime work there is a websocket endpoint where only the time the model spends generating counts against the concurrency limit. Determinism is best effort: a seed repeats a take, and previous_text or previous_request_ids keep prosody continuous across chunks.
Format choice is where the tier structure becomes visible. The default is mp3_44100_128; MP3 at 192 kbps needs Creator, and PCM or WAV at 44.1 kHz needs Pro, which is the tier most speech pipelines actually want. Concurrency is enforced through a queue rather than a rejection, and the docs say crossing the limit typically adds about 50 ms. The response headers current-concurrent-requests and maximum-concurrent-requests are the only telemetry provided for it.
Models and latency
The model page publishes latency figures with a footnote that matters: every one of them is marked as excluding application and network latency. Read them as inference time in isolation, which is useful for comparing Flash against Multilingual v2 and useless as a user-facing promise until the socket path has been measured on the region you will call.
| Model | Published latency | Languages | Characters per request |
|---|---|---|---|
| eleven_v4 | No figure published | 90+ | 10,000 |
| eleven_v4_turbo | About 100 ms median inference | 90+ | 10,000 |
| eleven_flash_v2_5 | About 75 ms | 32 | 40,000 |
| eleven_multilingual_v2 | No figure published | 29 | 10,000 |
| eleven_v3 | 280 ms in the Conversational variant | 70+ | 5,000 |
One behavioural detail decides model choice more often than quality does. Flash v2.5 leaves text normalisation off by default to protect latency, so phone numbers, dates and currency come out awkwardly, and switching it on with apply_text_normalization is an Enterprise-only option for the v2.5 models. Multilingual v2 normalises numbers properly, which is why the docs recommend having the LLM normalise the text before synthesis rather than paying for it inside the TTS call.
Getting started
The Python client installs with pip install elevenlabs and the key lives in the dashboard under Developers. The snippet below synthesises one clip with the default output format; swapping model_id to eleven_flash_v2_5 buys latency and a 40,000 character limit, and swapping it to eleven_v4 buys the newest model.
# pip install elevenlabs
from elevenlabs import ElevenLabs
client = ElevenLabs(api_key="xi-api-key")
audio = client.text_to_speech.convert(
voice_id="JBFqnCBsd6RMkjVDRZzb",
output_format="mp3_44100_128",
text="The first move is what sets everything in motion.",
model_id="eleven_multilingual_v2",
)
with open("clip.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)Two knobs change behaviour without changing price: the voice settings (stability, similarity boost, style, speed) and the language code. For long text, split into chunks and pass previous_request_ids so prosody continues across requests. For agents, stream over the websocket instead of buffering a complete MP3, because buffering puts the whole generation time in front of the first audible byte.
Pricing
There are two meters and they do not convert into each other. The Creative plans bill in credits, a single shared pool that every product draws from: text to speech costs about a credit per character, speech to text 330 credits a minute, music 900 a minute and dubbing 2,000 a minute. The ElevenAPI side is billed in US dollars instead of credits, which is the only version that can be lined up against another vendor's rate card.
- Free: $0, 10,000 credits a month, no commercial licence.
- Starter: $6, 30,000 credits, commercial licence, instant voice cloning.
- Creator: $22, 121,000 credits, professional voice cloning, 192 kbps audio.
- Pro: $99, 600,000 credits, 44.1 kHz PCM and WAV through the API.
- Scale: $299, 1,800,000 credits, 3 workspace seats.
- Business: $990, 6,000,000 credits, 10 seats, low-latency text to speech from 5 cents a minute.
- Annual billing costs ten months for twelve, which puts Starter at $5 a month and Pro at $82.50.
On the API, text to speech is $0.04 per 1,000 characters for Flash and v4 Turbo and $0.08 for v3 and Multilingual v2, with a promotional $0.022 for Eleven v4 until 12 October 2026. Speech to text is $0.22 an hour for Scribe v2 and $0.39 for Scribe v2 Realtime, the speech engine for agents is $0.08 a minute with burst pricing at $0.16, music is $0.15 a minute, and dubbing runs from $0.33 a minute with a watermark to $2.20 for Dubbing v2.
Credits roll over for up to two months while a paid subscription stays active, so the balance can reach three times the monthly quota, and pay-as-you-go top-ups sit outside that cap. Unused credits are lost on downgrade or cancellation, which turns the annual plan from a discount into a bet on usage.
Where it shingles
The weaknesses are structural rather than cosmetic. Nothing is self-hostable, so availability, retention and price are all vendor decisions, and zero retention exists only as a parameter on Enterprise. Per-request character limits cap long-form work at 10,000 characters on Eleven v4 and 40,000 on Flash, so a book is a loop of stitched requests. The credit pool makes cost forecasting harder than a plain character rate, useful formats sit behind plan gates rather than feature gates, and the free tier does not allow commercial use at all.
| Attribute | ElevenLabs | OpenAI TTS | Amazon Polly |
|---|---|---|---|
| Price per 1,000 characters | $0.04 for Flash and v4 Turbo, $0.08 for v3 and Multilingual v2 | $0.015 for tts-1 and $0.030 for tts-1-hd; gpt-4o-mini-tts bills audio tokens | $0.004 Standard, $0.016 Neural, $0.030 Generative |
| Billing unit | Characters on the API, one shared credit pool on the Creative plans | Characters for tts-1, audio tokens for gpt-4o-mini-tts | Characters, charged monthly |
| Free tier | 10,000 credits a month, commercial licence starts at Starter | None listed on the pricing page | 5 million Standard characters a month for the first 12 months |
| Where it fits | Cloned and expressive voices, 90+ languages, character work | Utility speech beside an existing OpenAI stack | Bulk narration inside AWS, with cached replay at no extra cost |
The lock-in is the voice itself. A cloned voice, the pronunciation dictionaries and the prosody decisions behind a product line have no export format, so an exit plan is really a re-recording plan. Treat the audio as the artefact and the API as a supplier that can be replaced only by re-synthesis, and keep the source text and the seed for every take worth reproducing.
Verdict
Adopt it when the voice is part of the product. ElevenLabs is the right choice for audiobooks, character work, multilingual narration and any interface where the listener can hear the difference; it is the wrong choice for bulk utility speech, where Polly at $0.016 per thousand characters does the same job. The pricing is the weak surface: legible only after normalising, gating formats rather than features, and rewarding whoever reads the FAQ before the architecture.
- Choose it when voice quality is audible to the user: narration, characters, marketing material, multilingual content.
- Take Flash or v4 Turbo for agents and keep the request on a websocket; treat 75 ms and 100 ms as inference figures, not as round-trip promises.
- Budget at Creator rather than Starter if cloned voices or 192 kbps audio matter, and at Pro if the pipeline needs 44.1 kHz PCM.
- Do not choose it for pure utility speech, for data residency without Enterprise terms, or for anything that must run inside your own VPC.
- Model the bill in characters per finished minute of audio rather than in credits, and check that unit against a competitor before committing.
The footnote attached to every latency figure in the model documentation reads excluding application and network latency. An architecture decision that quotes 75 ms without measuring the socket path is quoting a number that has already left out the part that usually dominates.
Sources
Frequently asked questions
How much does ElevenLabs cost per 1,000 characters?
On the API, $0.04 for Flash and v4 Turbo and $0.08 for v3 and Multilingual v2, with a promotional $0.022 for Eleven v4 until 12 October 2026. The Creative plans bill in credits instead, from $6 for 30,000 credits on Starter to $990 for 6 million on Business, and every product draws from one shared pool.
Which ElevenLabs model has the lowest latency?
Eleven Flash v2.5 publishes about 75 ms and Eleven v4 Turbo about 100 ms, both median inference latency excluding application and network latency. Flash v2.5 also accepts 40,000 characters per request, the largest limit in the lineup, against 10,000 for Eleven v4.
Can ElevenLabs output be used commercially on the free plan?
No. The commercial licence starts with the Starter plan at $6 a month, which also unlocks instant voice cloning. Professional voice cloning appears at Creator ($22) and 44.1 kHz PCM output through the API at Pro ($99).
Does ElevenLabs run inside my own cloud account?
No. It is a hosted API with regional endpoints in the US, EU, India and Singapore, and zero retention is an Enterprise option. There is no self-hosted or open-weight deployment, so leaving means re-synthesising or re-recording the audio.