ElevenLabs Dubbing Review: Great Voices, Watch the Meter
We measured ElevenLabs dubbing to the millisecond and read the Chinese output. Timing alignment is excellent, the translation underneath is not, and the pricing is the real story.
The short answer
ElevenLabs dubs 18 seconds of English into Chinese and lands the speech segments within 2 milliseconds of the original speaker’s timing. That part is genuinely impressive, and we have the measurements.
Two things temper it. The translation underneath the voice is ordinary machine translation, and in our test it produced one sentence that a Chinese listener cannot parse. And the pricing is in a different universe from subtitling: dubbing runs $2.23–$4.91 per minute against roughly $0.20/minute for AI transcription. The free tier buys you about 22 seconds.
If you need a voice track and your budget can absorb the per-minute cost, the voice quality is there. If you are choosing between dubbing and subtitles for reach, this review is mostly an argument for subtitles.
Disclosure: every measurement below was taken, written and published before we joined the ElevenLabs affiliate programme, which we did on 2026-07-31. Links to ElevenLabs and Happy Scribe now earn us a commission. Not one word of the findings changed when that happened, including the parts that argue you probably shouldn’t buy.
What dubbing actually costs
This is the finding that should drive your decision, so it goes first.
We uploaded a 2-minute clip and the interface quoted 54,000 credits. The free plan includes 10,000 credits per month. The job was 5.4× our entire monthly allowance before we had dubbed a single second.
Trimming to 18 seconds brought it to 8,100 credits, which pins the rate precisely at 27,000 credits per minute. Against a 10,000-credit free plan, that is 22 seconds of dubbing per month.
ElevenLabs’ own plan table agrees, and it is worth reading carefully:
| Dubbing v2 | Minutes included | Extra minutes |
|---|---|---|
| Free ($0) | 0.4 | $4.91/min |
| Starter ($6) | 2 | $2.70/min |
| Creator ($11) | 9 | $2.45/min |
| Pro ($99) | 44 | $2.23/min |
Nine minutes of dubbing on the $11 plan. A single 10-minute YouTube video does not fit.
There is a much cheaper path that the interface does not push you toward: legacy Dubbing v1 with a watermark gives the free plan 5 minutes instead of 0.4, and charges $0.36/minute overage on Creator instead of $2.45, roughly seven times cheaper. Dubbing v2 is the default and is currently in alpha. If cost matters more than being on the newest model, v1 deserves a look.
One thing ElevenLabs does right: the cost is displayed before you confirm. Nothing gets spent by surprise. That sounds like a low bar until you try a competitor: Maestra’s free tier let us upload, configure and click through before revealing we had zero usable credits.
Note also that free-tier dubs are watermarked automatically on v2, with no option to trade a watermark for a discount. That option exists only on the legacy v1 flow.
How we tested
We used the same source as our other tool tests: a LibriVox recording of Conan Doyle’s “The Adventure of Wisteria Lodge”, whose exact text is on Project Gutenberg, so the English is known-correct before any tool sees it.
The 18-second excerpt was chosen deliberately: it contains a sentence we have already put through two other translation engines, which lets us compare three systems on identical input.
We dubbed English → Chinese (Simplified) using Dubbing v2 Alpha, default settings. Timing was measured with ffprobe and ffmpeg’s silence detection; the Chinese was read by a native speaker.
What we did not test: other language pairs, multi-speaker content, lip-sync, Dubbing Studio’s manual editing, and the legacy v1 model. One clip, one language pair, one voice.
Timing: measured, and genuinely good
Total duration came back at 18.001 seconds against an original of 18.000, a one-millisecond difference.
But matching total length is trivial; anyone can pad with silence. The real question is whether each spoken segment starts when the original speaker started. We detected silence gaps in both files and compared the onsets:
| Speech segment | Original | Dubbed | Drift |
|---|---|---|---|
| 1 | 0.931s | 0.675s | −256 ms |
| 2 | 5.354s | 5.240s | −114 ms |
| 3 | 9.777s | 9.776s | −2 ms |
| 4 | 13.156s | 13.142s | −14 ms |
| 5 | 16.517s | 16.172s | −344 ms |
Two segments land within 15 milliseconds. That is not duration-filling. The model is genuinely re-anchoring each utterance to the original speaker’s rhythm, which is what makes a dub sit against picture instead of drifting away from it.
The cost of that alignment
Holding the segments in place has to come from somewhere, and it comes out of the pauses:
| Original | Dubbed | Change | |
|---|---|---|---|
| Speech | 12.56s | 14.35s | +14.3% |
| Silence | 5.44s | 3.65s | −33.0% |
The Chinese takes 14% longer to say than the English, and that overflow eats a third of the original silence. Our clip is literary narration with 5.4 seconds of breathing room in 18, so it absorbed the overflow comfortably, which matches the listening impression, that the result sounds unhurried and natural.
The practical warning is what happens when that headroom is absent. Fast-paced dialogue, tutorial voiceover, anything cut tight with minimal pauses gives the model nothing to absorb expansion into. Before committing a whole video, dub 30 seconds of your densest passage, not a comfortable one.
The translation underneath is the weak layer
Automatic dubbing is three systems stacked: transcribe, translate, synthesise. The voice is the part you notice; the translation is the part that decides whether the dub means anything.
Ours had problems at three levels of severity.
One sentence broke outright
The English “He shook his head at my definition” came back as 他听了玩的定义摇了摇头, “he shook his head at the definition of play”. The character for “my” (我, wǒ) has become “play” (玩, wán).
A Chinese listener hits this and simply loses the thread; there is no recovering the intended meaning from context. This is the error-stacking failure we described in our YouTube auto-translate article: an upstream mistake propagating into the translation and then getting spoken aloud in a confident, natural-sounding voice. The better the synthesis, the more convincingly the error is delivered.
Register and punctuation slipped
- “I suggested” became 我提议到, using the wrong homophone: the speech-attribution suffix is 道, not 到. Visible instantly in subtitles; inaudible in speech.
- “Strange—remarkable,” where Watson is offering two definitions, became 奇特?非同寻常?, turning statements into questions. We saw the identical failure in our Happy Scribe review when we degraded the audio: em-dashes and prosody are where these systems lose the difference between proposing and asking.
- The attribution “said he” drifted across a sentence boundary, attaching to Holmes’s next question rather than closing his previous line.
- “I suppose, Watson, …” became 我想华生 without the comma that marks Watson as being addressed, which reads as “I think about Watson”.
What it got right
Credit where due: man of letters → 文人墨客 is a genuinely idiomatic choice, not a literal gloss. The name 华生 is the standard Chinese rendering of Watson and stayed consistent. And the delivery itself (pacing, intonation, the sense of a person speaking rather than a machine reciting) is convincing.
Three engines, one identical mistake
Here is the finding we did not expect.
“A mischievous twinkle in his eye” came back as 眼里闪烁着顽皮的光芒. 顽皮 in Chinese describes a child being naughty. Holmes is a grown man about to tease Watson about his prose; the register calls for 狡黠 (sly) or 促狭. And 光芒 means radiance, far too bright for a twinkle.
We have now run this exact sentence through three unrelated systems:
| Engine | Output |
|---|---|
| Happy Scribe | 顽皮的光芒 |
| Google Translate (via our free SRT translator) | 顽皮的光芒 |
| ElevenLabs Dubbing v2 | 顽皮的光芒 |
Three independent pipelines, character-for-character identical, and identically wrong.
That rules out “this vendor’s translation is weaker.” Register, the difference between playful-child and sly-adult, is a structural blind spot of current machine translation, not a bug in any one product. Switching tools will not fix it. Only a human reader of the target language will catch it, which is exactly why we read our own dubs rather than trusting them.
Who this is for
Worth it if you need a voice track specifically: the voice quality and timing alignment are real, and no subtitle workflow substitutes for audio when your audience is driving, cooking, or visually impaired. Budget for the per-minute rate as a production cost, and have someone who reads the target language check the script before you publish. Try it on the free tier first: 22 seconds is enough to hear whether the voice suits your material.
Not worth it if your goal is international reach per dollar. At $2.23–$4.91/minute, dubbing a 20-minute video costs $45–100 in credits. Subtitling the same video costs a few dollars of transcription plus translation, and our free SRT translator handles the translation at no cost. For most creators asking “how do I reach viewers in other languages”, subtitles remain the answer by an order of magnitude.
If you do dub: test your densest 30 seconds first, seriously consider legacy v1 for the ~7× cost saving, and never ship a dub in a language nobody on your team reads. A fluent, confident voice reading a broken sentence is worse than an obvious subtitle error, because nothing signals to the listener that anything went wrong.
For how the subtitle-first path compares on cost and quality, see our Happy Scribe vs Maestra comparison and the AI subtitle translators roundup.
All measurements on this page are ours: credit costs observed in the product on 2026-07-31, timings from ffprobe/ffmpeg on the source and dubbed files, Chinese assessed by a native speaker. Plan figures verified on elevenlabs.io/pricing the same day. Nothing here was supplied or reviewed by the vendor. Links to ElevenLabs and Happy Scribe are affiliate links; we joined both programmes after testing, and the recommendations above are the ones we would give without them.