SubtitleLens

Happy Scribe Review: We Measured It Against Ground Truth

A measured Happy Scribe review: 0.37% word error rate on clean and degraded audio, where punctuation breaks down, and why the Chinese subtitles need reformatting.


The short answer

Happy Scribe’s English transcription is excellent: we measured 0.37% word error rate, one wrong word in 267, and degrading the audio to laptop-mic quality did not change that number at all.

What did degrade was punctuation. And the Chinese translation, while linguistically competent, produced subtitles that are not usable without reformatting: 71% of the cues exceed the standard line length for Chinese subtitles.

So: a strong transcription engine, a usable translator, and an output that still needs a human pass before it goes on screen.

Disclosure: SubtitleLens earns a commission if you subscribe through our Happy Scribe link. We paid nothing for this test (it ran inside the free tier), and the findings below are measurements, not impressions.


How we tested

Most tool reviews describe accuracy with adjectives. We wanted a number, which means you need text that is known to be correct before the tool sees it.

Our source was a LibriVox recording of “The Adventure of Wisteria Lodge” (Conan Doyle, public domain), because the exact text is on Project Gutenberg. That gives a ground truth to score against rather than a judgement call.

From it we made two clips:

ClipWhat it is
A (clean)2 minutes of the original studio-quality reading
B (degraded)The same two minutes, processed to laptop-mic conditions

Clip B was band-limited to 250–3800 Hz, given room reverb, and mixed with pink noise to a 12 dB signal-to-noise ratio, measured rather than eyeballed: the noise floor rises from −56 dB to −22 dB. Same words, same voice, same reading. Only the audio quality changes, so any difference in the output is attributable to audio quality and nothing else.

Both clips went through Happy Scribe’s AI transcription in English. Clip A’s transcript was then translated to Simplified Chinese with Happy Scribe’s own translation feature and reviewed by a native speaker.

Scoring was done by aligning the transcript to the reference text word by word and counting substitutions, deletions and insertions. British/American spellings (recognise/recognize) and number formats (five/5) are treated as equivalent, because those are localisation choices rather than recognition errors. Counting them, as our first pass did, overstated the error rate by 15 percentage points. Front matter read aloud by the narrator (book preface, chapter announcement) is excluded; only the narrative is scored.

What this test does not cover: the human-review tier (that costs real money, see pricing below), speaker diarisation on multi-person audio, languages other than English and Chinese, and long-file behaviour. Treat the numbers as what they are: one voice, clear diction, two minutes.


English accuracy: better than the audio deserved

Clip A (clean)Clip B (12 dB SNR)
Word error rate0.37%0.37%
Substitutions11
Deletions00
Insertions00

Both clips produced exactly one error across 267 reference words, and it was the same error in both: “a mischievous twinkle in his eyes” transcribed as “his eye”.

Update (2026-07-22): We later ran the same clip through Maestra for our Happy Scribe vs Maestra comparison, and it produced the identical “eye” reading. Two independent engines making the same single deviation strongly suggests the narrator actually read “eye”: this is the source audio, not a transcription error. Happy Scribe’s real-world error rate on this clip is likely 0.00%. We’re leaving the original number above and noting the correction here rather than quietly editing it.

This is the headline finding, and it is more interesting than a simple “it’s accurate”. Degrading the audio to the quality of a cheap laptop microphone in a noisy room (a 34 dB rise in noise floor) did not cost a single additional word. Whatever acoustic model is running underneath is robust well past the point where the recording sounds bad to a human.

The practical implication: if your audio is merely bad rather than genuinely broken, you probably do not need to re-record before transcribing.


Where degradation actually showed up: punctuation

Word accuracy held. Punctuation did not.

Clip A (clean)Clip B (degraded)
Quotation marks104
Em dashes10
Question marks46
Subtitle cues3336
Words per cue9.58.7

The degraded clip lost 60% of its quotation marks. In a text that is largely dialogue between two people, that is not cosmetic: quotation marks are how a reader tracks who is speaking.

The clearest example is a single line. Watson is offering definitions of the word “grotesque”:

Clean: "Strange—remarkable," I suggested.

Degraded: Strange? Remarkable? I suggested.

Same words, different meaning. In the clean version Watson proposes two synonyms. In the degraded version he appears to be asking questions. The two spurious question marks in clip B are both from this line. The model lost the prosodic cues that distinguish a statement from a question, and guessed wrong.

The degraded clip also fragmented more (36 cues instead of 33, fewer words each), which for subtitles means more cuts mid-thought.

The takeaway: noisy audio does not necessarily corrupt what was said, but it does corrupt the structure: punctuation, sentence boundaries, dialogue attribution. For a transcript you will read, that is a minor annoyance. For subtitles that must be parsed at reading speed, it is the difference between a usable file and an editing job.


Chinese translation: competent language, unusable subtitles

This is where a review written only by English speakers stops being useful. We build translation software and read Chinese natively, so we scored the translation the same way we scored the transcript: specifically.

What it got right

Proper nouns were handled correctly and consistently: 华生 (Watson), 福尔摩斯 (Holmes), 埃克尔斯 (Eccles), 查林十字 (Charing Cross), 紫藤小屋 (Wisteria Lodge). These are the established Chinese renderings, not transliterations invented on the spot, and every name is identical on each reappearance. Name drift across a long file is a classic machine-translation failure and it did not happen here.

Where it slipped

A missed context switch. Holmes receives a telegram and “scribbled a reply”. The translation renders the reply as 回, a letter. In a scene about telegrams (a reply-paid telegram is discussed a few lines later) the correct word is 回. The word “telegram” appears in the same subtitle cue, so the context was available and still not used.

A register error. “A mischievous twinkle in his eyes” became 眼中闪过一丝顽皮的光芒. 顽皮 in Chinese is used almost exclusively for children being naughty. Holmes here is a grown man about to tease Watson about his prose; the idiomatic rendering is 狡黠 (sly) or 促狭. On top of that, 光芒 means radiance, far too strong for a twinkle.

A tone shift that contradicts the next sentence. “Scribbled a reply” became 随手草草写了回信, “casually, carelessly dashed off a reply”. But the very next line says Holmes “made no remark, but the matter remained in his thoughts”. The English “scribbled” is about speed; the Chinese adds a dismissiveness that the passage then contradicts.

One genuine comprehension failure. The English reads:

“…those narratives with which you have afflicted a long-suffering public”

Holmes is teasing Watson that his stories torment his readers. The translation:

那些曾让你长期受苦的读者们饱受折磨的故事

The modifier structure has collapsed. It reads as something like “those stories that tormented the readers who long suffered you”: the relationship between Watson, the stories and the public is scrambled, and 受苦/饱受折磨 says “suffer” twice. This is the one place where a reader of the Chinese would not be able to recover the meaning.

The formatting problem is bigger than the translation problem

Chinese subtitle convention is roughly 15–20 full-width characters per line, with reading speed comfortable at 4–6 characters per second. The output:

Result
Cues over 20 characters15 of 21 (71%)
Longest single cue48 characters
Peak reading speed9.0 characters/second

The translator merged the 33 English cues into 21 Chinese ones and made no attempt to re-break them for the target language. Chinese is far more compact than English (one Chinese line often carries two English lines’ worth of meaning), so a good translation should produce shorter, not longer, cues. Instead the timings were inherited unchanged.

There is a typography bug too: in one cue an opening quotation mark 「“」 is left stranded at the end of a line, separated from the sentence it opens. Chinese line-breaking rules forbid this.

The result is a file that reads fine as a transcript and fails as subtitles. Anyone shipping Chinese subtitles from this output will be re-breaking most of the lines by hand.


Pricing

Verified on the pricing page in July 2026:

PlanAnnual (per month)MonthlyIncluded
Freen/an/a10 minutes, 45 min per recording
Basic$8.50$17120 min/month
Pro$19$29600 min/month
Business$59$896,000 min/month

Extra AI minutes are $0.20/min on any paid tier. Human-made transcription starts at $2.00/min ($1.90 on Business), covering 65+ languages with delivery from 4 hours.

We did not test the human tier: at $2.00/min that is a real purchase, and we would rather tell you it is untested than review something we did not run. What we can say is that the price sets the decision cleanly: human review costs about ten times the AI rate, so it is worth it exactly when a wrong word costs you more than ten minutes of your own proofreading time.

The free tier’s 10 minutes is enough to run this entire test, which is a fair way to evaluate before paying.


Who should use it

Good fit if you need English transcription that is accurate out of the box, if your source audio is imperfect (the degradation test suggests it copes), or if you want transcription and subtitle export in one place. The export formats (SRT, VTT, TXT, DOCX) cover the practical bases, and DOCX in particular makes it easy to hand a transcript to an editor who does not want to learn a new interface.

Think twice if your deliverable is subtitles in a CJK language. The translation is good enough to work from, but the line-breaking is not, and you will be reformatting. If that is your workflow, budget the editing time or plan to re-break the cues. Our free SRT tools handle the mechanical parts, and the subtitle shifter covers timing adjustments after the fact.

Not tested here: multi-speaker diarisation, the human-review tier, and languages beyond the English–Chinese pair we ran. If those are central to your decision, the free 10 minutes will tell you more than any review will.

For alternatives and how they compare on price and language coverage, see our AI subtitle translators roundup. If you already have a subtitle file and just need it in another language, our free SRT translator does that in your browser at no cost.

Try Happy Scribe →


Affiliate disclosure: links to Happy Scribe on this page are affiliate links and we earn a commission on subscriptions bought through them, at no extra cost to you. Every number above was measured by us; nothing on this page was supplied or reviewed by the vendor.