How to Blind-Test Local TTS Without Missing Words
A reproducible local TTS test for script completion, skipped words, hallucinations, naturalness, speed, and memory, without relying on polished demos.
Direct answer: test script completion before asking which local TTS voice sounds best. Use the same short, medium, and long passages on one Mac, hide the model names during listening, align every output against the source text, and record omissions, substitutions, repetitions, added speech, speed, and peak memory. A model should not win because of an expressive ten-second clip if it drops a date, changes a price, repeats a clause, or invents a sentence. The protocol below gives audiobook producers, video editors, course creators, accessibility teams, and developers a reproducible way to run that comparison without publishing invented benchmark scores.
Quick Answer
- Score exact script completion before asking listeners which voice sounds best.
- Use blind labels such as Model A and Model B so brand reputation cannot steer votes.
- Report omissions, substitutions, repetitions, extra speech, punctuation behavior, time to first audio, total generation time, and peak memory.
- Keep model version, runtime, seed, hardware, reference clip, and generation settings public.
- Do not publish the winner until the pending result table is replaced by measured values and the matching audio files are attached.
In This Local TTS Benchmark
- Why demo quality hides production failures
- The models and hardware disclosure
- The fixed completion and naturalness tests
- How transcription errors are scored
- The blind listening panel
- Charts and audio required before publication
- How to run the same test on your Mac
- Frequently asked questions
Why Script Completion Comes First
TTS comparisons often begin and end with a listener preference vote. That is useful for naturalness, but it misses failures that become expensive in production. A model can sound warm and confident while saying only 72 of 80 source words. A missing adjective may be harmless. A missing dosage, decimal, surname, product tier, or negation changes the meaning. Long-form projects make the problem worse because a producer may not notice the omission until after editing, captioning, or client review.
Recent community benchmark discussions have asked testers to expose model versions, licenses, parameters, release dates, skipped-text behavior, and full audio rather than a highlight reel. One recurring complaint is an output that sounds more natural yet speaks only part of the passage. That evidence shapes this protocol, but measured claims in the final article must come from Murmur's own run. The benchmark separates three questions: Did the model say the requested text? Did it sound natural? How costly was the generation?
Models Included and Why
| Model | Official size | Core reason to test | Primary source | Result status |
|---|---|---|---|---|
| Pocket TTS | 100M | CPU-first streaming and cloning | Kyutai repository | Pending measured run |
| Kokoro | 82M | Small preset-voice narration baseline | Hexgrad repository | Pending measured run |
| Chatterbox Multilingual V3 | 500M | Zero-shot multilingual cloning and reduced hallucination claim | Resemble AI repository | Pending measured run |
| MOSS-TTS-v1.5 | 8B | Long-form cloning, punctuation prosody, explicit pauses | OpenMOSS repository | Pending measured run |
| Qwen3-TTS Base | 0.6B | Multilingual reference-audio cloning | Qwen model card | Pending measured run |
The list spans small CPU models, compact narration models, and larger cloning systems. It is not a claim that every checkpoint runs through the same backend. The disclosure table must identify the exact runtime used for each model and whether it is an upstream checkpoint or an MLX conversion. Official facts should link to the Pocket TTS repository, Chatterbox repository, MOSS-TTS repository, and each model card.
Test Hardware and Controls
Run every model on one Apple Silicon Mac after a restart, with the same power mode and no unrelated foreground apps. Record the chip, CPU and GPU core counts, unified memory, macOS version, free storage, runtime version, model commit, checkpoint precision, and whether compilation or model loading is included in timing. Cold-start and warm-start numbers belong in separate columns. Peak memory should come from the same measurement method for every run.
- Generate three seeded runs per model and passage, then report the median timing.
- Use one consented 20-second reference recording for cloning models and a documented default voice for preset-only models.
- Normalize published web samples to the same integrated loudness without changing tempo, pitch, pauses, or noise.
- Export lossless WAV masters. Derive MP3 or AAC web copies only after scoring.
- Keep failures. A crash, empty file, truncated file, or endless continuation is a result, not a clip to hide.
The Fixed Audio Test
The primary 82-word passage mixes punctuation, a contraction, a hyphenated phrase, a decimal, a currency value, a date, initials, and a question. These details make completion measurable without turning the script into a tongue twister: At 7:45 a.m. on Tuesday, Dr. Rivera checked Room 3B and said, “The revised total is $49.95, not $94.95.” Maya paused, reread the note, and asked, “Should we ship twenty-one units today, or wait until July 18?” After a brief call, the team approved the smaller order. Nothing dramatic happened; every label, number, and name simply had to be spoken exactly once, in the right order.
A second 300-word passage tests paragraph boundaries and a third 2,000-word chapter tests drift, repetition, and truncation. The short passage is the only sample used for the first blind vote because listeners can compare five clips without fatigue. The longer outputs are scored by transcript alignment and spot-checked at the beginning, midpoint, and end. All generated clips must appear beside the source transcript.
How We Score Omissions and Hallucinations
Automatic speech recognition provides a first alignment pass, not the final judgment. A human reviewer compares the waveform and transcript against the source. Each missing source word counts as an omission. A changed word counts as a substitution. Repeated source text counts separately from new, invented speech. Numbers are scored by meaning, so “forty-nine dollars and ninety-five cents” matches “$49.95,” while “ninety-four ninety-five” does not.
| Metric | Definition | Better direction | Publication rule |
|---|---|---|---|
| Completion rate | Correct source words divided by source words | Higher | Show numerator and denominator |
| Critical-token accuracy | Correct names, dates, money, quantities, and negations | Higher | List every failed token |
| Extra speech | Words not present in the source | Lower | Publish the appended transcript |
| Repeated words | Unrequested repeated source words | Lower | Count every recurrence |
| Naturalness MOS | Blind 1 to 5 listener rating | Higher | Report panel size and interval |
| Warm generation speed | Audio seconds divided by compute seconds | Higher | Median of three runs |
| Peak memory | Highest observed resident or unified memory use | Lower | Use one measurement method |
Results Dashboard: Publication Gate
The table below is intentionally a gate, not a simulated result. Replace every pending cell after the audio run, then reveal the blind model labels. If an article generator fills these cells with plausible numbers, delete them. A benchmark becomes useful because a reader can audit the script, clips, settings, and arithmetic. Brand familiarity is not evidence.
| Blind label | Completion | Critical tokens | Extra words | Naturalness | Warm speed | Peak memory |
|---|---|---|---|---|---|---|
| Model A | Pending | Pending | Pending | Pending | Pending | Pending |
| Model B | Pending | Pending | Pending | Pending | Pending | Pending |
| Model C | Pending | Pending | Pending | Pending | Pending | Pending |
| Model D | Pending | Pending | Pending | Pending | Pending | Pending |
| Model E | Pending | Pending | Pending | Pending | Pending | Pending |
Blind Listening Without Brand Bias
Randomize model labels for each listener and hide filenames, waveforms, model sizes, and generation times. Ask for a 1 to 5 naturalness score, a 1 to 5 speaker-consistency score, and a forced choice for the clip the listener would use in a finished video. Give listeners the written source after the preference vote, then ask them to flag missing or added content. This order prevents transcript hunting from replacing the first impression of the voice.
Report how many people completed the panel, their listening equipment, whether they are native speakers of the tested language, and an uncertainty interval. Ten votes can reveal an obvious artifact, but they do not establish a universal ranking. Publish individual vote distributions as well as averages. A 4.2 average built from polarized 1s and 5s means something different from consistent 4s.
How to Reproduce the Benchmark
- Copy the fixed scripts without changing punctuation or typography.
- Record a clean, consented reference voice and keep it unchanged across cloning models.
- Generate three runs with logged seeds, versions, settings, and timestamps.
- Align each output transcript to the source and manually verify every mismatch.
- Loudness-normalize derived listening copies while preserving the masters.
- Randomize filenames and collect blind ratings before revealing models.
- Publish all clips, settings, failures, result data, and chart calculations.
If you are choosing a model before the result run is complete, start with the broader local TTS model guide, the Kokoro Mac setup guide, and the local voice-cloning app comparison. For a visual workflow instead of separate command-line environments, review the models and voices available in Murmur.
Frequently Asked Questions
Conclusion
A useful local TTS benchmark makes script completion visible. Run the complete protocol before naming a winner, publish every fixed-script sample, and show the transcript errors beside the naturalness vote. The method is useful before results exist because it tells buyers which evidence to demand and gives reviewers a repeatable way to produce that evidence. It also makes future updates comparable: rerun the same scripts after a model, runtime, or operating-system change, then report whether accuracy, speed, and memory improved or regressed.
Compare Local Voices in One Mac App
Download Murmur to try supported local TTS models on Apple Silicon, keep projects on your Mac, and export production-ready audio. Buy Murmur for $49 as a one-time purchase.
macOS 15+ · Apple Silicon required · 7-day refund policy