Benchmark

How to Blind-Test Local TTS Without Missing Words

A reproducible local TTS test for script completion, skipped words, hallucinations, naturalness, speed, and memory, without relying on polished demos.

·7 min read

Direct answer: test script completion before asking which local TTS voice sounds best. Use the same short, medium, and long passages on one Mac, hide the model names during listening, align every output against the source text, and record omissions, substitutions, repetitions, added speech, speed, and peak memory. A model should not win because of an expressive ten-second clip if it drops a date, changes a price, repeats a clause, or invents a sentence. The protocol below gives audiobook producers, video editors, course creators, accessibility teams, and developers a reproducible way to run that comparison without publishing invented benchmark scores.

Quick Answer

  • Score exact script completion before asking listeners which voice sounds best.
  • Use blind labels such as Model A and Model B so brand reputation cannot steer votes.
  • Report omissions, substitutions, repetitions, extra speech, punctuation behavior, time to first audio, total generation time, and peak memory.
  • Keep model version, runtime, seed, hardware, reference clip, and generation settings public.
  • Do not publish the winner until the pending result table is replaced by measured values and the matching audio files are attached.

In This Local TTS Benchmark

  1. Why demo quality hides production failures
  2. The models and hardware disclosure
  3. The fixed completion and naturalness tests
  4. How transcription errors are scored
  5. The blind listening panel
  6. Charts and audio required before publication
  7. How to run the same test on your Mac
  8. Frequently asked questions

Why Script Completion Comes First

TTS comparisons often begin and end with a listener preference vote. That is useful for naturalness, but it misses failures that become expensive in production. A model can sound warm and confident while saying only 72 of 80 source words. A missing adjective may be harmless. A missing dosage, decimal, surname, product tier, or negation changes the meaning. Long-form projects make the problem worse because a producer may not notice the omission until after editing, captioning, or client review.

Recent community benchmark discussions have asked testers to expose model versions, licenses, parameters, release dates, skipped-text behavior, and full audio rather than a highlight reel. One recurring complaint is an output that sounds more natural yet speaks only part of the passage. That evidence shapes this protocol, but measured claims in the final article must come from Murmur's own run. The benchmark separates three questions: Did the model say the requested text? Did it sound natural? How costly was the generation?

Models Included and Why

ModelOfficial sizeCore reason to testPrimary sourceResult status
Pocket TTS100MCPU-first streaming and cloningKyutai repositoryPending measured run
Kokoro82MSmall preset-voice narration baselineHexgrad repositoryPending measured run
Chatterbox Multilingual V3500MZero-shot multilingual cloning and reduced hallucination claimResemble AI repositoryPending measured run
MOSS-TTS-v1.58BLong-form cloning, punctuation prosody, explicit pausesOpenMOSS repositoryPending measured run
Qwen3-TTS Base0.6BMultilingual reference-audio cloningQwen model cardPending measured run

The list spans small CPU models, compact narration models, and larger cloning systems. It is not a claim that every checkpoint runs through the same backend. The disclosure table must identify the exact runtime used for each model and whether it is an upstream checkpoint or an MLX conversion. Official facts should link to the Pocket TTS repository, Chatterbox repository, MOSS-TTS repository, and each model card.

Test Hardware and Controls

Run every model on one Apple Silicon Mac after a restart, with the same power mode and no unrelated foreground apps. Record the chip, CPU and GPU core counts, unified memory, macOS version, free storage, runtime version, model commit, checkpoint precision, and whether compilation or model loading is included in timing. Cold-start and warm-start numbers belong in separate columns. Peak memory should come from the same measurement method for every run.

  • Generate three seeded runs per model and passage, then report the median timing.
  • Use one consented 20-second reference recording for cloning models and a documented default voice for preset-only models.
  • Normalize published web samples to the same integrated loudness without changing tempo, pitch, pauses, or noise.
  • Export lossless WAV masters. Derive MP3 or AAC web copies only after scoring.
  • Keep failures. A crash, empty file, truncated file, or endless continuation is a result, not a clip to hide.

The Fixed Audio Test

The primary 82-word passage mixes punctuation, a contraction, a hyphenated phrase, a decimal, a currency value, a date, initials, and a question. These details make completion measurable without turning the script into a tongue twister: At 7:45 a.m. on Tuesday, Dr. Rivera checked Room 3B and said, “The revised total is $49.95, not $94.95.” Maya paused, reread the note, and asked, “Should we ship twenty-one units today, or wait until July 18?” After a brief call, the team approved the smaller order. Nothing dramatic happened; every label, number, and name simply had to be spoken exactly once, in the right order.

A second 300-word passage tests paragraph boundaries and a third 2,000-word chapter tests drift, repetition, and truncation. The short passage is the only sample used for the first blind vote because listeners can compare five clips without fatigue. The longer outputs are scored by transcript alignment and spot-checked at the beginning, midpoint, and end. All generated clips must appear beside the source transcript.

How We Score Omissions and Hallucinations

Automatic speech recognition provides a first alignment pass, not the final judgment. A human reviewer compares the waveform and transcript against the source. Each missing source word counts as an omission. A changed word counts as a substitution. Repeated source text counts separately from new, invented speech. Numbers are scored by meaning, so “forty-nine dollars and ninety-five cents” matches “$49.95,” while “ninety-four ninety-five” does not.

MetricDefinitionBetter directionPublication rule
Completion rateCorrect source words divided by source wordsHigherShow numerator and denominator
Critical-token accuracyCorrect names, dates, money, quantities, and negationsHigherList every failed token
Extra speechWords not present in the sourceLowerPublish the appended transcript
Repeated wordsUnrequested repeated source wordsLowerCount every recurrence
Naturalness MOSBlind 1 to 5 listener ratingHigherReport panel size and interval
Warm generation speedAudio seconds divided by compute secondsHigherMedian of three runs
Peak memoryHighest observed resident or unified memory useLowerUse one measurement method

Results Dashboard: Publication Gate

The table below is intentionally a gate, not a simulated result. Replace every pending cell after the audio run, then reveal the blind model labels. If an article generator fills these cells with plausible numbers, delete them. A benchmark becomes useful because a reader can audit the script, clips, settings, and arithmetic. Brand familiarity is not evidence.

Blind labelCompletionCritical tokensExtra wordsNaturalnessWarm speedPeak memory
Model APendingPendingPendingPendingPendingPending
Model BPendingPendingPendingPendingPendingPending
Model CPendingPendingPendingPendingPendingPending
Model DPendingPendingPendingPendingPendingPending
Model EPendingPendingPendingPendingPendingPending
Five ordered local TTS checks from script completion to peak memory
Score accuracy first, then blind naturalness, generation speed, and memory on the same machine.

Blind Listening Without Brand Bias

Randomize model labels for each listener and hide filenames, waveforms, model sizes, and generation times. Ask for a 1 to 5 naturalness score, a 1 to 5 speaker-consistency score, and a forced choice for the clip the listener would use in a finished video. Give listeners the written source after the preference vote, then ask them to flag missing or added content. This order prevents transcript hunting from replacing the first impression of the voice.

Report how many people completed the panel, their listening equipment, whether they are native speakers of the tested language, and an uncertainty interval. Ten votes can reveal an obvious artifact, but they do not establish a universal ranking. Publish individual vote distributions as well as averages. A 4.2 average built from polarized 1s and 5s means something different from consistent 4s.

How to Reproduce the Benchmark

  1. Copy the fixed scripts without changing punctuation or typography.
  2. Record a clean, consented reference voice and keep it unchanged across cloning models.
  3. Generate three runs with logged seeds, versions, settings, and timestamps.
  4. Align each output transcript to the source and manually verify every mismatch.
  5. Loudness-normalize derived listening copies while preserving the masters.
  6. Randomize filenames and collect blind ratings before revealing models.
  7. Publish all clips, settings, failures, result data, and chart calculations.

If you are choosing a model before the result run is complete, start with the broader local TTS model guide, the Kokoro Mac setup guide, and the local voice-cloning app comparison. For a visual workflow instead of separate command-line environments, review the models and voices available in Murmur.

Frequently Asked Questions

Conclusion

A useful local TTS benchmark makes script completion visible. Run the complete protocol before naming a winner, publish every fixed-script sample, and show the transcript errors beside the naturalness vote. The method is useful before results exist because it tells buyers which evidence to demand and gives reviewers a repeatable way to produce that evidence. It also makes future updates comparable: rerun the same scripts after a model, runtime, or operating-system change, then report whether accuracy, speed, and memory improved or regressed.

Compare Local Voices in One Mac App

Download Murmur to try supported local TTS models on Apple Silicon, keep projects on your Mac, and export production-ready audio. Buy Murmur for $49 as a one-time purchase.

macOS 15+ · Apple Silicon required · 7-day refund policy