Skip to content
Guide

How to Choose a Kokoro Voice for Narration

Compare Kokoro presets with an original audition script and downloadable scorecard. Check pronunciation, pacing, and listening comfort before producing a long project.

Murmur6 min read

Choose a Kokoro voice by testing a short passage from your actual project with several compatible presets. Keep the script, speed, and playback volume consistent, then compare missing words, pronunciation, pauses, and listening comfort. A pleasant opening sentence is useful, but it does not tell you whether that voice can carry your whole lesson or chapter. This guide provides an original audition worksheet; it does not claim that one preset won a listening study.

Already generating speech? Start with the worksheet below. If you still need an installation path, use the Kokoro setup guide for Mac. Keeping setup and voice selection separate makes it easier to tell whether a problem comes from the runtime, the text, or the voice choice.

Start with a compatible shortlist

The upstream Kokoro voice catalogue groups presets by language and accent. Examples include af_heart and af_bella for American English, bf_emma for British English, ff_siwis for French, and jf_alpha for Japanese. Check the voice IDs exposed by your installed runtime; an application can offer a subset or use friendlier display names. The official catalogue also describes uneven support across languages. Do not assume that every preset handles every language equally well.

Shortlist three candidates that fit the audience and the language of the script. Write down why each is included: a familiar accent, a less forceful delivery, or clearer distinctions between speakers. Those are project requirements, not universal rankings. Keep a fourth option only if it tests a genuinely different direction. Auditioning dozens of voices without a clear brief makes the decision harder to explain and repeat.

Kokoro uses preset voices rather than cloning a speaker from a recording. If a client needs their own voice identity, changing Kokoro presets will not solve that requirement. Review voice cloning, design, and conversion before spending time tuning the wrong kind of system.

Use this original audition script

Read the following passage with each candidate. It is an original test fixture written for this guide, not an audio benchmark. First use the text unchanged. Then replace the example company, deadline, and amount with terms from your own project, without adding private information you cannot share with reviewers. Save the original and revised versions separately so a later comparison uses the same words.

Before you press Export, check the final paragraph. The Northbridge project has three parts: an introduction, a demonstration, and a short recap. We will meet on October ninth at nine thirty. The estimate is forty-nine dollars, and the first delivery contains twelve files. Is the preview ready? Almost. Listen once more, leave a short pause, and make sure the last word is complete.

Original Murmur audition fixture

This passage checks a question, a short answer, a list, and a sentence ending. It deliberately spells out numbers, which helps separate delivery problems from number interpretation. Add a second passage containing the raw numbers and abbreviations your project actually uses. If only the second version fails, investigate text normalization before rejecting an otherwise useful voice.

Also include a quiet transition, a difficult name, and one paragraph from the longest section you plan to produce. The synthetic fixture cannot represent every genre. A software tutorial, a first-person memoir, and a dialogue scene place different demands on the listener. Keep your real excerpt short enough that you can listen to every candidate completely.

Score errors before preferences

CheckWhat to recordDecision
CompletenessMissing, repeated, or cut-off wordsResolve before accepting the take
PronunciationExact word and expected readingTry a documented correction, then retest
PausesSentence or list boundary that feels wrongRevise punctuation or chunking separately
ComfortListener preference from 1 to 5Use only after content checks pass
Repair effortMinutes spent fixing the excerptCompare against the project deadline

Download the free CSV audition scorecard and open it in Numbers, Excel, or another spreadsheet. Each row leaves results blank on purpose. Enter the model version, runtime, voice ID, speed, and script version yourself. A blank cell means unmeasured, not zero errors. Keep the audio files alongside the sheet so another person can check what led to your choice.

Use one row per take. If you change punctuation after hearing a problem, record a new take rather than overwriting the old score. That separates a voice improvement from a script repair. Track critical pronunciation failures as counts, then use notes for the exact location. A single wrong product name may matter more than a lower average preference score, especially when the finished audio will explain a product to new customers.

Make the listening comparison fair

  • Generate every candidate from the same saved script and initial speed setting. Record any runtime-specific controls instead of assuming that identical numbers mean identical behavior across applications.
  • Listen at a comfortable, similar perceived volume. Do not pick the louder file automatically. If you adjust gain for listening, keep the original export as well.
  • Rename the review copies A, B, and C. Keep the name-to-voice mapping in your own notes until reviewers have written their observations.
  • Ask reviewers to mark concrete errors before discussing which voice they prefer. Then reveal the preset names and agree on the project requirements.

These steps reduce avoidable bias; they are not a controlled research study. For a solo project, take a break before the second listen and change the playback order. For client work, ask the person who will approve the final narration to review the shortlist early. Record that approval with the exact sample rather than relying on a vague description such as warm or professional.

Check a longer passage before committing

After choosing a candidate, generate one representative section using your intended production workflow. Listen to the beginning, transitions, and ending, then listen all the way through. Shortlist success does not establish long-form reliability. Keep the same voice and configuration while fixing text issues so you do not introduce a second variable halfway through the test.

If the voice rushes, clips an ending, or reads a short label awkwardly, change one thing at a time. Try clearer punctuation or different segment boundaries before large speed changes. Save a fresh take after each adjustment. For a systematic repair process, see the missing-words and punctuation guide. Stop if repairs take longer than recording or using another voice would.

Murmur provides public voice samples and a voice library overview so you can hear examples before buying. Its bundled Kokoro path supports basic narration; optional cloning and Voice Design models have additional download and memory requirements. Speech generation runs locally, while activation, periodic license validation, updates, and downloads still use the network. The audition sheet works with other Kokoro runtimes too.

Questions about Kokoro voice selection

Sources

Make a short narration project in Murmur

Murmur is a local voice studio for Apple Silicon Macs running macOS 15 or later. The website edition costs $49 one-time, with no free trial and a 7-day refund policy. Listen to the public samples before buying.

macOS 15+ · Apple Silicon required · 7-day refund policy