Comparison

Chatterbox Nano vs Turbo vs Multilingual V3: Which Model Should You Run?

A practical comparison of Chatterbox Nano, Turbo, Multilingual V3, and Single Language Packs by size, languages, compute, cloning, tags, and workflow fit.

·5 min read

Direct answer: choose Chatterbox Nano for an English CPU-first path with tight memory and latency budgets; choose Chatterbox Turbo for stronger low-latency English cloning and paralinguistic tags such as laughter; choose Chatterbox Multilingual V3 for cross-language cloning and 23-plus languages; and test a Single Language Pack when one supported language needs tighter regional quality control. Resemble AI lists Nano at 110M parameters, Turbo at 350M, and Multilingual V3 at 500M. Size and upstream claims do not replace a completion, pronunciation, and memory test on your own machine.

Quick decision table

ModelOfficial sizeLanguage scopeBest first use
Nano110MEnglishCPU and constrained on-device inference
Turbo350MEnglishLow-latency expressive cloning and agents
Multilingual V3500M23+Cross-language voice cloning and localization
Single Language Pack500M eachSix dedicated modelsLanguage-specific quality control
Original Chatterbox500MEnglishCreative CFG and exaggeration control

Nano is a deployment decision

The official Chatterbox table describes Nano as a 110M English model with the same broad architecture direction as Turbo, paralinguistic tags, and CPU performance of about three times real time on eight cores. That makes it interesting for assistants, games, previews, and accessibility tools where a GPU should remain available. The claim is an upstream reference point, not a Mac benchmark. Measure first audio, total real-time factor, and CPU use on the minimum device.

Nano's small size does not mean it is the correct audiobook narrator. Long-form stability, pronunciation, and acting range need separate tests. It may be ideal as a responsive draft voice while a larger model produces final lines.

Turbo is for expressive low-latency English

Turbo reduces the decoder path from ten steps to one according to Resemble AI and supports reaction tags such as [laugh]. It targets English agents and production where response time and expressive events matter. A zero-shot clone still depends heavily on clean reference audio and can reproduce room noise or performance artifacts.

Test tags in context. A laugh that sounds good alone may break timing in a final edit. Measure the spoken words around the event, the length of the inserted reaction, and whether repeated generations drift. Murmur includes Chatterbox in a broader local model decision guide rather than presenting tags as universal quality.

Multilingual V3 changes the evaluation

V3 is a 500M general-purpose multilingual model with 23-plus languages. Resemble AI says it improves speaker similarity, reduces hallucinations, and produces more natural multilingual speech compared with earlier versions. Verify all three claims on a fixed local evaluation. Cross-language cloning needs native-speaker review because a voice can remain recognizable while stress, names, and borrowed words sound wrong.

Do not assume every MLX conversion labeled Chatterbox is V3. Check the upstream commit or model card used for the conversion. Store that identity with results. For English and German methodology, see the bilingual voice-cloning scorecard.

Choose by production stage

A useful workflow can use more than one model. Nano can produce immediate previews. Turbo can generate an expressive English host. Multilingual V3 can preserve a speaker across localization. A language pack can handle a priority market. The production system should show which model generated each clip and should not silently substitute one family without updating saved metadata.

Compare the same 250-word passage, clone reference, sample rate, and export settings. Score exact completion, critical tokens, speaker similarity, native pronunciation, first audio, warm speed, memory, and license. Hear Murmur's current output on the samples page before downloading every checkpoint.

Evidence to collect before publication

Generate the same 250-word English passage with Nano and Turbo, then a bilingual passage with Multilingual V3. Use one consented reference, normalize playback level, hide model labels, and score completion before expression. Record first audio, real-time factor, peak memory, CPU and GPU use, reaction-tag timing, and repeated-run drift. Add the samples to the Murmur sample library only if each model and voice asset permits redistribution.

  • Same passage and consented reference
  • Completion and critical-token gate
  • Blind listening order
  • CPU, GPU, memory, and latency results
  • Exact upstream model revisions

Reader worksheet

Create a one-page worksheet before acting on this Chatterbox Nano vs Turbo vs Multilingual V3 guide. Write the job, current baseline, minimum supported Mac, privacy requirement, model and runtime revision, acceptable failure rate, and the one result that would change your decision. Add measured values only after the test, and keep upstream claims in a separate column. Save links to raw output and licenses beside the row they support. Review the worksheet after seven days of real use, not only after a successful demo. Browse the Murmur research library for the related memory, privacy, model, and production checks that belong in the same decision.

Practical checklist

  1. Pick English CPU, English expressive, multilingual, or language-specific intent.
  2. Verify the exact upstream version behind any conversion.
  3. Use the same consented reference and script across models.
  4. Score completion and pronunciation before expressiveness.
  5. Measure first audio, warm speed, peak memory, and CPU or GPU use.
  6. Save model family and version beside each clip.

Frequently asked questions

Sources

Turn the research into finished local audio

Murmur brings local models, reusable voices, projects, queueing, and export into one Mac production workspace. Hear the current models before choosing a workflow.

macOS 15+ · Apple Silicon required · 7-day refund policy