---
title: "Best Local TTS Models for Mac in 2026"
description: "Hear 12 TTS voice samples, compare the open-weight leaderboard, and check Mac setup and licenses for Kokoro, Qwen3, Breeze, Voxtral, Fish S2, and more."
canonical: "https://www.murmurtts.com/blog/best-local-tts-models-2026"
---
[Murmur](https://www.murmurtts.com/)/[Blog](https://www.murmurtts.com/blog)/Guide

Guide

# The Best Local TTS Models for Mac Creators in 2026

Listen to local TTS voice samples and compare Mac workflows, Artificial Analysis rankings, MLX support, and model licenses for 2026.

![](https://www.murmurtts.com/murmur-icon.webp)Murmur·April 11, 2026·14 min read

On this page

Quick VerdictIn This GuideListen: 12 TTS voice samplesOpen-weight TTS leaderboard: September 2026More leaderboard models to consider on a MacModels Currently Available in MurmurThe Best Model Changes With the JobKokoro: A Small Preset-Voice Starting PointQwen3-TTS: Cloning and Voice Design on MacChatterbox: Two Different RecommendationsFish Audio S2 Pro: Powerful, Large, and License-LimitedSparkTTS and OmniVoice: Specialized AlternativesWhat Apple Silicon ChangesA Five-Minute Model Selection ChecklistWhen Local TTS Is the Wrong ChoiceFrequently Asked Questions

[Back to the journal](https://www.murmurtts.com/blog)

Start with **Kokoro** for preset-voice narration on a Mac, or **Qwen3-TTS** for cloning and voice design. The wider open-weight field also includes Breeze TTS 2, Step Audio EditX, Voxtral TTS, NVIDIA Magpie, Maya1, and Higgs Audio V3. This guide compares their listening results with the setup and licensing choices a Mac creator faces. Choose for your script, target language, available memory, and whether you need a preset, cloned, or designed voice.

## Quick Verdict

- Choose Kokoro for quick drafts, steady narration, preset voices, and lower-memory Macs.
- Choose Qwen3-TTS Base for rapid reference-audio cloning across 10 supported languages.
- Choose Qwen3-TTS VoiceDesign when you want to describe a new voice without recording a reference clip.
- Choose Chatterbox Turbo for English cloning with native reaction tags such as [laugh] and [chuckle].
- Choose Chatterbox Multilingual for cross-language cloning across its documented 23-language set.
- Consider Fish Audio S2 Pro for fine-grained delivery control only after checking memory and commercial licensing.
- Treat SparkTTS and OmniVoice as specialized choices, not automatic recommendations for every production workflow.
- For more options beyond Murmur, compare the leaderboard below and the documented MLX paths for Breeze, Voxtral, Magpie, Maya1, and Higgs.

## In This Guide

1. Listen to 12 voice samples with model and source labels
2. The open-weight leaderboard and what its scores measure
3. Breeze, Step Audio EditX, Voxtral, Magpie, Maya1, and Higgs on Mac
4. Models currently available in Murmur
5. How to choose by project instead of hype
6. Kokoro, Qwen3-TTS, Chatterbox, Fish Audio, SparkTTS, and OmniVoice explained
7. Apple Silicon and model-download considerations
8. Commercial-use license checks
9. When local TTS is the wrong choice
10. Frequently asked questions

## Listen: 12 TTS voice samples

Listen for pronunciation, breath, pacing, and changes in emotion. The first three clips were recorded in Murmur. The remaining clips are published model-creator demos, including the OmniVoice recording already used on our language pages. External audio loads when you press Play and needs an internet connection.

**How to compare these recordings**

These clips use different scripts, speakers, recording settings, and runtimes. Creator demos are curated examples, not our Mac benchmarks or the audio used to calculate the leaderboard scores. Their sound does not establish local speed, memory use, or the quality of an MLX conversion.

### Recorded in Murmur

Kokoro · Heart · English narration0:00

**Kokoro full script:** Some ideas start as a scribble. A sentence in your notes. A story you cannot stop thinking about. Give yours a voice.

Qwen3-TTS · Ryan · English audiobook0:00

**Qwen3-TTS script excerpt:** The morning light crept through the curtains like a hesitant guest, unsure of its welcome. Margaret sat at the kitchen table…

Chatterbox Turbo · Comedy · English0:00

**Chatterbox Turbo script excerpt:** Okay so I tried that viral thirty day cold shower challenge and honestly? `[laugh]` The first week was absolute torture… Read the [Nano, Turbo, and V3 comparison](https://www.murmurtts.com/blog/chatterbox-nano-turbo-v3-comparison) for model-specific setup and support.

### Creator demos: design, cloning, and expression

Breeze TTS 2 · English · Creator demo0:00

**Script:** Every journey finds its meaning when someone dares to take the first step. [BreezeBlue’s TTS 2 demo](https://breezeblue.ai/breeze-tts-2).

Breeze TTS 2 · Mandarin Chinese · Creator demo0:00

The Chinese example from the same multilingual demo. Compare its phrasing with the English clip above. [BreezeBlue’s TTS 2 demo](https://breezeblue.ai/breeze-tts-2).

Fish Audio S2 Pro · James · English expression demo0:00

**Script:** `[clears throat]` Hey chat, how do I solve merge conflicts again? I can’t believe I forgot how to do it. [Fish Audio’s S2 demo](https://fish.audio/s2/).

Step Audio EditX · English · Voice cloning demo0:00

**Script:** Underneath the courtyard is a large underground exhibition room which connects the two buildings. This is the generated English-to-English output, not the reference recording. [Step Audio EditX paper demos](https://stepaudiollm.github.io/step-audio-editx/).

Voxtral TTS · Sanchit · British English cloning demo0:00

Mistral’s generated voice-cloning example using Sanchit’s reference voice. Listen for accent and sentence rhythm. [Mistral’s Voxtral TTS announcement](https://mistral.ai/news/voxtral-tts/).

Maya1 · English · Designed event-host voice0:00

An energetic female event-host voice from the model card. **Script excerpt:** Wow. This place looks even better than I imagined. [Maya1 model-card examples](https://huggingface.co/maya-research/maya1).

### Creator demos: narration

Zonos v0.1 · American English · Narration demo0:00

Zyphra’s first generated sample from the v0.1 announcement. The OmniVoice clip below uses the same passage about nature, although the generation settings differ. [Zyphra’s Zonos v0.1 demos](https://www.zyphra.com/post/beta-release-of-zonos-v0-1).

OmniVoice · English · Research demo0:00

**Shared script excerpt:** I don’t really care what you call me. I’ve been a silent spectator, watching species evolve, empires rise and fall. The authors source this prompt and text from Seed-TTS. [OmniVoice paper demos](https://zhu-han.github.io/omnivoice/).

StyleTTS 2 · English · LJSpeech narration demo0:00

The generated StyleTTS 2 output from the paper’s first LJSpeech example. **Script excerpt:** After the construction and action of the machine had been explained, the doctor asked the governor… [StyleTTS 2 paper demos](https://styletts2.github.io/).

Use these examples to shortlist models, then run the [same-script listening test](https://www.murmurtts.com/blog/local-tts-blind-benchmark-script-completion) with your own material. The model profiles below also link to NVIDIA’s interactive Magpie demo and include Boson’s Higgs TTS 3 video.

## Open-weight TTS leaderboard: September 2026

[Artificial Analysis’s Provider Voice leaderboard](https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice?open-weights=true), checked on **9 September 2026**, lists these 15 open-weight entries with all accents and categories selected. Scores reflect blind listening preferences using each provider’s own voices.

| Rank range | Model | Elo |
| --- | --- | --- |
| 1 | Breeze TTS 2 (provisional) | 1215 |
| 2 | Fish Audio S2 Pro | 1128 |
| 3 | Step Audio EditX (Mar 2026) | 1104 |
| 4 | Voxtral TTS | 1078 |
| 5–6 | Magpie-Multilingual 357M (Feb 2026) | 1065 |
| 5–6 | Kokoro 82M v1.0 | 1060 |
| 7–8 | Maya1 | 1043 |
| 7–8 | Higgs Audio V3 TTS | 1038 |
| 9 | Chatterbox | 1021 |
| 10 | Zonos-v0.1 | 1000 |
| 11 | VibeVoice 7B | 969 |
| 12 | OpenVoice v2 | 955 |
| 13 | XTTS v2 | 922 |
| 14 | StyleTTS 2 | 893 |
| 15 | MetaVoice v1 | 840 |

Breeze’s result is provisional. Rank ranges reflect uncertainty; see the source for confidence intervals and sample counts. These scores do not measure Mac speed, memory, or your cloned voice. The source’s prices are hosted API charges, not local-generation costs.

Keep checkpoint names exact: the Chatterbox row is not a Nano, Turbo, or V3 score. Qwen3-TTS, SparkTTS, and OmniVoice are absent from this filtered snapshot; absence is not a quality ranking. A community conversion needs its own listening test.

## More leaderboard models to consider on a Mac

**Mac setup and Murmur support**

The six models below are not in Murmur’s current model catalog. Community MLX weights and documented runtime support are starting points for separate setups. We have not run or benchmarked these six models for this guide.

### Breeze TTS 2: voice design with a non-commercial license

[Breeze TTS 2](https://huggingface.co/BreezeBlue/Breeze-TTS-2) supports English and Chinese cloning, voice design, and voice direction. Its official setup requires Linux and an NVIDIA GPU. An [unofficial MLX port](https://huggingface.co/mlx-community/Breeze-TTS-2-mlx) supplies bf16, 8-bit, and 4-bit weights, but the card does not provide a complete generation walkthrough. Check the matching runtime before downloading. The original weights, derivatives, and self-hosted outputs are restricted to research and non-commercial use; the Apache license on the code does not remove that restriction.

### Step Audio EditX: expressive editing with a CUDA setup

[Step Audio EditX](https://huggingface.co/stepfun-ai/Step-Audio-EditX) is a 3B model for voice cloning and editing emotion, delivery, and vocal events. The official instructions require NVIDIA CUDA and were tested on Linux. They describe 12 GB of GPU memory as a tight boundary and recommend 16 GB for more headroom. That is a different hardware path from an Apple Silicon Mac. The card identifies an Apache 2.0 code license; check the checkpoint’s own terms before commercial use.

### Voxtral TTS: a documented MLX path for nine languages

[Mistral’s Voxtral 4B TTS](https://huggingface.co/mistralai/Voxtral-4B-TTS-2603) offers 20 preset voices across nine languages. Its [MLX conversion](https://huggingface.co/mlx-community/Voxtral-4B-TTS-2603-mlx-bf16) includes MLX-Audio generation instructions and lists an approximately 8 GB bf16 download. Download size is not peak memory. The weights and supplied reference voices use CC BY-NC 4.0, so this is a non-commercial evaluation option under those terms. Its leaderboard score does not establish the quality or speed of a quantized Mac version.

### NVIDIA Magpie: check the checkpoint and runtime together

The leaderboard names the February 2026 [Magpie-Multilingual 357M](https://huggingface.co/nvidia/magpie_tts_multilingual_357m) release. NVIDIA’s current card also describes a newer v2607 checkpoint, so keep those versions separate. The official runtime targets NVIDIA systems. A [community INT8 MLX port](https://huggingface.co/aufklarer/Magpie-TTS-Multilingual-357M-MLX-8bit) points to a Swift SDK and requires tokenization and audio-generation integration. It is a developer path, not an installable Murmur model. Both cards identify the NVIDIA Open Model License; review its terms and the precise revision you download.

**Listen:** [NVIDIA’s interactive Magpie demo](https://huggingface.co/spaces/nvidia/magpie_tts_multilingual_demo) lets you enter a script and choose a speaker and language. It generates audio on demand; this guide does not have a verified fixed clip for the February leaderboard checkpoint.

### Maya1: English voice design with an MLX conversion

[Maya1](https://huggingface.co/maya-research/maya1) is a 3B English model with description-based voice design and emotion controls under Apache 2.0. The official deployment examples focus on GPUs, but the [community 4-bit conversion](https://huggingface.co/mlx-community/maya1-4bit) documents CLI and Python use with MLX-Audio. Its card names version 0.2.9 as the conversion toolchain. Verify compatibility with your installed runtime and test pronunciation and longer passages before choosing it for a project.

### Higgs Audio V3: conversational speech and creator-use terms

The leaderboard calls this model Higgs Audio V3 TTS; its linked repository now uses the name [Higgs TTS 3](https://huggingface.co/bosonai/higgs-tts-3-4b). It supports multilingual cloning and inline delivery controls, and [MLX-Audio lists Higgs Audio v3 support](https://github.com/Blaizzy/mlx-audio#supported-models). Pin the checkpoint and runtime together when evaluating it. Boson’s current license includes a creator-use grant for monetized videos, podcasts, audiobooks, and similar content with attribution. Hosting an API or embedding the model in a product requires a separate commercial license. Read the grant before treating all commercial uses alike.

*The multilingual demo embedded in Boson’s Higgs TTS 3 announcement. This is the creator’s video, not a Murmur recording. [Boson AI · Higgs TTS 3 announcement](https://www.boson.ai/blog/higgs-tts-3)*

## Models Currently Available in Murmur

| Model | Official size | Voice cloning | Language scope | License caution | Best first use |
| --- | --- | --- | --- | --- | --- |
| Kokoro | 82M parameters | No | 9 documented language or accent pipelines | Apache 2.0 | Fast narration and drafts |
| Qwen3-TTS Base | 0.6B checkpoint | Yes, from reference audio | 10 languages | Apache 2.0 | Multilingual voice cloning |
| Qwen3-TTS VoiceDesign | 1.7B checkpoint | No reference clip needed | 10 languages | Apache 2.0 | Voice design from a description |
| Chatterbox Turbo | 350M parameters | Yes | English | MIT | Expressive English reads |
| Chatterbox Multilingual | 500M parameters | Yes | 23 languages | MIT | Localization and cross-language cloning |
| Fish Audio S2 Pro | 4B Slow AR plus 400M Fast AR; card lists 5B | Yes | Trained across 80+ languages | Commercial license required | Fine-grained expressive control |
| SparkTTS | 0.5B parameters | Yes | English and Chinese | Weights are CC BY-NC-SA 4.0 | Non-commercial bilingual experiments |
| OmniVoice | 0.6B class | Yes | 646 documented languages | Review dependency licenses | Broad-language experimentation |

The numbers above come from the current official model repositories and cards, including [Hexgrad's Kokoro repository](https://github.com/hexgrad/kokoro), the [official Qwen3-TTS Base model card](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base), [Resemble AI's Chatterbox repository](https://github.com/resemble-ai/chatterbox), and the [Fish Audio S2 Pro model card](https://huggingface.co/fishaudio/s2-pro). Parameter count does not predict voice quality by itself. It mostly helps explain download, memory, and inference tradeoffs.

## The Best Model Changes With the Job

For an audiobook or a long course, consistency matters more than a dramatic ten-second demo. Kokoro is a sensible first test because its 82M-parameter footprint makes iteration practical, and its preset voices avoid managing consent and reference recordings. Qwen3-TTS becomes more attractive when the book crosses languages or needs a reusable cloned narrator. For either model, test a full representative chapter before committing to hundreds of pages.

For a YouTube explainer, speed and revision cost often win. A creator may regenerate three versions of the same paragraph while adjusting timing to an edit. Kokoro fits that loop. Chatterbox Turbo is worth the extra complexity when reactions and expressive delivery are part of the channel's style. Its official model includes native tags such as [cough], [laugh], and [chuckle], but it is English-only.

For localization, start with language support, not the English demo. Qwen3-TTS documents Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian. Chatterbox Multilingual lists 23 languages. OmniVoice documents 646 languages, but the amount of training data varies enormously by language, so 646 should not be read as 646 equally strong production voices. Review each target language with a native speaker.

For character dialogue and performance, Chatterbox Turbo and Fish Audio S2 Pro deserve a test. Fish accepts free-form inline directions such as [whisper in small voice] or [pitch up]. That control comes with a much larger model and a restrictive commercial-use condition. If the output is for a paying client, monetized channel, or product, licensing is a selection criterion, not a footnote.

## Kokoro: A Small Preset-Voice Starting Point

Kokoro is an open-weight 82M-parameter TTS model with Apache 2.0 licensed code and weights. The official pipeline covers American and British English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, and Mandarin Chinese. It emits 24 kHz audio in the official example and uses Misaki for text-to-phoneme processing. The current MLX-Audio documentation lists 54 preset voices and bf16, 8-bit, 6-bit, and 4-bit model variants.

Kokoro does not natively clone a reference voice. That limitation is also part of its appeal. You choose a preset, adjust speed, and generate. It is a strong fit for documentation, article narration, internal training, and draft voiceovers. Read the deeper [Kokoro TTS on Mac guide](https://www.murmurtts.com/blog/kokoro-tts-mac-guide) for the command-line and native-app paths.

## Qwen3-TTS: Cloning and Voice Design on Mac

Qwen3-TTS separates jobs across checkpoints. The 0.6B Base model accepts reference audio for rapid voice cloning. The 1.7B VoiceDesign model creates a voice from a natural-language description without a reference clip. Both sit inside a family trained on more than 5 million hours of speech across 10 languages. The official model cards are Apache 2.0.

The Qwen team reports end-to-end streaming latency as low as 97 ms, but that is a model-system result, not a promise for your Mac. Generation time changes with the checkpoint, quantization, text length, runtime, and Apple Silicon generation. Murmur's current catalog estimates a 1.7 GB download for its 0.6B bf16 Base checkpoint. VoiceDesign is estimated at 4.5 GB for bf16 or 2.3 GB for the 4-bit option. See the [Qwen3-TTS Mac guide](https://www.murmurtts.com/blog/qwen3-tts-for-multilingual-narration) for its creator workflow.

## Chatterbox: Two Different Recommendations

Chatterbox Turbo and Chatterbox Multilingual solve different problems. Turbo is a streamlined 350M-parameter English model with zero-shot cloning and native reaction tags. Multilingual V3 is a 500M-parameter model with 23 listed languages, improved speaker similarity, and reduced unwanted continuation compared with earlier versions, according to Resemble AI. Both official model cards use the MIT license.

One version detail matters. Murmur currently points to the MLX checkpoint named `mlx-community/chatterbox-fp16`. That should be described as Chatterbox Multilingual across 23 languages, not automatically as V3 unless the converted checkpoint is verified as V3. This kind of checkpoint-level precision prevents a current upstream feature from being accidentally promised in an older conversion.

## Fish Audio S2 Pro: Powerful, Large, and License-Limited

Fish Audio S2 Pro combines a 4B-parameter Slow AR model with a 400M-parameter Fast AR model. Its Hugging Face card displays the overall model as 5B parameters. The architecture uses 10 codebooks at roughly 21 Hz and supports free-form instructions inside the text. Fish says it trained the system on more than 10 million hours across more than 80 languages. Those facts make it attractive for nuanced performance, but not lightweight.

Murmur estimates 11 GB for its bf16 Fish checkpoint and 6.72 GB for the preferred low-memory 8-bit route. The app warns that S2 Pro requires at least 16 GB of unified memory. More important, the Fish Audio Research License permits research and non-commercial use without charge, while commercial use requires a separate written agreement. The license also includes notice and attribution requirements for distribution. Do not choose it for paid production without reviewing those terms.

## SparkTTS and OmniVoice: Specialized Alternatives

SparkTTS is a 0.5B English and Chinese model with zero-shot voice cloning and controllable virtual-speaker generation using gender, pitch, and speaking rate. Its repository code is Apache 2.0, but the official model weights were changed to CC BY-NC-SA 4.0 because of training-data terms. That means the weights are non-commercial and derivatives must follow the share-alike license. Spark is useful for research, education, and evaluation, not an uncomplicated commercial recommendation.

OmniVoice documents 646 languages, zero-shot cloning, voice design, and an Apple Silicon MPS path. Its voice-design training is concentrated in Chinese and English, and the maintainers warn that results can be unstable for some low-resource languages. The main repository is Apache 2.0, but an unresolved public issue questions the license of an underlying tokenizer. Buyers who need clear commercial provenance should review the complete dependency chain before shipping.

## What Apple Silicon Changes

Apple's [MLX framework](https://github.com/ml-explore/mlx) is designed for Apple Silicon and uses unified memory shared by CPU and GPU. MLX-Audio builds TTS runtimes on that foundation and currently requires Python 3.10 or newer plus an Apple Silicon Mac for its documented Mac path. WAV output works without ffmpeg. MP3, FLAC, OGG, Opus, and Vorbis encoding require ffmpeg.

Local does not mean zero setup or zero downloads. Model weights still need to reach the Mac, except when an app ships a model in its bundle. Larger checkpoints take more storage and unified memory. Download size also does not equal peak memory use. If you want the concepts without the command-line details, read [What Is MLX TTS?](https://www.murmurtts.com/blog/what-is-mlx-tts) and compare the available [Murmur voice workflows](https://www.murmurtts.com/voices).

## A Five-Minute Model Selection Checklist

1. Write down the exact target languages and have a native speaker review important output.
2. Decide whether a preset voice is enough or whether reference-audio cloning is essential.
3. Check the weights license, code license, and major runtime dependencies for commercial restrictions.
4. Compare model download estimates with your free storage and unified memory, not just the chip name.
5. Test a representative 300 to 500 word passage containing names, numbers, abbreviations, dialogue, and punctuation.
6. Use the same passage across shortlisted models and listen for pronunciation, consistency, pacing, and unwanted words.
7. Confirm that the workflow can export the format and number of files your editor or client expects.

## When Local TTS Is the Wrong Choice

A hosted service can be the better purchase when a team needs browser collaboration, centralized projects, a service-level agreement, or an API that scales without managing Macs. Cloud systems may also offer broader polished language coverage and faster server generation. Local TTS is strongest for private scripts, offline access after setup, predictable costs, and a single creator who owns the production machine. Compare those tradeoffs in [Murmur vs ElevenLabs](https://www.murmurtts.com/compare/murmur-vs-elevenlabs).

Murmur provides a Mac interface for Kokoro, Qwen3-TTS, Chatterbox, Fish Audio S2 Pro, SparkTTS, and OmniVoice workflows. Kokoro is bundled in the current app; larger models are downloaded when chosen. Murmur costs $49 one-time, has no free trial, and includes a 7-day refund policy. The app simplifies installation, projects, queueing, previews, and export, but it does not erase each model's license or hardware requirements.

## Frequently Asked Questions

## Use local TTS models in one Mac workflow

Murmur brings local voices, cloning, projects, queueing, and export into a $49 one-time Mac app. There is no free trial, and purchases include a 7-day refund policy.

[Buy Murmur · $49](https://murmur-licenses.tarunyadav9761.workers.dev/checkout)

macOS 15+ · Apple Silicon required · 7-day refund policy

MADE FOR YOUR MAC

## Less reading. *More making.*

[Explore Murmur](https://www.murmurtts.com/#product)
