AI SeedbankHelp preserve open and free AI for humanity's future

← All models

ResembleAI_chatterbox-turbo

ResembleAI · View on Hugging Face ↗

Get this model

Download TorrentMagnet Link

Seeders: 1 · Leechers: 0

Observed 2026-09-02T13:56:39Z via announce.aitorrent.org:7070.

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


license: mit language:

  • en pipeline_tag: text-to-speech tags:
  • text-to-speech
  • speech
  • speech-generation
  • voice-cloning

Chatterbox TTS

Made with ❤️ by

Chatterbox is a family of three state-of-the-art, open-source text-to-speech models by Resemble AI.

We are excited to introduce Chatterbox-Turbo, our most efficient model yet. Built on a streamlined 350M parameter architecture, Turbo delivers high-quality speech with less compute and VRAM than our previous models. We have also distilled the speech-token-to-mel decoder, previously a bottleneck, reducing generation from 10 steps to just one, while retaining high-fidelity audio output.

Paralinguistic tags are now native to the Turbo model, allowing you to use [cough], [laugh], [chuckle], and more to add distinct realism. While Turbo was built primarily for low-latency voice agents, it excels at narration and creative workflows.

If you like the model but need to scale or tune it for higher accuracy, check out our competitively priced TTS service (link). It delivers reliable performance with ultra-low latency of sub 200ms—ideal for production use in agents, applications, or interactive media.

⚡ Model Zoo

Choose the right model for your application.

Model Size Languages Key Features Best For 🤗 Examples
Chatterbox-Turbo 350M English Paralinguistic Tags ([laugh]), Lower Compute and VRAM Zero-shot voice agents, Production Demo Listen
Chatterbox-Multilingual (Language list) 500M 23+ Zero-shot cloning, Multiple Languages Global applications, Localization Demo Listen
Chatterbox (Tips and Tricks) 500M English CFG & Exaggeration tuning General zero-shot TTS with creative controls Demo Listen

Installation

pip install chatterbox-tts

Alternatively, you can install from source:

# conda create -yn chatterbox python=3.11
# conda activate chatterbox

git clone https://github.com/resemble-ai/chatterbox.git
cd chatterbox
pip install -e .

We developed and tested Chatterbox on Python 3.11 on Debian 11 OS; the versions of the dependencies are pinned in pyproject.toml to ensure consistency. You can modify the code or dependencies in this installation mode.

Usage

Chatterbox-Turbo
import torchaudio as ta
import torch
from chatterbox.tts_turbo import ChatterboxTurboTTS

# Load the Turbo model
model = ChatterboxTurboTTS.from_pretrained(device="cuda")

# Generate with Paralinguistic Tags
text = "Hi there, Sarah here from MochaFone calling you back [chuckle], have you got one minute to chat about the billing issue?"

# Generate audio (requires a reference clip for voice cloning)
wav = model.generate(text, audio_prompt_path="your_10s_ref_clip.wav")

ta.save("test-turbo.wav", wav, model.sr)
Chatterbox and Chatterbox-Multilingual

import torchaudio as ta
from chatterbox.tts import ChatterboxTTS
from chatterbox.mtl_tts import ChatterboxMultilingualTTS

# English example
model = ChatterboxTTS.from_pretrained(device="cuda")

text = "Ezreal and Jinx teamed up with Ahri, Yasuo, and Teemo to take down the enemy's Nexus in an epic late-game pentakill."
wav = model.generate(text)
ta.save("test-english.wav", wav, model.sr)

# Multilingual examples
multilingual_model = ChatterboxMultilingualTTS.from_pretrained(device=device)

french_text = "Bonjour, comment ça va? Ceci est le modèle de synthèse vocale multilingue Chatterbox, il prend en charge 23 langues."
wav_french = multilingual_model.generate(spanish_text, language_id="fr")
ta.save("test-french.wav", wav_french, model.sr)

chinese_text = "你好,今天天气真不错,希望你有一个愉快的周末。"
wav_chinese = multilingual_model.generate(chinese_text, language_id="zh")
ta.save("test-chinese.wav", wav_chinese, model.sr)

# If you want to synthesize with a different voice, specify the audio prompt
AUDIO_PROMPT_PATH = "YOUR_FILE.wav"
wav = model.generate(text, audio_prompt_path=AUDIO_PROMPT_PATH)
ta.save("test-2.wav", wav, model.sr)

See example_tts.py and example_vc.py for more examples.

Supported Languages

Arabic (ar) • Danish (da) • German (de) • Greek (el) • English (en) • Spanish (es) • Finnish (fi) • French (fr) • Hebrew (he) • Hindi (hi) • Italian (it) • Japanese (ja) • Korean (ko) • Malay (ms) • Dutch (nl) • Norwegian (no) • Polish (pl) • Portuguese (pt) • Russian (ru) • Swedish (sv) • Swahili (sw) • Turkish (tr) • Chinese (zh)

Original Chatterbox Tips

  • General Use (TTS and Voice Agents):

    • Ensure that the reference clip matches the specified language tag. Otherwise, language transfer outputs may inherit the accent of the reference clip’s language. To mitigate this, set cfg_weight to 0.
    • The default settings (exaggeration=0.5, cfg_weight=0.5) work well for most prompts across all languages.
    • If the reference speaker has a fast speaking style, lowering cfg_weight to around 0.3 can improve pacing.
  • Expressive or Dramatic Speech:

    • Try lower cfg_weight values (e.g. ~0.3) and increase exaggeration to around 0.7 or higher.
    • Higher exaggeration tends to speed up speech; reducing cfg_weight helps compensate with slower, more deliberate pacing.

Built-in PerTh Watermarking for Responsible AI

Every audio file generated by Chatterbox includes Resemble AI's Perth (Perceptual Threshold) Watermarker - imperceptible neural watermarks that survive MP3 compression, audio editing, and common manipulations while maintaining nearly 100% detection accuracy.

Watermark extraction

You can look for the watermark using the following script.

import perth
import librosa

AUDIO_PATH = "YOUR_FILE.wav"

# Load the watermarked audio
watermarked_audio, sr = librosa.load(AUDIO_PATH, sr=None)

# Initialize watermarker (same as used for embedding)
watermarker = perth.PerthImplicitWatermarker()

# Extract watermark
watermark = watermarker.get_watermark(watermarked_audio, sample_rate=sr)
print(f"Extracted watermark: {watermark}")
# Output: 0.0 (no watermark) or 1.0 (watermarked)

Official Discord

👋 Join us on Discord and let's build something awesome together!

Acknowledgements

  • Cosyvoice
  • Real-Time-Voice-Cloning
  • HiFT-GAN
  • Llama 3
  • S3Tokenizer

Citation

If you find this model useful, please consider citing.

@misc{chatterboxtts2025,
  author       = {{Resemble AI}},
  title        = {{Chatterbox-TTS}},
  year         = {2025},
  howpublished = {\url{https://github.com/resemble-ai/chatterbox}},
  note         = {GitHub repository}
}

Disclaimer

Don't use this model to do bad things. Prompts are sourced from freely available data on the internet.

Magnet link

Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:

magnet:?xt=urn:btih:a4e0b3e233cd070ade84f2dd6fb98b9cd5634218&dn=ResembleAI_chatterbox-turbo

Open magnet in torrent client · infohash a4e0b3e233cd070ade84f2dd6fb98b9cd5634218

Files & hashes

PathSizesha1sha256
README.md9.3 KB (9,544 B)7d61412fa834e0e8d9217e86a8ad0bf44de04b6dd007d8c2ed2093530d95eee9d623f8977c321f6f361c19dafa07c000375808d7
added_tokens.json418 B (418 B)fd85325cdfafc690469b6f0e8aeb5cf4649c145072e4ab6acb0d9309ac3df4b526ae5fd80a2da5bc5ab7bb02d85096a374f69193
conds.pt165.5 KB (169,454 B)4f89af14ff6ca313608336d0e2613ff8bb379f2ab1852099306fd6a7814eb9d0bd10186caba7249596cc23868f78a0eefbfa5033
merges.txt445.6 KB (456,318 B)226b0752cac7789c48f0cb3ec53eda48b7be36cc1ce1664773c50f3e0cc8842619a93edc4624525b728b188a9e0be33b7726adc5
s3gen.safetensors1007.5 MB (1,056,484,620 B)972a6da50a9f3e3410f36632dc58f4aea73c8c1c2b78103c654207393955e4900aac14a12de8ef25f4b09424f1ef91941f161d4e
s3gen_meanflow.safetensors1015.5 MB (1,064,875,036 B)a9e5e0ac30c088fa68635419a4cce02dcd24299bd65cb687a2ed581ee6cc297e919ffefa63386944f42364ae13b78a594945514f
special_tokens_map.json470 B (470 B)b2d2dbc845800c78fe98656af927d831ab4ff7b792ba8063bf40aa163eadebbfe0de07c2aebe44cf0d4a9e8726580b0781fd2640
t3_turbo_v1.safetensors1.78 GB (1,915,480,052 B)27d93d787840f9b445b9ff0daa528cd44b61d78cfcf1f8c1d651bb7e3acd69ee5be269b4ac10c02980b7708213d598bc9f7cdf87
t3_turbo_v1.yaml8.3 KB (8,457 B)ff62b05c534a082aa5c5399c98986b8450205cc857623237f5072148a138a47fd93da1241bc5069fd0df7f1850c00053391c50de
tokenizer_config.json3.8 KB (3,878 B)58bc759ea7dab3dd5442f0500e493f170eeba67fbca16a2ac1ddbd78b8d6228f0031884cc74b6ea54b967d6f6d2ebae9ccde23e6
ve.safetensors5.4 MB (5,695,784 B)cf554c3c3ab64211a8da2d0422b7d8a731763b3ff0921cab452fa278bc25cd23ffd59d36f816d7dc5181dd1bef9751a7fb61f63c
vocab.json975.8 KB (999,186 B)a15dd0028acd1dd2c1c2394c80e1de8f3f12a0e4f6bd25a65e4e63ca31360e9fb11c7e4f9a391a78385d640acd814092dd6eee4f

Cite this release

Canonical URL
https://aiseedbank.org/models/ResembleAI_chatterbox-turbo/
Slug
ResembleAI_chatterbox-turbo
Infohash
a4e0b3e233cd070ade84f2dd6fb98b9cd5634218
License
mit
Signing key fingerprint
85a3b32c3712427b

Every file carries a locally computed sha256 — verify a download against the signed sums: ResembleAI_chatterbox-turbo.SHA256SUMS (+ minisign signature).

Provenance

Upstream repositoryResembleAI/chatterbox-turbo
Revision (pinned)749d1c1a46eb10492095d68fbcf55691ccf137cd
Fetched at2026-09-02T12:06:57Z
License at fetchmit
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

✓ verified · rehash-vs-hf-metadata at 2026-09-02T12:07:46Z

mit3.77 GB (4,044,183,217 bytes)text-to-speechspeechspeech-generationvoice-cloning1 language (en)