stepfun-ai_Step-Audio-TTS-3B
stepfun-ai · View on Hugging Face ↗
Model card
The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.
license: apache-2.0 pipeline_tag: text-to-speech
Step-Audio-TTS-3B
Step-Audio-TTS-3B represents the industry's first Text-to-Speech (TTS) model trained on a large-scale synthetic dataset utilizing the LLM-Chat paradigm. It has achieved SOTA Character Error Rate (CER) results on the SEED TTS Eval benchmark. The model supports multiple languages, a variety of emotional expressions, and diverse voice style controls. Notably, Step-Audio-TTS-3B is also the first TTS model in the industry capable of generating RAP and Humming, marking a significant advancement in the field of speech synthesis.
This repository provides the model weights for StepAudio-TTS-3B, which is a dual-codebook trained LLM (Large Language Model) for text-to-speech synthesis. Additionally, it includes a vocoder trained using the dual-codebook approach, as well as a specialized vocoder specifically optimized for humming generation. These resources collectively enable high-quality speech synthesis and humming capabilities, leveraging the advanced dual-codebook training methodology.
Performance comparison of content consistency (CER/WER) between GLM-4-Voice and MinMo.
| Model | test-zh | test-en |
|---|---|---|
| CER (%) ↓ | WER (%) ↓ | |
| GLM-4-Voice | 2.19 | 2.91 |
| MinMo | 2.48 | 2.90 |
| Step-Audio | 1.53 | 2.71 |
Results of TTS Models on SEED Test Sets.
- StepAudio-TTS-3B-Single denotes dual-codebook backbone with single-codebook vocoder*
| Model | test-zh | test-en | ||
|---|---|---|---|---|
| CER (%) ↓ | SS ↑ | WER (%) ↓ | SS ↑ | |
| FireRedTTS | 1.51 | 0.630 | 3.82 | 0.460 |
| MaskGCT | 2.27 | 0.774 | 2.62 | 0.774 |
| CosyVoice | 3.63 | 0.775 | 4.29 | 0.699 |
| CosyVoice 2 | 1.45 | 0.806 | 2.57 | 0.736 |
| CosyVoice 2-S | 1.45 | 0.812 | 2.38 | 0.743 |
| Step-Audio-TTS-3B-Single | 1.37 | 0.802 | 2.52 | 0.704 |
| Step-Audio-TTS-3B | 1.31 | 0.733 | 2.31 | 0.660 |
| Step-Audio-TTS | 1.17 | 0.73 | 2.0 | 0.660 |
Performance comparison of Dual-codebook Resynthesis with Cosyvoice.
| Token | test-zh | test-en | ||
|---|---|---|---|---|
| CER (%) ↓ | SS ↑ | WER (%) ↓ | SS ↑ | |
| Groundtruth | 0.972 | - | 2.156 | - |
| CosyVoice | 2.857 | 0.849 | 4.519 | 0.807 |
| Step-Audio-TTS-3B | 2.192 | 0.784 | 3.585 | 0.742 |
More information
For more information, please refer to our repository: Step-Audio.
Magnet link
Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:
magnet:?xt=urn:btih:6e79801aec7f2b7b734eac2c44f8e46be20247a4&dn=stepfun-ai_Step-Audio-TTS-3BOpen magnet in torrent client · infohash 6e79801aec7f2b7b734eac2c44f8e46be20247a4
Files & hashes
| Path | Size | sha1 | sha256 |
|---|---|---|---|
| CosyVoice-300M-25Hz-Music/VERSION_Vq0206Vocoder_Sing_0106 | 133 B (133 B) | 2d468d86821058fbddf5fcbbf8faa67119ef83d9 | 34b05c07a3bc8c7d10faf763ec658328f37dc819c3e0a33ef80911a7247983b4 |
| CosyVoice-300M-25Hz-Music/cosyvoice.yaml | 3.1 KB (3,172 B) | 100c00f424dd6a250a0329596b4872229e4b9f89 | e67225adeca0bb3739e75f490d3d9e951615090ec5cbe5d276b6a03e2dc5e0cf |
| CosyVoice-300M-25Hz-Music/flow.pt | 402.6 MB (422,109,962 B) | b4fae4de7de362d2cfb28550c9100907724ea1b5 | acdae6101bc558903b68f506a7034f5fc582f16801e6f8e4a416dce4c509bae7 |
| CosyVoice-300M-25Hz-Music/hift.pt | 78.1 MB (81,896,716 B) | 09be04863939f6c56a4c0fee063202319384a00e | 91e679b6ca1eff71187ffb4f3ab0444935594cdcc20a9bd12afad111ef8d6012 |
| CosyVoice-300M-25Hz/VERSION_Vq0206Vocoder_1202 | 0 B (0 B) | e69de29bb2d1d6434b8b29ae775ad8c2e48c5391 | e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 |
| CosyVoice-300M-25Hz/cosyvoice.yaml | 3.1 KB (3,172 B) | 100c00f424dd6a250a0329596b4872229e4b9f89 | e67225adeca0bb3739e75f490d3d9e951615090ec5cbe5d276b6a03e2dc5e0cf |
| CosyVoice-300M-25Hz/flow.pt | 402.6 MB (422,117,160 B) | c6111792e2d5e5f7e2f2555bfc7f9e314a813939 | a0fefc7bc07c57cf4b8b13adec54585dd3ea4285b32f03b85fc9b7f2f34887f5 |
| CosyVoice-300M-25Hz/hift.pt | 78.1 MB (81,896,716 B) | 09be04863939f6c56a4c0fee063202319384a00e | 91e679b6ca1eff71187ffb4f3ab0444935594cdcc20a9bd12afad111ef8d6012 |
| README.md | 6.6 KB (6,795 B) | 857c40a1ed9d77703aadd9a2306a0db9050969d4 | dc3dfef376aef909b1e2c53ed38fb860d3b190a2d98d98f5a3c34005c1acc68b |
| config.json | 514 B (514 B) | 78619a35b594335317baddc69f5c134d36ce4f5f | 4a52e9b9397df4fdc0b9b3a89dddf3fc060ca33d27fd0760b90c1da9b37ddc9c |
| configuration_step1.py | 1.2 KB (1,251 B) | f6ad885955d74ee543130fa4e4d537321c94368e | ceba3d81a7f8fba186501b731bd9270ab17ddfe863a0921e1792f394d01dff20 |
| lib/liboptimus_ths-torch2.2-cu121.cpython-310-x86_64-linux-gnu.so | 29.8 MB (31,250,408 B) | 65f2f26bef5ade34d1d117ca2cc4c46cbf9dd9fc | e018916e5e93fb904be6b34af32e71d03ba9e888d8c086a43a5c9fcacda661a1 |
| lib/liboptimus_ths-torch2.3-cu121.cpython-310-x86_64-linux-gnu.so | 29.8 MB (31,250,472 B) | dc18894b9c6b634ac32f22e33c9031a87c88f01a | ee23bba95f7806364e101e285720892b755a176d603842fb4646822800ac2344 |
| lib/liboptimus_ths-torch2.5-cu124.cpython-310-x86_64-linux-gnu.so | 29.8 MB (31,258,792 B) | 3be6dc033f4f07c573bfeaa987093c081b60481a | 6fa1a77f035203ff90a071218f775381f705269ef454163474d22501684b7e1f |
| model-00001.safetensors | 6.57 GB (7,059,446,656 B) | ef628d8375e016e5a523694798a962bc68065128 | 99b74741d91dbe52844497fe5faf50cb195d28ffd51bb31777c632ae2eed3176 |
| model.safetensors.index.json | 19.7 KB (20,147 B) | 301504a7d89e5e54e18aa9c237d249fa91a35af6 | 94cadf7795a0e36dee76000aabf2804ed60a896718905d0f94a10908f95496fc |
| modeling_step1.py | 14.5 KB (14,855 B) | 4ecaa52f33c666982de2f491e18a31407d8f1d7d | 782f2a9eea0838acdb49cc6e067704fc546919d873b68a30ad109206043ad3a0 |
| tokenizer.model | 1.2 MB (1,264,044 B) | 7579f4a8ec1f12ff0e25949b5affa246b69a7716 | 25e122d9205d035033a9994c4d46a6a1b467a938654e4178fc0e5f4f5d610674 |
| tokenizer_config.json | 314 B (314 B) | f9bf4b6a00b757af0559e4c2aff443abeafc842a | 0db16fc3d1de979315e6bc1ab22ab0c8add09b50631f35cc551945e3d685382b |
Cite this release
- Canonical URL
- https://aiseedbank.org/models/stepfun-ai_Step-Audio-TTS-3B/
- Slug
- stepfun-ai_Step-Audio-TTS-3B
- Infohash
- 6e79801aec7f2b7b734eac2c44f8e46be20247a4
- License
- apache-2.0
- Signing key fingerprint
- 85a3b32c3712427b
Every file carries a locally computed sha256 — verify a download against the signed sums: stepfun-ai_Step-Audio-TTS-3B.SHA256SUMS (+ minisign signature).
Provenance
| Upstream repository | stepfun-ai/Step-Audio-TTS-3B |
|---|---|
| Revision (pinned) | 9ddb7cb28b97bbfceb429f1a2567b30256b7137c |
| Fetched at | 2026-09-04T06:18:22Z |
| License at fetch | apache-2.0 |
| Snapshot tool | huggingface · seedbank 0.1.0 |
Trackers
- udp://announce.aitorrent.org:6969/announce
- http://announce.aitorrent.org:7070/announce
- udp://announce2.aitorrent.org:6970/announce
- http://announce2.aitorrent.org:7071/announce
- udp://tracker.opentrackr.org:1337/announce
- udp://open.demonii.com:1337/announce
- udp://open.stealth.si:80/announce
- udp://exodus.desync.com:6969/announce
- udp://tracker.torrent.eu.org:451/announce
✓ verified · rehash-vs-hf-metadata at 2026-09-04T06:19:45Z
apache-2.07.60 GB (8,162,541,279 bytes)onnxsafetensorsstep1text-to-speechcustom_code