AI SeedbankHelp preserve open and free AI for humanity's future

← All models

fishaudio_s2-pro

fishaudio · View on Hugging Face ↗

Get this model

Download TorrentMagnet Link

Seeders: · Leechers:

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


language:

  • zh
  • en
  • ja
  • ko
  • es
  • pt
  • ar
  • ru
  • fr
  • de
  • sv
  • it
  • tr
  • 'no'
  • nl
  • cy
  • eu
  • ca
  • da
  • gl
  • ta
  • hu
  • fi
  • pl
  • et
  • hi
  • la
  • ur
  • th
  • vi
  • jw
  • bn
  • yo
  • sl
  • cs
  • sw
  • nn
  • he
  • ms
  • uk
  • id
  • kk
  • bg
  • lv
  • my
  • tl
  • sk
  • ne
  • fa
  • af
  • el
  • bo
  • hr
  • ro
  • sn
  • mi
  • yi
  • am
  • be
  • km
  • is
  • az
  • sd
  • br
  • sq
  • ps
  • mn
  • ht
  • ml
  • sr
  • sa
  • te
  • ka
  • bs
  • pa
  • lt
  • kn
  • si
  • hy
  • mr
  • as
  • gu
  • fo license: other license_name: fish-audio-research-license license_link: LICENSE.md pipeline_tag: text-to-speech tags:
  • text-to-speech
  • instruction-following
  • multilingual inference: false extra_gated_prompt: You agree to not use the model to generate contents that violate DMCA or local laws. extra_gated_fields: Country: country Specific date: date_picker I agree to use this model for non-commercial use ONLY: checkbox

Fish Audio S2 Pro

Technical Report | GitHub | Playground

Fish Audio S2 Pro is a leading text-to-speech (TTS) model with fine-grained inline control of prosody and emotion. Trained on over 10M+ hours of audio data across 80+ languages, the system combines reinforcement learning alignment with a dual-autoregressive architecture. The release includes model weights, fine-tuning code, and an SGLang-based streaming inference engine.

Architecture

S2 Pro builds on a decoder-only transformer combined with an RVQ-based audio codec (10 codebooks, ~21 Hz frame rate) using a Dual-Autoregressive (Dual-AR) architecture:

  • Slow AR (4B parameters): Operates along the time axis and predicts the primary semantic codebook.
  • Fast AR (400M parameters): Generates the remaining 9 residual codebooks at each time step, reconstructing fine-grained acoustic detail.

This asymmetric design keeps inference efficient while preserving audio fidelity. Because the Dual-AR architecture is structurally isomorphic to standard autoregressive LLMs, it inherits all LLM-native serving optimizations from SGLang — including continuous batching, paged KV cache, CUDA graph replay, and RadixAttention-based prefix caching.

Fine-Grained Inline Control

S2 Pro enables localized control over speech generation by embedding natural-language instructions directly within the text using [tag] syntax. Rather than relying on a fixed set of predefined tags, S2 Pro accepts free-form textual descriptions — such as [whisper in small voice], [professional broadcast tone], or [pitch up] — allowing open-ended expression control at the word level.

Common tags (15,000+ unique tags supported):

[pause] [emphasis] [laughing] [inhale] [chuckle] [tsk] [singing] [excited] [laughing tone] [interrupting] [chuckling] [excited tone] [volume up] [echo] [angry] [low volume] [sigh] [low voice] [whisper] [screaming] [shouting] [loud] [surprised] [short pause] [exhale] [delight] [panting] [audience laughter] [with strong accent] [volume down] [clearing throat] [sad] [moaning] [shocked]

Supported Languages

S2 Pro supports 80+ languages.

Tier 1: Japanese (ja), English (en), Chinese (zh)

Tier 2: Korean (ko), Spanish (es), Portuguese (pt), Arabic (ar), Russian (ru), French (fr), German (de)

Other supported languages: sv, it, tr, no, nl, cy, eu, ca, da, gl, ta, hu, fi, pl, et, hi, la, ur, th, vi, jw, bn, yo, xsl, cs, sw, nn, he, ms, uk, id, kk, bg, lv, my, tl, sk, ne, fa, af, el, bo, hr, ro, sn, mi, yi, am, be, km, is, az, sd, br, sq, ps, mn, ht, ml, sr, sa, te, ka, bs, pa, lt, kn, si, hy, mr, as, gu, fo, and more.

Production Streaming Performance

On a single NVIDIA H200 GPU:

  • Real-Time Factor (RTF): 0.195
  • Time-to-first-audio: ~100 ms
  • Throughput: 3,000+ acoustic tokens/s while maintaining RTF below 0.5

Links

  • Fish Speech GitHub
  • Fish Audio Playground
  • Blog & Tech Report

Technical Report

If you find our work useful, please consider citing our report:

@misc{liao2026fishaudios2technical,
      title={Fish Audio S2 Technical Report}, 
      author={Shijia Liao and Yuxuan Wang and Songting Liu and Yifan Cheng and Ruoyi Zhang and Tianyu Li and Shidong Li and Yisheng Zheng and Xingwei Liu and Qingzheng Wang and Zhizhuo Zhou and Jiahua Liu and Xin Chen and Dawei Han},
      year={2026},
      eprint={2603.08823},
      archivePrefix={arXiv},
      primaryClass={cs.SD},
      url={https://arxiv.org/abs/2603.08823},
}

License

This model is licensed under the Fish Audio Research License. Research and non-commercial use is permitted free of charge. Commercial use requires a separate license from Fish Audio — contact [email protected].

Magnet link

Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:

magnet:?xt=urn:btih:0c286deacb8f4328dae79e5eae5e15a8d33120c8&dn=fishaudio_s2-pro

Open magnet in torrent client · infohash 0c286deacb8f4328dae79e5eae5e15a8d33120c8

Files & hashes

PathSizesha1sha256
LICENSE.md10.1 KB (10,360 B)b469a1a983b3f961e839881fc0ea5c4eb82ec4beaa7d9206e9d710590987a3636934f643529c00cd490323594e6206aaa0c32d80
README.md5.0 KB (5,117 B)afb5f3a5d842294faf1f65dd0f93b38b65ad64cb3c2d78af0f3991ef6047272d4df18f204b4340e1e5fcc975d0e720d4312c4b29
chat_template.jinja4.0 KB (4,116 B)699ff8df401fe4788525e9c1f9b86a99eadd623087a2728cb8dc9fe424d624542f6060ec05a1d285ebbec578bb078900e33396b5
config.json1.8 KB (1,863 B)5214fbf56c44b20d00b1d5654d68b3b1f9ec9ea5261b519a2a9576710fc8533a77297fae007f0e7b3aa28a217f8352b7f32fe993
model-00001-of-00002.safetensors4.64 GB (4,986,872,984 B)02c822ea88c8d21a0b6b5ff2cfc58bdd7a8aa335c4218e8ac93be83b35eee30b4f94cb2e9b5ecff40f3e21611438d2f4f8804aad
model-00002-of-00002.safetensors3.85 GB (4,136,876,104 B)103d814940079f7ad28b6d243ca31ad10f4d5d6476738d23465deaac431433232c0762908cc99a6eddc3d49f67307d92680827be
model.safetensors.index.json31.9 KB (32,686 B)4a355d7d68ae0885db2c3ed3e2d9846420eef9a1c8cb9974d3d17663a95dba1d6c3ea531fc42d377c0dff946f3a01a0ec48d45a3
overview.png3.4 MB (3,544,504 B)5f137da47cfed5ac45f0b80f405acb38a7a28eabf77da18e7d3cf59182fb714fc6c1bc526e7877b538e6380d8f002e84e8c56ea9
special_tokens_map.json99.5 KB (101,864 B)7989e2cfbbfec23cfe75e17cb5494c2354a1769ec2ff18fde6e43b7408435bc8ed079af74531befba549358a97cbc59ce606bc6b
tokenizer.json11.7 MB (12,217,872 B)bccdea1b298536ea7eff7f08fd77159c0e4a3088f24e08099d45a8adf3f52f5f0b03276e433bb9d689bb15fcbcc48ce58744588b
tokenizer_config.json840.7 KB (860,832 B)47a7c4ee0e88f50904020eec33f723cec6933c61b8d149343ae425b0da67e6708686aceb51be7815d9792f265fc12ff04d5e9856

Cite this release

Canonical URL
https://aiseedbank.org/models/fishaudio_s2-pro/
Slug
fishaudio_s2-pro
Infohash
0c286deacb8f4328dae79e5eae5e15a8d33120c8
License
custom/other license
Signing key fingerprint
85a3b32c3712427b

Every file carries a locally computed sha256 — verify a download against the signed sums: fishaudio_s2-pro.SHA256SUMS (+ minisign signature).

Provenance

Upstream repositoryfishaudio/s2-pro
Revision (pinned)1de9996b6be38b745688de084d87a5633f714e4e
Fetched at2026-09-03T22:59:35Z
License at fetchother
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

✓ verified · rehash-vs-hf-metadata at 2026-09-03T23:02:28Z

custom/other license8.51 GB (9,140,528,302 bytes)safetensorsfish_qwen3_omnitext-to-speechinstruction-followingmultilingual83 languages (zh, en, ja …)paper: 2603.08823