AI SeedbankHelp preserve open and free AI for humanity's future

← All models

stepfun-ai_Step-Audio-2-mini

stepfun-ai · View on Hugging Face ↗

Get this model

Download TorrentMagnet Link

Seeders: · Leechers:

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


license: apache-2.0 library_name: transformers pipeline_tag: any-to-any language:

  • en
  • zh

Introduction

Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation, presented in the paper Step-Audio 2 Technical Report.

  • Advanced Speech and Audio Understanding: Promising performance in ASR and audio understanding by comprehending and reasoning semantic information, para-linguistic and non-vocal information.

  • Intelligent Speech Conversation: Achieving natural and intelligent interactions that are contextually appropriate for various conversational scenarios and paralinguistic information.

  • Tool Calling and Multimodal RAG: By leveraging tool calling and RAG to access real-world knowledge (both textual and acoustic), Step-Audio 2 can generate responses with fewer hallucinations for diverse scenarios, while also having the ability to switch timbres based on retrieved speech.

  • State-of-the-Art Performance: Achieving state-of-the-art performance on various audio understanding and conversational benchmarks compared to other open-source and commercial solutions. (See Evaluation and Technical Report).

Model Download

Huggingface

Models 🤗 Hugging Face
Step-Audio 2 mini stepfun-ai/Step-Audio-2-mini
Step-Audio 2 mini Base stepfun-ai/Step-Audio-2-mini-Base

Model Usage

🔧 Dependencies and Installation

  • Python >= 3.10
  • PyTorch >= 2.3-cu121
  • CUDA Toolkit
conda create -n stepaudio2 python=3.10
conda activate stepaudio2
pip install transformers==4.49.0 torchaudio librosa onnxruntime s3tokenizer diffusers hyperpyyaml

git clone https://github.com/stepfun-ai/Step-Audio2.git
cd Step-Audio2
git lfs install
git clone https://huggingface.co/stepfun-ai/Step-Audio-2-mini

🚀 Inference Scripts

python examples.py

🚀 Local web demonstration

pip install gradio
python web_demo.py

Online demonstration

StepFun Audio Studio

  • Both Step-Audio 2 and Step-Audio 2.3 are available in our StepFun Audio Studio.
  • You will need an API key from the StepFun Open Platform.

StepFun AI Assistant

  • Step-Audio 2 is also available in our StepFun AI Assistant mobile App with both web and audio search tools enabled.
  • Please scan the following QR code to download it from your app store then tap the phone icon in the top-right corner.

WeChat group

You can scan the following QR code to join our WeChat group for communication and discussion.

Evaluation

Automatic speech recognition

CER for Chinese, Cantonese and Japanese and WER for Arabian and English. N/A indicates that the language is not supported.

Category Test set Doubao LLM ASR GPT-4o Transcribe Kimi-Audio Qwen-Omni Step-Audio 2 Step-Audio 2 mini
English Common Voice 9.20 9.30 7.83 8.33 5.95 6.76
FLEURS English 7.22 2.71 4.47 5.05 3.03 3.05
LibriSpeech clean 2.92 1.75 1.49 2.93 1.17 1.33
LibriSpeech other 5.32 4.23 2.91 5.07 2.42 2.86
Average 6.17 4.50 4.18 5.35 3.14 3.50
Chinese AISHELL 0.98 3.52 0.64 1.17 0.63 0.78
AISHELL-2 3.10 4.26 2.67 2.40 2.10 2.16
FLEURS Chinese 2.92 2.62 2.91 7.01 2.68 2.53 2.53
KeSpeech phase1 6.48 26.80 5.11 6.45 3.63 3.97
WenetSpeech meeting 4.90 31.40 5.21 6.61 4.75 4.87
WenetSpeech net 4.46 15.71 5.93 5.24 4.67 4.82
Average 3.81 14.05 3.75 4.81 3.08 3.19
Multilingual FLEURS Arabian N/A 11.72 N/A 25.13 14.22 16.46
Common Voice yue 9.20 11.10 38.90 7.89 7.90 8.32
FLEURS Japanese N/A 3.27 N/A 10.49 3.18 4.67
In-house Anhui accent 8.83 50.55 22.17 18.73 10.61 11.65
Guangdong accent 4.99 7.83 3.76 4.03 3.81 4.44
Guangxi accent 3.37 7.09 4.29 3.35 4.11 3.51
Shanxi accent 20.26 55.03 34.71 25.95 12.44 15.60
Sichuan dialect 3.01 32.85 5.26 5.61 4.35 4.57
Shanghai dialect 47.49 89.58 82.90 58.74 17.77 19.30
Average 14.66 40.49 25.52 19.40 8.85 9.85

Paralinguistic information understanding

StepEval-Audio-Paralinguistic

Model Avg. Gender Age Timbre Scenario Event Emotion Pitch Rhythm Speed Style Vocal
GPT-4o Audio 43.45 18 42 34 22 14 82 40 60 58 64 44
Kimi-Audio 49.64 94 50 10 30 48 66 56 40 44 54 54
Qwen-Omni 44.18 40 50 16 28 42 76 32 54 50 50 48
Step-Audio-AQAA 36.91 70 66 18 14 14 40 38 48 54 44 0
Step-Audio 2 83.09 100 96 82 78 60 86 82 86 88 88 68
Step-Audio 2 mini 80.00 100 94 80 78 60 82 82 68 74 86 76

Audio understanding and reasoning

MMAU

Model Avg. Sound Speech Music
Audio Flamingo 3 73.1 76.9 66.1 73.9
Gemini 2.5 Pro 71.6 75.1 71.5 68.3
GPT-4o Audio 58.1 58.0 64.6 51.8
Kimi-Audio 69.6 79.0 65.5 64.4
Omni-R1 77.0 81.7 76.0 73.4
Qwen2.5-Omni 71.5 78.1 70.6 65.9
Step-Audio-AQAA 49.7 50.5 51.4 47.3
Step-Audio 2 78.0 83.5 76.9 73.7
Step-Audio 2 mini 73.2 76.6 71.5 71.6

Speech translation

Model CoVoST 2 (S2TT)
Avg. English-to-Chinese Chinese-to-English
GPT-4o Audio 29.61 40.20 19.01
Qwen2.5-Omni 35.40 41.40 29.40
Step-Audio-AQAA 28.57 37.71 19.43
Step-Audio 2 39.26 49.01 29.51
Step-Audio 2 mini 39.29 49.12 29.47
Model CVSS (S2ST)
Avg. English-to-Chinese Chinese-to-English
GPT-4o Audio 23.68 20.07 27.29
Qwen-Omni 15.35 8.04 22.66
Step-Audio-AQAA 27.36 30.74 23.98
Step-Audio 2 30.87 34.83 26.92
Step-Audio 2 mini 29.08 32.81 25.35

Tool calling

StepEval-Audio-Toolcall. Date and time tools have no parameter.

Model Objective Metric Audio search Date & Time Weather Web search
Qwen3-32B Trigger Precision / Recall 67.5 / 98.5 98.4 / 100.0 90.1 / 100.0 86.8 / 98.5
Type Accuracy 100.0 100.0 98.5 98.5
Parameter Accuracy 100.0 N/A 100.0 100.0
Step-Audio 2 Trigger Precision / Recall 86.8 / 99.5 96.9 / 98.4 92.2 / 100.0 88.4 / 95.5
Type Accuracy 100.0 100.0 90.5 98.4
Parameter Accuracy 100.0 N/A 100.0 100.0

Speech-to-speech conversation

URO-Bench. U. R. O. stands for understanding, reasoning, and oral conversation, respectively.

86.08 58.79
Model Language Basic Pro
Avg. U. R. O. Avg. U. R. O.
GPT-4o Audio Chinese 78.59 89.40 65.48 85.24 67.10 70.60 57.22 70.20
Kimi-Audio 73.59 79.34 64.66 79.75 66.07 60.44 59.29 76.21
Qwen-Omni 68.98 59.66 69.74 77.27 59.11 59.01 59.82 58.74
Step-Audio-AQAA 74.71 87.61 59.63 81.93 65.61 74.76 47.29 68.97
Step-Audio 2 83.32 91.05 75.45 68.25 74.78 63.18 65.10
Step-Audio 2 mini 77.81 89.19 64.53 84.12 69.57 76.84 58.90 69.42
GPT-4o Audio English 84.54 90.18 75.90 90.41 67.51 60.65 64.36 78.46
Kimi-Audio 60.04 83.36 42.31 60.36 49.79 50.32 40.59 56.04
Qwen-Omni 70.58 66.29 69.62 76.16 50.99 44.51 63.88 49.41
Step-Audio-AQAA 71.11 90.15 56.12 72.06 52.01 44.25 54.54 59.81
Step-Audio 2 83.90 92.72 76.51 84.92 66.07 64.86 67.75 66.33
Step-Audio 2 mini 74.36 90.07 60.12 77.65 61.2561.94 63.80

License

The model and code in the repository is licensed under Apache 2.0 License.

Citation

@misc{wu2025stepaudio2technicalreport,
      title={Step-Audio 2 Technical Report},
      author={Boyong Wu and Chao Yan and Chen Hu and Cheng Yi and Chengli Feng and Fei Tian and Feiyu Shen and Gang Yu and Haoyang Zhang and Jingbei Li and Mingrui Chen and Peng Liu and Wang You and Xiangyu Tony Zhang and Xingyuan Li and Xuerui Yang and Yayue Deng and Yechang Huang and Yuxin Li and Yuxin Zhang and Zhao You and Brian Li and Changyi Wan and Hanpeng Hu and Jiangjie Zhen and Siyu Chen and Song Yuan and Xuelin Zhang and Yimin Jiang and Yu Zhou and Yuxiang Yang and Bingxin Li and Buyun Ma and Changhe Song and Dongqing Pang and Guoqiang Hu and Haiyang Sun and Kang An and Na Wang and Shuli Gao and Wei Ji and Wen Li and Wen Sun and Xuan Wen and Yong Ren and Yuankai Ma and Yufan Lu and Bin Wang and Bo Li and Changxin Miao and Che Liu and Chen Xu and Dapeng Shi and Dingyuan Hu and Donghang Wu and Enle Liu and Guanzhe Huang and Gulin Yan and Han Zhang and Hao Nie and Haonan Jia and Hongyu Zhou and Jianjian Sun and Jiaoren Wu and Jie Wu and Jie Yang and Jin Yang and Junzhe Lin and Kaixiang Li and Lei Yang and Liying Shi and Li Zhou and Longlong Gu and Ming Li and Mingliang Li and Mingxiao Li and Nan Wu and Qi Han and Qinyuan Tan and Shaoliang Pang and Shengjie Fan and Siqi Liu and Tiancheng Cao and Wanying Lu and Wenqing He and Wuxun Xie and Xu Zhao and Xueqi Li and Yanbo Yu and Yang Yang and Yi Liu and Yifan Lu and Yilei Wang and Yuanhao Ding and Yuanwei Liang and Yuanwei Lu and Yuchu Luo and Yuhe Yin and Yumeng Zhan and Yuxiang Zhang and Zidong Yang and Zixin Zhang and Binxing Jiao and Daxin Jiang and Heung-Yeung Shum and Jiansheng Chen and Jing Li and Xiangyu Zhang and Yibo Zhu},
      year={2025},
      eprint={2507.16632},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2507.16632},
}

Magnet link

Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:

magnet:?xt=urn:btih:eb633d71656ad44695f25581f8e67a5ef779b4c6&dn=stepfun-ai_Step-Audio-2-mini

Open magnet in torrent client · infohash eb633d71656ad44695f25581f8e67a5ef779b4c6

Files & hashes

PathSizesha1sha256
README.md32.0 KB (32,766 B)e561da6b7f9caf7667727b78fc79cbe4f23c17e66b475ec1347dd660d9a50a8e421eb87c19227158841abc8bfb4b57e9c371d25e
added_tokens.json169.3 KB (173,344 B)3dc2bd80bb01fd6a02b7f3e755c212925c0204b1af6486629171010d08496ffd77a9ffa504cbe581731303fe7ed2ada203991a82
assets/architecture5.png1.5 MB (1,547,253 B)efd45065858fab56a22a5bb7dc5a691b98b4854b3c85a631baf8687bf35edaa22a76af8b641ece3cefc114befd6c9f27f72f73b0
assets/arxiv.svg1.1 KB (1,114 B)9e815e78e563907ba325f9c4eb553e18c75357eda67434015c97469d72035acbfabaac96f8baa69f73263fd8c00a1e7170264033
assets/logo.png6.7 KB (6,870 B)0b972779477f5707f49b9f4db9c920af87ed90d7bcfd1a80ed4c0485360a62a4ddea769e94d9f1c3746db2550b8e416719f4d526
assets/qrcode.jpg56.3 KB (57,678 B)d543543dea98c57c46f5e0597b8eb9fe126d8a67a10d8611679dcf7c1341a46b4bb041927ef4832c20155385d653bc644e6b2f4c
assets/radar.png1.7 MB (1,733,839 B)1b2015b6c0f99ac6f46c40715e2d9d4fa53a68eff7c16383ccd791312635793fb214aa53cce9a0b2135ab9e56672d873fee231ab
assets/wechat_group.png4.6 KB (4,751 B)d4abcf9af152c59c437ed8f23a295539127f1e4b84e9beee0a4f28aedd51ac533af0a786350ea3830d8a12365757079cd85c7e91
config.json1.0 KB (1,030 B)3434805f04fee7ddcdd6f7b57d273a4616466f27bb3bc7cbfc4f11f7e32c55709b66cf76eb8a80e72856944409fabe1e13a300f5
configuration_step_audio_2.py4.4 KB (4,555 B)95c636d22187de863d21fb5c01aa375cb4d841dc699df8e17a713b07ceccdd0d3a79f61b2e38361541a617c65a1dd6181ac96bb4
merges.txt1.6 MB (1,671,853 B)31349551d90c7606f325fe0f11bbb8bd5fa0d7c78831e4f1a044471340f7c0a83d7bd71306a5b867e95fd870f74d0c5308a904d5
model-00001-of-00004.safetensors4.59 GB (4,925,370,984 B)2218dc934b68008e193aced387dad89adb0c63797b88e02b0b8c643412ec68cae009b3952dbd8e27642d61626065a2c420a8b73c
model-00002-of-00004.safetensors4.59 GB (4,932,751,008 B)4c7b1ec2e1219058c5f9469c58ce926fa9c7ce5c3d412c8d2fc17ca3351751f3171d48ff5b139af623aa05749062f132ac2585f1
model-00003-of-00004.safetensors4.65 GB (4,988,307,424 B)583eecedc012b116dff27dd3050016ec679488f7135ae4a891350e8ebf9791ef073d310314e1f75192bece0971bfab7b86c5587c
model-00004-of-00004.safetensors1.66 GB (1,784,019,520 B)10f7fcdb828559b0dae8b0c1f5f7e6697333c12dd35bf0ec42ff9ec160dfc6c5cb20a65247f0f8ba1c6edc620398c2ef49a66295
model.safetensors.index.json63.1 KB (64,645 B)438fd67c13e0369bf05b6d11159f063dd412b55d300d35f38e13985c4a8899ceed2053d8dfcbf2120eab965da2b1bda2eae68bd8
modeling_step_audio_2.py16.3 KB (16,647 B)50dfd88aeb19f32e8488f4391cda9b5cc83e537c82b93db45a4bf37c2877660b321782aea4715e3e0d7be221e91e99e13815cbe3
special_tokens_map.json819 B (819 B)48237cf3928d1d2b90e2ac3018b8a694b15448bd48904d0bfbf462af87a7de5e963b266974a5bdc4b50068343e968597e6bcf73d
token2wav/flow.pt594.6 MB (623,466,603 B)9d94d97d7b57d25e9d78c406260839eadcb4685715ccff24256ff61537c7f8b51e025116b83405f3fb017b54b008fc97da115446
token2wav/flow.yaml1.1 KB (1,099 B)8638dca97881751e50c132498acbfa41ecf74d6f723295d37bf11f5f1b896ca4f2f4c81ebc2fbb3e51b753c2507ef8461c751486
token2wav/hift.pt79.5 MB (83,390,254 B)8f059cd7d06516c52a62dfce1b18ae25e79c70ec3386cc880324d4e98e05987b99107f49e40ed925b8ecc87c1f4939432d429879
tokenizer.json12.1 MB (12,684,784 B)4bdecce756b25645bd487df4afc3c37db5729dce529f599b059f0e73c30a25205b906a84add3988c964de0b7db090f34d71e2e6e
tokenizer_config.json1.1 MB (1,203,453 B)efec20d5b802e723549a85acd030bef77793a35d55ac36fac7a75a46cdb78fa14a7efabbc63429807c6f1d54ab82477b4b20fe8f
vocab.json2.6 MB (2,776,833 B)4783fe10ac3adce15ac8f358ef5462739852c569ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910

Cite this release

Canonical URL
https://aiseedbank.org/models/stepfun-ai_Step-Audio-2-mini/
Slug
stepfun-ai_Step-Audio-2-mini
Infohash
eb633d71656ad44695f25581f8e67a5ef779b4c6
License
apache-2.0
Signing key fingerprint
85a3b32c3712427b

Every file carries a locally computed sha256 — verify a download against the signed sums: stepfun-ai_Step-Audio-2-mini.SHA256SUMS (+ minisign signature).

Provenance

Upstream repositorystepfun-ai/Step-Audio-2-mini
Revision (pinned)e36fdd5d71e0ea22f09dd94bbab9bfc544ca1e36
Fetched at2026-09-04T06:15:11Z
License at fetchapache-2.0
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

✓ verified · rehash-vs-hf-metadata at 2026-09-04T06:18:22Z

apache-2.016.17 GB (17,359,289,126 bytes)transformersonnxsafetensorsstep_audio_2text-generationany-to-anycustom_code2 languages (en, zh)paper: 2507.16632