microsoft_VibeVoice-ASR
microsoft · View on Hugging Face ↗
Model card
The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.
language:
- en # English
- zh # Chinese
- es # Spanish
- pt # Portuguese
- de # German
- ja # Japanese
- ko # Korean
- fr # French
- ru # Russian
- id # Indonesian
- sv # Swedish
- it # Italian
- he # Hebrew
- nl # Dutch
- pl # Polish
- no # Norwegian
- tr # Turkish
- th # Thai
- ar # Arabic
- hu # Hungarian
- ca # Catalan
- cs # Czech
- da # Danish
- fa # Persian
- af # Afrikaans
- hi # Hindi
- fi # Finnish
- et # Estonian
- aa # Afar
- el # Greek
- ro # Romanian
- vi # Vietnamese
- bg # Bulgarian
- is # Icelandic
- sl # Slovenian
- sk # Slovak
- lt # Lithuanian
- sw # Swahili
- uk # Ukrainian
- kl # Kalaallisut
- lv # Latvian
- hr # Croatian
- ne # Nepali
- sr # Serbian
- tl # Filipino (ISO 639-1; 常见工程别名: fil)
- yi # Yiddish
- ms # Malay
- ur # Urdu
- mn # Mongolian
- hy # Armenian
- jv # Javanese license: mit pipeline_tag: automatic-speech-recognition tags:
- ASR
- Transcriptoin
- Diarization
- Speech-to-Text library_name: transformers
VibeVoice-ASR
VibeVoice-ASR is a unified speech-to-text model designed to handle 60-minute long-form audio in a single pass, generating structured transcriptions containing Who (Speaker), When (Timestamps), and What (Content), with support for Customized Hotwords and over 50 languages.
➡️ Code: microsoft/VibeVoice
➡️ Demo: VibeVoice-ASR-Demo
➡️ Report: VibeVoice-ASR Technical Report
➡️ Finetuning: Finetuning
➡️ vLLM: vLLM-VibeVoice-ASR
🔥 Key Features
🕒 60-minute Single-Pass Processing: Unlike conventional ASR models that slice audio into short chunks (often losing global context), VibeVoice ASR accepts up to 60 minutes of continuous audio input within 64K token length. This ensures consistent speaker tracking and semantic coherence across the entire hour.
👤 Customized Hotwords: Users can provide customized hotwords (e.g., specific names, technical terms, or background info) to guide the recognition process, significantly improving accuracy on domain-specific content.
📝 Rich Transcription (Who, When, What): The model jointly performs ASR, diarization, and timestamping, producing a structured output that indicates who said what and when.
🌍 Multilingual & Code-Switching Support: It supports over 50 languages, requires no explicit language setting, and natively handles code-switching within and across utterances. Language distribution can be found here.
Evaluation
Installation and Usage
Please refer to GitHub README.
Language Distribution
License
This project is licensed under the MIT License.
Contact
This project was conducted by members of Microsoft Research. We welcome feedback and collaboration from our audience. If you have suggestions, questions, or observe unexpected/offensive behavior in our technology, please contact us at [email protected]. If the team receives reports of undesired behavior or identifies issues independently, we will update this repository with appropriate mitigations.
Magnet link
Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:
magnet:?xt=urn:btih:4364903daa558d11dca44599471cdc17ea327650&dn=microsoft_VibeVoice-ASROpen magnet in torrent client · infohash 4364903daa558d11dca44599471cdc17ea327650
Files & hashes
| Path | Size | sha1 | sha256 |
|---|---|---|---|
| README.md | 4.2 KB (4,345 B) | 501e3de4efe413e63e399bc441c0658f3f802134 | 635e608dc781dd53ee002a7fddc4a4775b9e77d7c077a22be086dcf67fb5dad6 |
| config.json | 3.4 KB (3,520 B) | 2ecc9f450c501030c10bbccee77fb31b899dd651 | 1798906d016a625ffa0100182cad152e055bfee53fb228a45ffe25d8179b9b24 |
| figures/DER.jpg | 61.2 KB (62,653 B) | 6d87843aef4f48013b1db3952cad6e09eb511884 | 7472f1f44110da44ebfc49f7926096098c8598c0c63dd837cbace3748d9256d1 |
| figures/VibeVoice_ASR_archi.png | 146.0 KB (149,474 B) | 3313a88445ba76f7bab944a9af0bdd78ee823df2 | c5973094212c6c5c9c0bba6ee274e396ddbb272d8642f86c0d96c5b9e85dd52b |
| figures/cpWER.jpg | 66.9 KB (68,492 B) | d5ef00b0e971306a6a5430f73bb847d271d4ec25 | cd686c29637e4267ecb1119e68c973fe6782ef09e5b9fb136b5265b048255947 |
| figures/language_distribution_horizontal.png | 867.6 KB (888,409 B) | 6ff46228bc74731dec74fc1bf62d57ca9eec67a5 | 1151d7587c660c7417df0194bbb7aeefe02a6be5a3ba8a89203d52a3602fa700 |
| figures/tcpWER.jpg | 62.7 KB (64,251 B) | b0dcb013fdf109151f1ac2c27d97e8008c203c84 | 1586b7b5635a9be4f1f67f8e31db639ca4382cc6dea747adff1afe37d27d0dc0 |
| model-00001-of-00008.safetensors | 2.32 GB (2,488,346,272 B) | 4ff0233ee19396c7b814a6b42f832b0e73dba090 | 5548c67885d423ba184bc8c33f2e9f81b582a6d119cef79907e19a274b916637 |
| model-00002-of-00008.safetensors | 2.23 GB (2,389,315,976 B) | 6a69f0ae2cfe8bc3dbc9e0a65427129522f42d1a | 163023c61a3fb047745cbaf53ed41c1e27e515e9786a376e122bfac2ea6e687e |
| model-00003-of-00008.safetensors | 2.30 GB (2,466,376,368 B) | 41fd46fbcc9cddc85858b4d1375fd147ffeb35c5 | 4e021702dfac2c52e8fdd6688de82c118be7bb7ad9b5c7988725ec63c44a64fb |
| model-00004-of-00008.safetensors | 2.30 GB (2,466,376,400 B) | 6aee76d730e5b183dd60c73c4785fbd0582358bc | b17657bb151daa117a5a4671374ac1b248acb696691a2a67ac227a1115925e30 |
| model-00005-of-00008.safetensors | 2.33 GB (2,499,431,136 B) | 494dc89fda5af253b335b0d4d94bbb44635b40eb | 0ed4e457268f7b02dda5cffe16b3a32614ccc2ccfe5de2db39bdd79700836406 |
| model-00006-of-00008.safetensors | 2.31 GB (2,483,469,928 B) | cada4b067420f2fbd67727b8cc4860fcbe810776 | 6de8246bb042fd853b57d40995efd289ea44e4d1b611cec2e122570b8d2122bd |
| model-00007-of-00008.safetensors | 1.36 GB (1,464,887,482 B) | 44ab7acdf0d7a51981b446cefe148beb295ca14d | a2ba6960d994dc7598efc6796f85ab097da7708f4dd56095f7fccf4df8dc00e5 |
| model-00008-of-00008.safetensors | 1.02 GB (1,089,994,848 B) | 528434da9b683fc4212b676186c086a5d3648dff | 1b9d9b328f85a25b4efca712d31513c6eed9e178152cc8cf4a6f0c2cd2bb623f |
| model.safetensors.index.json | 117.3 KB (120,151 B) | 07c48a206a91ab04b8cf17cf6bcd4d881015dfd7 | 1468c7b7c74fe27831d8db57871fbf15efd270c747f3f99caf689119ace658ba |
Cite this release
- Canonical URL
- https://aiseedbank.org/models/microsoft_VibeVoice-ASR/
- Slug
- microsoft_VibeVoice-ASR
- Infohash
- 4364903daa558d11dca44599471cdc17ea327650
- License
- mit
- Signing key fingerprint
- 85a3b32c3712427b
Every file carries a locally computed sha256 — verify a download against the signed sums: microsoft_VibeVoice-ASR.SHA256SUMS (+ minisign signature).
Provenance
| Upstream repository | microsoft/VibeVoice-ASR |
|---|---|
| Revision (pinned) | d0c9efdb8d614685062c04425d91e01b6f37d944 |
| Fetched at | 2026-09-04T02:30:35Z |
| License at fetch | mit |
| Snapshot tool | huggingface · seedbank 0.1.0 |
Trackers
- udp://announce.aitorrent.org:6969/announce
- http://announce.aitorrent.org:7070/announce
- udp://announce2.aitorrent.org:6970/announce
- http://announce2.aitorrent.org:7071/announce
- udp://tracker.opentrackr.org:1337/announce
- udp://open.demonii.com:1337/announce
- udp://open.stealth.si:80/announce
- udp://exodus.desync.com:6969/announce
- udp://tracker.torrent.eu.org:451/announce
✓ verified · rehash-vs-hf-metadata at 2026-09-04T02:33:44Z
mit16.16 GB (17,349,559,705 bytes)transformerssafetensorsvibevoiceASRTranscriptoinDiarizationSpeech-to-Textautomatic-speech-recognitionendpoints_compatible51 languages (en, zh, es …)paper: 2601.18184