AI SeedbankHelp preserve open and free AI for humanity's future

← All models

microsoft_VibeVoice-ASR

microsoft · View on Hugging Face ↗

Get this model

Download TorrentMagnet Link

Seeders: · Leechers:

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


language:

  • en # English
  • zh # Chinese
  • es # Spanish
  • pt # Portuguese
  • de # German
  • ja # Japanese
  • ko # Korean
  • fr # French
  • ru # Russian
  • id # Indonesian
  • sv # Swedish
  • it # Italian
  • he # Hebrew
  • nl # Dutch
  • pl # Polish
  • no # Norwegian
  • tr # Turkish
  • th # Thai
  • ar # Arabic
  • hu # Hungarian
  • ca # Catalan
  • cs # Czech
  • da # Danish
  • fa # Persian
  • af # Afrikaans
  • hi # Hindi
  • fi # Finnish
  • et # Estonian
  • aa # Afar
  • el # Greek
  • ro # Romanian
  • vi # Vietnamese
  • bg # Bulgarian
  • is # Icelandic
  • sl # Slovenian
  • sk # Slovak
  • lt # Lithuanian
  • sw # Swahili
  • uk # Ukrainian
  • kl # Kalaallisut
  • lv # Latvian
  • hr # Croatian
  • ne # Nepali
  • sr # Serbian
  • tl # Filipino (ISO 639-1; 常见工程别名: fil)
  • yi # Yiddish
  • ms # Malay
  • ur # Urdu
  • mn # Mongolian
  • hy # Armenian
  • jv # Javanese license: mit pipeline_tag: automatic-speech-recognition tags:
  • ASR
  • Transcriptoin
  • Diarization
  • Speech-to-Text library_name: transformers

VibeVoice-ASR

VibeVoice-ASR is a unified speech-to-text model designed to handle 60-minute long-form audio in a single pass, generating structured transcriptions containing Who (Speaker), When (Timestamps), and What (Content), with support for Customized Hotwords and over 50 languages.

➡️ Code: microsoft/VibeVoice
➡️ Demo: VibeVoice-ASR-Demo
➡️ Report: VibeVoice-ASR Technical Report
➡️ Finetuning: Finetuning
➡️ vLLM: vLLM-VibeVoice-ASR

🔥 Key Features

  • 🕒 60-minute Single-Pass Processing: Unlike conventional ASR models that slice audio into short chunks (often losing global context), VibeVoice ASR accepts up to 60 minutes of continuous audio input within 64K token length. This ensures consistent speaker tracking and semantic coherence across the entire hour.

  • 👤 Customized Hotwords: Users can provide customized hotwords (e.g., specific names, technical terms, or background info) to guide the recognition process, significantly improving accuracy on domain-specific content.

  • 📝 Rich Transcription (Who, When, What): The model jointly performs ASR, diarization, and timestamping, producing a structured output that indicates who said what and when.

  • 🌍 Multilingual & Code-Switching Support: It supports over 50 languages, requires no explicit language setting, and natively handles code-switching within and across utterances. Language distribution can be found here.

Evaluation

Installation and Usage

Please refer to GitHub README.

Language Distribution

License

This project is licensed under the MIT License.

Contact

This project was conducted by members of Microsoft Research. We welcome feedback and collaboration from our audience. If you have suggestions, questions, or observe unexpected/offensive behavior in our technology, please contact us at [email protected]. If the team receives reports of undesired behavior or identifies issues independently, we will update this repository with appropriate mitigations.

Magnet link

Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:

magnet:?xt=urn:btih:4364903daa558d11dca44599471cdc17ea327650&dn=microsoft_VibeVoice-ASR

Open magnet in torrent client · infohash 4364903daa558d11dca44599471cdc17ea327650

Files & hashes

PathSizesha1sha256
README.md4.2 KB (4,345 B)501e3de4efe413e63e399bc441c0658f3f802134635e608dc781dd53ee002a7fddc4a4775b9e77d7c077a22be086dcf67fb5dad6
config.json3.4 KB (3,520 B)2ecc9f450c501030c10bbccee77fb31b899dd6511798906d016a625ffa0100182cad152e055bfee53fb228a45ffe25d8179b9b24
figures/DER.jpg61.2 KB (62,653 B)6d87843aef4f48013b1db3952cad6e09eb5118847472f1f44110da44ebfc49f7926096098c8598c0c63dd837cbace3748d9256d1
figures/VibeVoice_ASR_archi.png146.0 KB (149,474 B)3313a88445ba76f7bab944a9af0bdd78ee823df2c5973094212c6c5c9c0bba6ee274e396ddbb272d8642f86c0d96c5b9e85dd52b
figures/cpWER.jpg66.9 KB (68,492 B)d5ef00b0e971306a6a5430f73bb847d271d4ec25cd686c29637e4267ecb1119e68c973fe6782ef09e5b9fb136b5265b048255947
figures/language_distribution_horizontal.png867.6 KB (888,409 B)6ff46228bc74731dec74fc1bf62d57ca9eec67a51151d7587c660c7417df0194bbb7aeefe02a6be5a3ba8a89203d52a3602fa700
figures/tcpWER.jpg62.7 KB (64,251 B)b0dcb013fdf109151f1ac2c27d97e8008c203c841586b7b5635a9be4f1f67f8e31db639ca4382cc6dea747adff1afe37d27d0dc0
model-00001-of-00008.safetensors2.32 GB (2,488,346,272 B)4ff0233ee19396c7b814a6b42f832b0e73dba0905548c67885d423ba184bc8c33f2e9f81b582a6d119cef79907e19a274b916637
model-00002-of-00008.safetensors2.23 GB (2,389,315,976 B)6a69f0ae2cfe8bc3dbc9e0a65427129522f42d1a163023c61a3fb047745cbaf53ed41c1e27e515e9786a376e122bfac2ea6e687e
model-00003-of-00008.safetensors2.30 GB (2,466,376,368 B)41fd46fbcc9cddc85858b4d1375fd147ffeb35c54e021702dfac2c52e8fdd6688de82c118be7bb7ad9b5c7988725ec63c44a64fb
model-00004-of-00008.safetensors2.30 GB (2,466,376,400 B)6aee76d730e5b183dd60c73c4785fbd0582358bcb17657bb151daa117a5a4671374ac1b248acb696691a2a67ac227a1115925e30
model-00005-of-00008.safetensors2.33 GB (2,499,431,136 B)494dc89fda5af253b335b0d4d94bbb44635b40eb0ed4e457268f7b02dda5cffe16b3a32614ccc2ccfe5de2db39bdd79700836406
model-00006-of-00008.safetensors2.31 GB (2,483,469,928 B)cada4b067420f2fbd67727b8cc4860fcbe8107766de8246bb042fd853b57d40995efd289ea44e4d1b611cec2e122570b8d2122bd
model-00007-of-00008.safetensors1.36 GB (1,464,887,482 B)44ab7acdf0d7a51981b446cefe148beb295ca14da2ba6960d994dc7598efc6796f85ab097da7708f4dd56095f7fccf4df8dc00e5
model-00008-of-00008.safetensors1.02 GB (1,089,994,848 B)528434da9b683fc4212b676186c086a5d3648dff1b9d9b328f85a25b4efca712d31513c6eed9e178152cc8cf4a6f0c2cd2bb623f
model.safetensors.index.json117.3 KB (120,151 B)07c48a206a91ab04b8cf17cf6bcd4d881015dfd71468c7b7c74fe27831d8db57871fbf15efd270c747f3f99caf689119ace658ba

Cite this release

Canonical URL
https://aiseedbank.org/models/microsoft_VibeVoice-ASR/
Slug
microsoft_VibeVoice-ASR
Infohash
4364903daa558d11dca44599471cdc17ea327650
License
mit
Signing key fingerprint
85a3b32c3712427b

Every file carries a locally computed sha256 — verify a download against the signed sums: microsoft_VibeVoice-ASR.SHA256SUMS (+ minisign signature).

Provenance

Upstream repositorymicrosoft/VibeVoice-ASR
Revision (pinned)d0c9efdb8d614685062c04425d91e01b6f37d944
Fetched at2026-09-04T02:30:35Z
License at fetchmit
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

✓ verified · rehash-vs-hf-metadata at 2026-09-04T02:33:44Z

mit16.16 GB (17,349,559,705 bytes)transformerssafetensorsvibevoiceASRTranscriptoinDiarizationSpeech-to-Textautomatic-speech-recognitionendpoints_compatible51 languages (en, zh, es …)paper: 2601.18184