AI SeedbankHelp preserve open and free AI for humanity's future

← All models

kyutai_hibiki-zero-3b-pytorch-bf16

kyutai · View on Hugging Face ↗

Get this model

Download TorrentMagnet Link

Seeders: 1 · Leechers: 0

Observed 2026-09-02T13:56:39Z via announce.aitorrent.org:7070.

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


language:

  • fr
  • es
  • pt
  • de
  • en metrics:
  • bleu
  • comet base_model:
  • kyutai/hibiki-zero-3b-pytorch-bf16 pipeline_tag: audio-to-audio

Hibiki-Zero

Hibiki-Zero is a model for simultaneous speech translation. Traditional approaches for building simultaneous translation systems rely on supervised training with word-level aligned data between the source and the target content. Hibiki-Zero eliminates the need for word-level alignments entirely so that it fundamentally simplifies the training pipeline and enables seamless scaling to multiple languages with varying grammatical structures.

Hibiki-Zero supports translation from 🇫🇷 French, 🇪🇸 Spanish, 🇵🇹 Portuguese and 🇩🇪 German to 🇬🇧 English. At inference, Hibiki-Zero adapts its flow to accumulate just enough context so that it produces a real-time and natural speech translation with voice transfer along with a text translation. Hibiki-Zero can also be adapted to a new input language with less than 1000h of speech data.


Model Details

This is the model simply referred to as Hibiki-Zero in our paper, a 3B-parameter hierarchical Transformer producing speech and text tokens at a framerate of 12.5Hz, with audio being generated at a 2.2kbps bitrate.

Model Description

Hibiki-Zero is a decoder-only model that can receive and generate audio tokens produced by the the streaming neural audio codec Mimi. It leverages the same multistream architecture as Moshi or Hibiki to model source and target speech jointly. This allows Hibiki-Zero to continuously process the input stream while generating the target speech and text tokens at a constant framerate of 12.5Hz producing a continuous output audio stream, along with timestamped text translation. Hibiki-Zero consist of a main backbone of 3 billion parameters.

At inference, Hibiki-Zero continuously encodes the input user speech and produces real-time speech and text translation. Our model relies on simple temperature sampling and is thus compatible with batching unlike models using complex inference policies. It is also possible to run batched inference 3x faster than real-time on a single H100 GPU as demonstrated by our inference code. Hibiki-Zero only supports a single speaker in a single language per session. However, it shows zero-shot capabilities for translation with voice transfer of multiple speakers with different languages in the same audio.

  • Developed by: Kyutai
  • Model type: Simultaneous speech-to-speech and speech-to-text translation.
  • Languages: {French,Spanish,Portuguese,German}-to-English
  • License: CC BY-NC-SA 4.0

Model Sources


Usage

Direct Use

The model can be used for streaming translation from French, Spanish, Portuguese and German to English in real-time settings, or for batched simultaneous translation of many input sequences. It is robust to noisy conditions and is trained on sequences up to 120 seconds.

Downstream Use

Some components of the model can be used independently or repurposed relatively easily. For instance the Mimi codec is a state-of-the-art audio neural codec that combines semantic and acoustic information into audio tokens running at 12Hz and a bitrate of 1.1kbps, which make it particularly adapted to train speech language models or text-to-speech systems. Regarding the main Hibiki-Zero architecture, we demonstrated that it was possible to finetune it to adapt to a new input language with less than 1000h of speech and explicit the method in our paper.

Out-of-Scope Use

The model is not intended to be used to impersonate other people or any malicious use of any kind.

How to Get Started with the Model

See the README file for the inference code.


Training Details

Training Data

  • Textual data: The underlying text LLM model Helium-1-2B is trained on a mix of data including: Wikipedia, Stack Exchange, open-access scientific articles (from peS2o) and Common Crawl.

  • Audio data:

    • Unsupervised audio dataset: This dataset used for audio pretraining is a large collection readily available audio content in French, Spanish, Portuguese, German and English. Our data mixture contains approximately 12% of audio in each source language, 50% of English and less than 2% of Italian (see Section 4.2.2).
    • Speech translation dataset: This dataset used for speech translation training and reinforcement contains around 40k hours of real speech data for each source language with synthetic sentence-level aligned speech in English (see Sections 4.2.3 and 4.2.4).
    • Speech translation fine-tuning dataset: This dataset is a small 200h resynthesized subset of the speech translation dataset with natural pauses to improve audio quality and speech naturalness (see Section 4.2.5).

Training procedure and hyper-parameters

The different training stages along with the hyper-parameters are detailled in the paper.

Compute Infrastructure

The final model was trained on 48 H100 Nvidia GPUs.


Citation

If you use this model, please cite:

@unpublished{hibikizero2026,
  title={Simultaneous Speech-to-Speech Translation Without Aligned Data},
  author={Tom Labiausse and Romain Fabre and Yannick Estève and Alexandre Défossez and Neil Zeghidour},
  note={Preprint},
  year={2026},
  url={https://arxiv.org/abs/2602.11072v1}
}

Magnet link

Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:

magnet:?xt=urn:btih:328ed9bd1dfb30c05f8e7f150c4635880ff9ec05&dn=kyutai_hibiki-zero-3b-pytorch-bf16

Open magnet in torrent client · infohash 328ed9bd1dfb30c05f8e7f150c4635880ff9ec05

Files & hashes

PathSizesha1sha256
README.md6.0 KB (6,162 B)cb87b8993e605b700588c296c11e268bc270827cc0482e28804df5f881b747997dd83313d4a17bbd36a00fd06c7fbf1de0af3ca8
config.json1.6 KB (1,675 B)7388d8998eed7a648e196f664db815a40488d1e9a99f354a6131034b688fc9f91c889dc10e7eeff96ce65e94447be33d1be325a5
[email protected]5.83 GB (6,263,420,344 B)98f9550bb376b683fb858560bd5d7bb23daea6e5cd78e453b3b80299255bea02be439bcc2552b57c03cd82dbf0e9792e20100db8
[email protected]366.8 MB (384,644,900 B)612e137da20d8bbe47e4520227788a63d92a537509b782f0629851a271227fb9d36db65c041790365f11bbe5d3d59369cf863f50
tokenizer_spm_48k_multi6_2.model837.2 KB (857,314 B)60afa2fd517392a159d59be9dae966be7f3ec537c22110fb855aa049e17346ea2e88355bdd664f06cbfd09948380ab5e85b39697

Cite this release

Canonical URL
https://aiseedbank.org/models/kyutai_hibiki-zero-3b-pytorch-bf16/
Slug
kyutai_hibiki-zero-3b-pytorch-bf16
Infohash
328ed9bd1dfb30c05f8e7f150c4635880ff9ec05
License
no license recorded
Signing key fingerprint
85a3b32c3712427b

Every file carries a locally computed sha256 — verify a download against the signed sums: kyutai_hibiki-zero-3b-pytorch-bf16.SHA256SUMS (+ minisign signature).

Provenance

Upstream repositorykyutai/hibiki-zero-3b-pytorch-bf16
Revision (pinned)73175ce6243f8ad66b2138b0264a80044b35c1bd
Fetched at2026-09-02T05:11:09Z
License at fetchno license recorded
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

✓ verified · rehash-vs-hf-metadata at 2026-09-02T05:12:10Z

no license recorded6.19 GB (6,648,930,395 bytes)hibikiaudio-to-audio5 languages (fr, es, pt …)paper: 2410.00037paper: 2502.03382paper: 2602.11072