AI SeedbankHelp preserve open and free AI for humanity's future

← All models

nari-labs_Dia-1.6B

nari-labs · View on Hugging Face ↗

Get this model

Download TorrentMagnet Link

Seeders: 1 · Leechers: 0

Observed 2026-09-02T13:56:39Z via announce.aitorrent.org:7070.

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


license: apache-2.0 pipeline_tag: text-to-speech language:

  • en tags:
  • model_hub_mixin
  • pytorch_model_hub_mixin widget:
  • text: "[S1] Dia is an open weights text to dialogue model. [S2] You get full control over scripts and voices. [S1] Wow. Amazing. (laughs) [S2] Try it now on Git hub or Hugging Face." example_title: "Dia intro"
  • text: "[S1] Oh fire! Oh my goodness! What's the procedure? What to we do people? The smoke could be coming through an air duct! [S2] Oh my god! Okay.. it's happening. Everybody stay calm! [S1] What's the procedure... [S2] Everybody stay fucking calm!!!... Everybody fucking calm down!!!!! [S1] No! No! If you touch the handle, if its hot there might be a fire down the hallway!" example_title: "Panic protocol"

Dia is a 1.6B parameter text to speech model created by Nari Labs. It was pushed to the Hub using the PytorchModelHubMixin integration.

Dia directly generates highly realistic dialogue from a transcript. You can condition the output on audio, enabling emotion and tone control. The model can also produce nonverbal communications like laughter, coughing, clearing throat, etc.

To accelerate research, we are providing access to pretrained model checkpoints and inference code. The model weights are hosted on Hugging Face. The model only supports English generation at the moment.

We also provide a demo page comparing our model to ElevenLabs Studio and Sesame CSM-1B.

  • (Update) We have a ZeroGPU Space running! Try it now here. Thanks to the HF team for the support :)
  • Join our discord server for community support and access to new features.
  • Play with a larger version of Dia: generate fun conversations, remix content, and share with friends. 🔮 Join the waitlist for early access.

⚡️ Quickstart

This will open a Gradio UI that you can work on.

git clone https://github.com/nari-labs/dia.git
cd dia && uv run app.py

or if you do not have uv pre-installed:

git clone https://github.com/nari-labs/dia.git
cd dia
python -m venv .venv
source .venv/bin/activate
pip install uv
uv run app.py

Note that the model was not fine-tuned on a specific voice. Hence, you will get different voices every time you run the model. You can keep speaker consistency by either adding an audio prompt (a guide coming VERY soon - try it with the second example on Gradio for now), or fixing the seed.

Features

  • Generate dialogue via [S1] and [S2] tag
  • Generate non-verbal like (laughs), (coughs), etc.
    • Below verbal tags will be recognized, but might result in unexpected output.
    • (laughs), (clears throat), (sighs), (gasps), (coughs), (singing), (sings), (mumbles), (beep), (groans), (sniffs), (claps), (screams), (inhales), (exhales), (applause), (burps), (humming), (sneezes), (chuckle), (whistles)
  • Voice cloning. See example/voice_clone.py for more information.
    • In the Hugging Face space, you can upload the audio you want to clone and place its transcript before your script. Make sure the transcript follows the required format. The model will then output only the content of your script.

⚙️ Usage

As a Python Library

import soundfile as sf

from dia.model import Dia


model = Dia.from_pretrained("nari-labs/Dia-1.6B")

text = "[S1] Dia is an open weights text to dialogue model. [S2] You get full control over scripts and voices. [S1] Wow. Amazing. (laughs) [S2] Try it now on Git hub or Hugging Face."

output = model.generate(text)

sf.write("simple.mp3", output, 44100)

A pypi package and a working CLI tool will be available soon.

💻 Hardware and Inference Speed

Dia has been tested on only GPUs (pytorch 2.0+, CUDA 12.6). CPU support is to be added soon. The initial run will take longer as the Descript Audio Codec also needs to be downloaded.

On enterprise GPUs, Dia can generate audio in real-time. On older GPUs, inference time will be slower. For reference, on a A4000 GPU, Dia roughly generates 40 tokens/s (86 tokens equals 1 second of audio). torch.compile will increase speeds for supported GPUs.

The full version of Dia requires around 10GB of VRAM to run. We will be adding a quantized version in the future.

If you don't have hardware available or if you want to play with bigger versions of our models, join the waitlist here.

🪪 License

This project is licensed under the Apache License 2.0 - see the LICENSE file for details.

⚠️ Disclaimer

This project offers a high-fidelity speech generation model intended for research and educational use. The following uses are strictly forbidden:

  • Identity Misuse: Do not produce audio resembling real individuals without permission.
  • Deceptive Content: Do not use this model to generate misleading content (e.g. fake news)
  • Illegal or Malicious Use: Do not use this model for activities that are illegal or intended to cause harm.

By using this model, you agree to uphold relevant legal standards and ethical responsibilities. We are not responsible for any misuse and firmly oppose any unethical usage of this technology.

🔭 TODO / Future Work

  • Docker support.
  • Optimize inference speed.
  • Add quantization for memory efficiency.

🤝 Contributing

We are a tiny team of 1 full-time and 1 part-time research-engineers. We are extra-welcome to any contributions! Join our Discord Server for discussions.

🤗 Acknowledgements

  • We thank the Google TPU Research Cloud program for providing computation resources.
  • Our work was heavily inspired by SoundStorm, Parakeet, and Descript Audio Codec.
  • HuggingFace for providing the ZeroGPU Grant.
  • "Nari" is a pure Korean word for lily.
  • We thank Jason Y. for providing help with data filtering.

Magnet link

Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:

magnet:?xt=urn:btih:cff6543cf823e0a9b2004a01897b0c682458a350&dn=nari-labs_Dia-1.6B

Open magnet in torrent client · infohash cff6543cf823e0a9b2004a01897b0c682458a350

Files & hashes

PathSizesha1sha256
README.md6.4 KB (6,551 B)146916d420c1f14cf794a171811ac56e42d13dbdba8921486a0a601d65b9445a449b799c338a1a80943c8c0669eaba9ea4404b21
config.json941 B (941 B)0a586180c3246fefa312c5e3977a6e419a7a113d9140e85fd15b82d7f268c5681a41fb1de68573ec206b87cb3a5862fe5c598d6e
model.safetensors6.00 GB (6,444,682,848 B)16fa8b3b15032385c99b555ab501047788674e0acaba289b60f6d7d1e58fc744f4dc25aae88995fcca46be3d05e220b971486a26
preprocessor_config.json172 B (172 B)a812a82c392c511dc04417b8f8bcde9411347af01d47667309198c616ff39964928543718a180aaa0269ce38442418edc1f19f75

Cite this release

Canonical URL
https://aiseedbank.org/models/nari-labs_Dia-1.6B/
Slug
nari-labs_Dia-1.6B
Infohash
cff6543cf823e0a9b2004a01897b0c682458a350
License
apache-2.0
Signing key fingerprint
85a3b32c3712427b

Every file carries a locally computed sha256 — verify a download against the signed sums: nari-labs_Dia-1.6B.SHA256SUMS (+ minisign signature).

Provenance

Upstream repositorynari-labs/Dia-1.6B
Revision (pinned)257bc72f9b78182ccc6fa07675a9ae4c1a44e2cd
Fetched at2026-09-02T12:21:00Z
License at fetchapache-2.0
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

✓ verified · rehash-vs-hf-metadata at 2026-09-02T12:22:04Z

apache-2.06.00 GB (6,444,690,512 bytes)safetensorsmodel_hub_mixinpytorch_model_hub_mixintext-to-speech1 language (en)paper: 2305.09636