KBLab_wav2vec2-large-voxrex-swedish
KBLab · View on Hugging Face ↗
Swedish speech recognition (wav2vec 2.0 large, VoxRex data) from the KBLab national-model program.
✓ verified · rehash-vs-hf-metadata at 2026-08-24T02:51:05Z
cc0-1.02.35 GB (2,524,448,099 bytes)transformerspytorchsafetensorswav2vec2automatic-speech-recognitionaudiospeechhf-asr-leaderboardmodel-indexendpoints_compatible1 language (sv)paper: 2205.03026
Get this model
Download KBLab_wav2vec2-large-voxrex-swedish.torrent
Recommended — the .torrent carries the webseed url-list, so your client can fall back to plain HTTPS if the swarm is thin. See/verify for the full download + verification walkthrough.
Model card
The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.
language: sv arxiv: https://arxiv.org/abs/2205.03026 datasets:
- common_voice
- NST_Swedish_ASR_Database
- P4 metrics:
- wer tags:
- audio
- automatic-speech-recognition
- speech
- hf-asr-leaderboard license: cc0-1.0 model-index:
- name: Wav2vec 2.0 large VoxRex Swedish
results:
- task:
name: Speech Recognition
type: automatic-speech-recognition
dataset:
name: Common Voice
type: common_voice
args: sv-SE
metrics:
- name: Test WER type: wer value: 8.49
- task:
name: Speech Recognition
type: automatic-speech-recognition
dataset:
name: Common Voice
type: common_voice
args: sv-SE
metrics:
Wav2vec 2.0 large VoxRex Swedish (C)
Finetuned version of KBs VoxRex large model using Swedish radio broadcasts, NST and Common Voice data. Evalutation without a language model gives the following: WER for NST + Common Voice test set (2% of total sentences) is 2.5%. WER for Common Voice test set is 8.49% directly and 7.37% with a 4-gram language model.
When using this model, make sure that your speech input is sampled at 16kHz.
Update 2022-01-10: Updated to VoxRex-C version.
Update 2022-05-16: Paper is is here.
Performance*
*Chart shows performance without the additional 20k steps of Common Voice fine-tuningTraining
This model has been fine-tuned for 120000 updates on NST + CommonVoice and then for an additional 20000 updates on CommonVoice only. The additional fine-tuning on CommonVoice hurts performance on the NST+CommonVoice test set somewhat and, unsurprisingly, improves it on the CommonVoice test set. It seems to perform generally better though [citation needed].
Usage
The model can be used directly (without a language model) as follows:
import torch
import torchaudio
from datasets import load_dataset
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
test_dataset = load_dataset("common_voice", "sv-SE", split="test[:2%]").
processor = Wav2Vec2Processor.from_pretrained("KBLab/wav2vec2-large-voxrex-swedish")
model = Wav2Vec2ForCTC.from_pretrained("KBLab/wav2vec2-large-voxrex-swedish")
resampler = torchaudio.transforms.Resample(48_000, 16_000)
# Preprocessing the datasets.
# We need to read the aduio files as arrays
def speech_file_to_array_fn(batch):
speech_array, sampling_rate = torchaudio.load(batch["path"])
batch["speech"] = resampler(speech_array).squeeze().numpy()
return batch
test_dataset = test_dataset.map(speech_file_to_array_fn)
inputs = processor(test_dataset["speech"][:2], sampling_rate=16_000, return_tensors="pt", padding=True)
with torch.no_grad():
logits = model(inputs.input_values, attention_mask=inputs.attention_mask).logits
predicted_ids = torch.argmax(logits, dim=-1)
print("Prediction:", processor.batch_decode(predicted_ids))
print("Reference:", test_dataset["sentence"][:2])
Citation
https://arxiv.org/abs/2205.03026
@inproceedings{malmsten2022hearing,
title={Hearing voices at the national library : a speech corpus and acoustic model for the Swedish language},
author={Malmsten, Martin and Haffenden, Chris and B{\"o}rjeson, Love},
booktitle={Proceeding of Fonetik 2022 : Speech, Music and Hearing Quarterly Progress and Status Report, TMH-QPSR},
volume={3},
year={2022}
}
Magnet link (secondary — no webseeds)
Opens the swarm directly, but carries no webseed url-list. Prefer the.torrent download above — HTTP fallback seeds ride inside it.
magnet:?xt=urn:btih:bcc8082f0e6609714d1e07db418ede46e3f72334&dn=KBLab_wav2vec2-large-voxrex-swedishOpen magnet in torrent client · infohash bcc8082f0e6609714d1e07db418ede46e3f72334
Files & hashes
| Path | Size | Method | Hash |
|---|---|---|---|
| README.md | 3.3 KB (3,377 B) | sha1-git-blob | 002762cb2ded2ad4544b3b38212fce7881f3499f |
| chart_1.svg | 185.1 KB (189,589 B) | sha1-git-blob | 00737aa745a5aa86fd718a4b07e29615aa09f536 |
| comparison.png | 146.5 KB (150,062 B) | sha1-git-blob | bdcedebace678b88f689894c1406015176bcf683 |
| config.json | 1.7 KB (1,757 B) | sha1-git-blob | d56120374ea7409712e6a26c776d6f762fb44a5b |
| model.safetensors | 1.18 GB (1,261,996,032 B) | sha256-lfs | d09f24c48fdb07865246e3cf06fc47c881ce911f02a6983bce70595ba417e02b |
| preprocessor_config.json | 212 B (212 B) | sha1-git-blob | 36ebe8b7c1cc967b3059f0494ae8a1069dd67655 |
| pytorch_model.bin | 1.18 GB (1,262,106,353 B) | sha256-lfs | 7138b3f9c5700388ddbd8b08a194290529778aba95327e7445a621cbe24ef508 |
| special_tokens_map.json | 85 B (85 B) | sha1-git-blob | 25bc39604f72700b3b8e10bd69bb2f227157edd1 |
| tokenizer_config.json | 211 B (211 B) | sha1-git-blob | 92169da5ef839418d20bd11deb1e1db05e74bf35 |
| vocab.json | 421 B (421 B) | sha1-git-blob | 6c477f44f1c66dbdea44c2e0781ba0aeb4fa53f6 |
Provenance
| Upstream repository | KBLab/wav2vec2-large-voxrex-swedish |
|---|---|
| Revision (pinned) | ca70e31c06a2617bf7fe3b4fb5d387d2b19b2983 |
| Fetched at | 2026-08-24T02:48:31Z |
| License at fetch | cc0-1.0 |
| Snapshot tool | huggingface · seedbank 0.1.0 |
Trackers
- udp://announce.aitorrent.org:6969/announce
- http://announce.aitorrent.org:7070/announce
- udp://announce2.aitorrent.org:6970/announce
- http://announce2.aitorrent.org:7071/announce
- udp://tracker.opentrackr.org:1337/announce
- udp://open.demonii.com:1337/announce
- udp://open.stealth.si:80/announce
- udp://exodus.desync.com:6969/announce
- udp://tracker.torrent.eu.org:451/announce