Help preserve open and free AI for humanity's future

← All models

KBLab_wav2vec2-large-voxrex-swedish

KBLab · View on Hugging Face ↗

Swedish speech recognition (wav2vec 2.0 large, VoxRex data) from the KBLab national-model program.

✓ verified · rehash-vs-hf-metadata at 2026-08-24T02:51:05Z

cc0-1.02.35 GB (2,524,448,099 bytes)transformerspytorchsafetensorswav2vec2automatic-speech-recognitionaudiospeechhf-asr-leaderboardmodel-indexendpoints_compatible1 language (sv)paper: 2205.03026

Get this model

Download KBLab_wav2vec2-large-voxrex-swedish.torrent

Recommended — the .torrent carries the webseed url-list, so your client can fall back to plain HTTPS if the swarm is thin. See/verify for the full download + verification walkthrough.

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


language: sv arxiv: https://arxiv.org/abs/2205.03026 datasets:

  • common_voice
  • NST_Swedish_ASR_Database
  • P4 metrics:
  • wer tags:
  • audio
  • automatic-speech-recognition
  • speech
  • hf-asr-leaderboard license: cc0-1.0 model-index:
  • name: Wav2vec 2.0 large VoxRex Swedish results:
    • task: name: Speech Recognition type: automatic-speech-recognition dataset: name: Common Voice type: common_voice args: sv-SE metrics:
      • name: Test WER type: wer value: 8.49

Wav2vec 2.0 large VoxRex Swedish (C)

Finetuned version of KBs VoxRex large model using Swedish radio broadcasts, NST and Common Voice data. Evalutation without a language model gives the following: WER for NST + Common Voice test set (2% of total sentences) is 2.5%. WER for Common Voice test set is 8.49% directly and 7.37% with a 4-gram language model.

When using this model, make sure that your speech input is sampled at 16kHz.

Update 2022-01-10: Updated to VoxRex-C version.

Update 2022-05-16: Paper is is here.

Performance*

*Chart shows performance without the additional 20k steps of Common Voice fine-tuning

Training

This model has been fine-tuned for 120000 updates on NST + CommonVoice and then for an additional 20000 updates on CommonVoice only. The additional fine-tuning on CommonVoice hurts performance on the NST+CommonVoice test set somewhat and, unsurprisingly, improves it on the CommonVoice test set. It seems to perform generally better though [citation needed].

Usage

The model can be used directly (without a language model) as follows:

import torch
import torchaudio
from datasets import load_dataset
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
test_dataset = load_dataset("common_voice", "sv-SE", split="test[:2%]").
processor = Wav2Vec2Processor.from_pretrained("KBLab/wav2vec2-large-voxrex-swedish")
model = Wav2Vec2ForCTC.from_pretrained("KBLab/wav2vec2-large-voxrex-swedish")
resampler = torchaudio.transforms.Resample(48_000, 16_000)
# Preprocessing the datasets.
# We need to read the aduio files as arrays
def speech_file_to_array_fn(batch):
    speech_array, sampling_rate = torchaudio.load(batch["path"])
    batch["speech"] = resampler(speech_array).squeeze().numpy()
    return batch
test_dataset = test_dataset.map(speech_file_to_array_fn)
inputs = processor(test_dataset["speech"][:2], sampling_rate=16_000, return_tensors="pt", padding=True)
with torch.no_grad():
    logits = model(inputs.input_values, attention_mask=inputs.attention_mask).logits
predicted_ids = torch.argmax(logits, dim=-1)
print("Prediction:", processor.batch_decode(predicted_ids))
print("Reference:", test_dataset["sentence"][:2])

Citation

https://arxiv.org/abs/2205.03026

@inproceedings{malmsten2022hearing,
  title={Hearing voices at the national library : a speech corpus and acoustic model for the Swedish language},
  author={Malmsten, Martin and Haffenden, Chris and B{\"o}rjeson, Love},
  booktitle={Proceeding of Fonetik 2022 : Speech, Music and Hearing Quarterly Progress and Status Report, TMH-QPSR},
  volume={3},
  year={2022}
}

Magnet link (secondary — no webseeds)

Opens the swarm directly, but carries no webseed url-list. Prefer the.torrent download above — HTTP fallback seeds ride inside it.

magnet:?xt=urn:btih:bcc8082f0e6609714d1e07db418ede46e3f72334&dn=KBLab_wav2vec2-large-voxrex-swedish

Open magnet in torrent client · infohash bcc8082f0e6609714d1e07db418ede46e3f72334

Files & hashes

PathSizeMethodHash
README.md3.3 KB (3,377 B)sha1-git-blob002762cb2ded2ad4544b3b38212fce7881f3499f
chart_1.svg185.1 KB (189,589 B)sha1-git-blob00737aa745a5aa86fd718a4b07e29615aa09f536
comparison.png146.5 KB (150,062 B)sha1-git-blobbdcedebace678b88f689894c1406015176bcf683
config.json1.7 KB (1,757 B)sha1-git-blobd56120374ea7409712e6a26c776d6f762fb44a5b
model.safetensors1.18 GB (1,261,996,032 B)sha256-lfsd09f24c48fdb07865246e3cf06fc47c881ce911f02a6983bce70595ba417e02b
preprocessor_config.json212 B (212 B)sha1-git-blob36ebe8b7c1cc967b3059f0494ae8a1069dd67655
pytorch_model.bin1.18 GB (1,262,106,353 B)sha256-lfs7138b3f9c5700388ddbd8b08a194290529778aba95327e7445a621cbe24ef508
special_tokens_map.json85 B (85 B)sha1-git-blob25bc39604f72700b3b8e10bd69bb2f227157edd1
tokenizer_config.json211 B (211 B)sha1-git-blob92169da5ef839418d20bd11deb1e1db05e74bf35
vocab.json421 B (421 B)sha1-git-blob6c477f44f1c66dbdea44c2e0781ba0aeb4fa53f6

Provenance

Upstream repositoryKBLab/wav2vec2-large-voxrex-swedish
Revision (pinned)ca70e31c06a2617bf7fe3b4fb5d387d2b19b2983
Fetched at2026-08-24T02:48:31Z
License at fetchcc0-1.0
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

Webseeds