Harveenchadha_vakyansh-wav2vec2-tamil-tam-250
Harveenchadha · View on Hugging Face ↗
Tamil speech recognition (wav2vec 2.0, 250 hours) from the Vakyansh open-speech initiative.
✓ verified · rehash-vs-hf-metadata at 2026-08-24T03:03:00Z
mit360.2 MB (377,738,686 bytes)transformerspytorchwav2vec2automatic-speech-recognitionaudiospeechmodel-indexendpoints_compatible1 language (ta)paper: 2107.07402
Get this model
Download Harveenchadha_vakyansh-wav2vec2-tamil-tam-250.torrent
Recommended — the .torrent carries the webseed url-list, so your client can fall back to plain HTTPS if the swarm is thin. See/verify for the full download + verification walkthrough.
Model card
The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.
language: ta #datasets: #- Interspeech 2021 metrics:
- wer tags:
- audio
- automatic-speech-recognition
- speech license: mit model-index:
- name: Wav2Vec2 Vakyansh Tamil Model by Harveen Chadha
results:
- task:
name: Speech Recognition
type: automatic-speech-recognition
dataset:
name: Common Voice ta
type: common_voice
args: ta
metrics:
- name: Test WER type: wer value: 53.64
- task:
name: Speech Recognition
type: automatic-speech-recognition
dataset:
name: Common Voice ta
type: common_voice
args: ta
metrics:
Pretrained Model
Fine-tuned on Multilingual Pretrained Model CLSRIL-23. The original fairseq checkpoint is present here. When using this model, make sure that your speech input is sampled at 16kHz.
Note: The result from this model is without a language model so you may witness a higher WER in some cases.
Dataset
This model was trained on 4200 hours of Hindi Labelled Data. The labelled data is not present in public domain as of now.
Training Script
Models were trained using experimental platform setup by Vakyansh team at Ekstep. Here is the training repository.
In case you want to explore training logs on wandb they are here.
Colab Demo
Usage
The model can be used directly (without a language model) as follows:
import soundfile as sf
import torch
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
import argparse
def parse_transcription(wav_file):
# load pretrained model
processor = Wav2Vec2Processor.from_pretrained("Harveenchadha/vakyansh-wav2vec2-tamil-tam-250")
model = Wav2Vec2ForCTC.from_pretrained("Harveenchadha/vakyansh-wav2vec2-tamil-tam-250")
# load audio
audio_input, sample_rate = sf.read(wav_file)
# pad input values and return pt tensor
input_values = processor(audio_input, sampling_rate=sample_rate, return_tensors="pt").input_values
# INFERENCE
# retrieve logits & take argmax
logits = model(input_values).logits
predicted_ids = torch.argmax(logits, dim=-1)
# transcribe
transcription = processor.decode(predicted_ids[0], skip_special_tokens=True)
print(transcription)
Evaluation
The model can be evaluated as follows on the hindi test data of Common Voice.
import torch
import torchaudio
from datasets import load_dataset, load_metric
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
import re
test_dataset = load_dataset("common_voice", "ta", split="test")
wer = load_metric("wer")
processor = Wav2Vec2Processor.from_pretrained("Harveenchadha/vakyansh-wav2vec2-tamil-tam-250")
model = Wav2Vec2ForCTC.from_pretrained("Harveenchadha/vakyansh-wav2vec2-tamil-tam-250")
model.to("cuda")
resampler = torchaudio.transforms.Resample(48_000, 16_000)
chars_to_ignore_regex = '[\,\?\.\!\-\;\:\"\“]'
# Preprocessing the datasets.
# We need to read the aduio files as arrays
def speech_file_to_array_fn(batch):
batch["sentence"] = re.sub(chars_to_ignore_regex, '', batch["sentence"]).lower()
speech_array, sampling_rate = torchaudio.load(batch["path"])
batch["speech"] = resampler(speech_array).squeeze().numpy()
return batch
test_dataset = test_dataset.map(speech_file_to_array_fn)
# Preprocessing the datasets.
# We need to read the aduio files as arrays
def evaluate(batch):
inputs = processor(batch["speech"], sampling_rate=16_000, return_tensors="pt", padding=True)
with torch.no_grad():
logits = model(inputs.input_values.to("cuda")).logits
pred_ids = torch.argmax(logits, dim=-1)
batch["pred_strings"] = processor.batch_decode(pred_ids, skip_special_tokens=True)
return batch
result = test_dataset.map(evaluate, batched=True, batch_size=8)
print("WER: {:2f}".format(100 * wer.compute(predictions=result["pred_strings"], references=result["sentence"])))
Test Result: 53.64 %
Colab Evaluation
Credits
Thanks to Ekstep Foundation for making this possible. The vakyansh team will be open sourcing speech models in all the Indic Languages.
Magnet link (secondary — no webseeds)
Opens the swarm directly, but carries no webseed url-list. Prefer the.torrent download above — HTTP fallback seeds ride inside it.
magnet:?xt=urn:btih:76f041c846da6f72a55aaf64086f0d35aba128a4&dn=Harveenchadha_vakyansh-wav2vec2-tamil-tam-250Open magnet in torrent client · infohash 76f041c846da6f72a55aaf64086f0d35aba128a4
Files & hashes
| Path | Size | Method | Hash |
|---|---|---|---|
| README.md | 4.3 KB (4,366 B) | sha1-git-blob | f1eba299bb919388056a0a432c0030d9eb9e1c1e |
| config.json | 1.6 KB (1,658 B) | sha1-git-blob | 84640d660b095c75f6aba1c0ac4f9262da12d0a6 |
| preprocessor_config.json | 213 B (213 B) | sha1-git-blob | 8df8da1de6563b3f11638f4df5f2336f4ca94c04 |
| pytorch_model.bin | 360.2 MB (377,731,607 B) | sha256-lfs | 00adef7dc21c04f9014807f73c21c5fa12f9284aed6a5494a72bfc42ff3e55e7 |
| special_tokens_map.json | 85 B (85 B) | sha1-git-blob | 25bc39604f72700b3b8e10bd69bb2f227157edd1 |
| tokenizer_config.json | 181 B (181 B) | sha1-git-blob | 94d9657079dc7dd16de16f9abeb1878e5a592f47 |
| vocab.json | 576 B (576 B) | sha1-git-blob | 4fbee556e70fb8ef63e6bce2e53aed6f286dbba8 |
Provenance
| Upstream repository | Harveenchadha/vakyansh-wav2vec2-tamil-tam-250 |
|---|---|
| Revision (pinned) | 0bd7c7d87da18a71b246ce3e543244bdba983e36 |
| Fetched at | 2026-08-24T03:02:51Z |
| License at fetch | mit |
| Snapshot tool | huggingface · seedbank 0.1.0 |
Trackers
- udp://announce.aitorrent.org:6969/announce
- http://announce.aitorrent.org:7070/announce
- udp://announce2.aitorrent.org:6970/announce
- http://announce2.aitorrent.org:7071/announce
- udp://tracker.opentrackr.org:1337/announce
- udp://open.demonii.com:1337/announce
- udp://open.stealth.si:80/announce
- udp://exodus.desync.com:6969/announce
- udp://tracker.torrent.eu.org:451/announce