Help preserve open and free AI for humanity's future

← All models

Harveenchadha_vakyansh-wav2vec2-tamil-tam-250

Harveenchadha · View on Hugging Face ↗

Tamil speech recognition (wav2vec 2.0, 250 hours) from the Vakyansh open-speech initiative.

✓ verified · rehash-vs-hf-metadata at 2026-08-24T03:03:00Z

mit360.2 MB (377,738,686 bytes)transformerspytorchwav2vec2automatic-speech-recognitionaudiospeechmodel-indexendpoints_compatible1 language (ta)paper: 2107.07402

Get this model

Download Harveenchadha_vakyansh-wav2vec2-tamil-tam-250.torrent

Recommended — the .torrent carries the webseed url-list, so your client can fall back to plain HTTPS if the swarm is thin. See/verify for the full download + verification walkthrough.

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


language: ta #datasets: #- Interspeech 2021 metrics:

  • wer tags:
  • audio
  • automatic-speech-recognition
  • speech license: mit model-index:
  • name: Wav2Vec2 Vakyansh Tamil Model by Harveen Chadha results:
    • task: name: Speech Recognition type: automatic-speech-recognition dataset: name: Common Voice ta type: common_voice args: ta metrics:
      • name: Test WER type: wer value: 53.64

Pretrained Model

Fine-tuned on Multilingual Pretrained Model CLSRIL-23. The original fairseq checkpoint is present here. When using this model, make sure that your speech input is sampled at 16kHz.

Note: The result from this model is without a language model so you may witness a higher WER in some cases.

Dataset

This model was trained on 4200 hours of Hindi Labelled Data. The labelled data is not present in public domain as of now.

Training Script

Models were trained using experimental platform setup by Vakyansh team at Ekstep. Here is the training repository.

In case you want to explore training logs on wandb they are here.

Colab Demo

Usage

The model can be used directly (without a language model) as follows:

import soundfile as sf
import torch
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
import argparse

def parse_transcription(wav_file):
    # load pretrained model
    processor = Wav2Vec2Processor.from_pretrained("Harveenchadha/vakyansh-wav2vec2-tamil-tam-250")
    model = Wav2Vec2ForCTC.from_pretrained("Harveenchadha/vakyansh-wav2vec2-tamil-tam-250")

    # load audio
    audio_input, sample_rate = sf.read(wav_file)

    # pad input values and return pt tensor
    input_values = processor(audio_input, sampling_rate=sample_rate, return_tensors="pt").input_values

    # INFERENCE
    # retrieve logits & take argmax
    logits = model(input_values).logits
    predicted_ids = torch.argmax(logits, dim=-1)

    # transcribe
    transcription = processor.decode(predicted_ids[0], skip_special_tokens=True)
    print(transcription)

Evaluation

The model can be evaluated as follows on the hindi test data of Common Voice.


import torch
import torchaudio
from datasets import load_dataset, load_metric
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
import re

test_dataset = load_dataset("common_voice", "ta", split="test")
wer = load_metric("wer")

processor = Wav2Vec2Processor.from_pretrained("Harveenchadha/vakyansh-wav2vec2-tamil-tam-250")
model = Wav2Vec2ForCTC.from_pretrained("Harveenchadha/vakyansh-wav2vec2-tamil-tam-250")
model.to("cuda")

resampler = torchaudio.transforms.Resample(48_000, 16_000)

chars_to_ignore_regex = '[\,\?\.\!\-\;\:\"\“]'

# Preprocessing the datasets.
# We need to read the aduio files as arrays
def speech_file_to_array_fn(batch):
  batch["sentence"] = re.sub(chars_to_ignore_regex, '', batch["sentence"]).lower()
  speech_array, sampling_rate = torchaudio.load(batch["path"])
  batch["speech"] = resampler(speech_array).squeeze().numpy()
  return batch

test_dataset = test_dataset.map(speech_file_to_array_fn)

# Preprocessing the datasets.
# We need to read the aduio files as arrays
def evaluate(batch):
  inputs = processor(batch["speech"], sampling_rate=16_000, return_tensors="pt", padding=True)

  with torch.no_grad():
      logits = model(inputs.input_values.to("cuda")).logits

      pred_ids = torch.argmax(logits, dim=-1)
      batch["pred_strings"] = processor.batch_decode(pred_ids, skip_special_tokens=True)
      return batch

result = test_dataset.map(evaluate, batched=True, batch_size=8)

print("WER: {:2f}".format(100 * wer.compute(predictions=result["pred_strings"], references=result["sentence"])))

Test Result: 53.64 %

Colab Evaluation

Credits

Thanks to Ekstep Foundation for making this possible. The vakyansh team will be open sourcing speech models in all the Indic Languages.

Magnet link (secondary — no webseeds)

Opens the swarm directly, but carries no webseed url-list. Prefer the.torrent download above — HTTP fallback seeds ride inside it.

magnet:?xt=urn:btih:76f041c846da6f72a55aaf64086f0d35aba128a4&dn=Harveenchadha_vakyansh-wav2vec2-tamil-tam-250

Open magnet in torrent client · infohash 76f041c846da6f72a55aaf64086f0d35aba128a4

Files & hashes

PathSizeMethodHash
README.md4.3 KB (4,366 B)sha1-git-blobf1eba299bb919388056a0a432c0030d9eb9e1c1e
config.json1.6 KB (1,658 B)sha1-git-blob84640d660b095c75f6aba1c0ac4f9262da12d0a6
preprocessor_config.json213 B (213 B)sha1-git-blob8df8da1de6563b3f11638f4df5f2336f4ca94c04
pytorch_model.bin360.2 MB (377,731,607 B)sha256-lfs00adef7dc21c04f9014807f73c21c5fa12f9284aed6a5494a72bfc42ff3e55e7
special_tokens_map.json85 B (85 B)sha1-git-blob25bc39604f72700b3b8e10bd69bb2f227157edd1
tokenizer_config.json181 B (181 B)sha1-git-blob94d9657079dc7dd16de16f9abeb1878e5a592f47
vocab.json576 B (576 B)sha1-git-blob4fbee556e70fb8ef63e6bce2e53aed6f286dbba8

Provenance

Upstream repositoryHarveenchadha/vakyansh-wav2vec2-tamil-tam-250
Revision (pinned)0bd7c7d87da18a71b246ce3e543244bdba983e36
Fetched at2026-08-24T03:02:51Z
License at fetchmit
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

Webseeds