AI SeedbankHelp preserve open and free AI for humanity's future

← All models

microsoft_deberta-v3-large

microsoft · View on Hugging Face ↗

Get this model

Download TorrentMagnet Link

Seeders: · Leechers:

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


language: en tags:


DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing

DeBERTa improves the BERT and RoBERTa models using disentangled attention and enhanced mask decoder. With those two improvements, DeBERTa out perform RoBERTa on a majority of NLU tasks with 80GB training data.

In DeBERTa V3, we further improved the efficiency of DeBERTa using ELECTRA-Style pre-training with Gradient Disentangled Embedding Sharing. Compared to DeBERTa, our V3 version significantly improves the model performance on downstream tasks. You can find more technique details about the new model from our paper.

Please check the official repository for more implementation details and updates.

The DeBERTa V3 large model comes with 24 layers and a hidden size of 1024. It has 304M backbone parameters with a vocabulary containing 128K tokens which introduces 131M parameters in the Embedding layer. This model was trained using the 160GB data as DeBERTa V2.

Fine-tuning on NLU tasks

We present the dev results on SQuAD 2.0 and MNLI tasks.

Model Vocabulary(K) Backbone #Params(M) SQuAD 2.0(F1/EM) MNLI-m/mm(ACC)
RoBERTa-large 50 304 89.4/86.5 90.2
XLNet-large 32 - 90.6/87.9 90.8
DeBERTa-large 50 - 90.7/88.0 91.3
DeBERTa-v3-large 128 304 91.5/89.0 91.8/91.9

Fine-tuning with HF transformers

#!/bin/bash

cd transformers/examples/pytorch/text-classification/

pip install datasets
export TASK_NAME=mnli

output_dir="ds_results"

num_gpus=8

batch_size=8

python -m torch.distributed.launch --nproc_per_node=${num_gpus} \
  run_glue.py \
  --model_name_or_path microsoft/deberta-v3-large \
  --task_name $TASK_NAME \
  --do_train \
  --do_eval \
  --evaluation_strategy steps \
  --max_seq_length 256 \
  --warmup_steps 50 \
  --per_device_train_batch_size ${batch_size} \
  --learning_rate 6e-6 \
  --num_train_epochs 2 \
  --output_dir $output_dir \
  --overwrite_output_dir \
  --logging_steps 1000 \
  --logging_dir $output_dir

Citation

If you find DeBERTa useful for your work, please cite the following papers:

@misc{he2021debertav3,
      title={DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing}, 
      author={Pengcheng He and Jianfeng Gao and Weizhu Chen},
      year={2021},
      eprint={2111.09543},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}
@inproceedings{
he2021deberta,
title={DEBERTA: DECODING-ENHANCED BERT WITH DISENTANGLED ATTENTION},
author={Pengcheng He and Xiaodong Liu and Jianfeng Gao and Weizhu Chen},
booktitle={International Conference on Learning Representations},
year={2021},
url={https://openreview.net/forum?id=XPZIaotutsD}
}

Magnet link

Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:

magnet:?xt=urn:btih:1de4c35c612957fab6ed4747aaa53c9207d993e8&dn=microsoft_deberta-v3-large

Open magnet in torrent client · infohash 1de4c35c612957fab6ed4747aaa53c9207d993e8

Files & hashes

PathSizesha1sha256
README.md3.2 KB (3,266 B)5a967796d9eabc9eb149d951253d793ef8931771f6a88d8dba8986495c8274f7bc142c72919307982bb6b200a81fcf2acd509d17
config.json580 B (580 B)9d95ab0632a030f501c2e15336123041c556b8a7ddec8b81d079d218ce9e54fc0af5d1d5937d6d53b5d42e70c1f251a1cebc830d
generator_config.json560 B (560 B)97bd497a506de3bf306f52506fab62cdb954a26236956d8ce96c4a9f3acbb43501394c78756f75f491bb7288cff065ad7636e223
pytorch_model.bin833.2 MB (873,673,253 B)cf300f8c39fa569a2b7fa66a4918cf7e00c578fadd5b5d93e2db101aaf281df0ea1216c07ad73620ff59c5b42dccac4bf2eef5b5
pytorch_model.generator.bin544.8 MB (571,293,153 B)a01bbe664f95d38af5d28bc39b2f2473658cfdeaff85455c562822ea7001f810d026a68da8a24ffdae5a095081dfe7e84e27989d
spm.model2.4 MB (2,464,616 B)1993e578cb006883fd01014f831c6261e8136823c679fbf93643d19aab7ee10c0b99e460bdbc02fedf34b92b05af343b4af586fd
tokenizer_config.json52 B (52 B)acfd94e399c5659e4bed75f91b4ee24b111fc7a63f3978e0c036f2c2588cac34a6047cbb0af0b0dc1814254e291028529805496d

Cite this release

Canonical URL
https://aiseedbank.org/models/microsoft_deberta-v3-large/
Slug
microsoft_deberta-v3-large
Infohash
1de4c35c612957fab6ed4747aaa53c9207d993e8
License
mit
Signing key fingerprint
85a3b32c3712427b

Every file carries a locally computed sha256 — verify a download against the signed sums: microsoft_deberta-v3-large.SHA256SUMS (+ minisign signature).

Provenance

Upstream repositorymicrosoft/deberta-v3-large
Revision (pinned)64a8c8eab3e352a784c658aef62be1662607476f
Fetched at2026-09-04T02:42:17Z
License at fetchmit
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

✓ verified · rehash-vs-hf-metadata at 2026-09-04T02:42:34Z

mit1.35 GB (1,447,435,480 bytes)transformerspytorchdeberta-v2debertadeberta-v3fill-maskendpoints_compatible2 languages (tf, en)paper: 2006.03654paper: 2111.09543