AI SeedbankHelp preserve open and free AI for humanity's future

← All models

answerdotai_ModernBERT-large

answerdotai · View on Hugging Face ↗

Get this model

Download TorrentMagnet Link

Seeders: · Leechers:

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


library_name: transformers license: apache-2.0 language:

  • en tags:
  • fill-mask
  • masked-lm
  • long-context
  • modernbert pipeline_tag: fill-mask inference: false

ModernBERT

Table of Contents

  1. Model Summary
  2. Usage
  3. Evaluation
  4. Limitations
  5. Training
  6. License
  7. Citation

Model Summary

ModernBERT is a modernized bidirectional encoder-only Transformer model (BERT-style) pre-trained on 2 trillion tokens of English and code data with a native context length of up to 8,192 tokens. ModernBERT leverages recent architectural improvements such as:

  • Rotary Positional Embeddings (RoPE) for long-context support.
  • Local-Global Alternating Attention for efficiency on long inputs.
  • Unpadding and Flash Attention for efficient inference.

ModernBERT’s native long context length makes it ideal for tasks that require processing long documents, such as retrieval, classification, and semantic search within large corpora. The model was trained on a large corpus of text and code, making it suitable for a wide range of downstream tasks, including code retrieval and hybrid (text + code) semantic search.

It is available in the following sizes:

For more information about ModernBERT, we recommend our release blog post for a high-level overview, and our arXiv pre-print for in-depth information.

ModernBERT is a collaboration between Answer.AI, LightOn, and friends.

Usage

You can use these models directly with the transformers library starting from v4.48.0:

pip install -U transformers>=4.48.0

Since ModernBERT is a Masked Language Model (MLM), you can use the fill-mask pipeline or load it via AutoModelForMaskedLM. To use ModernBERT for downstream tasks like classification, retrieval, or QA, fine-tune it following standard BERT fine-tuning recipes.

⚠️ If your GPU supports it, we recommend using ModernBERT with Flash Attention 2 to reach the highest efficiency. To do so, install Flash Attention as follows, then use the model as normal:

pip install flash-attn

Using AutoModelForMaskedLM:

from transformers import AutoTokenizer, AutoModelForMaskedLM

model_id = "answerdotai/ModernBERT-base"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForMaskedLM.from_pretrained(model_id)

text = "The capital of France is [MASK]."
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)

# To get predictions for the mask:
masked_index = inputs["input_ids"][0].tolist().index(tokenizer.mask_token_id)
predicted_token_id = outputs.logits[0, masked_index].argmax(axis=-1)
predicted_token = tokenizer.decode(predicted_token_id)
print("Predicted token:", predicted_token)
# Predicted token:  Paris

Using a pipeline:

import torch
from transformers import pipeline
from pprint import pprint

pipe = pipeline(
    "fill-mask",
    model="answerdotai/ModernBERT-base",
    torch_dtype=torch.bfloat16,
)

input_text = "He walked to the [MASK]."
results = pipe(input_text)
pprint(results)

Note: ModernBERT does not use token type IDs, unlike some earlier BERT models. Most downstream usage is identical to standard BERT models on the Hugging Face Hub, except you can omit the token_type_ids parameter.

Evaluation

We evaluate ModernBERT across a range of tasks, including natural language understanding (GLUE), general retrieval (BEIR), long-context retrieval (MLDR), and code retrieval (CodeSearchNet and StackQA).

Key highlights:

  • On GLUE, ModernBERT-base surpasses other similarly-sized encoder models, and ModernBERT-large is second only to Deberta-v3-large.
  • For general retrieval tasks, ModernBERT performs well on BEIR in both single-vector (DPR-style) and multi-vector (ColBERT-style) settings.
  • Thanks to the inclusion of code data in its training mixture, ModernBERT as a backbone also achieves new state-of-the-art code retrieval results on CodeSearchNet and StackQA.

Base Models

Model IR (DPR) IR (DPR) IR (DPR) IR (ColBERT) IR (ColBERT) NLU Code Code
BEIR MLDR_OOD MLDR_ID BEIR MLDR_OOD GLUE CSN SQA
BERT 38.9 23.9 32.2 49.0 28.1 84.7 41.2 59.5
RoBERTa 37.7 22.9 32.8 48.7 28.2 86.4 44.3 59.6
DeBERTaV3 20.2 5.4 13.4 47.1 21.9 88.1 17.5 18.6
NomicBERT 41.0 26.7 30.3 49.9 61.3 84.0 41.6 61.4
GTE-en-MLM 41.4 34.3 44.4 48.2 69.3 85.6 44.9 71.4
ModernBERT 41.6 27.4 44.0 51.3 80.2 88.4 56.4 73.6

Large Models

Model IR (DPR) IR (DPR) IR (DPR) IR (ColBERT) IR (ColBERT) NLU Code Code
BEIR MLDR_OOD MLDR_ID BEIR MLDR_OOD GLUE CSN SQA
BERT 38.9 23.3 31.7 49.5 28.5 85.2 41.6 60.8
RoBERTa 41.4 22.6 36.1 49.8 28.8 88.9 47.3 68.1
DeBERTaV3 25.6 7.1 19.2 46.7 23.0 91.4 21.2 19.7
GTE-en-MLM 42.5 36.4 48.9 50.7 71.3 87.6 40.5 66.9
ModernBERT 44.0 34.3 48.6 52.4 80.4 90.4 59.5 83.9

Table 1: Results for all models across an overview of all tasks. CSN refers to CodeSearchNet and SQA to StackQA. MLDRID refers to in-domain (fine-tuned on the training set) evaluation, and MLDR_OOD to out-of-domain.

ModernBERT’s strong results, coupled with its efficient runtime on long-context inputs, demonstrate that encoder-only models can be significantly improved through modern architectural choices and extensive pretraining on diversified data sources.

Limitations

ModernBERT’s training data is primarily English and code, so performance may be lower for other languages. While it can handle long sequences efficiently, using the full 8,192 tokens window may be slower than short-context inference. Like any large language model, ModernBERT may produce representations that reflect biases present in its training data. Verify critical or sensitive outputs before relying on them.

Training

  • Architecture: Encoder-only, Pre-Norm Transformer with GeGLU activations.
  • Sequence Length: Pre-trained up to 1,024 tokens, then extended to 8,192 tokens.
  • Data: 2 trillion tokens of English text and code.
  • Optimizer: StableAdamW with trapezoidal LR scheduling and 1-sqrt decay.
  • Hardware: Trained on 8x H100 GPUs.

See the paper for more details.

License

We release the ModernBERT model architectures, model weights, training codebase under the Apache 2.0 license.

Citation

If you use ModernBERT in your work, please cite:

@misc{modernbert,
      title={Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference}, 
      author={Benjamin Warner and Antoine Chaffin and Benjamin Clavié and Orion Weller and Oskar Hallström and Said Taghadouini and Alexis Gallagher and Raja Biswas and Faisal Ladhak and Tom Aarsen and Nathan Cooper and Griffin Adams and Jeremy Howard and Iacopo Poli},
      year={2024},
      eprint={2412.13663},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2412.13663}, 
}

Magnet link

Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:

magnet:?xt=urn:btih:878b3e348330f997c6a6f4e0b0c0176eacbe467e&dn=answerdotai_ModernBERT-large

Open magnet in torrent client · infohash 878b3e348330f997c6a6f4e0b0c0176eacbe467e

Files & hashes

PathSizesha1sha256
README.md8.2 KB (8,427 B)8b9d77d5a470f7ef31f7495da1593826849b6fd640539466fa1117be4693b8f1fdfdac81fac3c5db472ff57bc05d013d48eafa82
config.json1.2 KB (1,194 B)f58291adedf9dc2fe5a5d80f8e2221bd1f135a2355dca68059e98f0debf8eb63ab355c451c65b1f5d90fdca4754113fdaadd2cca
model.safetensors1.47 GB (1,583,544,840 B)1748cc7e37d5a703343e0947e42ebae56e7e924e44510fec5d3a81a1877f225637b869495f18e55f6f23a09abb9be0acc030295f
pytorch_model.bin1.47 GB (1,583,581,182 B)017a40a93ba2732ce55efbb6d9abafcf4b6e5bc1fabb014a0b62d28ec5f2de530d2ade920b9293f170a4ec38acfa1b8096531adc
special_tokens_map.json694 B (694 B)6bf65ea14185a88eef7e24b41c86a953a58dcbadea97ecdbcc73713039d8d64dbb05e3689495c96657fbd9a18f5bed381be81049
tokenizer.json2.0 MB (2,132,967 B)de3f3e316b66f3885df49ab2c0fff3e477ff16809fd55248d51d33976b324fc11592e28071da7d41e0e9401dfb7082e30574b7b1
tokenizer_config.json20.3 KB (20,810 B)d78aa50c890c167d8700301e5580164466e51fa53cd2017ff46d0a527e5d39cae39272eccfa1f19bb9f89b05d166aab2e38354e2

Cite this release

Canonical URL
https://aiseedbank.org/models/answerdotai_ModernBERT-large/
Slug
answerdotai_ModernBERT-large
Infohash
878b3e348330f997c6a6f4e0b0c0176eacbe467e
License
apache-2.0
Signing key fingerprint
85a3b32c3712427b

Every file carries a locally computed sha256 — verify a download against the signed sums: answerdotai_ModernBERT-large.SHA256SUMS (+ minisign signature).

Provenance

Upstream repositoryanswerdotai/ModernBERT-large
Revision (pinned)45bb4654a4d5aaff24dd11d4781fa46d39bf8c13
Fetched at2026-09-03T20:55:24Z
License at fetchapache-2.0
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

✓ verified · rehash-vs-hf-metadata at 2026-09-03T20:56:01Z

apache-2.02.95 GB (3,169,290,114 bytes)transformerspytorchonnxsafetensorsmodernbertfill-maskmasked-lmlong-context1 language (en)paper: 2412.13663