AI SeedbankHelp preserve open and free AI for humanity's future

← All models

cardiffnlp_twitter-roberta-base-sentiment

cardiffnlp · View on Hugging Face ↗

Get this model

Download TorrentMagnet Link

Seeders: · Leechers:

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


datasets:

  • tweet_eval language:
  • en

Twitter-roBERTa-base for Sentiment Analysis

This is a roBERTa-base model trained on ~58M tweets and finetuned for sentiment analysis with the TweetEval benchmark. This model is suitable for English (for a similar multilingual model, see XLM-T).

  • Reference Paper: TweetEval (Findings of EMNLP 2020).
  • Git Repo: Tweeteval official repository.

Labels: 0 -> Negative; 1 -> Neutral; 2 -> Positive

New! We just released a new sentiment analysis model trained on more recent and a larger quantity of tweets. See twitter-roberta-base-sentiment-latest and TweetNLP for more details.

Example of classification

from transformers import AutoModelForSequenceClassification
from transformers import TFAutoModelForSequenceClassification
from transformers import AutoTokenizer
import numpy as np
from scipy.special import softmax
import csv
import urllib.request

# Preprocess text (username and link placeholders)
def preprocess(text):
    new_text = []
 
 
    for t in text.split(" "):
        t = '@user' if t.startswith('@') and len(t) > 1 else t
        t = 'http' if t.startswith('http') else t
        new_text.append(t)
    return " ".join(new_text)

# Tasks:
# emoji, emotion, hate, irony, offensive, sentiment
# stance/abortion, stance/atheism, stance/climate, stance/feminist, stance/hillary

task='sentiment'
MODEL = f"cardiffnlp/twitter-roberta-base-{task}"

tokenizer = AutoTokenizer.from_pretrained(MODEL)

# download label mapping
labels=[]
mapping_link = f"https://raw.githubusercontent.com/cardiffnlp/tweeteval/main/datasets/{task}/mapping.txt"
with urllib.request.urlopen(mapping_link) as f:
    html = f.read().decode('utf-8').split("\n")
    csvreader = csv.reader(html, delimiter='\t')
labels = [row[1] for row in csvreader if len(row) > 1]

# PT
model = AutoModelForSequenceClassification.from_pretrained(MODEL)
model.save_pretrained(MODEL)

text = "Good night 😊"
text = preprocess(text)
encoded_input = tokenizer(text, return_tensors='pt')
output = model(**encoded_input)
scores = output[0][0].detach().numpy()
scores = softmax(scores)

# # TF
# model = TFAutoModelForSequenceClassification.from_pretrained(MODEL)
# model.save_pretrained(MODEL)

# text = "Good night 😊"
# encoded_input = tokenizer(text, return_tensors='tf')
# output = model(encoded_input)
# scores = output[0][0].numpy()
# scores = softmax(scores)

ranking = np.argsort(scores)
ranking = ranking[::-1]
for i in range(scores.shape[0]):
    l = labels[ranking[i]]
    s = scores[ranking[i]]
    print(f"{i+1}) {l} {np.round(float(s), 4)}")

Output:

1) positive 0.8466
2) neutral 0.1458
3) negative 0.0076

BibTeX entry and citation info

Please cite the reference paper if you use this model.

@inproceedings{barbieri-etal-2020-tweeteval,
    title = "{T}weet{E}val: Unified Benchmark and Comparative Evaluation for Tweet Classification",
    author = "Barbieri, Francesco  and
      Camacho-Collados, Jose  and
      Espinosa Anke, Luis  and
      Neves, Leonardo",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020",
    month = nov,
    year = "2020",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2020.findings-emnlp.148",
    doi = "10.18653/v1/2020.findings-emnlp.148",
    pages = "1644--1650"
}

Magnet link

Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:

magnet:?xt=urn:btih:97aa220234dac3cca360aad5bd703ab33f30e8c4&dn=cardiffnlp_twitter-roberta-base-sentiment

Open magnet in torrent client · infohash 97aa220234dac3cca360aad5bd703ab33f30e8c4

Files & hashes

PathSizesha1sha256
README.md3.6 KB (3,723 B)47cd107247eba114e9abca9f158e4ad098b31e27068f1528bb7f006dbd7d95884387db548908b87ae26f8d21f130428b69009927
config.json747 B (747 B)f1e21eb05e56dc207001d6684f988a75a29e66b1de9ed9df04ecd8c267564e60b73a216f7cc7951abf85b797cb717216a5eeba50
merges.txt445.6 KB (456,318 B)226b0752cac7789c48f0cb3ec53eda48b7be36cc1ce1664773c50f3e0cc8842619a93edc4624525b728b188a9e0be33b7726adc5
pytorch_model.bin475.6 MB (498,679,497 B)72c1cdab9ecdb7d462093275c19c022dcb4e4a18c37a3484c55954cd75b336a85f1e0c023ae874f3a73b05d2418dd04828e293b1
special_tokens_map.json150 B (150 B)6cd1d9021e10d47aed59399af6b0e30312b46ca47638f5bbbe86ef6d604ef28ad3647dc690d6d117c81c0d63e885416be8da1150
vocab.json877.8 KB (898,822 B)0a39732b2d8be8e493cab3da68b68cc3e28221de06b4d46c8e752d410213d9548eb27a54db70fda0319b6271fb8d59dead5e1cab

Cite this release

Canonical URL
https://aiseedbank.org/models/cardiffnlp_twitter-roberta-base-sentiment/
Slug
cardiffnlp_twitter-roberta-base-sentiment
Infohash
97aa220234dac3cca360aad5bd703ab33f30e8c4
License
no license recorded
Signing key fingerprint
85a3b32c3712427b

Every file carries a locally computed sha256 — verify a download against the signed sums: cardiffnlp_twitter-roberta-base-sentiment.SHA256SUMS (+ minisign signature).

Provenance

Upstream repositorycardiffnlp/twitter-roberta-base-sentiment
Revision (pinned)daefdd1f6ae931839bce4d0f3db0a1a4265cd50f
Fetched at2026-09-03T21:11:16Z
License at fetchno license recorded
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

✓ verified · rehash-vs-hf-metadata at 2026-09-03T21:11:23Z

no license recorded476.9 MB (500,039,257 bytes)transformerspytorchjaxrobertatext-classificationendpoints_compatible2 languages (tf, en)paper: 2010.12421