cardiffnlp_twitter-roberta-base-sentiment
cardiffnlp · View on Hugging Face ↗
Model card
The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.
datasets:
- tweet_eval language:
- en
Twitter-roBERTa-base for Sentiment Analysis
This is a roBERTa-base model trained on ~58M tweets and finetuned for sentiment analysis with the TweetEval benchmark. This model is suitable for English (for a similar multilingual model, see XLM-T).
- Reference Paper: TweetEval (Findings of EMNLP 2020).
- Git Repo: Tweeteval official repository.
Labels: 0 -> Negative; 1 -> Neutral; 2 -> Positive
New! We just released a new sentiment analysis model trained on more recent and a larger quantity of tweets. See twitter-roberta-base-sentiment-latest and TweetNLP for more details.
Example of classification
from transformers import AutoModelForSequenceClassification
from transformers import TFAutoModelForSequenceClassification
from transformers import AutoTokenizer
import numpy as np
from scipy.special import softmax
import csv
import urllib.request
# Preprocess text (username and link placeholders)
def preprocess(text):
new_text = []
for t in text.split(" "):
t = '@user' if t.startswith('@') and len(t) > 1 else t
t = 'http' if t.startswith('http') else t
new_text.append(t)
return " ".join(new_text)
# Tasks:
# emoji, emotion, hate, irony, offensive, sentiment
# stance/abortion, stance/atheism, stance/climate, stance/feminist, stance/hillary
task='sentiment'
MODEL = f"cardiffnlp/twitter-roberta-base-{task}"
tokenizer = AutoTokenizer.from_pretrained(MODEL)
# download label mapping
labels=[]
mapping_link = f"https://raw.githubusercontent.com/cardiffnlp/tweeteval/main/datasets/{task}/mapping.txt"
with urllib.request.urlopen(mapping_link) as f:
html = f.read().decode('utf-8').split("\n")
csvreader = csv.reader(html, delimiter='\t')
labels = [row[1] for row in csvreader if len(row) > 1]
# PT
model = AutoModelForSequenceClassification.from_pretrained(MODEL)
model.save_pretrained(MODEL)
text = "Good night 😊"
text = preprocess(text)
encoded_input = tokenizer(text, return_tensors='pt')
output = model(**encoded_input)
scores = output[0][0].detach().numpy()
scores = softmax(scores)
# # TF
# model = TFAutoModelForSequenceClassification.from_pretrained(MODEL)
# model.save_pretrained(MODEL)
# text = "Good night 😊"
# encoded_input = tokenizer(text, return_tensors='tf')
# output = model(encoded_input)
# scores = output[0][0].numpy()
# scores = softmax(scores)
ranking = np.argsort(scores)
ranking = ranking[::-1]
for i in range(scores.shape[0]):
l = labels[ranking[i]]
s = scores[ranking[i]]
print(f"{i+1}) {l} {np.round(float(s), 4)}")
Output:
1) positive 0.8466
2) neutral 0.1458
3) negative 0.0076
BibTeX entry and citation info
Please cite the reference paper if you use this model.
@inproceedings{barbieri-etal-2020-tweeteval,
title = "{T}weet{E}val: Unified Benchmark and Comparative Evaluation for Tweet Classification",
author = "Barbieri, Francesco and
Camacho-Collados, Jose and
Espinosa Anke, Luis and
Neves, Leonardo",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020",
month = nov,
year = "2020",
address = "Online",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2020.findings-emnlp.148",
doi = "10.18653/v1/2020.findings-emnlp.148",
pages = "1644--1650"
}
Magnet link
Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:
magnet:?xt=urn:btih:97aa220234dac3cca360aad5bd703ab33f30e8c4&dn=cardiffnlp_twitter-roberta-base-sentimentOpen magnet in torrent client · infohash 97aa220234dac3cca360aad5bd703ab33f30e8c4
Files & hashes
| Path | Size | sha1 | sha256 |
|---|---|---|---|
| README.md | 3.6 KB (3,723 B) | 47cd107247eba114e9abca9f158e4ad098b31e27 | 068f1528bb7f006dbd7d95884387db548908b87ae26f8d21f130428b69009927 |
| config.json | 747 B (747 B) | f1e21eb05e56dc207001d6684f988a75a29e66b1 | de9ed9df04ecd8c267564e60b73a216f7cc7951abf85b797cb717216a5eeba50 |
| merges.txt | 445.6 KB (456,318 B) | 226b0752cac7789c48f0cb3ec53eda48b7be36cc | 1ce1664773c50f3e0cc8842619a93edc4624525b728b188a9e0be33b7726adc5 |
| pytorch_model.bin | 475.6 MB (498,679,497 B) | 72c1cdab9ecdb7d462093275c19c022dcb4e4a18 | c37a3484c55954cd75b336a85f1e0c023ae874f3a73b05d2418dd04828e293b1 |
| special_tokens_map.json | 150 B (150 B) | 6cd1d9021e10d47aed59399af6b0e30312b46ca4 | 7638f5bbbe86ef6d604ef28ad3647dc690d6d117c81c0d63e885416be8da1150 |
| vocab.json | 877.8 KB (898,822 B) | 0a39732b2d8be8e493cab3da68b68cc3e28221de | 06b4d46c8e752d410213d9548eb27a54db70fda0319b6271fb8d59dead5e1cab |
Cite this release
- Canonical URL
- https://aiseedbank.org/models/cardiffnlp_twitter-roberta-base-sentiment/
- Slug
- cardiffnlp_twitter-roberta-base-sentiment
- Infohash
- 97aa220234dac3cca360aad5bd703ab33f30e8c4
- License
- no license recorded
- Signing key fingerprint
- 85a3b32c3712427b
Every file carries a locally computed sha256 — verify a download against the signed sums: cardiffnlp_twitter-roberta-base-sentiment.SHA256SUMS (+ minisign signature).
Provenance
| Upstream repository | cardiffnlp/twitter-roberta-base-sentiment |
|---|---|
| Revision (pinned) | daefdd1f6ae931839bce4d0f3db0a1a4265cd50f |
| Fetched at | 2026-09-03T21:11:16Z |
| License at fetch | no license recorded |
| Snapshot tool | huggingface · seedbank 0.1.0 |
Trackers
- udp://announce.aitorrent.org:6969/announce
- http://announce.aitorrent.org:7070/announce
- udp://announce2.aitorrent.org:6970/announce
- http://announce2.aitorrent.org:7071/announce
- udp://tracker.opentrackr.org:1337/announce
- udp://open.demonii.com:1337/announce
- udp://open.stealth.si:80/announce
- udp://exodus.desync.com:6969/announce
- udp://tracker.torrent.eu.org:451/announce
✓ verified · rehash-vs-hf-metadata at 2026-09-03T21:11:23Z
no license recorded476.9 MB (500,039,257 bytes)transformerspytorchjaxrobertatext-classificationendpoints_compatible2 languages (tf, en)paper: 2010.12421