metadata
license: apache-2.0
base_model: distilbert/distilbert-base-uncased
datasets:
- cardiffnlp/tweet_eval
language:
- en
pipeline_tag: text-classification
tags:
- sentiment-analysis
- twitter
- social-media
metrics:
- accuracy
- f1
distilbert-base → TweetEval Sentiment
Small, fast LLM fine-tuned for social-media (tweet) sentiment analysis. 3 classes: negative / neutral / positive.
- Base model: distilbert/distilbert-base-uncased (~67M, generic)
- Dataset: cardiffnlp/tweet_eval (
sentimentconfig, 45.6K train) - Training: 3 epochs, lr 2e-5, batch 32, max_len 128, warmup 0.1, weight decay 0.01
Test-set results (TweetEval sentiment, 12,284 tweets)
| Metric | Score |
|---|---|
| Accuracy | 0.6888 |
| Macro-F1 | 0.6877 |
| Macro-Recall | 0.6978 |
| Speed (T4) | ~2897 tweets/s |
Comparison vs twitter-roberta-base
| Model | Size | Accuracy | Macro-F1 | tweets/s |
|---|---|---|---|---|
| twitter-roberta-base | 125M | 0.7155 | 0.7155 | 1600 |
| distilbert-base (this) | 67M | 0.6888 | 0.6877 | 2897 |
This model trades 2.7 pts accuracy for **1.8× faster inference and half the size** — a good fit for high-throughput or edge deployment.
Usage
from transformers import pipeline
clf = pipeline("text-classification", model="Ido-shraga/distilbert-base-tweeteval-sentiment")
clf("I can't believe how good this is 🔥")