File size: 4,759 Bytes
ad5cda1 06204ae ad5cda1 06204ae ad5cda1 06204ae ad5cda1 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 | ---
license: apache-2.0
base_model: ProsusAI/finbert
tags:
- financial-sentiment
- social-media
- reddit
- wallstreetbets
- text-classification
pipeline_tag: text-classification
language:
- en
---
# WSB-FinBERT
A FinBERT model fine-tuned on r/wallstreetbets text for three-class sentiment
classification (negative / neutral / positive).
Off-the-shelf financial sentiment models are trained on formal financial English —
earnings calls, analyst reports, newswire. Retail social-media text is a different
register: irony, slang, emoji, and self-deprecation. On the held-out set below,
vanilla FinBERT scores **below the majority-class floor**, i.e. worse than ignoring
the text entirely. This model is a domain-adaptation check on that gap.
This is the frozen checkpoint behind the results reported in *Beyond the Volume of Attention:
Domain-Adapted Sentiment and the Content of Retail Investor Discussion* (ICAIF '26) — not a
retrained copy. The replication package for that paper is distributed separately; this
repository is the model only.
## Labels
| id | label |
|---|---|
| 0 | negative |
| 1 | neutral |
| 2 | positive |
## Usage
You do not need to download anything by hand. `transformers` fetches the weights on first use
and caches them under `~/.cache/huggingface/`, so the first call takes a moment (~440 MB) and
every call after that is instant.
```
pip install transformers torch
```
```python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_id = "AnonymousResearchICAIF/wsb-finbert"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
text = "NVDA printing again, loaded calls for next week"
inputs = tok(text, return_tensors="pt", truncation=True, max_length=128)
with torch.no_grad():
probs = model(**inputs).logits.softmax(-1)[0]
print({model.config.id2label[i]: round(p.item(), 3) for i, p in enumerate(probs)})
```
A continuous sentiment score in [-1, 1] is formed as `P(positive) - P(negative)`. This is the
`sentiment_wsb` variable used throughout the paper — the sign gives the direction and the
magnitude gives the confidence. **Use `max_length=128`**: that is what the model was trained
with, and longer inputs are truncated to it.
For many texts at once, batch them rather than looping:
```python
texts = ["...", "...", "..."]
batch = tok(texts, return_tensors="pt", truncation=True, max_length=128, padding=True)
with torch.no_grad():
p = model(**batch).logits.softmax(-1)
scores = (p[:, 2] - p[:, 0]).tolist()
```
## Training data
2,503 r/wallstreetbets posts and comments mentioning a fixed universe of AI-related
tickers, labelled for sentiment by an LLM teacher (Claude) and split 70/15/15
stratified by label with seed 42 — 1,752 train / 375 validation / **376 test**.
The sample is stratified across tickers and balanced 50/50 between the first and
second halves of the collection period.
## Training procedure
| | |
|---|---|
| Base model | `ProsusAI/finbert` |
| Max sequence length | 128 |
| Learning rate | 2e-5 |
| Epochs | 4 |
| Train batch size | 8 |
| Weight decay | 0.01 |
| Warmup ratio | 0.1 |
| Seed | 42 |
## Evaluation
On the 376 held-out texts:
| Model | Accuracy | Macro F1 |
|---|---|---|
| **WSB-FinBERT (this model)** | **0.524** | **0.495** |
| Vanilla `ProsusAI/finbert` | 0.410 | 0.400 |
| Majority-class baseline ("always positive") | 0.434 | — |
Cohen's κ = 0.252.
Per-class (this model):
| label | precision | recall | f1 | support |
|---|---|---|---|---|
| negative | 0.448 | 0.312 | 0.368 | 96 |
| neutral | 0.513 | 0.513 | 0.513 | 117 |
| positive | 0.557 | 0.656 | 0.603 | 163 |
## Limitations
- **Absolute accuracy is modest.** 52.4% on a three-class problem is only ~9 points
above the majority-class floor. The gain over vanilla FinBERT (+11.4 points) is the
meaningful result; the model is not a strong standalone classifier.
- **Labels are LLM-generated**, not human gold standard. They inherit the teacher's
biases. They do correlate with *past* five-day returns, as a text-only annotator
should, and show no positive correlation with *forward* returns (no evidence of
look-ahead leakage).
- **Narrow domain.** Trained on r/wallstreetbets text about a small universe of
AI-related tickers over a specific period. Generalisation to other subreddits,
other sectors, other periods, or to formal financial text is untested and unlikely.
- **Positive skew.** The forum's labels run ~1.7 to 1 positive and the model inherits
this (51.1% of predictions positive vs 43.4% in truth). Use within-entity demeaning
if the level matters for your application.
- Not investment advice; not suitable for trading decisions on its own.
|