File size: 4,759 Bytes
ad5cda1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
06204ae
 
 
 
 
ad5cda1
 
 
 
 
 
 
 
 
 
06204ae
 
 
 
 
 
 
 
ad5cda1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
06204ae
 
 
 
 
 
 
 
 
 
 
 
 
 
ad5cda1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
---
license: apache-2.0
base_model: ProsusAI/finbert
tags:
- financial-sentiment
- social-media
- reddit
- wallstreetbets
- text-classification
pipeline_tag: text-classification
language:
- en
---

# WSB-FinBERT

A FinBERT model fine-tuned on r/wallstreetbets text for three-class sentiment
classification (negative / neutral / positive).

Off-the-shelf financial sentiment models are trained on formal financial English —
earnings calls, analyst reports, newswire. Retail social-media text is a different
register: irony, slang, emoji, and self-deprecation. On the held-out set below,
vanilla FinBERT scores **below the majority-class floor**, i.e. worse than ignoring
the text entirely. This model is a domain-adaptation check on that gap.

This is the frozen checkpoint behind the results reported in *Beyond the Volume of Attention:
Domain-Adapted Sentiment and the Content of Retail Investor Discussion* (ICAIF '26) — not a
retrained copy. The replication package for that paper is distributed separately; this
repository is the model only.

## Labels

| id | label |
|---|---|
| 0 | negative |
| 1 | neutral |
| 2 | positive |

## Usage

You do not need to download anything by hand. `transformers` fetches the weights on first use
and caches them under `~/.cache/huggingface/`, so the first call takes a moment (~440 MB) and
every call after that is instant.

```
pip install transformers torch
```

```python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

model_id = "AnonymousResearchICAIF/wsb-finbert"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)

text = "NVDA printing again, loaded calls for next week"
inputs = tok(text, return_tensors="pt", truncation=True, max_length=128)
with torch.no_grad():
    probs = model(**inputs).logits.softmax(-1)[0]

print({model.config.id2label[i]: round(p.item(), 3) for i, p in enumerate(probs)})
```

A continuous sentiment score in [-1, 1] is formed as `P(positive) - P(negative)`. This is the
`sentiment_wsb` variable used throughout the paper — the sign gives the direction and the
magnitude gives the confidence. **Use `max_length=128`**: that is what the model was trained
with, and longer inputs are truncated to it.

For many texts at once, batch them rather than looping:

```python
texts = ["...", "...", "..."]
batch = tok(texts, return_tensors="pt", truncation=True, max_length=128, padding=True)
with torch.no_grad():
    p = model(**batch).logits.softmax(-1)
scores = (p[:, 2] - p[:, 0]).tolist()
```

## Training data

2,503 r/wallstreetbets posts and comments mentioning a fixed universe of AI-related
tickers, labelled for sentiment by an LLM teacher (Claude) and split 70/15/15
stratified by label with seed 42 — 1,752 train / 375 validation / **376 test**.
The sample is stratified across tickers and balanced 50/50 between the first and
second halves of the collection period.

## Training procedure

| | |
|---|---|
| Base model | `ProsusAI/finbert` |
| Max sequence length | 128 |
| Learning rate | 2e-5 |
| Epochs | 4 |
| Train batch size | 8 |
| Weight decay | 0.01 |
| Warmup ratio | 0.1 |
| Seed | 42 |

## Evaluation

On the 376 held-out texts:

| Model | Accuracy | Macro F1 |
|---|---|---|
| **WSB-FinBERT (this model)** | **0.524** | **0.495** |
| Vanilla `ProsusAI/finbert` | 0.410 | 0.400 |
| Majority-class baseline ("always positive") | 0.434 | — |

Cohen's κ = 0.252.

Per-class (this model):

| label | precision | recall | f1 | support |
|---|---|---|---|---|
| negative | 0.448 | 0.312 | 0.368 | 96 |
| neutral | 0.513 | 0.513 | 0.513 | 117 |
| positive | 0.557 | 0.656 | 0.603 | 163 |

## Limitations

- **Absolute accuracy is modest.** 52.4% on a three-class problem is only ~9 points
  above the majority-class floor. The gain over vanilla FinBERT (+11.4 points) is the
  meaningful result; the model is not a strong standalone classifier.
- **Labels are LLM-generated**, not human gold standard. They inherit the teacher's
  biases. They do correlate with *past* five-day returns, as a text-only annotator
  should, and show no positive correlation with *forward* returns (no evidence of
  look-ahead leakage).
- **Narrow domain.** Trained on r/wallstreetbets text about a small universe of
  AI-related tickers over a specific period. Generalisation to other subreddits,
  other sectors, other periods, or to formal financial text is untested and unlikely.
- **Positive skew.** The forum's labels run ~1.7 to 1 positive and the model inherits
  this (51.1% of predictions positive vs 43.4% in truth). Use within-entity demeaning
  if the level matters for your application.
- Not investment advice; not suitable for trading decisions on its own.