WSB-FinBERT

A FinBERT model fine-tuned on r/wallstreetbets text for three-class sentiment classification (negative / neutral / positive).

Off-the-shelf financial sentiment models are trained on formal financial English โ€” earnings calls, analyst reports, newswire. Retail social-media text is a different register: irony, slang, emoji, and self-deprecation. On the held-out set below, vanilla FinBERT scores below the majority-class floor, i.e. worse than ignoring the text entirely. This model is a domain-adaptation check on that gap.

This is the frozen checkpoint behind the results reported in Beyond the Volume of Attention: Domain-Adapted Sentiment and the Content of Retail Investor Discussion (ICAIF '26) โ€” not a retrained copy. The replication package for that paper is distributed separately; this repository is the model only.

Labels

id label
0 negative
1 neutral
2 positive

Usage

You do not need to download anything by hand. transformers fetches the weights on first use and caches them under ~/.cache/huggingface/, so the first call takes a moment (~440 MB) and every call after that is instant.

pip install transformers torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

model_id = "AnonymousResearchICAIF/wsb-finbert"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)

text = "NVDA printing again, loaded calls for next week"
inputs = tok(text, return_tensors="pt", truncation=True, max_length=128)
with torch.no_grad():
    probs = model(**inputs).logits.softmax(-1)[0]

print({model.config.id2label[i]: round(p.item(), 3) for i, p in enumerate(probs)})

A continuous sentiment score in [-1, 1] is formed as P(positive) - P(negative). This is the sentiment_wsb variable used throughout the paper โ€” the sign gives the direction and the magnitude gives the confidence. Use max_length=128: that is what the model was trained with, and longer inputs are truncated to it.

For many texts at once, batch them rather than looping:

texts = ["...", "...", "..."]
batch = tok(texts, return_tensors="pt", truncation=True, max_length=128, padding=True)
with torch.no_grad():
    p = model(**batch).logits.softmax(-1)
scores = (p[:, 2] - p[:, 0]).tolist()

Training data

2,503 r/wallstreetbets posts and comments mentioning a fixed universe of AI-related tickers, labelled for sentiment by an LLM teacher (Claude) and split 70/15/15 stratified by label with seed 42 โ€” 1,752 train / 375 validation / 376 test. The sample is stratified across tickers and balanced 50/50 between the first and second halves of the collection period.

Training procedure

Base model ProsusAI/finbert
Max sequence length 128
Learning rate 2e-5
Epochs 4
Train batch size 8
Weight decay 0.01
Warmup ratio 0.1
Seed 42

Evaluation

On the 376 held-out texts:

Model Accuracy Macro F1
WSB-FinBERT (this model) 0.524 0.495
Vanilla ProsusAI/finbert 0.410 0.400
Majority-class baseline ("always positive") 0.434 โ€”

Cohen's ฮบ = 0.252.

Per-class (this model):

label precision recall f1 support
negative 0.448 0.312 0.368 96
neutral 0.513 0.513 0.513 117
positive 0.557 0.656 0.603 163

Limitations

  • Absolute accuracy is modest. 52.4% on a three-class problem is only ~9 points above the majority-class floor. The gain over vanilla FinBERT (+11.4 points) is the meaningful result; the model is not a strong standalone classifier.
  • Labels are LLM-generated, not human gold standard. They inherit the teacher's biases. They do correlate with past five-day returns, as a text-only annotator should, and show no positive correlation with forward returns (no evidence of look-ahead leakage).
  • Narrow domain. Trained on r/wallstreetbets text about a small universe of AI-related tickers over a specific period. Generalisation to other subreddits, other sectors, other periods, or to formal financial text is untested and unlikely.
  • Positive skew. The forum's labels run ~1.7 to 1 positive and the model inherits this (51.1% of predictions positive vs 43.4% in truth). Use within-entity demeaning if the level matters for your application.
  • Not investment advice; not suitable for trading decisions on its own.
Downloads last month
8
Safetensors
Model size
0.1B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for AnonymousResearchICAIF/wsb-finbert

Base model

ProsusAI/finbert
Finetuned
(113)
this model