wsb-finbert / README.md
AnonymousResearchICAIF's picture
Expand model card: auto-download note, batching, score definition, provenance
06204ae verified
|
Raw
History Blame Contribute Delete
4.76 kB
metadata
license: apache-2.0
base_model: ProsusAI/finbert
tags:
  - financial-sentiment
  - social-media
  - reddit
  - wallstreetbets
  - text-classification
pipeline_tag: text-classification
language:
  - en

WSB-FinBERT

A FinBERT model fine-tuned on r/wallstreetbets text for three-class sentiment classification (negative / neutral / positive).

Off-the-shelf financial sentiment models are trained on formal financial English — earnings calls, analyst reports, newswire. Retail social-media text is a different register: irony, slang, emoji, and self-deprecation. On the held-out set below, vanilla FinBERT scores below the majority-class floor, i.e. worse than ignoring the text entirely. This model is a domain-adaptation check on that gap.

This is the frozen checkpoint behind the results reported in Beyond the Volume of Attention: Domain-Adapted Sentiment and the Content of Retail Investor Discussion (ICAIF '26) — not a retrained copy. The replication package for that paper is distributed separately; this repository is the model only.

Labels

id label
0 negative
1 neutral
2 positive

Usage

You do not need to download anything by hand. transformers fetches the weights on first use and caches them under ~/.cache/huggingface/, so the first call takes a moment (~440 MB) and every call after that is instant.

pip install transformers torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

model_id = "AnonymousResearchICAIF/wsb-finbert"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)

text = "NVDA printing again, loaded calls for next week"
inputs = tok(text, return_tensors="pt", truncation=True, max_length=128)
with torch.no_grad():
    probs = model(**inputs).logits.softmax(-1)[0]

print({model.config.id2label[i]: round(p.item(), 3) for i, p in enumerate(probs)})

A continuous sentiment score in [-1, 1] is formed as P(positive) - P(negative). This is the sentiment_wsb variable used throughout the paper — the sign gives the direction and the magnitude gives the confidence. Use max_length=128: that is what the model was trained with, and longer inputs are truncated to it.

For many texts at once, batch them rather than looping:

texts = ["...", "...", "..."]
batch = tok(texts, return_tensors="pt", truncation=True, max_length=128, padding=True)
with torch.no_grad():
    p = model(**batch).logits.softmax(-1)
scores = (p[:, 2] - p[:, 0]).tolist()

Training data

2,503 r/wallstreetbets posts and comments mentioning a fixed universe of AI-related tickers, labelled for sentiment by an LLM teacher (Claude) and split 70/15/15 stratified by label with seed 42 — 1,752 train / 375 validation / 376 test. The sample is stratified across tickers and balanced 50/50 between the first and second halves of the collection period.

Training procedure

Base model ProsusAI/finbert
Max sequence length 128
Learning rate 2e-5
Epochs 4
Train batch size 8
Weight decay 0.01
Warmup ratio 0.1
Seed 42

Evaluation

On the 376 held-out texts:

Model Accuracy Macro F1
WSB-FinBERT (this model) 0.524 0.495
Vanilla ProsusAI/finbert 0.410 0.400
Majority-class baseline ("always positive") 0.434

Cohen's κ = 0.252.

Per-class (this model):

label precision recall f1 support
negative 0.448 0.312 0.368 96
neutral 0.513 0.513 0.513 117
positive 0.557 0.656 0.603 163

Limitations

  • Absolute accuracy is modest. 52.4% on a three-class problem is only ~9 points above the majority-class floor. The gain over vanilla FinBERT (+11.4 points) is the meaningful result; the model is not a strong standalone classifier.
  • Labels are LLM-generated, not human gold standard. They inherit the teacher's biases. They do correlate with past five-day returns, as a text-only annotator should, and show no positive correlation with forward returns (no evidence of look-ahead leakage).
  • Narrow domain. Trained on r/wallstreetbets text about a small universe of AI-related tickers over a specific period. Generalisation to other subreddits, other sectors, other periods, or to formal financial text is untested and unlikely.
  • Positive skew. The forum's labels run ~1.7 to 1 positive and the model inherits this (51.1% of predictions positive vs 43.4% in truth). Use within-entity demeaning if the level matters for your application.
  • Not investment advice; not suitable for trading decisions on its own.