File size: 4,341 Bytes
f4d360c 4b03c17 f4d360c 4b03c17 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 | ---
language: multilingual
license: mit
tags:
- text-classification
- political-science
- stance-detection
- support
- parliamentary-speech
base_model: jhu-clsp/mmBERT-base
---
# Support Detector
Binary classifier for parliamentary sentences: among sentences that are
**not** Opposition, does the sentence express **Support** toward the
European Union (label `1`) or is it **Neutral** (label `0`)?
This is the second stage of a two-step stance-detection cascade. It is
applied only to sentences that the upstream Opposition detector has
classified as Non-Opposition. Both stages use a 0.5 decision threshold.
Fine-tuned from `jhu-clsp/mmBERT-base` on hand-annotated parliamentary speeches
from AUS, CZE, DEU, DNK, ESP, GBR, NLD, and SWE.
## Labels
- `0` — Neutral
- `1` — Support
## Training data
- Source: hand-annotated parliamentary sentences labelled `Neutral`,
`Support`, or `Opposition`.
- For this model, restricted to gold non-Opposition rows
(`Neutral` ∪ `Support`) and binarised as Support vs Neutral.
- File: `Stance_Retrain_undersampled.csv` (undersampled to address class imbalance).
- Split: leakage-safe `StratifiedGroupKFold` (n_splits=10) on
country × speech_ID, so no speech appears in more than one fold.
Realised allocation: 8 folds train / 1 fold val / 1 fold test
(~80/10/10). Shares the same underlying stance split as the
Opposition detector for consistent cascade evaluation.
## Hyperparameters
- Base model: `jhu-clsp/mmBERT-base`
- Max sequence length: 320
- Learning rate: 4e-05
- Epochs: 4
- Batch size: 32 (with gradient accumulation if large model)
- Warmup ratio: 0.2
- Weight decay: 0.05
- LR scheduler: cosine
- Optimizer: AdamW (HF Trainer default)
- Mixed precision: fp16
- Early stopping patience: 2 (monitoring `f1_positive` on val)
- Class weights: balanced (sklearn `compute_class_weight`)
- Focal loss: disabled (plain weighted cross-entropy)
- Random seed: 123
- Model selection: best checkpoint by validation `f1_positive` (minority-class F1)
## Input format
Sentence-only input (no surrounding context window). Truncation to
320 tokens.
## Usage (standalone — Support vs Neutral)
```python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
tok = AutoTokenizer.from_pretrained("LBenoit/support-detector-mmbert")
mdl = AutoModelForSequenceClassification.from_pretrained("LBenoit/support-detector-mmbert")
text = "European cooperation has brought decades of peace and prosperity."
enc = tok(text, truncation=True, max_length=320, return_tensors="pt")
with torch.no_grad():
prob_support = torch.softmax(mdl(**enc).logits, dim=-1)[0, 1].item()
print("P(Support | Non-Opposition) =", prob_support)
```
## Usage (cascade — full 3-way stance)
```python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
OPP_REPO = "LBenoit/opposition-detector-mmbert"
SUP_REPO = "LBenoit/support-detector-mmbert"
tok_o = AutoTokenizer.from_pretrained(OPP_REPO)
mdl_o = AutoModelForSequenceClassification.from_pretrained(OPP_REPO)
tok_s = AutoTokenizer.from_pretrained(SUP_REPO)
mdl_s = AutoModelForSequenceClassification.from_pretrained(SUP_REPO)
def predict_stance(text, thresh=0.5):
enc = tok_o(text, truncation=True, max_length=320, return_tensors="pt")
p_opp = torch.softmax(mdl_o(**enc).logits, dim=-1)[0, 1].item()
if p_opp >= thresh:
return "Opposition"
enc = tok_s(text, truncation=True, max_length=320, return_tensors="pt")
p_sup = torch.softmax(mdl_s(**enc).logits, dim=-1)[0, 1].item()
return "Support" if p_sup >= thresh else "Neutral"
```
## Intended use
Research on parliamentary stance toward the EU. Designed as the second
stage of an Opposition → Support cascade. Using it standalone on
arbitrary text (without first filtering out Opposition sentences) is
out of distribution and not recommended.
## Limitations
- Trained only on non-Opposition rows; applying it to Opposition
sentences without the upstream filter will produce unreliable
predictions.
- Trained on parliamentary register; performance on social media,
journalism, or other domains is not guaranteed.
- Coverage limited to the eight countries listed above; generalisation
to other parliaments is untested.
- Sentence-level only; longer-range discourse context is not modelled.
|