File size: 8,568 Bytes
39ddcbc
 
 
 
 
 
 
 
 
 
 
 
 
 
f76461b
39ddcbc
 
 
 
 
 
 
 
 
 
 
f76461b
659559c
f76461b
4320afc
 
 
 
 
 
 
39ddcbc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f76461b
 
 
39ddcbc
 
 
f76461b
 
 
 
 
39ddcbc
 
 
 
 
f76461b
 
 
 
 
 
 
 
 
 
 
 
39ddcbc
 
 
 
 
f76461b
39ddcbc
 
 
 
 
f76461b
 
 
 
 
 
 
 
 
39ddcbc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f76461b
39ddcbc
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
---
license: mit
base_model: FacebookAI/xlm-roberta-large
pipeline_tag: text-classification
language:
- multilingual
tags:
- news
- topic-classification
- multi-label
- xlm-roberta
library_name: transformers
---

# SNT News Classifier v0.6

Multi-label news topic classifier: **12 top-level (L1) and 71 sub-level (L2)
categories**, built on `xlm-roberta-large` with two independent sigmoid heads. Both levels are
genuinely multi-label — an article about a trade deal can be `world` + `money_and_business` +
`politics` at the same time. Per-class decision thresholds (tuned on a held-out validation split)
ship inside `config.json`; `predict_labels()` applies them and falls back to argmax so no article
is ever left unlabeled.

Built by [Sweenk](https://sweenk.com) to categorize its news feed; released so others can use and
scrutinize it.

**Links:** [GitHub — model & training code](https://github.com/sweenk/snt-model) ·
[GitHub — data pipeline](https://github.com/sweenk/snt-data)

## At a glance

- **Multi-label at both levels** — a story can be `world` + `money_and_business` + `politics` at once, each above its own per-class tuned threshold.
- **0.847 L1 macro-F1** on 26,412 held-out articles; all 12 top-level categories clear a 0.65 per-class floor.
- **Multilingual encoder**`xlm-roberta-large` (100 languages), fine-tuned on ~264K news articles.
- **Honest about its labels** — the gold is LLM-teacher-generated, not human-annotated, and the evaluation section says exactly what that means.

## Quick start

```python
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("sweenk/snt-classifier", trust_remote_code=True)
tok = AutoTokenizer.from_pretrained("sweenk/snt-classifier")

enc = tok("OpenAI raises $6.6B. The startup announced its latest funding round...",
          return_tensors="pt", truncation=True, max_length=512)
print(model.predict_labels(**enc))
# [{'l1': [{'key': 'money_and_business', 'p': 0.99}, {'key': 'tech_and_ai', 'p': 0.98}],
#    'primary_l1': 'money_and_business',
#    'l2': [{'key': 'companies_and_industries', 'p': 0.92}, ...]}]
```

Input convention: `"{title}\n\n{body}"`, truncated at 512 tokens. The classifier was trained
on title+body; titles alone work but body text improves routing (the training prompt explicitly
prioritizes body over headline).

## Taxonomy — 12 L1 / 71 L2

| L1 category | # L2 | L2 sub-categories |
|---|---|---|
| `sports` | 12 | `american_football`, `baseball`, `basketball`, `college_sports`, `combat_sports`, `golf`, `hockey`, `motorsports`, `olympics`, `other_sports`, `soccer`, `tennis` |
| `politics` | 6 | `elections_and_campaigns`, `government_and_policy`, `immigration_and_borders`, `political_figures_and_scandals`, `social_issues_and_activism`, `state_and_local_politics` |
| `world` | 4 | `geopolitics_and_diplomacy`, `humanitarian_crises`, `terrorism_and_security`, `war_and_conflict` |
| `entertainment_and_pop_culture` | 7 | `books_and_arts`, `celebrities_and_gossip`, `gaming`, `internet_culture_and_creators`, `media_and_journalism`, `movies_and_tv`, `music` |
| `money_and_business` | 8 | `companies_and_industries`, `cost_of_living`, `crypto_and_fintech`, `housing_and_real_estate`, `macro_economy_and_rates`, `markets_and_investing`, `personal_finance`, `work_and_careers` |
| `crime_and_justice` | 4 | `courts_and_trials`, `crime_and_policing`, `scams_and_fraud`, `true_crime` |
| `tech_and_ai` | 5 | `artificial_intelligence`, `big_tech_and_startups`, `cybersecurity_and_privacy`, `gadgets_and_apps`, `screen_time_and_digital_life` |
| `science_and_space` | 4 | `archaeology_and_history`, `psychology_and_behavior`, `scientific_discoveries`, `space_and_astronomy` |
| `health_and_wellness` | 5 | `fitness_and_exercise`, `medical_and_public_health`, `mental_health`, `nutrition_and_diet`, `sleep_and_longevity` |
| `lifestyle` | 8 | `education_and_schools`, `faith_and_spirituality`, `fashion_and_beauty`, `food_and_drink`, `home_and_garden`, `parenting_and_family`, `relationships_and_dating`, `travel_and_places` |
| `weather_and_environment` | 5 | `climate_change`, `disasters_and_accidents`, `energy_and_climate_solutions`, `nature_and_wildlife`, `severe_weather` |
| `human_stories` | 3 | `animals_and_pets`, `good_news_and_kindness`, `offbeat_and_unusual` |

The full machine-readable taxonomy (`l1_keys`, `l2_keys`, `l2_parent`, per-class thresholds) is in
`config.json`.

## Evaluation — our own numbers, stated plainly

Held-out test split: **26,412 articles** (a 10% slice of the labeled corpus,
stratified by primary L1). "Tuned" = per-class thresholds optimized on the *validation* split,
then applied unchanged to test.

| Metric | @0.5 threshold | tuned thresholds |
|---|---|---|
| L1 macro F1 | 0.817 | **0.847** |
| L1 micro F1 | 0.831 | **0.857** |
| L1 primary accuracy (argmax) | 0.833 | — |
| L2 macro F1 | 0.647 | 0.685 |
| L2 micro F1 | 0.753 | 0.754 |

### Per-class L1 F1 (test)

| L1 | F1 @0.5 | F1 tuned | threshold |
|---|---|---|---|
| `sports` | 0.941 | 0.954 | 0.850 |
| `politics` | 0.845 | 0.859 | 0.750 |
| `world` | 0.783 | 0.827 | 0.850 |
| `entertainment_and_pop_culture` | 0.887 | 0.898 | 0.850 |
| `money_and_business` | 0.786 | 0.825 | 0.950 |
| `crime_and_justice` | 0.823 | 0.857 | 0.950 |
| `tech_and_ai` | 0.802 | 0.857 | 0.950 |
| `science_and_space` | 0.802 | 0.833 | 0.950 |
| `health_and_wellness` | 0.821 | 0.860 | 0.900 |
| `lifestyle` | 0.865 | 0.867 | 0.700 |
| `weather_and_environment` | 0.816 | 0.857 | 0.950 |
| `human_stories` | 0.632 | 0.671 | 0.900 |

### What you should know before trusting these numbers

Read this section — it is the honest part.

- **The gold labels are model-assisted, not human-annotated.** The corpus (~278K articles from
  HuffPost archives, CommonCrawl News, daily.dev, and Sweenk production) was labeled by a
  mechanical migration from an earlier taxonomy plus multiple passes of a Claude Sonnet teacher
  with a rule-based prompt, spot-audited by humans (QA gates at 70–87% agreement on sampled
  batches). Test F1 therefore measures agreement with an LLM teacher, not with human ground truth.
- **Class imbalance is real (~21x).** `politics`/`lifestyle`/`entertainment` have ~44-47K training
  rows; `science_and_space` ~2.2K and `tech_and_ai` ~3.7K. Training compensates with per-class
  `pos_weight` (clamped at 10) and caps the majority classes at 25K primary-label train rows so
  they don't swamp the rare ones. The weakest class here is `human_stories` (F1 0.671) —
  the fuzziest category by construction.
- **Labels were corrected over time, so F1 is not comparable across releases.** v0.5.1 re-labeled
  ~1,500 systematically mislabeled rows with route-by-cause rules (accidents by cause, terror →
  crime, pharma earnings → money); on those rows agreement with the corrected gold went
  25.9% → 75.3%. Because the gold labels themselves changed between releases, aggregate F1 deltas
  are not like-for-like — treat each release's numbers as self-referential.
- **Multilingual ability is inherited, not measured.** The encoder is XLM-R, but nearly all
  training articles are English. Expect degraded (unquantified) quality on non-English news.
- **L3 (named topics / entities) is not part of this model** — Sweenk handles that downstream with
  a separate extraction step.

## Architecture

`xlm-roberta-large` encoder → CLS pooling → dropout(0.1) → two parallel linear heads
(L1: 12 logits, L2: 71 logits), both sigmoid. Trained 3 epochs, BCE loss with
per-class `pos_weight` (L1) and loss weights L1:1.0 / L2:2.0, bf16 autocast, gradient checkpointing.
Inference upcasts logits to fp32 before sigmoid (bf16 sigmoid saturates above logit ~6.2, which
collapses co-confident multi-label pairs).

## Versions

| Version | What changed |
|---|---|
| v0.5 | First multi-label release (12 L1 / 71 L2, dual sigmoid heads) |
| v0.5.1 | Corrective retrain: route-accidents-by-cause, terror→crime, earnings→money, govt-personnel→politics, wildlife→weather, body-over-headline; ~1,500 corrected labels; re-tuned thresholds |
| v0.6 | Retrain on the corrected corpus + ~23K newly teacher-labeled articles (Sweenk production + daily.dev science/tech); majority-class capping (25K/class) on top of `pos_weight`; re-tuned thresholds |

## License & attribution

Model weights: MIT. Base model: [FacebookAI/xlm-roberta-large](https://huggingface.co/FacebookAI/xlm-roberta-large) (MIT).
The training corpus contains article text from public news sources and is **not** redistributed
with this model.