Instructions to use pinthoz/gus-net-bert-large with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pinthoz/gus-net-bert-large with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="pinthoz/gus-net-bert-large")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("pinthoz/gus-net-bert-large") model = AutoModelForTokenClassification.from_pretrained("pinthoz/gus-net-bert-large", device_map="auto") - Notebooks
- Google Colab
- Kaggle
GUS-Net (BERT Large)
Token-level social-bias detector built on bert-large-uncased. Given a
sentence, it tags each token with one of four bias categories following a 7-label
BIO scheme, highlighting which words carry bias.
Part of the Attention Atlas project (a master's thesis on interpretable bias
detection through transformer attention). This is the larger-capacity BERT
variant of pinthoz/gus-net-bert.
- Base model:
bert-large-uncased - Task: multi-label token classification (per-token sigmoid, thresholded)
- Language: English
- Related models:
pinthoz/gus-net-bert,pinthoz/gus-net-gpt2,pinthoz/gus-net-gpt2-medium
Label scheme
| Index | Label | Category |
|---|---|---|
| 0 | O | none |
| 1 | B-STEREO | Stereotype (span start) |
| 2 | I-STEREO | Stereotype (span inside) |
| 3 | B-GEN | Generalisation (span start) |
| 4 | I-GEN | Generalisation (span inside) |
| 5 | B-UNFAIR | Unfair language (span start) |
| 6 | I-UNFAIR | Unfair language (span inside) |
- GEN — a blanket generalisation about a group.
- UNFAIR — unfair / disparaging language toward a group.
- STEREO — a stereotype attributed to a group.
Important: multi-label + per-label thresholds
Outputs are per-token sigmoid probabilities (multi-label), not a softmax.
F1-optimised thresholds (order [O, B-STEREO, I-STEREO, B-GEN, I-GEN, B-UNFAIR, I-UNFAIR]):
[0.5194, 0.4882, 0.5640, 0.4915, 0.4429, 0.5264, 0.5000]
Use the values above rather than a flat 0.5. They are calibrated against these specific weights.
Revised August 2026. The thresholds previously published here were calibrated on a different training run of the same architecture, and applying them to these weights collapses
UNFAIRprecision to 0.16. If you pinned the earlier values, update them.
Usage
import torch
from transformers import BertTokenizerFast, BertForTokenClassification
model_id = "pinthoz/gus-net-bert-large"
tok = BertTokenizerFast.from_pretrained("bert-large-uncased")
model = BertForTokenClassification.from_pretrained(model_id).eval()
CATEGORY_INDICES = {"STEREO": [1, 2], "GEN": [3, 4], "UNFAIR": [5, 6]}
THRESHOLDS = [0.5194, 0.4882, 0.5640, 0.4915, 0.4429, 0.5264, 0.5000]
text = "Women are naturally worse at driving."
enc = tok(text, return_tensors="pt")
with torch.no_grad():
probs = torch.sigmoid(model(input_ids=enc["input_ids"],
attention_mask=enc["attention_mask"]).logits)[0]
tokens = tok.convert_ids_to_tokens(enc["input_ids"][0])
for i, tokn in enumerate(tokens):
if tokn in ("[CLS]", "[SEP]", "[PAD]"):
continue
fired = {cat: float(probs[i, idxs].max())
for cat, idxs in CATEGORY_INDICES.items()
if any(probs[i, j] > THRESHOLDS[j] for j in idxs)}
if fired:
print(f"{tokn:15s} -> {fired}")
Training data
Fine-tuned on the GUS-Net dataset — a token-level social-bias corpus
annotated for Generalisations, Unfairness and Stereotypes
(ethical-spectacle/gus-dataset-v1).
Difference from the original GUS-Net dataset and models: in the original data
punctuation is almost always fused to the preceding word rather than tokenised
separately (only 159 standalone punctuation tokens across the corpus, against
5,879 after cleaning), so a comma or full stop falling inside a labelled span
inherits that span's categories — the
sentence-final mark carries a bias label in 1,942 of the 3,739 sentences, and an
in-span comma in 270. The data used here splits each mark into a token of its
own and labels it non-bias O, repairing the BIO sequence where the split
interrupts a span, since punctuation is not a social-bias carrier. Bias spans
predicted by these models therefore exclude leading/trailing punctuation.
Evaluation
StereoSet (intersentence split, 2123 examples)
| Metric | Score |
|---|---|
| LMS (language-modeling score, higher is better) | 69.15 |
| SS (stereotype score, 50 = ideal) | 54.64 |
| ICAT (bias-adjusted quality) | 62.73 |
Per-category SS: gender 54.96 · race 52.05 · religion 56.41 · profession 57.44.
Token classification (GUS-Net held-out test set)
Held-out partition (747 sentences) of the stratified cross-validation fold this
checkpoint was trained against — StratifiedKFold(n_splits=5, shuffle=True, random_state=42) over the cleaned corpus (see Training data), fold 4 —
scored with the per-label thresholds above. Each category aggregates its B-/I-
labels; the micro average covers the three bias categories and excludes the
majority O class. Label alignment mirrors training (secondary subtokens masked).
| Category | Precision | Recall | F1 |
|---|---|---|---|
| O (non-bias) | 0.929 | 0.955 | 0.942 |
| GEN | 0.813 | 0.791 | 0.802 |
| UNFAIR | 0.641 | 0.606 | 0.623 |
| STEREO | 0.859 | 0.827 | 0.843 |
| Micro-avg | 0.820 | 0.790 | 0.805 |
The operating point matters as much as the weights here. Scored with the
thresholds published before August 2026, which came from a different training
run, the same checkpoint returns UNFAIR precision 0.16 at recall 0.98 and a
micro F1 of 0.630, which reads as a detector that flags almost every token. The
weights were never the problem. For a balanced encoder at base size use
pinthoz/gus-net-bert;
for the strongest detector overall use
pinthoz/gus-net-gpt2-medium.
Limitations & intended use
- Research / auditing tool, not a content-moderation oracle. Predictions reflect a specific operationalisation of bias; subtle or context-dependent bias may be missed.
- English only.
- Labels are not error-free; treat spans as evidence to review, not ground truth.
- Do not use for automated decisions about individuals.
Citation
If you use these models, please cite the GUS-Net dataset and benchmark:
@article{powers2024gusnet,
title = {GUS-Net: Social Bias Classification in Text with Generalizations, Unfairness, and Stereotypes},
author = {Powers, Maximus and Raza, Shaina and Chang, Alex and Riaz, Rehana and Mavani, Umang and Jonala, Harshitha Reddy and Tiwari, Ansh and Wei, Hua},
journal = {arXiv preprint arXiv:2410.08388},
year = {2024}
}
License
Weights released under Apache-2.0 (matching the bert-large-uncased base
model). The Attention Atlas code is MIT-licensed.
- Downloads last month
- 13
Model tree for pinthoz/gus-net-bert-large
Base model
google-bert/bert-large-uncased