Instructions to use nishparadox/atomizer-gliner-small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER
How to use nishparadox/atomizer-gliner-small with GLiNER:
from gliner import GLiNER model = GLiNER.from_pretrained("nishparadox/atomizer-gliner-small") text = "Cristiano Ronaldo dos Santos Aveiro was born on 5 February 1985 in Funchal, Madeira, Portugal." labels = ["person", "date", "location"] entities = model.predict_entities(text, labels) for entity in entities: print(entity["text"], "=>", entity["label"]) - Notebooks
- Google Colab
- Kaggle
atomizer-gliner-small
A single-pass span atomizer with a critic. It splits a text into short, self-contained claims ("atoms") and scores each one:
verifiable: is this a factual claim that can be checked?checkworthy: is it worth fact-checking?
It runs in one forward pass of a 163M-parameter model, taking a few ms per text on a GPU and tens of ms on a laptop.
The model does not generate text. It selects words from the input, replaces references ("She" → "Marie Curie", "The company" → "Tesla") with earlier spans, and adds closed-class glue words ("is", "the", "'s"). An atom cannot contain a name or number that isn't in the input, and values are copied byte for byte ("8846m").
Status: research prototype (v0.2). v0.2 handles chat and agent text: first person, greetings, hedges ("I think …", which are unwrapped) and opinions (emitted as atoms with low
verifiable).
Usage
The loading code is the atomizer package at NISH1001/atomizer. That repo is
private for now, so access is on request. The weights are a plain PyTorch state dict (model.pt) plus atomizer.json
(backbone, settings, glue and insertion vocabularies).
uv add "atomizer @ git+ssh://git@github.com/NISH1001/atomizer" # needs access to the repo
from atomizer import Atomizer
atomizer = Atomizer.load("nishparadox/atomizer-gliner-small", revision="v0.2")
for atom in atomizer.atomize("I am good. My name is Nish. I think the height of mount everest is 8846m."):
print(atom.text, round(atom.verifiable, 2), round(atom.checkworthy, 2))
# I am good. 0.0 0.0
# My name is Nish. 0.0 0.0
# The height of mount everest is 8846m. 0.99 0.93
atomizer.atomize(text, min_verifiable=0.5) # only checkable facts
print(atomizer.explain(text).show()) # every head's predictions
atomizer.critic_scores(["We should invest more in education."]) # score claims you already have
atomize returns every atom by default, opinions and greetings included, each with its scores. Filter with
min_verifiable / min_checkworthy.
Model
GLiNER small v2.5 (DeBERTa-v3-small encoder, BiLSTM, span layer) plus four heads:
| Head | Predicts |
|---|---|
| anchor | per word: does an atom start here |
| pair | per (atom, word): KEEP / SUB / DROP, and glue after the word |
| antecedent | per reference: which earlier span it means |
| critic | per atom: verifiable, checkworthy |
Assembly from these predictions is deterministic. Texts longer than 1,024 tokens are atomized in sentence windows with read-only context. Details: ARCHITECTURE.md.
Training
Three stages, each continuing from the previous checkpoint. No API calls were made to create training data. Every
run is logged in W&B nasa-impact/atomizer (team project). The full log, with commands, curves and timings, is in
EXPERIMENTS.md.
Datasets
| Dataset | Role | Labels |
|---|---|---|
chentong00/propositionizer-wiki-data |
atoms: 42,857 Wikipedia passages, ~7.7 propositions each | GPT-4, by the dataset authors (Chen et al., Dense X Retrieval) |
Nithiwat/claimbuster |
critic: 23,533 US-debate sentences (check-worthy fact / unimportant fact / non-factual) | human |
iai-group/clef2024_checkthat_task1_en |
critic: CLEF CheckThat! 2024 task 1, check-worthiness yes/no (train 22,501, dev 1,032, test 318) | human |
nishparadox/atomizer-chat-synthetic@v0.1 |
atoms + per-atom critic labels on chat, agent, science and news text, 12 types (train 995, test 204) | AI-written (Claude); see its disclaimer |
chat templates (scripts/make_chat.py) |
hedged facts, opinions, insults, personal and social statements, questions; with noise | rule-based, generated locally |
sihaochen/propsegment (PropSegmEnt) |
test only (its atoms are fragments) | human |
ClaimBuster and CheckThat! come from the same debate transcripts. Every training sentence that also appears in a validation or test split was removed (3,279 sentences), so the critic is not scored on sentences it trained on.
Stages
| Stage 1: atomizer | Stage 2: critic → v0.1 | Stage 3: chat → v0.2 | |
|---|---|---|---|
| starts from | gliner-community/gliner_small-v2.5 |
stage 1, final | v0.1 |
| data per epoch | Propositionizer, 42,857 | Propositionizer 15,000 + ClaimBuster + CheckThat! 2024 (58,681) | Propositionizer 27,076 (30%) + critic 27,076 (30%) + templates 5,415 (6%) + synthetic 995 (all) = 60,562 |
| epochs / steps | 3 / 8,037 | 1 / 3,668 (released: step 1,000) | 2 / 3,780 (released: last) |
| batch | 16 | 16 | 32 |
| learning rates (encoder / GLiNER layers / heads) | 3e-5 / 1e-4 / 5e-4 | 2e-5 / 5e-5 / 5e-4 | 2e-5 / 5e-5 / 5e-4 |
| hardware, time | Apple M3 Max (MPS), 2 h 10 min | Apple M3 Max, ~40 min | 1× A100 80GB, 16 min |
| W&B run | gliner-small-v2.5-propwiki-3ep |
gliner-small-v2.5-propwiki-critic |
m1-mix-syn |
Shared by all stages:
- AdamW (fused), weight decay 0.01, gradient clip 1.0
- cosine schedule, 5% warmup, bf16 autocast, dropout 0.1
- insertion vocabulary of 300 words, 2,000 glue phrases
- anchor threshold 0.35, tuned on Propositionizer validation
Stage 3 also uses bucketed shuffling: each epoch, batches are rebuilt from shuffled chunks sorted by length, so they mix datasets (70% of batches) at 5% padding.
How labels are made
An aligner turns each gold atom into per-word targets:
- an anchor: the atom's first word no other atom uses
- KEEP / SUB / DROP for each word
- glue for each gap
- an antecedent span for each reference
Gold atoms that cannot be built this way are skipped, and the anchor loss is masked on their words. That covers reordering, verb changes and words not in the text: ~26–30% of GPT-4 propositions, and ~0% of the templates and synthetic data. The critic learns from gold atoms (per-atom labels, or "verifiable" for propositions) and from whole labelled sentences, with the atomization losses switched off for those.
Stage 3 ablations
Same data and settings, one difference each; test sets below:
| Run | Difference | synthetic soft F1 | critic AUC (synthetic atoms) | chat held-out | Propositionizer | PropSegmEnt | CheckThat! F1 |
|---|---|---|---|---|---|---|---|
| M1 = v0.2 | — | 0.897 | 0.994 | 0.988 | 0.733 | 0.741 | 0.866 |
| M2 | + glue penalty 0.5 | 0.897 | 0.994 | 0.989 | 0.733 | 0.742 | 0.866 |
| M3 | no synthetic data | 0.872 | 0.947 | 0.989 | 0.728 | 0.730 | 0.872 |
| M1 at step 500 | 0.26 epochs instead of 2 | 0.891 | 0.980 | 0.975 | 0.744 | 0.760 | 0.851 |
- The synthetic data mainly improves the critic on realistic text (+0.047 AUC).
- The glue penalty has no effect.
- Longer training improves chat and the critic but costs a little on Wikipedia-style text. The release favours chat and agent inputs.
Evaluation
Test splits only, none used for training or checkpoint selection. Atom sets are scored with soft F1 (best-match token F1 between predicted and reference atoms).
| Test set | v0.1 | v0.2 |
|---|---|---|
| synthetic chat/agent texts (204), soft F1 | 0.774 | 0.897 |
…critic verifiable on their gold atoms, AUC |
0.936 | 0.994 |
| chat templates with held-out patterns, facts and names, soft F1 | 0.525 | 0.988 |
| known v0.1 failure cases (8), soft F1 | 0.565 | 1.000 |
| Propositionizer test (500), soft F1 | 0.737 | 0.733 |
| PropSegmEnt test (500, human-labelled), soft F1 | 0.666 | 0.741 |
CLEF CheckThat! 2024 task 1 EN official test (315), checkworthy F1 / AUC |
0.827 / 0.974 | 0.866 / 0.981 |
fact-assessor claim cases (19), verifiable F1 |
1.00 | 1.00 |
| words not in the input (invented) | 0 | 0 |
The share of reference atoms a span model can express at all is about 70% on GPT-4 propositions, so word-exact match on that set is low.
Limitations
- The chat gains are measured on generated data. The templates and the AI-written synthetic set are cleaner than real user text. The 8 failure cases are real but few.
- Over-generates on long scientific passages: about 15 atoms per passage where an LLM atomizer writes about 8. This was measured on v0.1 and not re-measured.
- One reference resolves to one span. "The study" cannot become a description assembled from several places, and event references ("That loss raised sea level…") stay unresolved.
- No reordering and no verb changes. "carrying" cannot become "carried". An opinion word placed before the fact ("the ridiculously expensive X costs…") cannot be separated.
- The anchor threshold (0.35) controls how many atoms come out. Lower it for more atoms.
Versions
| Tag | Content |
|---|---|
v0.2 |
stage 3: chat, hedges, opinions; better critic |
v0.1 |
stages 1–2: Wikipedia atomizer + critic |
License
The weights are released under CC-BY-4.0. The atomizer code is MIT.
The training data has its own terms, which you should check for your use:
- Propositionizer labels were generated with GPT-4 by the dataset's authors (OpenAI terms).
- ClaimBuster is CC-BY-SA 4.0.
- CLEF CheckThat! 2024 is under the CLEF lab's terms.
atomizer-chat-syntheticwas written by an AI model (Claude, by Anthropic).
Model tree for nishparadox/atomizer-gliner-small
Base model
gliner-community/gliner_small-v2.5