Token Classification
GLiNER
atomizer
atomic-claims
claim-extraction
decontextualization
fact-checking
check-worthiness

atomizer-gliner-small

A single-pass span atomizer with a critic. It splits a text into short, self-contained claims ("atoms") and scores each one:

  • verifiable: is this a factual claim that can be checked?
  • checkworthy: is it worth fact-checking?

It runs in one forward pass of a 163M-parameter model, taking a few ms per text on a GPU and tens of ms on a laptop.

The model does not generate text. It selects words from the input, replaces references ("She" → "Marie Curie", "The company" → "Tesla") with earlier spans, and adds closed-class glue words ("is", "the", "'s"). An atom cannot contain a name or number that isn't in the input, and values are copied byte for byte ("8846m").

Status: research prototype (v0.2). v0.2 handles chat and agent text: first person, greetings, hedges ("I think …", which are unwrapped) and opinions (emitted as atoms with low verifiable).

Usage

The loading code is the atomizer package at NISH1001/atomizer. That repo is private for now, so access is on request. The weights are a plain PyTorch state dict (model.pt) plus atomizer.json (backbone, settings, glue and insertion vocabularies).

uv add "atomizer @ git+ssh://git@github.com/NISH1001/atomizer"     # needs access to the repo
from atomizer import Atomizer

atomizer = Atomizer.load("nishparadox/atomizer-gliner-small", revision="v0.2")

for atom in atomizer.atomize("I am good. My name is Nish. I think the height of mount everest is 8846m."):
    print(atom.text, round(atom.verifiable, 2), round(atom.checkworthy, 2))
# I am good.                                0.0   0.0
# My name is Nish.                          0.0   0.0
# The height of mount everest is 8846m.     0.99  0.93

atomizer.atomize(text, min_verifiable=0.5)                        # only checkable facts
print(atomizer.explain(text).show())                              # every head's predictions
atomizer.critic_scores(["We should invest more in education."])   # score claims you already have

atomize returns every atom by default, opinions and greetings included, each with its scores. Filter with min_verifiable / min_checkworthy.

Model

GLiNER small v2.5 (DeBERTa-v3-small encoder, BiLSTM, span layer) plus four heads:

Head Predicts
anchor per word: does an atom start here
pair per (atom, word): KEEP / SUB / DROP, and glue after the word
antecedent per reference: which earlier span it means
critic per atom: verifiable, checkworthy

Assembly from these predictions is deterministic. Texts longer than 1,024 tokens are atomized in sentence windows with read-only context. Details: ARCHITECTURE.md.

Training

Three stages, each continuing from the previous checkpoint. No API calls were made to create training data. Every run is logged in W&B nasa-impact/atomizer (team project). The full log, with commands, curves and timings, is in EXPERIMENTS.md.

Datasets

Dataset Role Labels
chentong00/propositionizer-wiki-data atoms: 42,857 Wikipedia passages, ~7.7 propositions each GPT-4, by the dataset authors (Chen et al., Dense X Retrieval)
Nithiwat/claimbuster critic: 23,533 US-debate sentences (check-worthy fact / unimportant fact / non-factual) human
iai-group/clef2024_checkthat_task1_en critic: CLEF CheckThat! 2024 task 1, check-worthiness yes/no (train 22,501, dev 1,032, test 318) human
nishparadox/atomizer-chat-synthetic@v0.1 atoms + per-atom critic labels on chat, agent, science and news text, 12 types (train 995, test 204) AI-written (Claude); see its disclaimer
chat templates (scripts/make_chat.py) hedged facts, opinions, insults, personal and social statements, questions; with noise rule-based, generated locally
sihaochen/propsegment (PropSegmEnt) test only (its atoms are fragments) human

ClaimBuster and CheckThat! come from the same debate transcripts. Every training sentence that also appears in a validation or test split was removed (3,279 sentences), so the critic is not scored on sentences it trained on.

Stages

Stage 1: atomizer Stage 2: critic → v0.1 Stage 3: chat → v0.2
starts from gliner-community/gliner_small-v2.5 stage 1, final v0.1
data per epoch Propositionizer, 42,857 Propositionizer 15,000 + ClaimBuster + CheckThat! 2024 (58,681) Propositionizer 27,076 (30%) + critic 27,076 (30%) + templates 5,415 (6%) + synthetic 995 (all) = 60,562
epochs / steps 3 / 8,037 1 / 3,668 (released: step 1,000) 2 / 3,780 (released: last)
batch 16 16 32
learning rates (encoder / GLiNER layers / heads) 3e-5 / 1e-4 / 5e-4 2e-5 / 5e-5 / 5e-4 2e-5 / 5e-5 / 5e-4
hardware, time Apple M3 Max (MPS), 2 h 10 min Apple M3 Max, ~40 min 1× A100 80GB, 16 min
W&B run gliner-small-v2.5-propwiki-3ep gliner-small-v2.5-propwiki-critic m1-mix-syn

Shared by all stages:

  • AdamW (fused), weight decay 0.01, gradient clip 1.0
  • cosine schedule, 5% warmup, bf16 autocast, dropout 0.1
  • insertion vocabulary of 300 words, 2,000 glue phrases
  • anchor threshold 0.35, tuned on Propositionizer validation

Stage 3 also uses bucketed shuffling: each epoch, batches are rebuilt from shuffled chunks sorted by length, so they mix datasets (70% of batches) at 5% padding.

How labels are made

An aligner turns each gold atom into per-word targets:

  • an anchor: the atom's first word no other atom uses
  • KEEP / SUB / DROP for each word
  • glue for each gap
  • an antecedent span for each reference

Gold atoms that cannot be built this way are skipped, and the anchor loss is masked on their words. That covers reordering, verb changes and words not in the text: ~26–30% of GPT-4 propositions, and ~0% of the templates and synthetic data. The critic learns from gold atoms (per-atom labels, or "verifiable" for propositions) and from whole labelled sentences, with the atomization losses switched off for those.

Stage 3 ablations

Same data and settings, one difference each; test sets below:

Run Difference synthetic soft F1 critic AUC (synthetic atoms) chat held-out Propositionizer PropSegmEnt CheckThat! F1
M1 = v0.2 — 0.897 0.994 0.988 0.733 0.741 0.866
M2 + glue penalty 0.5 0.897 0.994 0.989 0.733 0.742 0.866
M3 no synthetic data 0.872 0.947 0.989 0.728 0.730 0.872
M1 at step 500 0.26 epochs instead of 2 0.891 0.980 0.975 0.744 0.760 0.851
  • The synthetic data mainly improves the critic on realistic text (+0.047 AUC).
  • The glue penalty has no effect.
  • Longer training improves chat and the critic but costs a little on Wikipedia-style text. The release favours chat and agent inputs.

Evaluation

Test splits only, none used for training or checkpoint selection. Atom sets are scored with soft F1 (best-match token F1 between predicted and reference atoms).

Test set v0.1 v0.2
synthetic chat/agent texts (204), soft F1 0.774 0.897
…critic verifiable on their gold atoms, AUC 0.936 0.994
chat templates with held-out patterns, facts and names, soft F1 0.525 0.988
known v0.1 failure cases (8), soft F1 0.565 1.000
Propositionizer test (500), soft F1 0.737 0.733
PropSegmEnt test (500, human-labelled), soft F1 0.666 0.741
CLEF CheckThat! 2024 task 1 EN official test (315), checkworthy F1 / AUC 0.827 / 0.974 0.866 / 0.981
fact-assessor claim cases (19), verifiable F1 1.00 1.00
words not in the input (invented) 0 0

The share of reference atoms a span model can express at all is about 70% on GPT-4 propositions, so word-exact match on that set is low.

Limitations

  • The chat gains are measured on generated data. The templates and the AI-written synthetic set are cleaner than real user text. The 8 failure cases are real but few.
  • Over-generates on long scientific passages: about 15 atoms per passage where an LLM atomizer writes about 8. This was measured on v0.1 and not re-measured.
  • One reference resolves to one span. "The study" cannot become a description assembled from several places, and event references ("That loss raised sea level…") stay unresolved.
  • No reordering and no verb changes. "carrying" cannot become "carried". An opinion word placed before the fact ("the ridiculously expensive X costs…") cannot be separated.
  • The anchor threshold (0.35) controls how many atoms come out. Lower it for more atoms.

Versions

Tag Content
v0.2 stage 3: chat, hedges, opinions; better critic
v0.1 stages 1–2: Wikipedia atomizer + critic

License

The weights are released under CC-BY-4.0. The atomizer code is MIT.

The training data has its own terms, which you should check for your use:

  • Propositionizer labels were generated with GPT-4 by the dataset's authors (OpenAI terms).
  • ClaimBuster is CC-BY-SA 4.0.
  • CLEF CheckThat! 2024 is under the CLEF lab's terms.
  • atomizer-chat-synthetic was written by an AI model (Claude, by Anthropic).
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nishparadox/atomizer-gliner-small

Finetuned
(2)
this model

Dataset used to train nishparadox/atomizer-gliner-small