gliner2_agents_v3

Fine-tune of fastino/gliner2-large-v1 (GLiNER2, span architecture) for data-use mention extraction (dataset / survey / census / registry mentions in economics research papers).

Annotation workflow

Labels were produced with a locally run LLM through the passage-annotation harness, with human-in-the-loop passage review and adjudication under the project doctrine. Holdout F1 is therefore agreement with the annotating agent, not owner accuracy.

  • dataset: rafmacalaba/datause-agents-v3 (config gliner2)
  • holdout: 2,108 passages, 1,916 spans
  • origins scored: fcv_pads_east_africa, prwp, reliefweb, umar_pads

Labels

  • NAMED_DATA — a proper name, title, or acronym of a specific data source
  • DESCRIPTIVE_DATA — a source described in words but not named
  • VAGUE_DATA — generic data wording with no identifiable source

Training

  • base model: fastino/gliner2-large-v1
  • dataset: rafmacalaba/datause-agents-v3 (gliner2 config)
  • epochs: 5
  • encoder LR: 1e-05
  • task LR: 0.0005
  • batch size: 8
  • precision: bf16

Evaluation (holdout, label-agnostic)

2 gold mention(s) in the holdout sit inside a run the word splitter glues into one token (a URL, EM-DAT_), so no GLiNER2-family model can address them as a span; they are removed from every split before training and scoring. Listed in holdout_metrics.json.

thr tp fp fn precision recall f0.5 f1
0.10 1752 1279 161 0.5780 0.9158 0.6241 0.7087
0.20 1673 909 240 0.6479 0.8745 0.6834 0.7444
0.30 1582 682 331 0.6988 0.8270 0.7211 0.7575
0.40 1466 479 447 0.7537 0.7663 0.7562 0.7600
0.50 1330 302 583 0.8150 0.6952 0.7878 0.7504
0.60 1122 196 791 0.8513 0.5865 0.7808 0.6945
0.70 925 125 988 0.8810 0.4835 0.7566 0.6244
0.80 703 59 1210 0.9226 0.3675 0.7085 0.5256
0.90 345 18 1568 0.9504 0.1803 0.5126 0.3032

Best F0.5: 0.7878 (thr=0.5) Best F1: 0.7600 (thr=0.4)

Where it fails (holdout, score threshold 0.50)

Entity-level counts at the reported threshold, matched with the same Jaccard >= 0.5 rule as the table above — these are the spans those numbers were computed from.

tp fp fn redundant (same entity found twice)
1330 302 583 222

False positives, by closeness to a gold span: near_miss 6 · partial 40 · spurious 293

Misses, by span length: 1-2 tokens 196 · 3-5 tokens 283 · 6+ tokens 104

Misses, by cause: 161 never emitted at all (proposal / abstention — no threshold moves these) · 422 emitted but scored below the cutoff (calibration).

Label errors on matched spans: 176 of 1330 — boundary right, class wrong.

Suppression: 941 passages carry no gold span at all; the model still emitted 168 spans on 518 of them.

label gold clusters recalled recall
DESCRIPTIVE_DATA 548 244 0.4453
NAMED_DATA 1355 1084 0.8000
VAGUE_DATA 10 2 0.2000

Per-passage detail is in holdout_predictions.jsonl — one record per holdout passage with the passage text, every prediction with its score and label (kept marks the ones above the threshold above), and which of them matched. Re-derive any threshold from it; gold and predictions are both span text plus character offsets.

Downloads last month
1
Safetensors
Model size
0.5B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support