gliner_agents_v3

Fine-tune of urchade/gliner_large-v2.1 for data-use mention extraction from document text.

Annotation workflow

Labels were produced with a locally run LLM through the passage-annotation harness, with human-in-the-loop passage review and adjudication under the project doctrine.

  • dataset: rafmacalaba/datause-agents-v3 (config gliner)
  • holdout: 2,108 passages, 1,918 spans
  • origins scored: fcv_pads_east_africa, prwp, reliefweb, umar_pads

Labels

  • NAMED_DATA — a proper name, title, or acronym of a specific data source
  • DESCRIPTIVE_DATA — a source described in words but not named
  • VAGUE_DATA — generic data wording with no identifiable source

Entity matching is on span text, so these classes never enter the metrics above: they describe what the model was asked to separate, not what it is scored on.

Training

  • base model: urchade/gliner_large-v2.1
  • dataset: rafmacalaba/datause-agents-v3 (config gliner)
  • label mode: keep (prompt: NAMED_DATA, DESCRIPTIVE_DATA, VAGUE_DATA)
  • corpus: all
  • epochs: 5
  • learning rate: 5e-06
  • batch size: 16
  • precision: bf16

Evaluation (holdout)

thr tp fp fn precision recall f0.5 f1
0.10 1884 1955 29 0.4908 0.9848 0.5455 0.6551
0.20 1863 1444 50 0.5634 0.9739 0.6152 0.7138
0.30 1838 1101 75 0.6254 0.9608 0.6723 0.7576
0.40 1799 759 114 0.7033 0.9404 0.7406 0.8047
0.50 1725 385 188 0.8175 0.9017 0.8331 0.8576
0.60 1446 125 467 0.9204 0.7559 0.8820 0.8301
0.70 912 53 1001 0.9451 0.4767 0.7899 0.6338

Best F0.5: 0.8820 (thr=0.6) Best F1: 0.8576 (thr=0.5)

Evaluation breakdown (holdout)

group examples spans thr precision recall f0.5 f1
overall 2108 1918 0.60 0.9204 0.7559 0.8820 0.8301
None 0 0 0.60 0.9204 0.7559 0.8820 0.8301
fcv_pads_east_africa 49 7 0.70 1.0000 0.2857 0.6667 0.4444
prwp 670 510 0.60 0.9335 0.7199 0.8812 0.8129
reliefweb 112 56 0.60 0.8636 0.6786 0.8190 0.7600
umar_pads 1277 1345 0.60 0.9195 0.7736 0.8861 0.8403

Per-label (overall)

label examples spans thr precision recall f0.5 f1
NAMED_DATA 2108 1360 0.60 0.8120 0.7970 0.8090 0.8045
DESCRIPTIVE_DATA 2108 548 0.60 0.7125 0.3120 0.5670 0.4340
VAGUE_DATA 2108 10 0.10 0.0000 0.0000 0.0000 0.0000
Origin breakdown (per-origin metrics)
origin examples spans thr precision recall f0.5 f1
fcv_pads_east_africa 49 7 0.70 1.0000 0.2857 0.6667 0.4444
prwp 670 510 0.60 0.9335 0.7199 0.8812 0.8129
reliefweb 112 56 0.60 0.8636 0.6786 0.8190 0.7600
umar_pads 1277 1345 0.60 0.9195 0.7736 0.8861 0.8403
Per-label details
label examples spans thr precision recall f0.5 f1
NAMED_DATA 2108 1360 0.60 0.8120 0.7970 0.8090 0.8045
DESCRIPTIVE_DATA 2108 548 0.60 0.7125 0.3120 0.5670 0.4340
VAGUE_DATA 2108 10 0.10 0.0000 0.0000 0.0000 0.0000

Where it fails (holdout, score threshold 0.60)

Entity-level counts at the best-F0.5 threshold, matched with the same Jaccard >= 0.5 rule as the tables above — so these are the spans those numbers were computed from.

tp fp fn redundant (same entity found twice)
1447 125 466 8

False positives, by how close each comes to a gold span: spurious 112 · partial 10 · near_miss 3

Misses, by span length: 1-2 tokens 159 · 3-5 tokens 234 · 6+ tokens 73

Misses, by cause: 8 never emitted at all (proposal — no threshold moves these) · 458 emitted but scored below the cutoff (calibration).

Label errors on matched spans: 196 of 1447 (13.5%) — boundary right, class wrong.

Suppression: 940 passages carry no gold span at all; the model still emitted 67 spans on 940 of them.

label gold clusters recalled recall
DESCRIPTIVE_DATA 548 313 0.5712
NAMED_DATA 1355 1131 0.8347
VAGUE_DATA 10 3 0.3
Most frequent false positives
span passages
Demographic and Health Survey 2
Administrative data 2
World Bank Open Data 2
FERTIMAP 2
Crop Growth Monitoring system 2
2021 data as per Atlas methodology 2
UETCL BNG monitoring database 1
nutrition data 1
Bangladesh demographic and health survey 2007 1
Malawi De velopmental Assessment Test 1
Most frequent misses
span passages
Human Development Index 4
Gender Inequality Index 3
GHS 2
household survey data 2
Ministry of Finance 2
DHS 2
Statistical Capacity Index 2
Gender Development Index 2
NEGU Statistics 2
enquête EDTIC 2

Per-passage detail is in holdout_predictions.jsonl — one record per holdout passage with the passage text, every prediction and its score (kept marks the ones above the threshold above), and which of them matched. Re-derive any threshold from it. Gold spans are token indices; prediction spans are character offsets into text (GLiNER's units).

Downloads last month
157
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support