BERT NER on CoNLL-2003

Intended use

English newswire named entity recognition (PER, ORG, LOC, MISC) for a course experiment. Not a general-domain or safety-critical extractor.

Model and training

  • Base: bert-base-cased
  • Dataset: lhoestq/conll2003; splits train/validation/test; original shared task: CoNLL-2003.
  • Adaptation: full; initial seed 42; 3 epochs; batch 16; maximum sequence length 256.
  • Two compared alternatives: frozen BERT with trained token head and full fine-tuning. Head learning rate 1e-3; encoder learning rate 2e-5 when trained.
  • Labels: O, B-PER, I-PER, B-ORG, I-ORG, B-LOC, I-LOC, B-MISC, I-MISC. Supervise the first subtoken only; other positions -100.

Results

Entity-level strict IOB2 span F1 on held-out test: 0.9120; precision 0.9113; recall 0.9127; token accuracy 0.9824. Validation F1: 0.9455. Single seed; close differences may be noise.

Limitations

English Reuters news from 1996; domain and time shift can hurt performance. One label per word; truncated sentences beyond 256 subtokens lose supervised words. Named entities and noisy labels may differ across datasets. Review the upstream dataset terms before reuse or redistribution.

References

Downloads last month
10
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for Perry-DLC/upy-tds-bert-ner-conll2003