ModernBERT-large - Medical Dataset Name Extraction
A ModernBERT-large token classification model trained to extract dataset names from medical/scientific research papers. Uses answerdotai/ModernBERT-large as the encoder with a linear classification head and cross-entropy loss. Trained with an 8192-token context window — 16x longer than the SciBERT/CRF baselines (512 tokens).
Names of datasets, corpus, collections, databases, benchmarks used in scientific research
Performance
Validation Set
Metric
Precision
Recall
F1
seqeval (entity)
-
-
0.9025
Exact Match (chunk)
0.8689
0.9149
0.8913
Partial Match (chunk)
0.9379
0.9877
0.9622
Test Set (In-Distribution)
Metric
Precision
Recall
F1
seqeval (entity)
-
-
0.9124
Exact Match (chunk)
0.8815
0.9222
0.9014
Partial Match (chunk)
0.9551
0.9992
0.9766
OOD Set (Out-of-Distribution)
Metric
Precision
Recall
F1
seqeval (entity)
-
-
0.6673
Exact Match (chunk)
0.6584
0.7204
0.6880
Partial Match (chunk)
0.8290
0.9071
0.8663
Training Details
Parameter
Value
Base model
answerdotai/ModernBERT-large (395M)
Max sequence length
8192 tokens
Primary chunk size
6000 chars (~1800 BPE tokens)
Chunk overlap
500 chars
Negative sampling ratio
0.15
Entity-centered augmentation
window=4000 chars, max 3/doc
Training documents
542 (+ 151 OOD held out)
Training examples
2324
Effective batch size
4 (batch 1 x accum 4)
Epochs
10
Precision
bf16
GPU used
NVIDIA GeForce RTX 5070 Ti
Comparison vs Baselines
All three models (CRF, SciBERT, ModernBERT-large) were trained on the same 542 manually
annotated medical-dataset-mention documents using the same random seed (42), and evaluated
on the same 151-document OOD set. All pipelines apply negative sampling (ratio 0.15) and
entity-centered augmentation for class balance. ModernBERT uses 6000-char primary
chunks and 4000-char augmentation windows to exploit its 8K context window — so chunk
counts differ numerically from the 512-token baselines, but the underlying documents, seed,
and balancing techniques are identical. Comparison is meaningful at the document / entity F1
level (see eval_metrics.json).
Usage
from inference import load_model, predict
model, tokenizer, id2label, config = load_model(".", device="cuda")
text = "We evaluated our model on the MIMIC-III dataset and the PhysioNet challenge corpus."
entities = predict(text, model, tokenizer, id2label, device="cuda")
for ent in entities:
print(ent)
Dependencies
torch>=2.1
transformers>=4.48 # required for native ModernBERT support
seqeval
Limitations
Recognizes only dataset / corpus / database / benchmark names, not other biomedical entities
Although the encoder supports 8192 tokens, inputs are chunked at 6000 chars during inference to match training; cross-chunk entities are deduplicated by span
Trained on manually annotated medical research abstracts and full-text sections; generalization to other scientific domains is not guaranteed
Long-document inference is memory-heavy on small GPUs — use a quantized build or CPU-offload for the largest inputs