haiderAI's picture
Upload README.md
be36fec verified
|
Raw
History Blame Contribute Delete
4.06 kB
metadata
tags:
  - modernbert
  - ner
  - medical
  - pytorch
  - dataset-extraction
  - token-classification
  - long-context
license: apache-2.0

ModernBERT-large - Medical Dataset Name Extraction

A ModernBERT-large token classification model trained to extract dataset names from medical/scientific research papers. Uses answerdotai/ModernBERT-large as the encoder with a linear classification head and cross-entropy loss. Trained with an 8192-token context window — 16x longer than the SciBERT/CRF baselines (512 tokens).

Architecture

  • Encoder: ModernBERT-large (answerdotai/ModernBERT-large), 395M params, hidden size 1024, 8192-token context
  • Classification head: Linear(1024, 3) with CrossEntropyLoss
  • Mixed precision: bf16 on Blackwell/Ampere GPUs, fp16 + GradScaler fallback on Turing/Volta
  • Gradient checkpointing enabled to fit ModernBERT-large at 8K context on 16 GB GPUs
  • Differential LR: 5e-5 encoder / 1e-3 head, warmup 10%, AdamW

Entity Types

Entity Type Description
Dataset Names of datasets, corpus, collections, databases, benchmarks used in scientific research

Performance

Validation Set

Metric Precision Recall F1
seqeval (entity) - - 0.9025
Exact Match (chunk) 0.8689 0.9149 0.8913
Partial Match (chunk) 0.9379 0.9877 0.9622

Test Set (In-Distribution)

Metric Precision Recall F1
seqeval (entity) - - 0.9124
Exact Match (chunk) 0.8815 0.9222 0.9014
Partial Match (chunk) 0.9551 0.9992 0.9766

OOD Set (Out-of-Distribution)

Metric Precision Recall F1
seqeval (entity) - - 0.6673
Exact Match (chunk) 0.6584 0.7204 0.6880
Partial Match (chunk) 0.8290 0.9071 0.8663

Training Details

Parameter Value
Base model answerdotai/ModernBERT-large (395M)
Max sequence length 8192 tokens
Primary chunk size 6000 chars (~1800 BPE tokens)
Chunk overlap 500 chars
Negative sampling ratio 0.15
Entity-centered augmentation window=4000 chars, max 3/doc
Training documents 542 (+ 151 OOD held out)
Training examples 2324
Effective batch size 4 (batch 1 x accum 4)
Epochs 10
Precision bf16
GPU used NVIDIA GeForce RTX 5070 Ti

Comparison vs Baselines

All three models (CRF, SciBERT, ModernBERT-large) were trained on the same 542 manually annotated medical-dataset-mention documents using the same random seed (42), and evaluated on the same 151-document OOD set. All pipelines apply negative sampling (ratio 0.15) and entity-centered augmentation for class balance. ModernBERT uses 6000-char primary chunks and 4000-char augmentation windows to exploit its 8K context window — so chunk counts differ numerically from the 512-token baselines, but the underlying documents, seed, and balancing techniques are identical. Comparison is meaningful at the document / entity F1 level (see eval_metrics.json).

Usage

from inference import load_model, predict

model, tokenizer, id2label, config = load_model(".", device="cuda")
text = "We evaluated our model on the MIMIC-III dataset and the PhysioNet challenge corpus."
entities = predict(text, model, tokenizer, id2label, device="cuda")
for ent in entities:
    print(ent)

Dependencies

torch>=2.1
transformers>=4.48   # required for native ModernBERT support
seqeval

Limitations

  • Recognizes only dataset / corpus / database / benchmark names, not other biomedical entities
  • Although the encoder supports 8192 tokens, inputs are chunked at 6000 chars during inference to match training; cross-chunk entities are deduplicated by span
  • Trained on manually annotated medical research abstracts and full-text sections; generalization to other scientific domains is not guaranteed
  • Long-document inference is memory-heavy on small GPUs — use a quantized build or CPU-offload for the largest inputs