federated-ehr-fm / README.md
oelemento's picture
Withhold per-site ensemble members (membership-inference caution); available on request
f1b8d8c verified
|
Raw
History Blame Contribute Delete
4.51 kB
metadata
license: other
license_name: federated-ehr-fm-research-license
license_link: LICENSE
extra_gated_prompt: >-
  These checkpoints are derivative works of credentialed intensive-care datasets
  (MIMIC-IV, eICU-CRD, HiRID via PhysioNet; SICdb). Access is granted only to
  users who already hold the relevant credentialed-access approvals and agree to
  the same data-use terms. Each request is reviewed individually by the authors.
extra_gated_fields:
  Full name: text
  Institution / affiliation: text
  Intended research use: text
  PhysioNet credentialed (MIMIC-IV / eICU / HiRID) and SICdb access held: checkbox
  I agree to the data-use terms and will not redistribute the weights or attempt re-identification: checkbox
tags:
  - electronic-health-records
  - foundation-model
  - federated-learning
  - clinical-prediction
  - ICU
pipeline_tag: other

Federated EHR Foundation Model (autoregressive, multi-institutional)

Research checkpoints accompanying Multi-Hospital Electronic Health Record Foundation Models Without Pooling Patient Data: A Comparison of Federated Learning and Inference-Time Ensembling (O. Elemento, Weill Cornell Medicine). Research artifact — not a medical device; not cleared or approved for clinical use.

Model details

  • Type: GPT-style decoder-only transformer (~31.6M parameters) trained with a next-token autoregressive objective (strict causal attention, cross-entropy over the full vocabulary) on tokenized patient sequences (diagnosis, procedure, medication, laboratory tokens). RoPE indexed by inter-visit day delta.
  • Architecture: 8 layers, 8 heads, d_model 512, d_ff 2048. Regularization (reg2): dropout 0.4, weight decay 0.5, label smoothing 0.1 (train-time only); the four-cohort cross-database FedAvg checkpoint uses a lighter recipe (dropout 0.3, wd 0.3, no label smoothing), which would not converge otherwise.
  • Released arms (MIMIC-IV vocabulary): centralized, fedavg, fedper, and fedper_lora, plus the tokenizer / vocabulary. The five per-site ensemble members are not included in this repository: they are the only arms that show a measurable membership signal under white-box attack, so they are withheld and provided to credentialed researchers on direct request (individual review). The cross-database (Table 1 / Figure 4) results are reproducible from the released code and the aggregate per-run metrics in the code repository.
  • Developers: Englander Institute for Precision Medicine & Department of Systems and Computational Biomedicine, Weill Cornell Medicine.

Intended use

Reproduction and verification of the manuscript's results, and downstream research on multi-institutional EHR modeling under appropriate governance. Not for clinical decision-making, patient-level deployment, or re-identification.

Training data

MIMIC-IV (Beth Israel Deaconess, US), eICU-CRD (US multi-hospital), HiRID (Switzerland) — all PhysioNet-credentialed — and SICdb (Austria). Diagnosis/ procedure tokens use ICD-10/ICD-9 chapter prefixes; labs mapped to LOINC; medications to ATC where available. Held-out evaluation is leakage-clean (target admission cut at the second-to-last separator; single-stay mortality scored from the first 12-hour block).

Performance (reg2, held-out test)

On MIMIC-IV, per-condition AUROC 0.84–0.91; FedAvg and inference-time ensembling are statistically indistinguishable from centralized training. Across four real ICU databases, the inference-time ensemble tracks centralized pooling while FedAvg degrades by 0.02–0.16 AUROC under cross-database heterogeneity (see manuscript Table 1, Figures 2–4).

Limitations & risks

~31.6M parameters (small vs frontier EHR models); simulated non-IID partitions for the homogeneous-federation experiments; gated release reflects the restricted, credentialed nature of the source data. A membership-inference audit of these checkpoints is reported in the manuscript: the released centralized/FedAvg models show near-chance membership signal (AUROC 0.52-0.55). The individual per-site members can leak under white-box attack (AUROC up to 0.78, an upper bound confounded by per-site distribution shift), so they are withheld from this repository and shared only on direct request to credentialed researchers. The deployed ensemble exposes only its averaged output, whose membership signal is near-random (AUROC 0.55).

Citation

See the accompanying manuscript. Code: https://github.com/oelemento/gpt-ehr-poc