MELP ECG encoder (reproduction)

The ECG tower of a MELP model trained from this fork: HKU-MedAI/MELP#4. Same architecture and packaging as fuyingw/MELP_Encoder, so it is a drop-in replacement. Original paper: From Token to Rhythm: A Multi-Scale Approach for ECG-Language Pretraining (ICML 2025), Wang, Xu & Yu — all credit for the method belongs to the authors.

import torch
from transformers import AutoModel

model = AutoModel.from_pretrained("xjc1022/MELP-Encoder-repro", trust_remote_code=True).eval()
ecg = torch.rand(1, 12, 5000)          # 12 leads, 10 s at 500 Hz, scaled to [0, 1]
with torch.no_grad():
    out = model(ecg)
# out["proj_ecg_emb"]    (B, 256)     rhythm level, the zero-shot embedding
# out["ecg_beat_emb"]    (B, 12, 256) beat level
# out["ecg_token_emb"]   (B, 128, 768) token level

Lead order is I, II, III, aVR, aVF, aVL, V1-V6, and each record is min-max scaled to [0, 1] over the whole 12x5000 array — the same preprocessing the training data used.

Zero-shot results

Six-dataset mean AUROC from scripts/zeroshot/test_zeroshot.py (the paper's protocol), test splits:

Rhythm Form Sub Super CPSC2018 CSN Average
This run 86.43 70.32 77.65 77.22 84.18 74.58 78.40
Paper (Table 3) 85.4 69.1 81.2 76.2 84.2 77.6 79.0

Those numbers describe the full MELP model this encoder came from: zero-shot needs the text tower as well, which is not part of this repo. What is published here is the ECG tower alone, for use as a feature extractor.

How it was trained

Text tower fuyingw/heart_bert, ECG tower initialised from an ECG-FM wav2vec2-CMSC checkpoint, trained on MIMIC-IV-ECG report pairs. 4x RTX 3090, batch 64/device, lr 1e-4, n_queries_contrast=12, loss weights 1.0 / 2.0 / 0.2. Best checkpoint by validation zero-shot AUROC, epoch 4.

Three things mattered more than any hyperparameter here, and all three are fixes in the linked PR:

  1. The shipped LR scheduler has no warmup, and without it both towers collapse within ~25 steps — every zero-shot AUROC comes out at exactly 0.5.
  2. Training on denoised waveforms while evaluating on raw ones opens a domain gap worth 5.6 AUROC points (70.75 -> 78.40 once the raw wfdb records are used).
  3. The peak is at epoch 4. Runs stopped earlier, or validated only at epoch boundaries, miss it.

Data

Pretrained on MIMIC-IV-ECG, which is PhysioNet credentialed-access data. These are model weights rather than data, but check the PhysioNet data use agreement before redistributing anything derived from them.

Downloads last month
24
Safetensors
Model size
65.6M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for xjc1022/MELP-Encoder-repro