timeagent / RUN.md
roh8exe's picture
Upload folder using huggingface_hub
60b21d3 verified
|
Raw
History Blame Contribute Delete
2.42 kB

TimeAgent β€” Llama-3.2-1B baseline: finish the ECG eval

This bundle lets another server complete the ONLY remaining step for the Llama-3.2-1B OpenTSLM-SoftPrompt baseline: the stage-5 ECG-QA-CoT evaluation. Stages 1-4 (TSQA, M4, HAR, Sleep) are fully complete, and ECG training is done (early-stopped @ epoch 14). Only the ECG generation eval remains β€” it kept OOM-crashing on our shared/contended GPUs.

Contents

  • results/Llama_3_2_1B/ β€” all stage checkpoints + metrics (stages 1-4 have metrics.json so they're skipped; stage5_ecg_cot has the trained checkpoint, no metrics.json β†’ the eval will run).
  • data/ecg_qa_cot/ β€” ECG-QA-CoT CSV splits (train/val/test).
  • data/ptbxl.tar.gz β€” PTB-XL records, tarred. On target: cd data && tar xzf ptbxl.tar.gz before running.
  • code/OpenTSLM/ β€” the code (includes our fixes: greedy eval do_sample=False, sdpa attention, left-padded batched generation, EVAL_BATCH_SIZE env).
  • requirements-key.txt β€” key package versions we ran with.

Setup on the target server

# 1) python env (Python 3.13; a venv or conda env)
python3 -m venv venv && source venv/bin/activate
pip install torch transformers peft wfdb datasets accelerate huggingface_hub numpy scipy pandas einops open_flamingo
#   (match requirements-key.txt if you hit version issues)

# 2) wire the checkpoints + data into the code tree (symlinks)
cd code/OpenTSLM
ln -sfn "$(realpath ../../results)"      results
ln -sfn "$(realpath ../../data)"         src/data

# 3) HF token (Llama-3.2-1B is gated; the base model auto-downloads)
export HF_TOKEN=<your_hf_token_with_llama_access>
export HF_HOME=/path/with/space/hf_cache

Run the ECG eval

cd code/OpenTSLM
# On a CLEAN 95GB GPU (no other tenants) you can batch the eval for speed:
EVAL_BATCH_SIZE=4 python3 -u curriculum_learning.py --model OpenTSLMSP --llm_id meta-llama/Llama-3.2-1B
#   If it OOMs, drop to EVAL_BATCH_SIZE=2, or omit it for batch=1 (slow but always fits).

It will: skip stages 1-4 (metrics exist) β†’ load the stage5 ECG checkpoint β†’ skip training (done) β†’ run the ECG eval β†’ write results/Llama_3_2_1B/OpenTSLMSP/stage5_ecg_cot/results/metrics.json. When that file appears, the 5-stage baseline is COMPLETE.

Note: eval is greedy (deterministic). ECG sequences are ~7.7k tokens, so batched generation is memory-heavy β€” batch=4 needs a mostly-free 95GB GPU.