ts-align-benchmarks / README.md
fengxr93's picture
README: document the caption GLM-extraction step (metric key + statistical-property explanation)
3b8f190
|
Raw
History Blame Contribute Delete
6.54 kB
---
tags:
- time-series
- time-series-question-answering
- captioning
- multimodal
- grpo
base_model:
- Qwen/Qwen3-4B-Instruct-2507
- AutonLab/MOMENT-1-base
---
# TS-Align
**TS-Align** aligns multivariate time series with a language model for **numeric metric
question-answering and captioning**. A frozen [MOMENT](https://huggingface.co/AutonLab/MOMENT-1-base)
time-series encoder feeds a fully fine-tuned
[Qwen3-4B-Instruct](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507) through a per-channel
("2-tunnel") projection with attention pooling, trained end-to-end with GRPO reinforcement
learning against verifiable metric and caption rewards.
This repository is a **self-contained benchmark bundle**: the trained TS-Align checkpoints,
the canonical evaluation data, and a reproduction harness for every baseline we compare against.
## Contents
| Path | What |
|---|---|
| `Models/4level_grpo_min2/` | TS-Align checkpoint — metric-QA + captioning |
| `Models/joint_grpo_min2/` | TS-Align checkpoint — metric-QA + captioning + TSQA (jointly trained) |
| `Models/<peer>/` | third-party baselines — **placeholder only** (weights not redistributed) |
| `Benchmark_eval/` | canonical evaluation inputs + scorers (`TS_Caption_test`, `TSQA`, `ChatTS_test`) |
| `Benchmarks/` | source datasets and their provenance — see [`Benchmarks/README.md`](Benchmarks/README.md) |
| `results/` | one reproduction bundle per evaluated model (runner + `run.sh` + `README`) |
## Trained checkpoints
Both checkpoints share the same architecture — **MOMENT-1-base** (frozen encoder) +
**Qwen3-4B-Instruct-2507** (fully fine-tuned), joined by a 2-tunnel per-channel projection
(1-layer, 8-head channel transformer) with attention pooling gated at
`channel_pool_attn_min = 2`. Trained with GRPO (2000 steps, PATRA-balanced caption reward,
split-wise Gaussian tolerance σ = {univar 0.5, bivar 0.5, multivar 1.0}).
| checkpoint | training tasks |
|---|---|
| `4level_grpo_min2` | metric-QA + captioning |
| `joint_grpo_min2` | metric-QA + captioning + TSQA (joint) |
Each `best-checkpoint/` holds:
- `model-state-0000{1..6}.pt` — the trainable weights (Qwen + projection + special-token
embeddings). The MOMENT encoder is frozen and **not** shipped; load it from its base repo.
- `tokenizer/` — the Qwen tokenizer with 4 added special tokens: `<ts>` `</ts>` `<scale>` `</scale>`.
- `training.pt`**metadata only** (≈1.6 KB: shard manifest, special-token ids, trainer
scalars). The optimizer state is intentionally not included.
### How to load
You need the two base models plus the TS-Align inference code:
- MOMENT-1-base — <https://huggingface.co/AutonLab/MOMENT-1-base>
- Qwen3-4B-Instruct-2507 — <https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507>
- TS-Align loader — `infer/infer_local.py` in the companion `ts-align-rl` code repository
`load_checkpoint` reads `training.pt` for the shard manifest and special-token ids, then loads
the `model-state-*.pt` shards onto a freshly built MOMENT + Qwen3-4B model:
```bash
python infer/infer_local.py \
--moment-path <path>/MOMENT-1-base \
--llm-path <path>/Qwen3-4B-Instruct-2507 \
--tokenizer-path Models/joint_grpo_min2/best-checkpoint/tokenizer \
--checkpoint-path Models/joint_grpo_min2/best-checkpoint/training.pt \
--channel-pool-attn-min-channels 2 \
--series "10,12,11,13,12,14,13,15" \
--prompt "<ts></ts> What is the mean of this series?"
```
## Baseline placeholders
`Models/<peer>/` — ChatTS-14B, GPT-OSS-20B, Qwen2.5-7B-Instruct, Qwen3-8B, Qwen3-14B,
TimeOmni-1-7B, ITFormer-7B, MOMENT-1-large, Mistral-7B-v0.3, Time-MQA-Mistral-7B — each contain
only a `PLACE_MODEL_FILE_HERE.txt`. These are third-party models we do not redistribute:
download each from its official source and drop the weights in place to reproduce.
## Reproducing the benchmarks
Every evaluated model has a bundle under `results/<model>/` with a `run.sh`, an `_infer/`
harness, and its own README. Inputs are read from `Benchmark_eval/`; scoring uses a multi-kernel
tolerance metric for numeric QA, precision/recall for captions, and RougeL/accuracy for TSQA.
Predictions and scores are regenerated by the runners and are not checked in.
Caption precision/recall uses an NLP-extraction step: **GLM-5.1** (temperature=0, the same un-tuned prompt for all models) reads each caption and extracts the numeric value it predicts for each mentioned metric. The prompt provides GLM with the metric key **and an explanation of its statistical property**, so it can align free-text mentions to the correct ground-truth metrics; extracted values are then scored against the ground truth under the same tolerance kernel (precision = correct / extracted, recall = ground-truth metrics correctly covered). Extractor: `Benchmark_eval/TS_Caption_test/extract_metrics_withdef.py` with the key→explanation map `metric_defs.json`.
Provenance for the third-party benchmarks is documented in
[`Benchmarks/README.md`](Benchmarks/README.md); `TS_Caption_test` is our own out-of-sample set.
## Notes
- Large files (checkpoint shards, evaluation data) are stored with **Git LFS**.
- The `training.pt.full_bak` optimizer states are kept locally only and are not published.
## License
This repository's **own content** — the fine-tuned TS-Align checkpoints
(`Models/4level_grpo_min2/`, `Models/joint_grpo_min2/`) and the `TS_Caption_test` evaluation set
— is **not yet released under a license** (all rights reserved by the authors for now).
All **third-party components** keep their original licenses:
### Models
| Model | Original license |
|---|---|
| Qwen3-4B-Instruct-2507, Qwen2.5-7B-Instruct, Qwen3-8B, Qwen3-14B | Apache-2.0 |
| MOMENT-1-base, MOMENT-1-large | MIT |
| GPT-OSS-20B | Apache-2.0 |
| ChatTS-14B | Apache-2.0 |
| TimeOmni-1-7B | Apache-2.0 |
| Mistral-7B-v0.3 (and its bnb-4bit quantization) | Apache-2.0 |
| Time-MQA-Mistral-7B (LoRA adapter) | see the Time-MQA project |
| ITFormer-7B | see the model's original repository |
### Datasets
| Dataset | Original license |
|---|---|
| TSQA (Time-MQA) | Apache-2.0 |
| ChatTS_test (ChatTS) | composite, per source component (e.g. NAB: Apache-2.0, Weather: CC-BY-4.0); see the ChatTS release (Zenodo `14349206`) |
| MTBench (Yale GGLab) | see the dataset's Hugging Face page |
| TimeSeriesExam (CMU Auton Lab) | see the dataset's original source |
Sources and links for every dataset are in [`Benchmarks/README.md`](Benchmarks/README.md).