--- tags: - time-series - time-series-question-answering - captioning - multimodal - grpo base_model: - Qwen/Qwen3-4B-Instruct-2507 - AutonLab/MOMENT-1-base --- # TS-Align **TS-Align** aligns multivariate time series with a language model for **numeric metric question-answering and captioning**. A frozen [MOMENT](https://huggingface.co/AutonLab/MOMENT-1-base) time-series encoder feeds a fully fine-tuned [Qwen3-4B-Instruct](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507) through a per-channel ("2-tunnel") projection with attention pooling, trained end-to-end with GRPO reinforcement learning against verifiable metric and caption rewards. This repository is a **self-contained benchmark bundle**: the trained TS-Align checkpoints, the canonical evaluation data, and a reproduction harness for every baseline we compare against. ## Contents | Path | What | |---|---| | `Models/4level_grpo_min2/` | TS-Align checkpoint — metric-QA + captioning | | `Models/joint_grpo_min2/` | TS-Align checkpoint — metric-QA + captioning + TSQA (jointly trained) | | `Models//` | third-party baselines — **placeholder only** (weights not redistributed) | | `Benchmark_eval/` | canonical evaluation inputs + scorers (`TS_Caption_test`, `TSQA`, `ChatTS_test`) | | `Benchmarks/` | source datasets and their provenance — see [`Benchmarks/README.md`](Benchmarks/README.md) | | `results/` | one reproduction bundle per evaluated model (runner + `run.sh` + `README`) | ## Trained checkpoints Both checkpoints share the same architecture — **MOMENT-1-base** (frozen encoder) + **Qwen3-4B-Instruct-2507** (fully fine-tuned), joined by a 2-tunnel per-channel projection (1-layer, 8-head channel transformer) with attention pooling gated at `channel_pool_attn_min = 2`. Trained with GRPO (2000 steps, PATRA-balanced caption reward, split-wise Gaussian tolerance σ = {univar 0.5, bivar 0.5, multivar 1.0}). | checkpoint | training tasks | |---|---| | `4level_grpo_min2` | metric-QA + captioning | | `joint_grpo_min2` | metric-QA + captioning + TSQA (joint) | Each `best-checkpoint/` holds: - `model-state-0000{1..6}.pt` — the trainable weights (Qwen + projection + special-token embeddings). The MOMENT encoder is frozen and **not** shipped; load it from its base repo. - `tokenizer/` — the Qwen tokenizer with 4 added special tokens: `` `` `` ``. - `training.pt` — **metadata only** (≈1.6 KB: shard manifest, special-token ids, trainer scalars). The optimizer state is intentionally not included. ### How to load You need the two base models plus the TS-Align inference code: - MOMENT-1-base — - Qwen3-4B-Instruct-2507 — - TS-Align loader — `infer/infer_local.py` in the companion `ts-align-rl` code repository `load_checkpoint` reads `training.pt` for the shard manifest and special-token ids, then loads the `model-state-*.pt` shards onto a freshly built MOMENT + Qwen3-4B model: ```bash python infer/infer_local.py \ --moment-path /MOMENT-1-base \ --llm-path /Qwen3-4B-Instruct-2507 \ --tokenizer-path Models/joint_grpo_min2/best-checkpoint/tokenizer \ --checkpoint-path Models/joint_grpo_min2/best-checkpoint/training.pt \ --channel-pool-attn-min-channels 2 \ --series "10,12,11,13,12,14,13,15" \ --prompt " What is the mean of this series?" ``` ## Baseline placeholders `Models//` — ChatTS-14B, GPT-OSS-20B, Qwen2.5-7B-Instruct, Qwen3-8B, Qwen3-14B, TimeOmni-1-7B, ITFormer-7B, MOMENT-1-large, Mistral-7B-v0.3, Time-MQA-Mistral-7B — each contain only a `PLACE_MODEL_FILE_HERE.txt`. These are third-party models we do not redistribute: download each from its official source and drop the weights in place to reproduce. ## Reproducing the benchmarks Every evaluated model has a bundle under `results//` with a `run.sh`, an `_infer/` harness, and its own README. Inputs are read from `Benchmark_eval/`; scoring uses a multi-kernel tolerance metric for numeric QA, precision/recall for captions, and RougeL/accuracy for TSQA. Predictions and scores are regenerated by the runners and are not checked in. Caption precision/recall uses an NLP-extraction step: **GLM-5.1** (temperature=0, the same un-tuned prompt for all models) reads each caption and extracts the numeric value it predicts for each mentioned metric. The prompt provides GLM with the metric key **and an explanation of its statistical property**, so it can align free-text mentions to the correct ground-truth metrics; extracted values are then scored against the ground truth under the same tolerance kernel (precision = correct / extracted, recall = ground-truth metrics correctly covered). Extractor: `Benchmark_eval/TS_Caption_test/extract_metrics_withdef.py` with the key→explanation map `metric_defs.json`. Provenance for the third-party benchmarks is documented in [`Benchmarks/README.md`](Benchmarks/README.md); `TS_Caption_test` is our own out-of-sample set. ## Notes - Large files (checkpoint shards, evaluation data) are stored with **Git LFS**. - The `training.pt.full_bak` optimizer states are kept locally only and are not published. ## License This repository's **own content** — the fine-tuned TS-Align checkpoints (`Models/4level_grpo_min2/`, `Models/joint_grpo_min2/`) and the `TS_Caption_test` evaluation set — is **not yet released under a license** (all rights reserved by the authors for now). All **third-party components** keep their original licenses: ### Models | Model | Original license | |---|---| | Qwen3-4B-Instruct-2507, Qwen2.5-7B-Instruct, Qwen3-8B, Qwen3-14B | Apache-2.0 | | MOMENT-1-base, MOMENT-1-large | MIT | | GPT-OSS-20B | Apache-2.0 | | ChatTS-14B | Apache-2.0 | | TimeOmni-1-7B | Apache-2.0 | | Mistral-7B-v0.3 (and its bnb-4bit quantization) | Apache-2.0 | | Time-MQA-Mistral-7B (LoRA adapter) | see the Time-MQA project | | ITFormer-7B | see the model's original repository | ### Datasets | Dataset | Original license | |---|---| | TSQA (Time-MQA) | Apache-2.0 | | ChatTS_test (ChatTS) | composite, per source component (e.g. NAB: Apache-2.0, Weather: CC-BY-4.0); see the ChatTS release (Zenodo `14349206`) | | MTBench (Yale GGLab) | see the dataset's Hugging Face page | | TimeSeriesExam (CMU Auton Lab) | see the dataset's original source | Sources and links for every dataset are in [`Benchmarks/README.md`](Benchmarks/README.md).