| # Benchmarks β dataset sources |
|
|
| This directory holds the **third-party time-series benchmarks** used to evaluate |
| TS-Align against peer models, together with our own OOS test set. The table below |
| documents the source of every dataset **except `TS_Caption_test`** (that one is our |
| own β see the repo root / `results/ts-align/` README). |
|
|
| | Dataset | Source | What it is | |
| |---|---|---| |
| | **TSQA** | **Time-MQA** β arXiv:[2503.01875](https://arxiv.org/abs/2503.01875) Β· HF: [Time-MQA](https://huggingface.co/Time-MQA) (ACL 2025 Findings) | The TSQA dataset from Time-MQA: ~200k QA pairs over 12 real-world domains Γ 5 task types (forecasting, imputation, anomaly detection, classification, open-ended reasoning). We use the open-ended reasoning split (true/false, multiple-choice, open-ended). | |
| | **ChatTS_test** | **ChatTS** β arXiv:[2412.03104](https://arxiv.org/abs/2412.03104) Β· GitHub: [NetManAIOps/ChatTS](https://github.com/NetManAIOps/ChatTS) Β· data: Zenodo record `14349206` (VLDB 2025, Tsinghua Γ ByteDance) | The official `dataset_a` / `dataset_b` evaluation sets released with ChatTS (aligning time series with LLMs via synthetic data). | |
| | **MTBench** | **MTBench** β arXiv:[2503.16858](https://arxiv.org/abs/2503.16858) Β· HF: [GGLabYale/MTBench_*](https://huggingface.co/datasets/GGLabYale/MTBench_finance_aligned_pairs_short) (Yale GGLab) | A multimodal TS benchmark pairing time series with text: **finance** (news + stock movements) and **weather** (reports + temperature). Tasks: forecasting, trend analysis, news-driven QA. Used here as an OOD / cross-domain probe. | |
| | **TimeSeriesExam** | **TimeSeriesExam** β arXiv:[2410.14752](https://arxiv.org/abs/2410.14752) (Cai, Choudhry, Goswami, Dubrawski β Auton Lab, CMU; NeurIPS 2024 workshop) | A configurable multiple-choice exam of time-series understanding across 5 categories: pattern recognition, noise, similarity, anomaly detection, causality. Encoder-only style, no textual leakage. | |
| |
| ## Not third-party |
| - **`TS_Caption_test/`** β our own OOS test set (metric-QA + captioning), de-duplicated. Source/provenance documented in the project's `TS_Caption` dataset release, not here. |
| - **`TS_Caption_test_duped/`** β the pre-de-duplication version of the above, kept locally for reference only; **git-ignored, not published**. |
|
|
| ## Notes |
| - Each dataset is consumed through its `benchmark_eval/adapters/<name>_adapter.py` (requests built locally, data path passed via CLI). See the bundle `run.sh` / `_infer/` for the exact invocation. |
| - MTBench caveat: most finance/weather QA tasks embed long news text β they are largely *text* questions; only `weather_rain` is near-pure time series. |
|
|