ChatTS eval requests — rebuilt inference-ready from raw
Builds the ChatTS eval request files (inference-ready for run_ts_align.py) from the
official ChatTS eval datasets. bash build_chatts.sh — deterministic (adapters have no
seed/shuffle), so re-running is bit-identical; there is no stored "canonical" to diff
against because the requests were historically ephemeral (built to /tmp), so this rebuild
is the reference.
Source
../Benchmarks/ChatTS_test/14349206/{dataset_a.json, dataset_b.json} — ChatTS official
Evaluation Datasets, Zenodo DOI 10.5281/zenodo.14349206 (dir named after the record).
md5 dataset_a 1978d2ae… / dataset_b ffbd5722… (byte-identical to replication's copy).
dataset_a = 159 univariate samples · dataset_b = 400 multivariate (7-channel) samples.
6 request files (all inference-ready: raw_ts + prompt + gt + system_prompt)
Reference md5 (deterministic — build_chatts.sh reproduces these bit-for-bit):
| file | n | md5 | from | task | gt / scorer |
|---|---|---|---|---|---|
chatts_a_qa.jsonl |
159 | e5a34635b1c4… |
dataset_a | free-form QA (text) | scored by chatts_bridge (official RAGAS evaluate_batch_qa) |
chatts_a_align.jsonl |
283 | ba977bf8a720… |
dataset_a | numeric | gt = ChatTS generator attributes → dispatch rel_acc |
chatts_a_tsametrics.jsonl |
789 | b5ba8a6d83b7… |
dataset_a | numeric | gt = TS-Align own metric defs (fairest) → dispatch |
chatts_b_qa.jsonl |
400 | 4e910372f0b3… |
dataset_b | free-form QA (text) | chatts_bridge RAGAS |
chatts_b_mv.jsonl |
275 | 002f2ba2da2d… |
dataset_b | multivariate relation (text) | dispatch rougeL |
chatts_b_mvmetrics.jsonl |
1100 | f901fe17f633… |
dataset_b | multivariate numeric | dispatch rel_acc (channels anonymized Series-k) |
(combined md5sum *.jsonl | md5sum = 67430f546377583d0559804e563f5ed6)
Pipeline
Each _build/adapters/chatts_*_adapter.py parses dataset_a/b → requests (deterministic,
from schema import build_request). Needs numpy + scipy + statsmodels (for
chatts_tsalign_metrics). QA files are scored by chatts_bridge.py (needs the ChatTS repo
- OpenAI/GLM key); numeric/text files by
external_eval/dispatch.py.
Files
build_chatts.sh— driver_build/{schema.py, adapters/chatts_{adapter,align_adapter,mv_adapter,mv_metrics_adapter,tsalign_metrics_adapter}.py}- raw source:
../Benchmarks/ChatTS_test/14349206/