fengxr93's picture
TS-Align benchmark reproduction bundles + canonical eval data + dataset sources
c3efe57
|
Raw
History Blame Contribute Delete
1.89 kB

TSQA 90/10 (tsqa_9010) — faithful rebuild from raw

Reproduces the canonical tsqa_9010 split byte-for-byte from the raw Time-MQA Open_Ended_QA/open_ended_QA.csv. Verified: test_requests.jsonl md5 = 659c50537a809a74215e0a56f00c036e (3519 rows), identical to the set champion / joint / all TSQA evals were run on.

One command

bash build_tsqa_9010.sh

Override raw CSV via CSV=.... Outputs into this dir: all_requests.jsonl (35758) · test_requests.jsonl (3519) · train/{sft_pairs,raw_values}.jsonl (31556).

Pipeline (3 steps — this is exactly how the canonical was built, 2026-06-25)

  1. _build/adapters/timemqa_adapter.py — raw open_ended_QA.csvall_requests.jsonl (35758 requests; auto task_type: open_ended→text, MCQ/TF→choice; skips 1871 malformed).
  2. build_9010.pyrandom.Random(42).shuffle → 10% test / 90% train (3576/32182), then de-leak prompts (numeric arrays → <ts></ts>, strip trailing "Time series:" suffix).
  3. fix_ts.py — drop rows whose prompt has ≠1 <ts> (multi-series questions) → 3519/31556. This second filter is what distinguishes the canonical from a naive build_9010-only split (which stops at 3576).

Files

  • build_tsqa_9010.sh — driver (runs all 3 steps + md5 check)
  • build_9010.py — step 2 (seed=42 split + de-leak)
  • fix_ts.py — step 3 (drop multi-<ts>)
  • _build/{schema.py, adapters/timemqa_adapter.py} — adapter + its schema dep
  • raw source: ../Benchmarks/TSQA/Open_Ended_QA/open_ended_QA.csv (md5 489063300176fb89a2cd365afa6070eb)

Provenance

Recovered from __former_session_archive/2026-06-25 (session 3e3f5433): build_9010 @ 01:52, fix_ts (multi-<ts> removal) @ 03:43. The middle all_requests.jsonl = timemqa_adapter over the single Open_Ended_QA CSV (which itself contains open-ended + MCQ + T/F formats).