tensorlink-dev's picture
Upload README.md with huggingface_hub
c86bdde verified
|
Raw
History Blame Contribute Delete
1.83 kB
---
license: apache-2.0
tags: [time-series, forecasting, foundation-model, tokenizer-ablation, toto-2-recipe]
---
# TSFM tokenizer ablation — Toto-2-4m-recipe clone (v10)
VQ-VAE OPTIMIZATION ablation (codebook health + next-code objective vs raw patches) on a ~3.6M-param decoder-only patched transformer
trained on the OFFICIAL TempoPFN synthetic prior generators (cloned from
automl/TempoPFN; GP/KernelSynth families capped at 6,144 steps, so the long
pool uses the cheap families only) with contiguous patch
masking, a 9-level quantile head, NorMuon+AdamW, and index RoPE at
time-scaled positions. Horizon decoding is fixed patch-32 in every arm; only
the CONTEXT tokenizer varies:
| arm | tokenizer | history | ctx tokens |
|-----|-----------|---------|-----------|
| T0 | fixed-32 (control) | 4,096 | 128 |
| T1 | pyramid, iso-context | 4,096 | 44 |
| T2 | pyramid, iso-token | 16,384 | 128 |
| T3 | adaptive equal-surprise | 16,384 | 128 |
Each subfolder is one (arm, seed) run: `model.pt` (final), `ckpt_15000.pt`
(rank-stability snapshot), `config.json`, `results.json` (dev-GIFT CRPS +
long-horizon probe). Dev metrics use a fixed 14-task GIFT-Eval subset — NOT
the full leaderboard; treat numbers as ablation-internal, not comparable to
published GIFT scores. Generated by the v10 experiment notebook.
## Results
```
=== transfer curve: GM-CRPS by checkpoint (dev-14, mean over seeds) ===
gm@30000 gm@100000 gm@200000 gm@300000 gm@final gm_final_std long_season params_m perp probe_amp n
arm
V0_champion_400k 0.2373 0.205 0.2126 0.2246 0.1991 NaN 0.2693 3.878 552.512 0.307 1
--- tripwire checks ---
```