File size: 1,831 Bytes
c86bdde
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
---
license: apache-2.0
tags: [time-series, forecasting, foundation-model, tokenizer-ablation, toto-2-recipe]
---
# TSFM tokenizer ablation — Toto-2-4m-recipe clone (v10)

VQ-VAE OPTIMIZATION ablation (codebook health + next-code objective vs raw patches) on a ~3.6M-param decoder-only patched transformer
trained on the OFFICIAL TempoPFN synthetic prior generators (cloned from
automl/TempoPFN; GP/KernelSynth families capped at 6,144 steps, so the long
pool uses the cheap families only) with contiguous patch
masking, a 9-level quantile head, NorMuon+AdamW, and index RoPE at
time-scaled positions. Horizon decoding is fixed patch-32 in every arm; only
the CONTEXT tokenizer varies:

| arm | tokenizer | history | ctx tokens |
|-----|-----------|---------|-----------|
| T0  | fixed-32 (control) | 4,096 | 128 |
| T1  | pyramid, iso-context | 4,096 | 44 |
| T2  | pyramid, iso-token | 16,384 | 128 |
| T3  | adaptive equal-surprise | 16,384 | 128 |

Each subfolder is one (arm, seed) run: `model.pt` (final), `ckpt_15000.pt`
(rank-stability snapshot), `config.json`, `results.json` (dev-GIFT CRPS +
long-horizon probe). Dev metrics use a fixed 14-task GIFT-Eval subset — NOT
the full leaderboard; treat numbers as ablation-internal, not comparable to
published GIFT scores. Generated by the v10 experiment notebook.

## Results
```
=== transfer curve: GM-CRPS by checkpoint (dev-14, mean over seeds) ===
                  gm@30000  gm@100000  gm@200000  gm@300000  gm@final  gm_final_std  long_season  params_m     perp  probe_amp  n
arm                                                                                                                              
V0_champion_400k    0.2373      0.205     0.2126     0.2246    0.1991           NaN       0.2693     3.878  552.512      0.307  1

--- tripwire checks ---

```