hackhackhack66666 commited on
Commit
d58168b
Β·
verified Β·
1 Parent(s): 4194805

Add Blockwise-OAT baseline artifacts (README.md)

Browse files
Files changed (1) hide show
  1. README.md +130 -0
README.md CHANGED
@@ -1,3 +1,133 @@
1
  ---
2
  license: mit
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: mit
3
+ tags:
4
+ - robotics
5
+ - manipulation
6
+ - oat
7
+ - libero
8
+ - blockwise-decoding
9
  ---
10
+
11
+ # Blockwise-OAT β€” strict original-OAT baseline (LIBERO-10)
12
+
13
+ Paired evaluation of **autoregressive (AR)** vs **blockwise parallel tail** action-token
14
+ generation on a frozen [OAT](https://arxiv.org/abs/2602.04215) policy.
15
+
16
+ **HF repo:** [hackhackhack66666/Blockwise-OAT](https://huggingface.co/hackhackhack66666/Blockwise-OAT)
17
+ **Code branch:** `Blockwise-OAT` on [GadzhiAskhabaliev/OAT-BLT-Dense](https://github.com/GadzhiAskhabaliev/OAT-BLT-Dense)
18
+
19
+ ## Summary
20
+
21
+ | Metric | AR baseline | Blockwise (P=4, r=1) |
22
+ |--------|-------------|----------------------|
23
+ | LIBERO-10 mean SR | **58.73% Β± 0.18%** | **52.33% Β± 1.04%** |
24
+ | Ξ” vs AR | β€” | **-6.40 pp** |
25
+ | Decoder speedup (bs=1) | β€” | **1.16Γ—** |
26
+ | E2E `predict_action` (bs=1) | 36.4 ms | 30.1 ms (**1.21Γ—**) |
27
+ | Tail train epochs | β€” | 15 (final CE 3.0607) |
28
+
29
+ AR baseline exceeds the published OAT8 reference (~56.3%) on our cluster stack.
30
+ Blockwise achieves **>1Γ— decoder speedup** but **βˆ’6.4 pp** SR at 15 tail epochs (resume training planned).
31
+
32
+ ## Baseline artifacts (frozen)
33
+
34
+ | Component | Source |
35
+ |-----------|--------|
36
+ | Policy | [Mirageinv/oat β€” policy_ep-0250_sr-0.596.ckpt](https://huggingface.co/Mirageinv/oat) |
37
+ | Tokenizer | [Mirageinv/oat β€” tokenizer_ep-0950_mse-0.002.ckpt](https://huggingface.co/Mirageinv/oat) |
38
+ | Tail decoder | `checkpoints/original_oat_tail_p4_r1.pt` (this repo) |
39
+
40
+ ## Architecture & data flow
41
+
42
+ OAT encodes observations and generates **8 action tokens** `z₁…zβ‚ˆ`. Blockwise-OAT splits decoding:
43
+
44
+ ```
45
+ Obs (RGB + proprio) ──► Vision encoder ──► cond [B, T_o, d]
46
+ β”‚
47
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
48
+ β”‚ AR path (baseline) β”‚
49
+ β”‚ BOS ──► AutoregressiveModel.generate (8 steps) ──► z₁…zβ‚ˆ β”‚
50
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
51
+ β”‚
52
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
53
+ β”‚ Blockwise path β”‚
54
+ β”‚ BOS ──► generate_prefix (P=4 AR steps) ──► z₁…zβ‚„, h_prefix β”‚
55
+ β”‚ (z₁…zβ‚„, h_prefix) ──► ParallelTailDecoder (1 pass) ──► z₅…zβ‚ˆβ”‚
56
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
57
+ β–Ό
58
+ cat(z_prefix, z_tail) ──► OATTok.detokenize ──► action chunk
59
+ ```
60
+
61
+ **Inputs:** multi-view RGB, robot state, task id (same as OAT).
62
+ **Outputs:** `action` / `action_pred` tensors (identical shapes for AR and Blockwise).
63
+ **Trainable in this run:** only `ParallelTailDecoder` (~4.5M params, 0.90Γ— AR size).
64
+
65
+ ### Generation schedule
66
+
67
+ | Mode | AR forward passes | Tail passes |
68
+ |------|-------------------|-------------|
69
+ | Full AR | 8 | 0 |
70
+ | Blockwise P=4 | 4 | 1 |
71
+
72
+ ## Experiment protocol
73
+
74
+ 1. Download Mirageinv/oat policy + tokenizer.
75
+ 2. Train `ParallelTailDecoder` on `libero10_N500` with frozen policy (15 epochs, bs=64, lr=1e-4).
76
+ 3. Paired sim-eval: `50` episodes/task Γ— `3` seeds (`test_start_seed=1000`).
77
+ 4. Benchmarks: dataset / training / policy verification + wall-clock speed.
78
+
79
+ Cluster launcher: `scripts/cluster/run_blockwise_original_oat_baseline.sh` (`PHASE=B NUM_EXP=3`).
80
+
81
+ ## Visualizations
82
+
83
+ | Figure | Description |
84
+ |--------|-------------|
85
+ | ![AR eval](benchmarks/ar_sim_eval_dashboard.png) | AR per-task SR |
86
+ | ![Blockwise eval](benchmarks/blockwise_sim_eval_dashboard.png) | Blockwise per-task SR |
87
+ | ![Paired](benchmarks/paired_sr_comparison_dashboard.png) | Side-by-side per-task comparison |
88
+ | ![Speed](benchmarks/speed_benchmark_dashboard.png) | Decoder + E2E latency |
89
+ | ![Tail train](benchmarks/tail_training_dashboard.png) | Tail CE loss curve |
90
+ | ![Verify](benchmarks/verification_summary_dashboard.png) | Verification kit |
91
+
92
+ ## Repository layout
93
+
94
+ ```
95
+ checkpoints/original_oat_tail_p4_r1.pt # trained tail decoder
96
+ eval/ar_eval_log.json # AR sim metrics
97
+ eval/blockwise_eval_log.json # Blockwise sim metrics
98
+ benchmarks/*.json # verification + speed raw logs
99
+ benchmarks/*_dashboard.png # plots above
100
+ ```
101
+
102
+ ## Reproduce inference
103
+
104
+ ```bash
105
+ python scripts/eval_policy_sim.py \
106
+ -c output/baselines/original_oat/hf/policy_ep-0250_sr-0.596.ckpt \
107
+ -o output/eval/blockwise/ar \
108
+ --tokenizer-checkpoint output/baselines/original_oat/hf/tokenizer_ep-0950_mse-0.002.ckpt
109
+
110
+ python scripts/eval_policy_sim.py \
111
+ -c output/baselines/original_oat/hf/policy_ep-0250_sr-0.596.ckpt \
112
+ -o output/eval/blockwise/bw \
113
+ --use-blockwise --blockwise-prefix-len 4 --blockwise-refine-iters 1 \
114
+ --blockwise-tail-checkpoint checkpoints/original_oat_tail_p4_r1.pt \
115
+ --tokenizer-checkpoint output/baselines/original_oat/hf/tokenizer_ep-0950_mse-0.002.ckpt
116
+ ```
117
+
118
+ ## Citation
119
+
120
+ ```bibtex
121
+ @misc{liu2026oatorderedactiontokenization,
122
+ title={OAT: Ordered Action Tokenization},
123
+ author={Chaoqi Liu and Xiaoshen Han and Jiawei Gao and Yue Zhao and Haonan Chen and Yilun Du},
124
+ year={2026},
125
+ eprint={2602.04215},
126
+ archivePrefix={arXiv},
127
+ primaryClass={cs.RO}}
128
+ ```
129
+
130
+ ## Next steps
131
+
132
+ - Resume tail training (target 30+ epochs) and re-run paired eval.
133
+ - Ablate `refine_iters`, prefix length `P`, and learning-rate schedule.