Paper-ready README: results table, quick start, structure, citation

#1
by nottygian - opened
Files changed (1) hide show
  1. README.md +130 -1
README.md CHANGED
@@ -1,3 +1,132 @@
1
  ---
2
- license: apache-2.0
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ license: other
3
+ license_name: mixed-apache-2.0-and-mit
4
+ tags:
5
+ - robotics
6
+ - world-model
7
+ - model-based-control
8
+ - jepa
9
+ - reinforcement-learning
10
+ - latent-planning
11
+ - iterative-control
12
  ---
13
+
14
+ # LePlanner: Amortized Iterative Control in Frozen Latent World Models
15
+
16
+ This repository contains the trained controllers, frozen world models, latent caches, and evaluation results for **LePlanner**:
17
+
18
+ > **LePlanner: Amortized Iterative Control in Frozen Latent World Models**
19
+ > [Authors], [Year].
20
+ > [Paper link β€” arXiv / venue]
21
+
22
+ LePlanner is an amortized iterative controller that plans through a **frozen JEPA world model** (LeWM) over a small number of learned refinement steps, trained with an **arrival–hold** objective. It matches or exceeds test-time search planners (CEM) while requiring an order-of-magnitude fewer predictor invocations and 8–15Γ— lower wall-clock per decision.
23
+
24
+ ## Results
25
+
26
+ | Environment | Method | Success | Notes |
27
+ |---|---|---|---|
28
+ | PushT | LePlanner (rh=1) | **94–96%** | arrival+hold checkpoint |
29
+ | PushT | LePlanner (rh=5) | **88–90%** | |
30
+ | PushT | CEM (rh=5) | 90% | baseline search planner |
31
+ | Reacher | LePlanner (rh=1) | **100%** | 50 episodes, seed 42 |
32
+ | TwoRooms | LePlanner (rh=1) | **100%** | 50 episodes, seed 42 |
33
+
34
+ LePlanner uses **~16Γ— fewer predictor rows per decision** than CEM and **8–15Γ— lower wall-clock** per decision.
35
+
36
+ ## Quick Start
37
+
38
+ ```python
39
+ from huggingface_hub import hf_hub_download
40
+
41
+ # Load a controller checkpoint
42
+ ckpt_path = hf_hub_download(
43
+ "SaltedLemon/lejepa-control-pusht",
44
+ "checkpoints/environment_controllers/reacher/controller.pt",
45
+ repo_type="model",
46
+ )
47
+
48
+ # Load a frozen world model
49
+ wm_path = hf_hub_download(
50
+ "SaltedLemon/lejepa-control-pusht",
51
+ "checkpoints/base_world_models/reacher/weights.pt",
52
+ repo_type="model",
53
+ )
54
+ ```
55
+
56
+ The training code is available at [`github.com/SaltedLemon/lejepa_control`](https://github.com/SaltedLemon/lejepa_control).
57
+
58
+ ## Repository Structure
59
+
60
+ ```
61
+ checkpoints/
62
+ β”œβ”€β”€ environment_controllers/ # Trained LePlanner controllers (per-environment)
63
+ β”‚ β”œβ”€β”€ reacher/
64
+ β”‚ β”‚ β”œβ”€β”€ controller.pt # 100% success, rh=1
65
+ β”‚ β”‚ └── density.pt # Behavior density model
66
+ β”‚ └── tworooms/
67
+ β”‚ β”œβ”€β”€ controller.pt # 100% success, rh=1
68
+ β”‚ └── density.pt
69
+ β”œβ”€β”€ controller/ # PushT iterative controller (phase-1)
70
+ β”œβ”€β”€ ablations/ # PushT training ablations
71
+ β”‚ └── ah_hold0.5/
72
+ β”‚ └── controller.pt # Headline PushT checkpoint (94-96%)
73
+ β”œβ”€β”€ base_world_models/ # Frozen LeWM world models
74
+ β”‚ β”œβ”€β”€ pusht/ # quentinll/lewm-pusht (MIT)
75
+ β”‚ β”œβ”€β”€ reacher/ # quentinll/lewm-reacher (MIT)
76
+ β”‚ └── tworooms/ # quentinll/lewm-tworooms (MIT)
77
+ β”œβ”€β”€ planner_*/ # Recursive & cross-attention planners
78
+ β”œβ”€β”€ manifold_transfer/ # Latent transfer experiments
79
+ └── density/ # PushT behavior density model
80
+
81
+ code/ # Training & evaluation scripts
82
+ β”œβ”€β”€ planner.py # LePlanner controller
83
+ β”œβ”€β”€ solver.py # Training/eval harness
84
+ β”œβ”€β”€ losses.py # Arrival-hold, action-Gaussian, support losses
85
+ └── scripts/ # train_planner.py, eval_planner.py, ...
86
+
87
+ latents/ # Pre-computed latent caches
88
+ β”œβ”€β”€ reacher/ # Latents, actions, episode metadata
89
+ └── tworoom_cls/
90
+
91
+ results/ # Raw evaluation logs
92
+ β”œβ”€β”€ eval/results.jsonl # PushT eval across planners
93
+ └── eval_planner/results.jsonl
94
+ ```
95
+
96
+ ## Key Checkpoints
97
+
98
+ | File | Description |
99
+ |---|---|
100
+ | `checkpoints/ablations/ah_hold0.5/controller.pt` | **Headline PushT controller** β€” 94–6% at rh=1 |
101
+ | `checkpoints/environment_controllers/reacher/controller.pt` | Reacher controller β€” 100% |
102
+ | `checkpoints/environment_controllers/tworooms/controller.pt` | TwoRooms controller β€” 100% |
103
+ | `checkpoints/base_world_models/pusht/weights.pt` | Frozen PushT world model (LeWM) |
104
+ | `checkpoints/base_world_models/reacher/weights.pt` | Frozen Reacher world model (LeWM) |
105
+ | `checkpoints/base_world_models/tworooms/weights.pt` | Frozen TwoRooms world model (LeWM) |
106
+
107
+ ## Training & Evaluation
108
+
109
+ The controller and planner checkpoints are trained in this project; the LeWM world model is **always frozen** while those planners are trained.
110
+
111
+ Training code: [`github.com/SaltedLemon/lejepa_control`](https://github.com/SaltedLemon/lejepa_control)
112
+
113
+ Each controller checkpoint embeds its training arguments, action statistics, step, and validation profile.
114
+
115
+ ## Citation
116
+
117
+ If you use this code or these checkpoints, please cite:
118
+
119
+ ```bibtex
120
+ @article{leplanner,
121
+ title = {LePlanner: Amortized Iterative Control in Frozen Latent World Models},
122
+ author = {[Authors]},
123
+ journal = {[Venue / arXiv]},
124
+ year = {[Year]},
125
+ url = {https://huggingface.co/SaltedLemon/lejepa-control-pusht}
126
+ }
127
+ ```
128
+
129
+ ## License
130
+
131
+ - **Controller/planner code and checkpoints:** as specified in the linked `lejepa_control` repository.
132
+ - **LeWM world model weights** (`base_world_models/pusht`, `reacher`, `tworooms`): MIT licensed upstream from `quentinll/lewm-*`.