Paper-ready README: results table, quick start, structure, citation
#1
by nottygian - opened
README.md
CHANGED
|
@@ -1,3 +1,132 @@
|
|
| 1 |
---
|
| 2 |
-
license:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: mixed-apache-2.0-and-mit
|
| 4 |
+
tags:
|
| 5 |
+
- robotics
|
| 6 |
+
- world-model
|
| 7 |
+
- model-based-control
|
| 8 |
+
- jepa
|
| 9 |
+
- reinforcement-learning
|
| 10 |
+
- latent-planning
|
| 11 |
+
- iterative-control
|
| 12 |
---
|
| 13 |
+
|
| 14 |
+
# LePlanner: Amortized Iterative Control in Frozen Latent World Models
|
| 15 |
+
|
| 16 |
+
This repository contains the trained controllers, frozen world models, latent caches, and evaluation results for **LePlanner**:
|
| 17 |
+
|
| 18 |
+
> **LePlanner: Amortized Iterative Control in Frozen Latent World Models**
|
| 19 |
+
> [Authors], [Year].
|
| 20 |
+
> [Paper link β arXiv / venue]
|
| 21 |
+
|
| 22 |
+
LePlanner is an amortized iterative controller that plans through a **frozen JEPA world model** (LeWM) over a small number of learned refinement steps, trained with an **arrivalβhold** objective. It matches or exceeds test-time search planners (CEM) while requiring an order-of-magnitude fewer predictor invocations and 8β15Γ lower wall-clock per decision.
|
| 23 |
+
|
| 24 |
+
## Results
|
| 25 |
+
|
| 26 |
+
| Environment | Method | Success | Notes |
|
| 27 |
+
|---|---|---|---|
|
| 28 |
+
| PushT | LePlanner (rh=1) | **94β96%** | arrival+hold checkpoint |
|
| 29 |
+
| PushT | LePlanner (rh=5) | **88β90%** | |
|
| 30 |
+
| PushT | CEM (rh=5) | 90% | baseline search planner |
|
| 31 |
+
| Reacher | LePlanner (rh=1) | **100%** | 50 episodes, seed 42 |
|
| 32 |
+
| TwoRooms | LePlanner (rh=1) | **100%** | 50 episodes, seed 42 |
|
| 33 |
+
|
| 34 |
+
LePlanner uses **~16Γ fewer predictor rows per decision** than CEM and **8β15Γ lower wall-clock** per decision.
|
| 35 |
+
|
| 36 |
+
## Quick Start
|
| 37 |
+
|
| 38 |
+
```python
|
| 39 |
+
from huggingface_hub import hf_hub_download
|
| 40 |
+
|
| 41 |
+
# Load a controller checkpoint
|
| 42 |
+
ckpt_path = hf_hub_download(
|
| 43 |
+
"SaltedLemon/lejepa-control-pusht",
|
| 44 |
+
"checkpoints/environment_controllers/reacher/controller.pt",
|
| 45 |
+
repo_type="model",
|
| 46 |
+
)
|
| 47 |
+
|
| 48 |
+
# Load a frozen world model
|
| 49 |
+
wm_path = hf_hub_download(
|
| 50 |
+
"SaltedLemon/lejepa-control-pusht",
|
| 51 |
+
"checkpoints/base_world_models/reacher/weights.pt",
|
| 52 |
+
repo_type="model",
|
| 53 |
+
)
|
| 54 |
+
```
|
| 55 |
+
|
| 56 |
+
The training code is available at [`github.com/SaltedLemon/lejepa_control`](https://github.com/SaltedLemon/lejepa_control).
|
| 57 |
+
|
| 58 |
+
## Repository Structure
|
| 59 |
+
|
| 60 |
+
```
|
| 61 |
+
checkpoints/
|
| 62 |
+
βββ environment_controllers/ # Trained LePlanner controllers (per-environment)
|
| 63 |
+
β βββ reacher/
|
| 64 |
+
β β βββ controller.pt # 100% success, rh=1
|
| 65 |
+
β β βββ density.pt # Behavior density model
|
| 66 |
+
β βββ tworooms/
|
| 67 |
+
β βββ controller.pt # 100% success, rh=1
|
| 68 |
+
β βββ density.pt
|
| 69 |
+
βββ controller/ # PushT iterative controller (phase-1)
|
| 70 |
+
βββ ablations/ # PushT training ablations
|
| 71 |
+
β βββ ah_hold0.5/
|
| 72 |
+
β βββ controller.pt # Headline PushT checkpoint (94-96%)
|
| 73 |
+
βββ base_world_models/ # Frozen LeWM world models
|
| 74 |
+
β βββ pusht/ # quentinll/lewm-pusht (MIT)
|
| 75 |
+
β βββ reacher/ # quentinll/lewm-reacher (MIT)
|
| 76 |
+
β βββ tworooms/ # quentinll/lewm-tworooms (MIT)
|
| 77 |
+
βββ planner_*/ # Recursive & cross-attention planners
|
| 78 |
+
βββ manifold_transfer/ # Latent transfer experiments
|
| 79 |
+
βββ density/ # PushT behavior density model
|
| 80 |
+
|
| 81 |
+
code/ # Training & evaluation scripts
|
| 82 |
+
βββ planner.py # LePlanner controller
|
| 83 |
+
βββ solver.py # Training/eval harness
|
| 84 |
+
βββ losses.py # Arrival-hold, action-Gaussian, support losses
|
| 85 |
+
βββ scripts/ # train_planner.py, eval_planner.py, ...
|
| 86 |
+
|
| 87 |
+
latents/ # Pre-computed latent caches
|
| 88 |
+
βββ reacher/ # Latents, actions, episode metadata
|
| 89 |
+
βββ tworoom_cls/
|
| 90 |
+
|
| 91 |
+
results/ # Raw evaluation logs
|
| 92 |
+
βββ eval/results.jsonl # PushT eval across planners
|
| 93 |
+
βββ eval_planner/results.jsonl
|
| 94 |
+
```
|
| 95 |
+
|
| 96 |
+
## Key Checkpoints
|
| 97 |
+
|
| 98 |
+
| File | Description |
|
| 99 |
+
|---|---|
|
| 100 |
+
| `checkpoints/ablations/ah_hold0.5/controller.pt` | **Headline PushT controller** β 94β6% at rh=1 |
|
| 101 |
+
| `checkpoints/environment_controllers/reacher/controller.pt` | Reacher controller β 100% |
|
| 102 |
+
| `checkpoints/environment_controllers/tworooms/controller.pt` | TwoRooms controller β 100% |
|
| 103 |
+
| `checkpoints/base_world_models/pusht/weights.pt` | Frozen PushT world model (LeWM) |
|
| 104 |
+
| `checkpoints/base_world_models/reacher/weights.pt` | Frozen Reacher world model (LeWM) |
|
| 105 |
+
| `checkpoints/base_world_models/tworooms/weights.pt` | Frozen TwoRooms world model (LeWM) |
|
| 106 |
+
|
| 107 |
+
## Training & Evaluation
|
| 108 |
+
|
| 109 |
+
The controller and planner checkpoints are trained in this project; the LeWM world model is **always frozen** while those planners are trained.
|
| 110 |
+
|
| 111 |
+
Training code: [`github.com/SaltedLemon/lejepa_control`](https://github.com/SaltedLemon/lejepa_control)
|
| 112 |
+
|
| 113 |
+
Each controller checkpoint embeds its training arguments, action statistics, step, and validation profile.
|
| 114 |
+
|
| 115 |
+
## Citation
|
| 116 |
+
|
| 117 |
+
If you use this code or these checkpoints, please cite:
|
| 118 |
+
|
| 119 |
+
```bibtex
|
| 120 |
+
@article{leplanner,
|
| 121 |
+
title = {LePlanner: Amortized Iterative Control in Frozen Latent World Models},
|
| 122 |
+
author = {[Authors]},
|
| 123 |
+
journal = {[Venue / arXiv]},
|
| 124 |
+
year = {[Year]},
|
| 125 |
+
url = {https://huggingface.co/SaltedLemon/lejepa-control-pusht}
|
| 126 |
+
}
|
| 127 |
+
```
|
| 128 |
+
|
| 129 |
+
## License
|
| 130 |
+
|
| 131 |
+
- **Controller/planner code and checkpoints:** as specified in the linked `lejepa_control` repository.
|
| 132 |
+
- **LeWM world model weights** (`base_world_models/pusht`, `reacher`, `tworooms`): MIT licensed upstream from `quentinll/lewm-*`.
|