Instructions to use Ngseo/stage3-ur5-currentview with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use Ngseo/stage3-ur5-currentview with LeRobot:
- Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| base_model: Qwen/Qwen2.5-VL-3B-Instruct | |
| datasets: | |
| - Ngseo/ur5_teleop_multitask | |
| tags: | |
| - robotics | |
| - vision-language-action | |
| - vla | |
| - qwen2_5_vl | |
| - lerobot | |
| - spurious-correlation | |
| # Stage 3 — UR5 teleop VLA (Qwen2.5-VL-3B), current-view variant | |
| Training run **complete**: 60,000 steps. This is the ablation partner of | |
| [`Ngseo/stage3-ur5-actheavy`](https://huggingface.co/Ngseo/stage3-ur5-actheavy): | |
| same objective, same data, same hyper-parameters — the frozen V-JEPA2 targets | |
| are computed on the **current** frame instead of on **future** frames. | |
| ``` | |
| act-heavy : V-JEPA sees frames t+4 … t+32 (1.07 s ahead) | |
| current-view : V-JEPA sees frame t, repeated 8x ← this model | |
| ``` | |
| Both branches move together: the task branch that `z_a` is aligned to *and* the | |
| domain branch that `z_b` is decorrelated from. Nothing in this model's loss | |
| looks ahead in time. | |
| **The question it answers:** does the V-JEPA target have to predict the future, | |
| or is shaping the representation against the current frame enough? | |
| ## The setup this was trained for | |
| Each of the 7 tasks in [`Ngseo/ur5_teleop_multitask`](https://huggingface.co/datasets/Ngseo/ur5_teleop_multitask) | |
| is only ever shown from **one** of the 4 cameras, so viewpoint alone almost | |
| determines the task. Object colour is deliberately crossed between the two task | |
| families so that "colour ⇒ camera" is not a valid shortcut on its own: | |
| | camera | tasks | | |
| |---|---| | |
| | `camera_0` | Point at the **red** cup · Pick up the **blue** die → basket | | |
| | `camera_1` | Pull a tissue out of the box · Close the laptop · Stand the shoe upright | | |
| | `camera_2` | Point at the **blue** cup · Pick up the **red** die → basket | | |
| ## Architecture | |
| ``` | |
| Qwen2.5-VL-3B + LoRA ─ forward_full_hidden ─┬─ Head A → z_a ─┐ | |
| └─ Head B → z_b ─┴─ concat(8192) → ResNetActionHead → 30×7 | |
| L = 1.0·L1(action) + 0.02·InfoNCE(z_a, z_task_CURRENT) + 0.002·SIGReg([z_domain_CURRENT ; z_b]) | |
| ``` | |
| Frozen target encoder: [`Ngseo/stage1`](https://huggingface.co/Ngseo/stage1) | |
| disentangled V-JEPA2 ViT-L. Both heads are `AttentiveLatentHead` | |
| (proj 4096, 8 queries, depth 2, 167.8M each). | |
| | | | | |
| |---|---| | |
| | LoRA | r=32, α=64 — LLM `q_proj`/`v_proj` (7.4M) **and** vision tower `qkv` (5.2M) | | |
| | Trainable | 386M of 3.77B | | |
| | Inputs | 1 RGB frame @224 + task string + 7-D joint state | | |
| | Output | 30-step action chunk (absolute joint positions, 1.0 s @ 30 fps) | | |
| | Optimiser | AdamW, lr 1e-4, 1000-step warmup + cosine to 0, batch 32, bf16 | | |
| | Augmentation | ColorJitter/SharpnessJitter ×2 + DomainRandomization p=0.7 | | |
| ## Final metrics vs. the future-frame variant | |
| | | current-view (this) | act-heavy (future) | | |
| |---|---|---| | |
| | action L1 @ 10k | 0.0826 | 0.0836 | | |
| | @ 30k | 0.0438 | 0.0442 | | |
| | @ 50k | 0.0326 | 0.0305 | | |
| | **@ 60k (final)** | **0.0319** | **0.0300** | | |
| | InfoNCE (chance 3.466) | 1.86 | 1.79 | | |
| | cos(z_a, z_target) | 0.233 | 0.292 | | |
| | cos(z_b, z_domain) | −0.002 | −0.000 | | |
| On in-distribution action accuracy the two are within ~6% of each other, i.e. | |
| looking ahead buys almost nothing here. That is expected: the InfoNCE term | |
| carries weight 0.02, so it barely competes with the action loss. | |
| Reference points on the same normalised scale, none of which use vision or | |
| language: dataset mean **0.834**, copying the input state across all 30 steps | |
| **0.155**. The per-step copy error grows from 0.019 at k=0 to 0.291 at k=29. | |
| > **Not evaluated on a robot.** Everything above is a training-set loss. The | |
| > question this study is actually about — what happens when a task is requested | |
| > from a camera it was never trained on — is not answered by these numbers, and | |
| > is exactly where the two variants might diverge. | |
| > | |
| > **Caveat specific to this variant:** the current-view clip is built from the | |
| > *same augmented context frame* the VLM is shown, so its InfoNCE aligns two | |
| > encodings of identical pixels. The future-frame variant aligns against a | |
| > different, independently-augmented clip. The two therefore differ in more | |
| > than just "future vs current", and the comparison should be read with that in | |
| > mind. | |
| > | |
| > The 7-D joint state is also an input, and a probe on that state alone recovers | |
| > which of the 7 tasks is running with 94% accuracy (chance 14%). | |
| ## Contents | |
| `epoch_6.pt` holds `model_state_dict` (full VLM incl. LoRA), `latent_head_state_dict` | |
| (Head A), `free_latent_head_state_dict` (Head B), `stage2_action_head_state_dict`, | |
| optimiser state, and the run config. `config.yaml` is the exact training config. | |