--- license: apache-2.0 base_model: Qwen/Qwen2.5-VL-3B-Instruct datasets: - Ngseo/ur5_teleop_multitask tags: - robotics - vision-language-action - vla - qwen2_5_vl - lerobot - spurious-correlation --- # Stage 3 — UR5 teleop VLA (Qwen2.5-VL-3B), current-view variant Training run **complete**: 60,000 steps. This is the ablation partner of [`Ngseo/stage3-ur5-actheavy`](https://huggingface.co/Ngseo/stage3-ur5-actheavy): same objective, same data, same hyper-parameters — the frozen V-JEPA2 targets are computed on the **current** frame instead of on **future** frames. ``` act-heavy : V-JEPA sees frames t+4 … t+32 (1.07 s ahead) current-view : V-JEPA sees frame t, repeated 8x ← this model ``` Both branches move together: the task branch that `z_a` is aligned to *and* the domain branch that `z_b` is decorrelated from. Nothing in this model's loss looks ahead in time. **The question it answers:** does the V-JEPA target have to predict the future, or is shaping the representation against the current frame enough? ## The setup this was trained for Each of the 7 tasks in [`Ngseo/ur5_teleop_multitask`](https://huggingface.co/datasets/Ngseo/ur5_teleop_multitask) is only ever shown from **one** of the 4 cameras, so viewpoint alone almost determines the task. Object colour is deliberately crossed between the two task families so that "colour ⇒ camera" is not a valid shortcut on its own: | camera | tasks | |---|---| | `camera_0` | Point at the **red** cup · Pick up the **blue** die → basket | | `camera_1` | Pull a tissue out of the box · Close the laptop · Stand the shoe upright | | `camera_2` | Point at the **blue** cup · Pick up the **red** die → basket | ## Architecture ``` Qwen2.5-VL-3B + LoRA ─ forward_full_hidden ─┬─ Head A → z_a ─┐ └─ Head B → z_b ─┴─ concat(8192) → ResNetActionHead → 30×7 L = 1.0·L1(action) + 0.02·InfoNCE(z_a, z_task_CURRENT) + 0.002·SIGReg([z_domain_CURRENT ; z_b]) ``` Frozen target encoder: [`Ngseo/stage1`](https://huggingface.co/Ngseo/stage1) disentangled V-JEPA2 ViT-L. Both heads are `AttentiveLatentHead` (proj 4096, 8 queries, depth 2, 167.8M each). | | | |---|---| | LoRA | r=32, α=64 — LLM `q_proj`/`v_proj` (7.4M) **and** vision tower `qkv` (5.2M) | | Trainable | 386M of 3.77B | | Inputs | 1 RGB frame @224 + task string + 7-D joint state | | Output | 30-step action chunk (absolute joint positions, 1.0 s @ 30 fps) | | Optimiser | AdamW, lr 1e-4, 1000-step warmup + cosine to 0, batch 32, bf16 | | Augmentation | ColorJitter/SharpnessJitter ×2 + DomainRandomization p=0.7 | ## Final metrics vs. the future-frame variant | | current-view (this) | act-heavy (future) | |---|---|---| | action L1 @ 10k | 0.0826 | 0.0836 | | @ 30k | 0.0438 | 0.0442 | | @ 50k | 0.0326 | 0.0305 | | **@ 60k (final)** | **0.0319** | **0.0300** | | InfoNCE (chance 3.466) | 1.86 | 1.79 | | cos(z_a, z_target) | 0.233 | 0.292 | | cos(z_b, z_domain) | −0.002 | −0.000 | On in-distribution action accuracy the two are within ~6% of each other, i.e. looking ahead buys almost nothing here. That is expected: the InfoNCE term carries weight 0.02, so it barely competes with the action loss. Reference points on the same normalised scale, none of which use vision or language: dataset mean **0.834**, copying the input state across all 30 steps **0.155**. The per-step copy error grows from 0.019 at k=0 to 0.291 at k=29. > **Not evaluated on a robot.** Everything above is a training-set loss. The > question this study is actually about — what happens when a task is requested > from a camera it was never trained on — is not answered by these numbers, and > is exactly where the two variants might diverge. > > **Caveat specific to this variant:** the current-view clip is built from the > *same augmented context frame* the VLM is shown, so its InfoNCE aligns two > encodings of identical pixels. The future-frame variant aligns against a > different, independently-augmented clip. The two therefore differ in more > than just "future vs current", and the comparison should be read with that in > mind. > > The 7-D joint state is also an input, and a probe on that state alone recovers > which of the 7 tasks is running with 94% accuracy (chance 14%). ## Contents `epoch_6.pt` holds `model_state_dict` (full VLM incl. LoRA), `latent_head_state_dict` (Head A), `free_latent_head_state_dict` (Head B), `stage2_action_head_state_dict`, optimiser state, and the run config. `config.yaml` is the exact training config.