File size: 5,994 Bytes
f2f43ac
 
 
 
 
 
 
 
712128b
f2f43ac
f291db4
 
 
986c58b
712128b
f2f43ac
712128b
 
 
 
 
 
 
 
 
 
 
 
 
 
f2f43ac
712128b
f2f43ac
712128b
 
 
 
f2f43ac
712128b
 
 
 
 
 
 
 
 
f2f43ac
712128b
f2f43ac
712128b
f2f43ac
712128b
 
 
 
 
 
 
 
 
f2f43ac
712128b
f2f43ac
712128b
f2f43ac
712128b
 
 
 
 
 
f2f43ac
712128b
f2f43ac
712128b
f2f43ac
712128b
 
f2f43ac
712128b
f2f43ac
712128b
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
---
license: apache-2.0
tags:
  - robotics
  - vision-language-action
  - grounding
---

# Galahad — 7-axis merged grounded VLA (single weight)

> [Project page](https://xn--7xa.monster/Galahad/) · [Code + battery](https://github.com/phi-monster/Galahad) · [Datasets](https://huggingface.co/phi-monster) · [Dual-output weight](https://huggingface.co/phi-monster/Galahad-object-unified)
>
> Paper: **Instruction Blindness in Vision–Language–Action Policies: Diagnosis and a Low-Rank Data Cure** — [PDF](https://xn--7xa.monster/Galahad/static/galahad.pdf)

A single merged weight (base + folded LoRA, no adapter to attach) that follows the **referring expression** in the instruction rather than a memorised position or scene prior. Seven referent axes, all measured **on this exact artifact**.

---

## 🔴 READ THIS BEFORE LOADING — the bundled `norm_stats.json` is WRONG for this model

This repository ships `norm_stats.json` inherited from the uncured base `allenai/MolmoAct2-Think-LIBERO`. It is **byte-identical to the base's** (`md5 03084dfec9bb21f716100fea20f3d2c9`) and **must not be used to denormalise this model's actions.**

The merge wrote HF-transformers format and dropped the lerobot processor files that carry the trained normalisation statistics. Those files are now included in this repo and **are the only correct source**:

```
policy_preprocessor.json
policy_preprocessor_step_3_molmoact2_masked_normalizer.safetensors
policy_postprocessor.json
policy_postprocessor_step_1_molmoact2_masked_unnormalizer.safetensors
```

Load with the lerobot processor path:

```python
from lerobot.policies.factory import make_pre_post_processors
pre, post = make_pre_post_processors(cfg, pretrained_path="<path to this repo>")
```

**Do NOT** use `make_molmoact2_pre_post_processors(cfg)` — it reads `norm_stats.json` and silently applies the base's action quantiles.

**Failure signature if you get this wrong** (measured, so you can recognise it): this corpus has roll/pitch **identically zero** and a yaw span of 0.0116, while the base's LIBERO spans are 0.233 / 0.335 / 0.444. Denormalising through the base's statistics injects a **constant roll −0.103 and pitch −0.155 on every step** plus a 38× yaw gain — a continuously tumbling wrist. Because x/y/z spans differ by only 5-12%, **the arm still reaches within ~2 mm of the correct object and then never closes the grasp.** It looks like a motor failure; it is a normalisation failure.

### Second requirement: `ZERO_STATE=1`

Serve with `observation.state` zeroed **and** the `state_encoder` output layer zeroed. This model's generation runs away on real proprioception. Every number below was measured under this contract.

---

## Main table — all seven axes, measured on this weight

Contract for every cell: this weight · the processor files above · `ZERO_STATE=1`.

| axis | **SWAP-OBEY** (told to fetch a different in-scene object, does the EEF go to the newly-named one?) | **OCC** (both cameras blanked) | TASK (completion) |
|---|---|---|---|
| **object** | **198/200 = 99.0%** · true-target **0/200** | 0/8 = 0% | 160/200 = 80.0% |
| **spatial** | **OBEYED 187 / WENT_TRUE 0** | 0/40 = 0% | 110/200 = 55.0% |
| **goal** | **VARY 40/48 = 83.3%** (referent displaced 7-9 cm) | 0/24 = 0% | — |
| **colour** | **32/36 = 88.9%** [74.7, 95.6] | 0/36 = 0% | 8/36 = 22.2% |
| **negation** | **43/48 = 89.6%** [77.8, 95.5] | 0/48 = 0% | 8/48 = 16.7% |
| **category** | **39/50 = 78.0%** [64.8, 87.2] | 1/50 = 2.0% | 10/50 = 20.0% |
| **compositional** | **36/48 = 75.0%** [61.2, 85.1] | 0/48 = 0% | 1/48 = 2.1% |

**All seven axes ground (75-99%) and all seven are vision-gated (0-2%).**

### How to read these numbers honestly

- **SWAP-OBEY is the grounding measurement; TASK is not.** Under a swapped instruction, "true-target ≈ 0" alone is ambiguous — a drop could mean "followed the new name" *or* "just failed". The pair (true-target 0 **and** obeyed-name high) is what rules out the second reading.
- **compositional decomposes, and the split is the useful part:** `flip_colour` **21/24 = 87.5%** vs `flip_size` **15/24 = 62.5%**. The colour attribute is followed as well as the single-attribute axes; **size is the short leg.**
- **TASK is low on the RoboCasa axes by bench design, not by regression.** The published single-axis anchors on the same benches are TASK 26-51% against SWAP-OBEY 77-95%. This weight sits ~10-15 pp below those TASK anchors (a motor cost of unifying seven axes into one weight) while its grounding sits at or above them.
- **spatial's OBEY count is backed by the raw distance distribution** (n=200), because its two referents are identical bowls placed close together: `d_named` median **20.0 mm** vs `d_true` median **124.0 mm** = **6.20× separation**, only 4.5% of trials ambiguous.
- **The OCC zeros are not a stuck counter.** The same scorers emit non-zero under other conditions in the same runs (category's own OCC reads 1/50), so the collapse is a real measurement.
- **The low TASK is not a serving artifact.** Paired control, same weight / harness / cells, normaliser the only variable: an independently regenerated normaliser gives **8/36 vs 8/36 — identical**.

## Baselines for context

The uncured base is not a weak model — it is strong *and* shortcut-driven, which is the point:

- **spatial:** uncured base scores TASK 99% but follows a swapped spatial referent only **50% = exact chance for two bowls**. Its 99% is a canonical-bowl shortcut.
- **goal:** uncured base **31.2%** on displaced referents (it goes to the remembered position) vs **83.3%** here.

## Limitations

- Closed vocabulary. Referents outside the training vocabulary are not expected to work.
- TASK completion on the RoboCasa axes is motor-limited (see above); this weight trades some motor precision for grounding.
- `compositional`'s size attribute is measurably weaker than its colour attribute.
- Sim-trained (LIBERO-PRO / RoboCasa). No real-robot transfer claim.