AAyano's picture
Add text-token-gate OpenVLA-OFT setting2 (10k fresh on 20k OFT backbone)
f4bdec1 verified
|
Raw
History Blame Contribute Delete
1.85 kB
---
license: mit
base_model: AAyano/oft_setting2_chunksize25_batch32_20k
tags:
- robotics
- vla
- openvla
- openvla-oft
- xarm
- task-conditioned-gate
library_name: transformers
---
# gate_setting2_chunksize25_batch32_from20000
Task-conditioned **text-token-gate** OpenVLA-OFT model for real-world XArm, **setting 2: cup stacking**.
Trained fresh for **10000 steps** on top of the [oft_setting2_chunksize25_batch32_20k](https://huggingface.co/AAyano/oft_setting2_chunksize25_batch32_20k) OFT backbone (action head / proprio projector / FiLM re-trained; cosine LR 5e-4 with 1000 warmup, effective batch 32, seed 6).
## Gate configuration (baked into `config.json`)
- Gazing: `self_attention` (dim 512, 8 heads, 1 layer), gate MLP hidden 512 / depth 1
- Text pooling: `contrastive_alignment_score` (visual τ 0.1, text τ 0.05), instruction-only pooling, stopword filtering
- Layer gate: `threshold` mode, threshold 0.3, strength 0.5
- Budget loss: target 0.4, weight 0.01
## ⚠️ Serving requires the Task-conditioned_Gate codebase
The gate modules live in this repo's `modeling_prismatic.py` (loaded via `trust_remote_code`), but the openvla-oft serve flow syncs the local repo's modeling files INTO the checkpoint before loading. Serve with `REPO_DIR` pointing at the [Task-conditioned_Gate](https://github.com/Kaixi66/Task-conditioned_Gate) repo — serving with vanilla openvla-oft silently drops the gate weights.
## Contents
- Merged full model (`model-*.safetensors`, gate config in `config.json`)
- `lora_adapter/` — LoRA + gate weights (`modules_to_save=["text_token_gate"]`), standalone copy
- `action_head--10000_checkpoint.pt`, `proprio_projector--10000_checkpoint.pt`, `vision_backbone--10000_checkpoint.pt` (FiLM)
- `dataset_statistics.json`, `oft_training_config.json` (chunk 25, action dim 7, proprio dim 6, BOUNDS_Q99)