| --- |
| license: apache-2.0 |
| pipeline_tag: robotics |
| tags: |
| - robotics |
| - gr00t |
| - vla |
| - bimanual |
| - deformable-object-manipulation |
| - cloth |
| --- |
| |
| # GR00T N1.7 — flatten_tshirt (bimanual cloth flattening) |
| |
| An [Isaac GR00T](https://github.com/NVIDIA/Isaac-GR00T) N1.7-3B vision-language-action policy |
| fine-tuned on the `flatten_tshirt` task of a bimanual deformable-object (cloth / bag) |
| manipulation benchmark. The robot is a dual-arm Piper; the task is to flatten a crumpled |
| t-shirt on a table. |
|
|
| Simulation uses a GPU cloth solver co-simulated with the robot in a single model, and |
| observations are rendered with a photorealistic renderer. |
|
|
| ## ⚠️ Requires access to a gated backbone |
|
|
| This checkpoint holds the fine-tuned GR00T weights only. Its vision-language backbone is |
| **[`nvidia/Cosmos-Reason2-2B`](https://huggingface.co/nvidia/Cosmos-Reason2-2B), a gated |
| repo** — `config.json` references it by repo id. You must request access to that repo and be |
| authenticated before this model will load. Those weights are NVIDIA's and are not redistributed |
| here. |
|
|
| ## Model |
|
|
| | | | |
| |---|---| |
| | Architecture | `Gr00tN1d7` (GR00T N1.7-3B) | |
| | Backbone | `nvidia/Cosmos-Reason2-2B` (gated, see above) | |
| | Precision | bfloat16 | |
| | Observation | 3 × RGB (`static_cam`, `left_hand_cam`, `right_hand_cam`) + 14-D joint state | |
| | Action | 14-D (left 6 joints + gripper, right 6 joints + gripper) | |
| | Action horizon | 40 | |
| | Inference timesteps | 4 | |
| | Image target size | 256 × 256 (crop fraction 0.95) | |
| | Embodiment tag | `new_embodiment` | |
|
|
| ## Training |
|
|
| | | | |
| |---|---| |
| | Dataset | `flatten_tshirt_200` — 200 episodes / 41,464 frames, LeRobot v3.0, 25 fps | |
| | Steps | 20,000 | |
| | Global batch size | 32 | |
| | Image augmentation | GR00T defaults (brightness 0.3, contrast 0.4, saturation 0.5, hue 0.08, + crop) | |
|
|
| Unlike the other policies in this benchmark, GR00T is trained with its author-default |
| augmentation recipe rather than a tuned one: each policy is trained with its own best-known |
| recipe, while the data and the evaluation protocol are held identical across policies. |
|
|
| 20,000 steps at batch 32 is the same sample budget (~640k) as the official 10,000 × 64 |
| configuration. |
|
|
| ## What is in this repo |
|
|
| Inference weights only. The optimizer state, HF Trainer arguments and W&B run config from the |
| original checkpoint have been removed — they carry no inference value and embedded internal |
| infrastructure paths. `trainer_state.json` is kept for its loss history. |
|
|
| ## Status |
|
|
| ⚠️ **This checkpoint has not been evaluated.** It has not been through the benchmark's |
| closed-loop protocol at all. Success-rate numbers are deliberately not published here; they |
| will be added once evaluation has run. Treat this as a training artifact, not a reported |
| result. |
|
|
| ## License |
|
|
| Apache-2.0 for these fine-tuned weights. The Cosmos backbone is covered by NVIDIA's own |
| license terms. |
|
|