Instructions to use kabilanKB/cosmos_edge_policy_so101 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Cosmos
How to use kabilanKB/cosmos_edge_policy_so101 with Cosmos:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- LeRobot
How to use kabilanKB/cosmos_edge_policy_so101 with LeRobot:
- Notebooks
- Google Colab
- Kaggle
cosmos_edge_policy_so101
A 6-DOF SO-101 bin-placement action policy, post-trained from
nvidia/Cosmos3-Edge-Policy-DROID.
Each folder holds one training checkpoint with its LoRA adapters merged into the weights, exported
as consolidated safetensors.
| Folder | Iteration | Epoch (approx.) | Benchmark |
|---|---|---|---|
iter_6000/ |
6000 | 3.5 | not evaluated yet |
iter_6500/ |
6500 | 3.8 | 5 / 98 (5.1%), the best checkpoint |
iter_7000/ |
7000 (final) | 4.0 | 1 / 100 (1.0%) |
Scores are on the benchmark below with a 25 s episode limit. Earlier checkpoints of the same run scored 0 / 50 (iteration 1500) and 0 / 100 (iteration 2500). For comparison, the Cosmos3-Nano SO-101 policy scores 3.1% on the same benchmark.
Runs on a Jetson Thor: see thor_deploy/ for serving iter_6500
natively on a Thor and driving the Isaac Lab benchmark from a PC through a hardware-in-the-loop web UI.
Training
| Setting | Value |
|---|---|
| Base model | nvidia/Cosmos3-Edge-Policy-DROID |
| Action-head warm start | action2llm / llm2action keep a separate weight row per embodiment. The DROID row (domain 8) was copied into the SO-101 row (domain 22) before training, so SO-101 starts from DROID's trained action mapping instead of random init. |
| Experiment | action_policy_so101_edge_focus5_multi (cosmos-framework 5e67049 plus local SO-101 support) |
| Data | so101_bench_sim_6: 5 single-object "Place the X in the plastic bin" instructions (97 episodes) plus the multi-object "Place each object in the plastic bin", capped at 45 episodes. 142 episodes, 10% held out. |
| Action space | Absolute joint_pos, 6-D (5 arm joints plus gripper, LeRobot .pos units). Row 0 is the current state. Chunk of 32 steps at 30 fps. |
| Normalization | minmax against the calibration bounds: joints [-100, 100], gripper [0, 100] (so101_lerobot_stats.json) |
| Video | concat_view: front wrist camera stacked on top of the overhead camera, 480p |
| Method | LoRA rank 64 / alpha 128 on q/k/v/o_proj_moe_gen, plus full training of the action heads |
| Schedule | Global batch 32, learning rate 1e-4, 200 warm-up steps, linear decay over 7000 iterations |
| Hardware | Iterations 0–5500 on 1× RTX PRO 6000 (batch 16 × 2 accumulation steps); 5500–7000 resumed on 2× RTX PRO 6000 (batch 16 per GPU) |
| Embodiment domain id | 22 (so101) |
Benchmark setup
so101_bench Isaac Lab digital twin, So101Bench-Bin-v0: 100 single-object episodes
(tasks/focus5.jsonl), 25 s per episode, 32 actions executed per inference call.
Serving
Served with cosmos_framework.scripts.action_policy_server_robolab, an openpi websocket server.
Every SO-101 flag below is required, and a missing one fails silently: the server starts but
returns wrong actions.
huggingface-cli download kabilanKB/cosmos_edge_policy_so101 --include "iter_7000/*" "so101_lerobot_stats.json" \
--local-dir cosmos_edge_policy_so101
python -m cosmos_framework.scripts.action_policy_server_robolab \
--checkpoint-path cosmos_edge_policy_so101/iter_7000 \
--port 8000 \
--domain-name so101 \
--action-dim 6 \
--arm-joint-dim 5 \
--action-space joint_pos \
--conditioning-fps 30 \
--no-flip-gripper \
--action-normalization minmax \
--normalizer-stats-path cosmos_edge_policy_so101/so101_lerobot_stats.json \
--view-description 'The top half is from the front-facing wrist camera. The bottom half is from the fixed overhead camera.' \
--no-guardrails
These flags depend on SO-101 support in the policy server (--arm-joint-dim, --no-flip-gripper,
--action-normalization, --view-description), which is not in upstream cosmos-framework 5e67049.
The pipeline for training, merge, export, serving and evaluation is in
kabilankb/so101-cosmos-nano-policy.
Request format: prompt (the instruction), observation/image (the concatenated view),
observation/joint_position (5 values), observation/gripper_position (1 value).
Response: action, 32 × 6 absolute joint targets in raw .pos units.
Run on a Jetson Thor
The policy server runs natively on an NVIDIA Jetson Thor (JetPack 7, CUDA 13); the ~9 GB of BF16
weights fit in the Thor's memory. thor_deploy/ has one-command setup and
serve scripts for the Thor, a setup script for the PC that runs the so101_bench simulation
(ported to Isaac Lab 3.0 / Isaac Sim 6.0), and a hardware-in-the-loop web UI that takes the Thor's IP
address, launches the policy server on the Thor over SSH, runs the benchmark and shows each camera
frame and action chunk.
# Thor
uvx hf@latest download kabilanKB/cosmos_edge_policy_so101 --include "thor_deploy/*" --local-dir ~/so101_deploy
bash ~/so101_deploy/thor_deploy/thor/setup_thor.sh && bash ~/so101-edge-thor/serve_thor.sh
On the Thor a 32-step chunk takes ~2.1 s with cuDNN attention and torch.compile (details and the
two fixes that make this possible are in the deploy README), versus 0.42 s on an RTX PRO 6000.
Limitations
- Trained and evaluated only in simulation. It has not been tested on a physical SO-101.
- Covers five objects plus one multi-object instruction, with 18–45 demonstrations per instruction.
- Sampling is stochastic: the server draws a new seed for every request.
License and attribution
This model is a derivative of NVIDIA Cosmos3-Edge-Policy-DROID, released under the
OpenMDW License 1.1, and is distributed under the same license.
The bundled vision encoder, processor and tokenizer files come from nvidia/Cosmos3-Edge.
Built with NVIDIA Cosmos. Post-trained by Kabilan KB.
- Downloads last month
- -
Model tree for kabilanKB/cosmos_edge_policy_so101
Base model
nvidia/Cosmos3-Edge-Policy-DROID