File size: 5,035 Bytes
0c69e19 dafcb05 1fbc7a6 0c69e19 1fbc7a6 dafcb05 1fbc7a6 dafcb05 1fbc7a6 dafcb05 1fbc7a6 dafcb05 1fbc7a6 dafcb05 1fbc7a6 dafcb05 1fbc7a6 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 | ---
license: cc-by-nc-sa-4.0
library_name: n0_twam
pipeline_tag: robotics
tags:
- robotics
- manipulation
- vision-tactile-action
- world-action-model
- diffusion
- flow-matching
- mixture-of-transformers
---
# N<sub>0</sub>-TWAM β A Tactile-Native World Action Model
$N_0$-TWAM is a VisionβTactileβAction world-action model. Vision, tactile, and action
are jointly modeled by a Mixture-of-Transformers (MoT) under a single rectified-flow /
flow-matching objective: the model predicts the visual and tactile future and generates
the low-level action that realizes it.
This repository holds the **pretrained weights** as a ready-to-use bundle. The training,
inference-server, and post-training **code** lives at
π **https://github.com/neoteai/N0-TWAM**
## Contents
A self-contained bundle β everything the model needs to load and serve:
```
.
βββ transformer/ # the trained N0-TWAM weights (MoT, 20-d action, tactile)
βββ vae/ # Wan2.2 AutoencoderKLWan (48-channel)
βββ text_encoder/ # umT5-xxl (frozen)
βββ tokenizer/
βββ norm_stat_pretrain.json # action q01/q99 the checkpoint was trained with β
β # the server de-normalizes actions with these
β # (auto-loaded from the bundle root by twam_server)
βββ empty_emb.pt # empty-prompt text embedding (used by the
# post-training dataloader's CFG text drop)
```
## Usage
### 1. Install the code
```bash
git clone https://github.com/neoteai/N0-TWAM.git
cd N0-TWAM
pip install .
pip install flash-attn --no-build-isolation
```
### 2. Download this bundle
```python
from huggingface_hub import snapshot_download
bundle = snapshot_download("NeoteAI/n0-twam-base")
# `bundle` now contains transformer/ vae/ text_encoder/ tokenizer/
# plus norm_stat_pretrain.json and empty_emb.pt
```
### 3a. Load the model
```python
import torch
from n0_twam.models.utils import load_mot_checkpoint
model = load_mot_checkpoint(f"{bundle}/transformer",
torch_dtype=torch.bfloat16, torch_device="cuda")
print(f"{sum(p.numel() for p in model.parameters()) / 1e9:.2f} B params")
# -> 7.16 B params
```
### 3b. Or serve it (observation β action)
Point the inference-server config at this downloaded bundle β it already contains
every component the server needs, including the training-time action norm stats
(`norm_stat_pretrain.json`, auto-loaded from the bundle root) β then launch the
websocket server:
```python
# in n0_twam/configs/twam_server_cfg.py
twam_server_cfg.wan22_pretrained_model_name_or_path = "<path to the downloaded bundle>"
```
```bash
export PYTHONPATH=$PWD:$PWD/n0_twam CUDA_VISIBLE_DEVICES=0
export RANK=0 LOCAL_RANK=0 WORLD_SIZE=1 MASTER_ADDR=127.0.0.1 MASTER_PORT=29988
python -m n0_twam.n0_twam_server --config-name twam_server --port 29601
```
Query it from a client (a reset with a language prompt, then per-step observations):
```python
import numpy as np
from n0_twam.utils.Simple_Remote_Infer.deploy.websocket_client_policy import WebsocketClientPolicy
client = WebsocketClientPolicy("127.0.0.1", 29601)
client.infer({"reset": True, "prompt": "pick up the object"})
cams = ["observation.images.third_view",
"observation.images.left_wrist_view",
"observation.images.right_wrist_view"]
frame = {k: np.zeros((256, 256, 3), np.uint8) for k in cams} # your RGB frames
state = np.zeros(20, np.float32) # your current EE state
action = client.infer({"obs": [frame], "current_state": state})["action"] # (20, 2, 16)
```
See [`DEPLOY.md`](https://github.com/neoteai/N0-TWAM/blob/main/docs/DEPLOY.md)
for the full serving guide, and
[`POST_TRAINING.md`](https://github.com/neoteai/N0-TWAM/blob/main/docs/POST_TRAINING.md)
to fine-tune the model on your own robot.
## Model details
| | |
|---|---|
| Architecture | 3-expert Mixture-of-Transformers (video / tactile / action) with shared attention |
| Parameters | ~7.16 B (bf16) |
| Backbone | WAN2.2 TI2V-5B video diffusion transformer |
| Video VAE | Wan2.2 `AutoencoderKLWan` (`z_dim=48`, 4Γ temporal / 16Γ spatial) |
| Text encoder | umT5-xxl (4096-d), frozen |
| Objective | Rectified-flow / flow-matching, per-frame timesteps |
| Action space | 20-dim dual-arm end-effector, Ο0.5-style horizon delta |
| Tactile | global (co-generated diffusion target) + optional local (observed input) |
## License
Released under the CC-BY-NC-SA-4.0 license (see [LICENSE](LICENSE)). The `vae/`,
`text_encoder/` and `tokenizer/` components are redistributed from
[Wan2.2](https://github.com/Wan-Video/Wan2.2) and keep their original Apache 2.0
license and notices.
## Acknowledgments
Builds upon [LingBot-VA](https://github.com/robbyant/lingbot-va),
[Wan2.2](https://github.com/Wan-Video/Wan2.2),
[FastWAM](https://github.com/yuantianyuan01/FastWAM) (MoT design), and
[LeRobot](https://github.com/huggingface/lerobot).
|