File size: 5,035 Bytes
0c69e19
 
dafcb05
1fbc7a6
 
 
 
 
 
 
 
 
0c69e19
1fbc7a6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
dafcb05
1fbc7a6
 
 
dafcb05
 
1fbc7a6
 
 
 
 
 
 
 
 
 
dafcb05
1fbc7a6
 
 
 
dafcb05
1fbc7a6
dafcb05
1fbc7a6
 
 
 
 
 
dafcb05
1fbc7a6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
---
license: cc-by-nc-sa-4.0
library_name: n0_twam
pipeline_tag: robotics
tags:
  - robotics
  - manipulation
  - vision-tactile-action
  - world-action-model
  - diffusion
  - flow-matching
  - mixture-of-transformers
---

# N<sub>0</sub>-TWAM β€” A Tactile-Native World Action Model

$N_0$-TWAM is a Vision–Tactile–Action world-action model. Vision, tactile, and action
are jointly modeled by a Mixture-of-Transformers (MoT) under a single rectified-flow /
flow-matching objective: the model predicts the visual and tactile future and generates
the low-level action that realizes it.

This repository holds the **pretrained weights** as a ready-to-use bundle. The training,
inference-server, and post-training **code** lives at
πŸ‘‰ **https://github.com/neoteai/N0-TWAM**

## Contents

A self-contained bundle β€” everything the model needs to load and serve:

```
.
β”œβ”€β”€ transformer/              # the trained N0-TWAM weights (MoT, 20-d action, tactile)
β”œβ”€β”€ vae/                      # Wan2.2 AutoencoderKLWan (48-channel)
β”œβ”€β”€ text_encoder/             # umT5-xxl (frozen)
β”œβ”€β”€ tokenizer/
β”œβ”€β”€ norm_stat_pretrain.json   # action q01/q99 the checkpoint was trained with β€”
β”‚                             #   the server de-normalizes actions with these
β”‚                             #   (auto-loaded from the bundle root by twam_server)
└── empty_emb.pt              # empty-prompt text embedding (used by the
                              #   post-training dataloader's CFG text drop)
```

## Usage

### 1. Install the code

```bash
git clone https://github.com/neoteai/N0-TWAM.git
cd N0-TWAM
pip install .
pip install flash-attn --no-build-isolation
```

### 2. Download this bundle

```python
from huggingface_hub import snapshot_download
bundle = snapshot_download("NeoteAI/n0-twam-base")
# `bundle` now contains transformer/ vae/ text_encoder/ tokenizer/
# plus norm_stat_pretrain.json and empty_emb.pt
```

### 3a. Load the model

```python
import torch
from n0_twam.models.utils import load_mot_checkpoint

model = load_mot_checkpoint(f"{bundle}/transformer",
                            torch_dtype=torch.bfloat16, torch_device="cuda")
print(f"{sum(p.numel() for p in model.parameters()) / 1e9:.2f} B params")
# -> 7.16 B params
```

### 3b. Or serve it (observation β†’ action)

Point the inference-server config at this downloaded bundle β€” it already contains
every component the server needs, including the training-time action norm stats
(`norm_stat_pretrain.json`, auto-loaded from the bundle root) β€” then launch the
websocket server:

```python
# in n0_twam/configs/twam_server_cfg.py
twam_server_cfg.wan22_pretrained_model_name_or_path = "<path to the downloaded bundle>"
```

```bash
export PYTHONPATH=$PWD:$PWD/n0_twam CUDA_VISIBLE_DEVICES=0
export RANK=0 LOCAL_RANK=0 WORLD_SIZE=1 MASTER_ADDR=127.0.0.1 MASTER_PORT=29988
python -m n0_twam.n0_twam_server --config-name twam_server --port 29601
```

Query it from a client (a reset with a language prompt, then per-step observations):

```python
import numpy as np
from n0_twam.utils.Simple_Remote_Infer.deploy.websocket_client_policy import WebsocketClientPolicy

client = WebsocketClientPolicy("127.0.0.1", 29601)
client.infer({"reset": True, "prompt": "pick up the object"})

cams = ["observation.images.third_view",
        "observation.images.left_wrist_view",
        "observation.images.right_wrist_view"]
frame = {k: np.zeros((256, 256, 3), np.uint8) for k in cams}   # your RGB frames
state = np.zeros(20, np.float32)                                # your current EE state
action = client.infer({"obs": [frame], "current_state": state})["action"]  # (20, 2, 16)
```

See [`DEPLOY.md`](https://github.com/neoteai/N0-TWAM/blob/main/docs/DEPLOY.md)
for the full serving guide, and
[`POST_TRAINING.md`](https://github.com/neoteai/N0-TWAM/blob/main/docs/POST_TRAINING.md)
to fine-tune the model on your own robot.

## Model details

| | |
|---|---|
| Architecture | 3-expert Mixture-of-Transformers (video / tactile / action) with shared attention |
| Parameters | ~7.16 B (bf16) |
| Backbone | WAN2.2 TI2V-5B video diffusion transformer |
| Video VAE | Wan2.2 `AutoencoderKLWan` (`z_dim=48`, 4Γ— temporal / 16Γ— spatial) |
| Text encoder | umT5-xxl (4096-d), frozen |
| Objective | Rectified-flow / flow-matching, per-frame timesteps |
| Action space | 20-dim dual-arm end-effector, Ο€0.5-style horizon delta |
| Tactile | global (co-generated diffusion target) + optional local (observed input) |

## License

Released under the CC-BY-NC-SA-4.0 license (see [LICENSE](LICENSE)). The `vae/`,
`text_encoder/` and `tokenizer/` components are redistributed from
[Wan2.2](https://github.com/Wan-Video/Wan2.2) and keep their original Apache 2.0
license and notices.

## Acknowledgments

Builds upon [LingBot-VA](https://github.com/robbyant/lingbot-va),
[Wan2.2](https://github.com/Wan-Video/Wan2.2),
[FastWAM](https://github.com/yuantianyuan01/FastWAM) (MoT design), and
[LeRobot](https://github.com/huggingface/lerobot).