--- license: cc-by-nc-sa-4.0 library_name: n0_twam pipeline_tag: robotics tags: - robotics - manipulation - vision-tactile-action - world-action-model - diffusion - flow-matching - mixture-of-transformers --- # N0-TWAM β€” A Tactile-Native World Action Model $N_0$-TWAM is a Vision–Tactile–Action world-action model. Vision, tactile, and action are jointly modeled by a Mixture-of-Transformers (MoT) under a single rectified-flow / flow-matching objective: the model predicts the visual and tactile future and generates the low-level action that realizes it. This repository holds the **pretrained weights** as a ready-to-use bundle. The training, inference-server, and post-training **code** lives at πŸ‘‰ **https://github.com/neoteai/N0-TWAM** ## Contents A self-contained bundle β€” everything the model needs to load and serve: ``` . β”œβ”€β”€ transformer/ # the trained N0-TWAM weights (MoT, 20-d action, tactile) β”œβ”€β”€ vae/ # Wan2.2 AutoencoderKLWan (48-channel) β”œβ”€β”€ text_encoder/ # umT5-xxl (frozen) β”œβ”€β”€ tokenizer/ β”œβ”€β”€ norm_stat_pretrain.json # action q01/q99 the checkpoint was trained with β€” β”‚ # the server de-normalizes actions with these β”‚ # (auto-loaded from the bundle root by twam_server) └── empty_emb.pt # empty-prompt text embedding (used by the # post-training dataloader's CFG text drop) ``` ## Usage ### 1. Install the code ```bash git clone https://github.com/neoteai/N0-TWAM.git cd N0-TWAM pip install . pip install flash-attn --no-build-isolation ``` ### 2. Download this bundle ```python from huggingface_hub import snapshot_download bundle = snapshot_download("NeoteAI/n0-twam-base") # `bundle` now contains transformer/ vae/ text_encoder/ tokenizer/ # plus norm_stat_pretrain.json and empty_emb.pt ``` ### 3a. Load the model ```python import torch from n0_twam.models.utils import load_mot_checkpoint model = load_mot_checkpoint(f"{bundle}/transformer", torch_dtype=torch.bfloat16, torch_device="cuda") print(f"{sum(p.numel() for p in model.parameters()) / 1e9:.2f} B params") # -> 7.16 B params ``` ### 3b. Or serve it (observation β†’ action) Point the inference-server config at this downloaded bundle β€” it already contains every component the server needs, including the training-time action norm stats (`norm_stat_pretrain.json`, auto-loaded from the bundle root) β€” then launch the websocket server: ```python # in n0_twam/configs/twam_server_cfg.py twam_server_cfg.wan22_pretrained_model_name_or_path = "" ``` ```bash export PYTHONPATH=$PWD:$PWD/n0_twam CUDA_VISIBLE_DEVICES=0 export RANK=0 LOCAL_RANK=0 WORLD_SIZE=1 MASTER_ADDR=127.0.0.1 MASTER_PORT=29988 python -m n0_twam.n0_twam_server --config-name twam_server --port 29601 ``` Query it from a client (a reset with a language prompt, then per-step observations): ```python import numpy as np from n0_twam.utils.Simple_Remote_Infer.deploy.websocket_client_policy import WebsocketClientPolicy client = WebsocketClientPolicy("127.0.0.1", 29601) client.infer({"reset": True, "prompt": "pick up the object"}) cams = ["observation.images.third_view", "observation.images.left_wrist_view", "observation.images.right_wrist_view"] frame = {k: np.zeros((256, 256, 3), np.uint8) for k in cams} # your RGB frames state = np.zeros(20, np.float32) # your current EE state action = client.infer({"obs": [frame], "current_state": state})["action"] # (20, 2, 16) ``` See [`DEPLOY.md`](https://github.com/neoteai/N0-TWAM/blob/main/docs/DEPLOY.md) for the full serving guide, and [`POST_TRAINING.md`](https://github.com/neoteai/N0-TWAM/blob/main/docs/POST_TRAINING.md) to fine-tune the model on your own robot. ## Model details | | | |---|---| | Architecture | 3-expert Mixture-of-Transformers (video / tactile / action) with shared attention | | Parameters | ~7.16 B (bf16) | | Backbone | WAN2.2 TI2V-5B video diffusion transformer | | Video VAE | Wan2.2 `AutoencoderKLWan` (`z_dim=48`, 4Γ— temporal / 16Γ— spatial) | | Text encoder | umT5-xxl (4096-d), frozen | | Objective | Rectified-flow / flow-matching, per-frame timesteps | | Action space | 20-dim dual-arm end-effector, Ο€0.5-style horizon delta | | Tactile | global (co-generated diffusion target) + optional local (observed input) | ## License Released under the CC-BY-NC-SA-4.0 license (see [LICENSE](LICENSE)). The `vae/`, `text_encoder/` and `tokenizer/` components are redistributed from [Wan2.2](https://github.com/Wan-Video/Wan2.2) and keep their original Apache 2.0 license and notices. ## Acknowledgments Builds upon [LingBot-VA](https://github.com/robbyant/lingbot-va), [Wan2.2](https://github.com/Wan-Video/Wan2.2), [FastWAM](https://github.com/yuantianyuan01/FastWAM) (MoT design), and [LeRobot](https://github.com/huggingface/lerobot).