Instructions to use tokimoa/pi0-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use tokimoa/pi0-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir pi0-mlx tokimoa/pi0-mlx
- LeRobot
How to use tokimoa/pi0-mlx with LeRobot:
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
File size: 2,816 Bytes
5cc2511 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 | ---
license: apache-2.0
base_model: lerobot/pi0
pipeline_tag: robotics
tags:
- mlx
- robotics
- vla
- lerobot
- pi0
---
# pi0-mlx
Physical Intelligenceのロボット基盤モデル [π0](https://huggingface.co/lerobot/pi0)(PaliGemma 3B + アクションエキスパート・Vision-Language-Action)のApple Silicon(MLX)移植です。カメラ画像・言語指示・関節状態から50手先までのアクションチャンクをflow matchingで生成します。
| 実行系 | 1チャンク(50手)生成 | ピークメモリ |
|---|---|---|
| **本移植(MLX・bf16)** | **522ms** | 8.0GB |
| PyTorch MPS(参照実装) | 660ms | |
| PyTorch CPU(参照実装) | 2,395ms | |
同一入力・同一ノイズでPyTorch参照実装(lerobot main / openpi直系)とコサイン類似度0.99993(bf16、fp32重みでは1.00000)の出力一致を検証済みです。プレフィル→エキスパートがキャッシュへ毎層アテンションする二相構造、Gemma固有の正規化((1+w)RMSNorm)・GeGLU・言語埋め込みの√widthスケーリングまで参照実装を忠実に再現しています。
## 使い方
トークナイザは上流(lerobot)と同じくgoogle/paligemma-3b-pt-224を参照します。Hugging FaceでGemma利用規約に同意し、`hf auth login`してから実行してください。
```bash
pip install mlx-vlm pillow transformers
hf download tokimoa/pi0-mlx --local-dir pi0-mlx
```
```python
from pi0_mlx import Pi0MLX
model = Pi0MLX.from_pretrained("pi0-mlx")
actions = model.predict(
images=[cam0, cam1, cam2], # HWC uint8(1〜3カメラ)
instruction="pick up the cube",
state=[0.1, -0.2, 0.3, 0.0, 0.5, 0.0],
) # -> (50, len(state)) アクションチャンク
```
CLIでも動きます。
```bash
python pi0-mlx/pi0_mlx.py --images cam0.png cam1.png cam2.png \
--instruction "pick up the cube" --state 0,0,0,0,0,0
```
## 位置づけ
π0はベースモデルであり、実タスクへの適用には手元のロボットでのファインチューニングが前提です。学習は[LeRobot](https://github.com/huggingface/lerobot)で行い、Mac上での推論・検証・デモに本移植を使う構成を想定しています。同一アーキテクチャのFT済み重みは`model.safetensors`を差し替えれば動きます。前処理(224pxアスペクト維持リサイズ・言語トークナイズ・状態パディング)はランタイムに内蔵しています。
同シリーズ: [smolvla-mlx](https://huggingface.co/tokimoa/smolvla-mlx)(450M・軽量版のVLA移植)
## ライセンス
Apache-2.0(ベースモデルlerobot/pi0のライセンスを継承)
---
Developed by [tokimoa](https://tokimoa.jp)
|