Instructions to use ZibinDong/ActionCodec2-Extended-1mm-8k with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ZibinDong/ActionCodec2-Extended-1mm-8k with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ZibinDong/ActionCodec2-Extended-1mm-8k", device_map="auto") - Notebooks
- Google Colab
- Kaggle
ActionCodec2 Extended First-Order — 1mm, 8k
Trained on curated data with lower-quality datasets and most stationary-arm segments removed.
| Setting | Value |
|---|---|
| Motion order | 1 |
| Precision tier | 1mm |
| Codebook budget | 8192 per profile |
| Fitted profiles | joint, eef |
| Registered layouts | 13 |
| Target rate | 15 Hz |
| Gripper | Binary, zero-order; open when ≥ 0.8 |
The budget applies separately to each profile; joint and EEF tokens use disjoint namespaces. Precision tiers describe physical quantization settings, not an end-to-end policy accuracy guarantee.
Usage
pip install numpy scipy torch "transformers>=4.57,<6" huggingface-hub pyyaml
import numpy as np
from transformers import AutoProcessor
codec = AutoProcessor.from_pretrained("ZibinDong/ActionCodec2-Extended-1mm-8k", trust_remote_code=True)
actions = np.zeros((12, 7), dtype=np.float32)
actions[:, 6] = 1.0
tokens = codec.encode(actions, fps=15)
decoded_actions = codec.decode(tokens, fps=15)
The default layout is single_eef_delta: position increments in metres (columns 0–2),
body-frame rotation-vector increments in radians (3–5), and an absolute gripper command (6).
Rotation follows R_next = R_previous @ Exp(rotvec). Supply physical values and the actual
recording rate through fps; the tokenizer resamples to 15 Hz.
Inputs may have shape (T, D) or (B, T, D). Decoding returns a CPU float32 tensor
of shape (B, T_out, D); resampling can change the length. Quantization is lossy.
Inspect and select layouts explicitly:
codec.print_action_spaces()
codec = codec.for_action_space("single_eef_abs")
Absolute layouts also need current_state for absolute reconstruction.
Joint, EEF, single-arm and dual-arm layouts are included.
Native acceleration
The artifact bundles C++17 sources for physical quantization, Set-BPE, resampling and pose operations. Build them once to enable accelerated, parallel batch encoding and decoding:
hf download ZibinDong/ActionCodec2-Extended-1mm-8k --local-dir ./actioncodec2-tokenizer
pip install ./actioncodec2-tokenizer
A C++17 compiler is required. The processor automatically uses the installed
kernels; installing the ActionCodec2 source
also supplies them. Without kernels, it uses the bundled Python runtime.
ACTIONCODEC2_NUM_THREADS controls the per-process thread budget;
ACTIONCODEC2_NO_NATIVE=1 selects the Python reference implementation.
Keep the complete repository when downloading or saving: router_config.yaml, profiles/,
runtime/, the Hugging Face entry point, metadata, native sources and license.
- Downloads last month
- -