Instructions to use hqfang/think-new with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use hqfang/think-new with Transformers:
# Load model directly from transformers import AutoModelForImageTextToText model = AutoModelForImageTextToText.from_pretrained("hqfang/think-new", trust_remote_code=True, device_map="auto") - LeRobot
How to use hqfang/think-new with LeRobot:
- Notebooks
- Google Colab
- Kaggle
YAMBox2 groceries — MolmoAct2 Think (step 20k)
This is the depth-reasoning MolmoAct2-Think checkpoint fine-tuned on
Jiafei1224/as for a bimanual
YAM robot putting groceries into a container. This repository contains the
20,000-step checkpoint from the 50,000-step training run.
Adaptive depth is enabled by default. config.json contains
"enable_depth_reasoning": true, and calling predict_action without an
enable_depth_reasoning argument reads that checkpoint default. Pass the
returned depth_cache into the next frame to reuse unchanged depth regions and
regenerate changed regions adaptively. No dataset depth_updated_mask is
required at inference time.
Policy contract
- Base checkpoint:
allenai/MolmoAct2-Think - Cameras, in order:
observation.images.top,observation.images.left,observation.images.right - State:
observation.state, 14 dimensions - Output: 14-dimensional absolute joint targets
- Action horizon / returned steps: 30 / 30
- Normalization tag:
yam_dual_molmoact2 - Grippers are not separately normalized
- Training sequence length: 896
- Depth representation: 10 x 10 discrete bins (100 codes)
- Default output style:
robot_depth_action
The converted action expert has a padded internal width of 32, but
norm_stats.json defines the real robot action width as 14 and predict_action
returns only those 14 dimensions.
Install
pip install "torch" "transformers==5.14.1" accelerate pillow numpy safetensors
The model uses custom Transformers code included in this repository, so loading
requires trust_remote_code=True.
Run with adaptive depth
import numpy as np
import torch
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor
model_id = "hqfang/think-new"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16,
).to("cuda").eval()
assert model.config.enable_depth_reasoning is True
depth_cache = None
def predict(top_path, left_path, right_path, state):
global depth_cache
# Keep this exact camera order. Adaptive change detection uses the first image.
images = [
Image.open(top_path).convert("RGB"),
Image.open(left_path).convert("RGB"),
Image.open(right_path).convert("RGB"),
]
with torch.inference_mode():
result = model.predict_action(
processor=processor,
images=images,
task="put the groceries into the container",
state=np.asarray(state, dtype=np.float32),
norm_tag="yam_dual_molmoact2",
expected_action_representation="absolute",
inference_action_mode="continuous",
num_steps=10,
n_action_steps=30,
depth_cache=depth_cache,
# No enable_depth_reasoning argument: the checkpoint default is True.
)
depth_cache = result.depth_cache
actions = result.actions[0].float().cpu().numpy()
assert actions.shape == (30, 14)
assert result.depth_bins.shape[-1] == 100
return actions
actions = predict("top.png", "left.png", "right.png", [0.0] * 14)
The first frame generates the complete depth representation. On later frames,
passing depth_cache enables adaptive selective regeneration. To explicitly
disable depth for an ablation, pass enable_depth_reasoning=False and omit the
cache.
actions[t] is an absolute 14-dimensional joint target, not a delta. Validate
joint ordering, limits, timing, and emergency-stop behavior before commanding
physical hardware.
Open-loop check
The fixed-random open-loop evaluation used episode 2 from Jiafei1224/as
(1,488 frames, continuous inference, adaptive depth, 10 flow steps). Raw action
MSE was 0.0010854247 at step 20,000. This is an offline reconstruction metric,
not a robot task-success measurement.
Citation
@misc{fang2026molmoact2actionreasoningmodels,
title={MolmoAct2: Action Reasoning Models for Real-world Deployment},
author={Haoquan Fang and Jiafei Duan and Donovan Clay and Sam Wang and Shuo Liu and Weikai Huang and Xiang Fan and Wei-Chuan Tsai and Shirui Chen and Yi Ru Wang and Shanli Xing and Jaemin Cho and Jae Sung Park and Ainaz Eftekhar and Peter Sushko and Karen Farley and Angad Wadhwa and Cole Harrison and Winson Han and Ying-Chun Lee and Eli VanderBilt and Rose Hendrix and Suveen Ellawela and Lucas Ngoo and Joyce Chai and Zhongzheng Ren and Ali Farhadi and Dieter Fox and Ranjay Krishna},
year={2026},
eprint={2605.02881},
archivePrefix={arXiv},
primaryClass={cs.RO}
}
- Downloads last month
- -
Model tree for hqfang/think-new
Base model
allenai/MolmoAct2-Think