SmolVLA: Optimized for AMD ROCm

SmolVLA (vision-language-action) is a behavior-cloning policy from Hugging Face LeRobot for 6-DOF robot arm control. This repository packages inference for robot arm action prediction using PyTorch, exported and validated for AMD ROCm so it runs efficiently on AMD GPUs and CPUs.

This is based on the implementation of SmolVLA found here. This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the smolVLA AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.


Task Overview

Task: Robot arm action prediction (vision-language-action)

Dataset: BlankHead/so101_redcube_greencloth_3cams (LeRobot format)

Output metrics: MAE, RMSE (per-joint and per-episode)

PyTorch note: CPU runs FP32; GPU runs BF16. No NPU (VitisAI) path is available — make setup-npu, make benchmark-npu, and make evaluate-npu print an informational note and exit cleanly.


AMD ROCm Optimization

This model export has been adapted and validated for AMD Instinct™ / Radeon™ GPUs running ROCm, as well as AMD CPUs. Key points:

  • Validated backend: PyTorch (native ROCm HIP kernels), FP32 on CPU, BF16 on GPU.
  • No code changes required versus the upstream SmolVLA implementation — only environment/runtime configuration differs.
  • No NPU (VitisAI) fallback path is available for this model.
Runtime Precision Backend Hardware Notes
PyTorch FP32 HIP (ROCm) AMD CPU
PyTorch BF16 HIP (ROCm) AMD Instinct™ / Radeon™ GPU No NPU path available

Getting Started

For setup instructions, evaluation scripts, and custom configuration options, see the smolVLA on GitHub.


Model Details

Model Type: Vision-language-action policy for robot arm control

Base Model: lerobot/smolvla_base

Model Stats:

  • Precision tested: FP32 (CPU), BF16 (GPU)
  • Configurable runtime knobs: rtc_config.enabled (Real-Time Chunking), num_steps (flow-matching denoising passes per chunk), n_action_steps (actions consumed per chunk)

Accuracy Pipeline

Open-loop offline evaluation is fully implemented: make evaluate-<device> runs inference on recorded dataset episodes and computes per-joint and per-episode MAE / RMSE against the recorded ground-truth actions. Lower is better for both metrics.

Metrics Explained

Metric Description
MAE Mean Absolute Error — average absolute difference between predicted and ground-truth joint positions across all timesteps. Lower is better.
RMSE Root Mean Squared Error — penalizes large deviations more heavily than MAE. Lower is better.

Accuracy Results

Published Results — Dataset: BlankHead/so101_redcube_greencloth_3cams (13 episodes, chunked_rtc mode):

Metric Value
Average MAE 3.9521
Average RMSE 7.7251

Per-joint breakdown:

Joint Avg MAE Avg RMSE
shoulder_pan 3.3654 5.1317
shoulder_lift 7.8224 14.0630
elbow_flex 4.9989 8.7409
wrist_flex 2.4138 3.4653
wrist_roll 2.6728 3.9475
gripper 2.4394 4.5340

Dig Deeper

Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?

📂 View the full project on GitHub

The GitHub repository includes:

  • Setup and prerequisites for ROCm environments
  • Open-loop dataset evaluation with trajectory plots and comparison videos
  • Latency benchmarking with Chrome trace output
  • PEFT/LoRA-aware checkpoint loading and runtime chunking configuration
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading