Spaces:
Sleeping
Sleeping
A newer version of the Gradio SDK is available: 6.22.0
metadata
title: LabVLA
emoji: 🧪
colorFrom: indigo
colorTo: blue
sdk: gradio
sdk_version: 6.15.1
app_file: app.py
short_description: VLA model for scientific laboratory robotics
python_version: '3.12'
startup_duration_timeout: 1h
LabVLA: Vision-Language-Action Model for Scientific Laboratories
LabVLA is the first VLA foundation model designed specifically for scientific laboratory environments. It combines a Qwen3-VL-4B vision-language backbone with a DiT flow-matching action expert trained with the π0.5 recipe: FAST action-token pre-training → flow-matching post-training with knowledge isolation.
This demo lets you provide a laboratory camera view and a natural language instruction, and the model predicts a chunk of 50 future robot actions (7 arm joint angles + 1 gripper width per step) for a Franka Panda manipulator.
Model
- Model: zjunlp/LabVLA-5B-Base
- Paper: LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories
- GitHub: zjunlp/LabVLA
Inputs
- Camera View: An image of the laboratory workspace
- Task Instruction: Natural language instruction (e.g., "Pick up the beaker")
- Robot State: 7 arm joint angles + 1 gripper width (Franka Panda 7-DOF)
Outputs
- Action Trajectory: 50-step predicted action chunk visualized as per-joint trajectory plots
- Action Details: JSON with the raw action values and timing information