rynn-value-8b / README.md
multimodalart's picture
multimodalart HF Staff
Align SDK/python to the org ZeroGPU baseline (gradio 5.49.1 / py3.10)
5f161c0 verified
|
Raw
History Blame Contribute Delete
2.98 kB

A newer version of the Gradio SDK is available: 6.24.0

Upgrade
metadata
title: RynnValue-8B
emoji: 🦾
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: 5.49.1
app_file: app.py
short_description: How much longer will this robot take?
python_version: '3.10'
startup_duration_timeout: 1h
pinned: false
license: apache-2.0
models:
  - Alibaba-DAMO-Academy/RynnValue-8B
tags:
  - robotics
  - value-model
  - video

RynnValue-8B

Demo of Alibaba-DAMO-Academy/RynnValue-8B, a ~9.6B robotic value model built on Qwen3-VL. Given a robot manipulation video and the task instruction, it predicts the remaining time to task completion (seconds) at every point along the clip — rendered as a curve synchronised with the video — plus a short Analysis block (video description, instruction Match, task Success).

The app follows the official reference implementation (rynn_infer/inference.py): prefix-uniform sampling, where evaluation step i resamples frames[0:i] to num_frames frames and reads the model's last prediction slot, so each score only conditions on frames seen so far; the custom pred_slot_isolated_eager attention implementation; and a greedy Analysis generation pass over the full-video prefix. plot_utils.py is vendored from the same repo (Apache-2.0) with the trend plot memoised per prediction step (it is otherwise re-rendered once per output video frame).

Deviations from the reference, forced by the ZeroGPU time budget: fewer evaluation steps (16 vs one-per-frame), fewer frames per step (32 vs 64), smaller frames (384 px vs 640 px), and the rendered video is temporally subsampled to ≤320 frames (playback fps scaled to match, so durations and the ground-truth reference curve are unchanged). All are adjustable in Advanced settings.

transformers is pinned to 4.57.x (what the checkpoint's remote code targets), which requires huggingface-hub<1.0 — hence the Gradio 5.x SDK version.

Example assets & attribution

  • examples/franka_box_into_drawer.mp4 — the demo clip bundled with alibaba-damo-academy/RynnValue (Apache-2.0), re-encoded to 640 px.
  • examples/soar_put_green_stick_in_brown_bowl.mp4, examples/berkeley_rpt_stack_cup.mp4, examples/jaco_play_pick_up_green_cup.mp4 — RoboMeter benchmark clips bundled in the same repository (MIT), originating from Open X-Embodiment (SOAR / Berkeley RPT / Jaco Play), CC BY 4.0. The task strings are the ones used in the RoboMeter README.
  • examples/so101_lego_into_box.mp4 — episode 1 of lerobot/svla_so101_pickplace (Apache-2.0), side camera.

The last example row pairs the Franka video with an unrelated instruction to show the video/instruction matching behaviour (Match: No).