futurefantasy commited on
Commit
dcbfddc
·
verified ·
1 Parent(s): 87894ff

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +5 -4
README.md CHANGED
@@ -24,9 +24,10 @@ The model is designed for long-horizon robot manipulation videos containing fail
24
 
25
  ## Highlights
26
 
27
- - **Video-level progress understanding:** reasons over temporally ordered robot videos rather than isolated images or image pairs.
28
  - **Non-monotonic process modeling:** recognizes advancement, stagnation, regression, and recovery.
29
- - **Generative output:** produces timestamped progress predictions through the standard autoregressive language-model interface, without a task-specific regression head.
 
30
 
31
  ## Model at a Glance
32
 
@@ -35,13 +36,13 @@ The model is designed for long-horizon robot manipulation videos containing fail
35
  | Base model | `Qwen/Qwen3-VL-30B-A3B-Instruct` |
36
  | Input | Task instruction, optional task plan, and sampled video frames |
37
  | Output | Timestamped task progress |
38
- | Default video sampling rate | `2 Hz` |
39
 
40
  ## Model Design
41
 
42
  VLAC-Cut is fine-tuned from Qwen3-VL-30B-A3B-Instruct and retains its original multimodal generation interface. The model receives a task instruction, an optional task plan, and sampled video frames, then generates task progress at different timestamps as text.
43
 
44
- The training data also include supervision for robot behavior description, failure analysis, and correction planning, helping the model understand realistic robot execution processes.
45
 
46
  This repository contains the model checkpoint and loading assets only. Inference scripts, examples, and VPB evaluation code are maintained in the GitHub repository.
47
 
 
24
 
25
  ## Highlights
26
 
27
+ - **Video-level progress understanding:** reasons over robot videos rather than isolated images or image pairs, and cuts the video into good segments and bad segments.
28
  - **Non-monotonic process modeling:** recognizes advancement, stagnation, regression, and recovery.
29
+ - **General understanding capability:** General understanding capability, enabling zero-shot generalization across tasks, scenes, and viewpoints.
30
+ - **Coarse/Fine-grained flexible:** The ability to adjust the frequency of video understanding, achieving progress analysis from coarse-grained to fine-grained.
31
 
32
  ## Model at a Glance
33
 
 
36
  | Base model | `Qwen/Qwen3-VL-30B-A3B-Instruct` |
37
  | Input | Task instruction, optional task plan, and sampled video frames |
38
  | Output | Timestamped task progress |
39
+ | flexiable video sampling rate | `2 Hz~20Hz`(Default: `2.0`) |
40
 
41
  ## Model Design
42
 
43
  VLAC-Cut is fine-tuned from Qwen3-VL-30B-A3B-Instruct and retains its original multimodal generation interface. The model receives a task instruction, an optional task plan, and sampled video frames, then generates task progress at different timestamps as text.
44
 
45
+ The training data also includes supervision for robot behavior description, failure analysis, and correction planning, helping the model understand realistic robot execution processes.
46
 
47
  This repository contains the model checkpoint and loading assets only. Inference scripts, examples, and VPB evaluation code are maintained in the GitHub repository.
48