futurefantasy commited on
Commit
986931d
·
verified ·
1 Parent(s): 81cb25c

Update VLAC-Cut citation and model card text

Browse files
Files changed (1) hide show
  1. README.md +13 -6
README.md CHANGED
@@ -24,9 +24,10 @@ The model is designed for long-horizon robot manipulation videos containing fail
24
 
25
  ## Highlights
26
 
27
- - **Video-level progress understanding:** reasons over temporally ordered robot videos rather than isolated images or image pairs.
28
  - **Non-monotonic process modeling:** recognizes advancement, stagnation, regression, and recovery.
29
- - **Generative output:** produces timestamped progress predictions through the standard autoregressive language-model interface, without a task-specific regression head.
 
30
 
31
  ## Model at a Glance
32
 
@@ -35,7 +36,7 @@ The model is designed for long-horizon robot manipulation videos containing fail
35
  | Base model | `Qwen/Qwen3-VL-30B-A3B-Instruct` |
36
  | Input | Task instruction, optional task plan, and sampled video frames |
37
  | Output | Timestamped task progress |
38
- | Default video sampling rate | `2 Hz` |
39
 
40
  ## Model Design
41
 
@@ -105,10 +106,16 @@ python scripts/utils/render_prediction_video.py \
105
 
106
  ## Citation
107
 
108
- Please replace the placeholder below with the final paper BibTeX before public release.
109
-
110
  ```bibtex
111
-
 
 
 
 
 
 
 
 
112
  ```
113
 
114
  ## License
 
24
 
25
  ## Highlights
26
 
27
+ - **Video-level progress understanding:** reasons over robot videos rather than isolated images or image pairs, and cuts the video into good segments and bad segments.
28
  - **Non-monotonic process modeling:** recognizes advancement, stagnation, regression, and recovery.
29
+ - **General understanding capability:** enables zero-shot generalization across tasks, scenes, and viewpoints.
30
+ - **Coarse/Fine-grained flexible:** adjusts the frequency of video understanding, achieving progress analysis from coarse-grained to fine-grained.
31
 
32
  ## Model at a Glance
33
 
 
36
  | Base model | `Qwen/Qwen3-VL-30B-A3B-Instruct` |
37
  | Input | Task instruction, optional task plan, and sampled video frames |
38
  | Output | Timestamped task progress |
39
+ | Flexible video sampling rate | `2 Hz`-`20 Hz` (default: `2.0`) |
40
 
41
  ## Model Design
42
 
 
106
 
107
  ## Citation
108
 
 
 
109
  ```bibtex
110
+ @misc{zhai2026maximizinghumanefficiencylargescale,
111
+ title={Maximizing Human Efficiency in Large-Scale Robot Post-Training via VLAC-Cut Guided Pipeline},
112
+ author={Shaopeng Zhai and Qi Zhang and Tianyi Zhang and Haoran Zhang and Fuxian Huang and Zhanhui Lin and Zijun Xu},
113
+ year={2026},
114
+ eprint={2607.09776},
115
+ archivePrefix={arXiv},
116
+ primaryClass={cs.RO},
117
+ url={https://arxiv.org/abs/2607.09776},
118
+ }
119
  ```
120
 
121
  ## License