Image-Text-to-Text
Transformers
Safetensors
qwen3_5
vllm
video
multimodal
reinforcement-learning
temporal-grounding
object-tracking
video-segmentation
visual-question-answering
spatial-reasoning
qwen3.5
conversational
Instructions to use OraRL/Video-ORA-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OraRL/Video-ORA-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="OraRL/Video-ORA-4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("OraRL/Video-ORA-4B") model = AutoModelForMultimodalLM.from_pretrained("OraRL/Video-ORA-4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use OraRL/Video-ORA-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "OraRL/Video-ORA-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OraRL/Video-ORA-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/OraRL/Video-ORA-4B
- SGLang
How to use OraRL/Video-ORA-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "OraRL/Video-ORA-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OraRL/Video-ORA-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "OraRL/Video-ORA-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OraRL/Video-ORA-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use OraRL/Video-ORA-4B with Docker Model Runner:
docker model run hf.co/OraRL/Video-ORA-4B
File size: 4,646 Bytes
0185029 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 | <p align="right"><a href="training.md">English</a></p>
# 使用 OraRL 训练
本流程将已获许可的源数据整理为可审计的 GRPO 或 OraRL 训练任务。请先
[安装固定版本软件栈](environment_zh.md#安装固定版本软件栈)。
`orarl-train` 是公开的训练入口:它会检查路径与覆盖项、解析发布配置,再启动
本仓库自带的训练器,无需准备第二份运行时源码。
## 1. 构建训练清单
整理后的训练数据将稍后单独上传,目前尚未包含在本 Git 仓库中。在发布完成前,
请按上游许可证自行获取标注和媒体文件,然后复制示例清单:
```bash
cp configs/data_sources.example.yaml ./data_sources.local.yaml
```
将所有 `../local_data` 占位符替换为已获许可的本地路径。相对路径以清单文件所在
目录为基准。每个数据源需要声明标注文件、任务、任务族、配额和媒体根目录,并可
记录许可证与来源页面。
公开示例对应论文使用的 100,032 条训练混合:
| 任务族 | 条数 |
| --- | ---: |
| 时序定位 | 20,096 |
| 跟踪 | 13,952 |
| 分割 | 12,032 |
| 空间定位 | 7,040 |
| 时空定位 | 9,536 |
| 视频问答 | 20,288 |
| 空间智能 | 17,088 |
构建确定性的训练集和 canary 集:
```bash
orarl-prepare \
--config ./data_sources.local.yaml \
--output ./prepared/train.jsonl \
--require-media
```
输出结构为:
```text
prepared/
├── train.jsonl
├── train.canary.jsonl
└── train.manifest.json
```
构建器会检查本地媒体、规范化任务记录、执行数据源配额和单媒体上限、去除重复
prompt、排除给定评测集中的样本,并保证训练集与 canary 集媒体互斥。审计清单会
记录计数、缺额、来源信息和 SHA-256 校验值。
## 2. 选择 GRPO 或 OraRL
| 配置 | 模型规模 | 方法 |
| --- | --- | --- |
| `grpo_4b.yaml` | 4B | GRPO baseline |
| `grpo_9b.yaml` | 9B | GRPO baseline |
| `orarl_4b.yaml` | 4B | OraRL |
| `orarl_9b.yaml` | 9B | OraRL |
论文默认每个 rollout/update batch 使用 64 个 prompt,每个 prompt 采样 8 个策略
回复,因此 100,032 条混合数据在一个 epoch 中对应 1,563 步。
## 3. 预览并启动
设置兼容的本地基础模型和已准备的数据路径:
```bash
MODEL_DIR=/path/to/local/base-model
OUTPUT_DIR="$PWD/runs/orarl-4b"
orarl-train \
--config orarl_4b.yaml \
--model "$MODEL_DIR" \
--train-data "$PWD/prepared/train.jsonl" \
--val-data "$PWD/prepared/train.canary.jsonl" \
--output "$OUTPUT_DIR" \
--nodes 1 \
--gpus-per-node 8
```
命令默认只执行 dry run。检查解析后的调用后,增加 `--run` 开始训练。可通过
`--set KEY=VALUE` 显式覆盖配置,并应随运行产物保留全部覆盖项。
单步 smoke test:
```bash
orarl-train \
--config orarl_4b.yaml \
--model "$MODEL_DIR" \
--train-data "$PWD/prepared/train.jsonl" \
--val-data "$PWD/prepared/train.canary.jsonl" \
--output "$OUTPUT_DIR" \
--nodes 1 \
--gpus-per-node 8 \
--set trainer.max_steps=1 \
--run
```
再使用 `grpo_4b.yaml` 和独立输出目录验证 baseline。9B 任务需使用匹配的模型与
配置。
## 4. 扩展到多节点
所有节点必须看到相同的源码、模型、数据和输出路径:
```bash
HOSTS=node-a,node-b \
bash scripts/launch_multinode.sh \
--gpus-per-node 8 \
-- \
--config "$PWD/configs/orarl_4b.yaml" \
--model "$MODEL_DIR" \
--train-data "$PWD/prepared/train.jsonl" \
--val-data "$PWD/prepared/train.canary.jsonl" \
--output "$OUTPUT_DIR"
```
启动器同样默认为 dry run;需要在 `--` 分隔符前加入其自身的 `--run`。SSH
主机密钥检查默认使用严格模式。
## 5. 验收一次运行
`scripts/smoke_training.sh` 会以较小的 batch 分别跑一步 GRPO 和一步 OraRL,
并各保存一个 checkpoint:
```bash
bash scripts/smoke_training.sh \
--model "$MODEL_DIR" \
--train-data "$PWD/prepared/train.jsonl" \
--val-data "$PWD/prepared/train.canary.jsonl" \
--size 4b \
--gpus-per-node 8
```
加上 `--dry-run` 可在不占用 GPU 的情况下检查解析后的命令。
正式实验前确认:
- GRPO 和 OraRL 均可完成一步更新,reward、loss、梯度范数和选择指标均为有限值。
- checkpoint 可保存、重新加载并继续完成一步更新。
- 多节点任务能组成预期的 Ray 集群并完成一步更新。
- 保存源码版本、配置、命令、数据清单校验值、环境版本、加速卡类型和所有覆盖项。
使用[评测文档](evaluation_zh.md)评估导出的 checkpoint。
|