Instructions to use OraRL/Video-ORA-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OraRL/Video-ORA-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="OraRL/Video-ORA-9B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("OraRL/Video-ORA-9B") model = AutoModelForMultimodalLM.from_pretrained("OraRL/Video-ORA-9B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use OraRL/Video-ORA-9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "OraRL/Video-ORA-9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OraRL/Video-ORA-9B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/OraRL/Video-ORA-9B
- SGLang
How to use OraRL/Video-ORA-9B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "OraRL/Video-ORA-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OraRL/Video-ORA-9B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "OraRL/Video-ORA-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OraRL/Video-ORA-9B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use OraRL/Video-ORA-9B with Docker Model Runner:
docker model run hf.co/OraRL/Video-ORA-9B
Evaluate Video-ORA
OraRL evaluates direct, task-native answers with canonical profiles for the
seven released task families. The corresponding runtime lives under
eval/task/; first
validate the evaluation node.
1. Materialize model and data
Install the Hugging Face CLI and download the released assets:
python -m pip install -U huggingface_hub
hf download OraRL/Video-ORA-9B \
--local-dir "$PWD/models/Video-ORA-9B"
hf download OraRL/OraRL-Data \
--repo-type dataset \
--include "OraRL-eval-data/**" \
--local-dir "$PWD/OraRL-Data"
The complete evaluation release is large. Downloads are resumable, and
OraRL-eval-data/assets.jsonl is the authoritative file inventory.
The local dataset layout is:
OraRL-Data/OraRL-eval-data/
โโโ datasets.jsonl # benchmark, prompt, parser, metric, and sampling profiles
โโโ assets.jsonl # released-file inventory
โโโ annotations/ # canonical JSONL rows
โโโ media/ # raw images, videos, and subtitles
Derived preprocessing caches are intentionally excluded. Evaluators decode the declared raw media when no compatible cache is present.
2. Run the canonical paper profile
Preview the resolved command first:
orarl-eval \
--model "$PWD/models/Video-ORA-9B" \
--tasks paper \
--dataset "$PWD/OraRL-Data/OraRL-eval-data" \
--summary "$PWD/outputs/Video-ORA-9B/evaluation.json"
orarl-eval is a dry run by default. After checking the model, dataset,
evaluator, task profiles, and output paths, add --run.
To validate the pipeline with a bounded smoke test:
orarl-eval \
--model "$PWD/models/Video-ORA-9B" \
--tasks videomme \
--dataset "$PWD/OraRL-Data/OraRL-eval-data" \
--max-samples 8 \
--summary "$PWD/outputs/Video-ORA-9B/videomme-smoke.json" \
--run
Smoke scores only validate execution and must not be reported as benchmark results.
3. Compose a task suite
--tasks paper selects all released tasks. --tasks video_qa selects the
seven Video QA benchmarks. Individual task names may be comma-separated.
| Family | Task names |
|---|---|
| Video QA | videomme, videommev2, mvbench, mmvu, videoholmes, longvideobench, mlvu |
| Spatial intelligence | vsi, mmsi, mindcube, revsi |
| Temporal grounding | temporal_grounding |
| Spatial grounding | spatial_grounding |
| Tracking | tracking |
| Spatial-temporal grounding | stvg |
| Segmentation | segmentation |
The canonical profiles in datasets.jsonl pin frame sampling, resolution,
prompts, parsers, and metrics. Changing them defines a different evaluation
setting. ReVSI is reported separately from the three-benchmark
spatial-intelligence average.
4. Add segmentation post-processing
Segmentation inference runs without SAM2 post-processing by default. Enabling
--segmentation-run-sam2 additionally requires:
- SAM2 weights (
SEGMENTATION_SAM2_CKPT) - the matching Hydra config (
SEGMENTATION_SAM2_CFG) - the official OneThinker
seg_post_sam2.py(SEGMENTATION_POSTPROCESSOR_PATH) - the
sam2Python package
OraRL validates all three paths before launch.
5. Preserve reportable outputs
Each task evaluator writes its native summary, then orarl-eval creates the
requested aggregate JSON with requested, completed, and missing tasks, the
return code, and official metrics. Outputs are grouped by paper task family
under outputs/<model>/.
For a reportable run, retain:
- the OraRL source revision
- the model and data revisions
- the exact command and aggregate summary
- software versions and accelerator type/count
- every non-default profile or CLI override
See ../eval/README.md for the evaluator map, video
decoding backend, and checkpoint-format notes.