Image-Text-to-Text
Transformers
Safetensors
qwen3_5
vllm
video
multimodal
reinforcement-learning
temporal-grounding
object-tracking
video-segmentation
visual-question-answering
spatial-reasoning
qwen3.5
conversational
Instructions to use OraRL/Video-ORA-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OraRL/Video-ORA-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="OraRL/Video-ORA-4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("OraRL/Video-ORA-4B") model = AutoModelForMultimodalLM.from_pretrained("OraRL/Video-ORA-4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use OraRL/Video-ORA-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "OraRL/Video-ORA-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OraRL/Video-ORA-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/OraRL/Video-ORA-4B
- SGLang
How to use OraRL/Video-ORA-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "OraRL/Video-ORA-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OraRL/Video-ORA-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "OraRL/Video-ORA-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OraRL/Video-ORA-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use OraRL/Video-ORA-4B with Docker Model Runner:
docker model run hf.co/OraRL/Video-ORA-4B
| # Copyright 2024 Bytedance Ltd. and/or its affiliates | |
| # | |
| # Licensed under the Apache License, Version 2.0 (the "License"); | |
| # you may not use this file except in compliance with the License. | |
| # You may obtain a copy of the License at | |
| # | |
| # http://www.apache.org/licenses/LICENSE-2.0 | |
| # | |
| # Unless required by applicable law or agreed to in writing, software | |
| # distributed under the License is distributed on an "AS IS" BASIS, | |
| # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. | |
| # See the License for the specific language governing permissions and | |
| # limitations under the License. | |
| """ | |
| Rollout config | |
| """ | |
| from dataclasses import asdict, dataclass, field | |
| from typing import Any, Optional | |
| class RolloutConfig: | |
| name: str = "vllm" | |
| n: int = 1 | |
| temperature: float = 1.0 | |
| top_p: float = 1.0 | |
| top_k: int = -1 | |
| seed: int = 1 | |
| limit_images: int = 0 | |
| dtype: str = "bf16" | |
| gpu_memory_utilization: float = 0.6 | |
| ignore_eos: bool = False | |
| enforce_eager: bool = False | |
| enable_chunked_prefill: bool = False # only for v0 engine | |
| tensor_parallel_size: int = 2 | |
| max_model_len: Optional[int] = None | |
| max_num_batched_tokens: int = 8192 | |
| disable_log_stats: bool = True | |
| disable_tqdm: bool = False | |
| val_override_config: dict[str, Any] = field(default_factory=dict) | |
| # vLLM KV cache dtype. Decode is HBM-bandwidth-bound on long-prompt RL | |
| # rollouts; switching from "auto" (matches model dtype, bf16=2 bytes/elt) | |
| # to "fp8" (1 byte/elt) halves KV traffic ⇒ ~2× decode speedup on Hopper. | |
| # Quality impact is typically < 0.5 IoU on temporal-grounding tasks | |
| # because attention is a low-rank op; H20 has native FP8 path so no | |
| # software emulation overhead. Choices: "auto" / "fp8" / "fp8_e5m2" / | |
| # "fp8_e4m3". Use "fp8" (vLLM picks the best variant on Hopper). | |
| kv_cache_dtype: str = "auto" | |
| # Return the per-token logprob of every sampled token in | |
| # ``rollout_log_probs`` and let the trainer reuse it as ``old_log_probs`` | |
| # instead of recomputing under FSDP, so the first PPO mini-batch update does | |
| # not start from ratio == 1. Incompatible with oracle rows, whose tokens | |
| # differ from what the rollout engine sampled. | |
| calculate_log_probs: bool = False | |
| # Emit the per-sequence mean logprob into | |
| # ``non_tensor_batch["seq_logprob_for_filter"]`` only. It never becomes | |
| # ``old_log_probs``, so it stays out of the PPO ratio path and remains | |
| # compatible with oracle rows. | |
| collect_seq_logprob_for_filter: bool = False | |
| # below are auto keys | |
| prompt_length: int = field(default=-1, init=False) | |
| response_length: int = field(default=-1, init=False) | |
| trust_remote_code: bool = field(default=False, init=False) | |
| def to_dict(self): | |
| return asdict(self) | |