Instructions to use OmniJev/OneJev-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OmniJev/OneJev-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="OmniJev/OneJev-4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("OmniJev/OneJev-4B") model = AutoModelForMultimodalLM.from_pretrained("OmniJev/OneJev-4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use OmniJev/OneJev-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "OmniJev/OneJev-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OmniJev/OneJev-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/OmniJev/OneJev-4B
- SGLang
How to use OmniJev/OneJev-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "OmniJev/OneJev-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OmniJev/OneJev-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "OmniJev/OneJev-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OmniJev/OneJev-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use OmniJev/OneJev-4B with Docker Model Runner:
docker model run hf.co/OmniJev/OneJev-4B
OneJev-4B is the 4B model of OneJev, a multimodal System One decision model. Give it a screenshot, a photo, a video or plain text along with a few typed questions, and it returns a calibrated probability for every option in a single forward pass. It is a full fine-tune of Qwen/Qwen3.5-4B on 99,193 questions drawn from real agent runs, videos and images.
Results
Accuracy in percent. The OneJev test set holds questions of the training kinds that no model saw in training. Jev 1.13's scores are its published ones, and it reads text only.
On one H200 with a 1280x720 screenshot, OneJev-4B answers 1 question in 64 ms and 10 questions in one request in 104 ms (10.4 ms per question).
Quick start
pip install git+https://github.com/OmniJev/OneJev.git
qev serve --model OmniJev/OneJev-4B
from qev import Client, Choice, Noul
from qev.media import data_uri
r = Client("http://localhost:8000").system_one(
state={"task": "Pay the open invoice from ACME", "screen": "<image:1>"},
media=[{"type": "image", "data": data_uri("screenshot.png")}],
questions={"done": Noul("The invoice has been paid"),
"next": Choice("What should the agent do next?", {"click": "click an element", "stop": "stop"})},
)
The server speaks TypeSafe's System One API plus a media field for images and video. More examples, the latency
benchmark and the code are on GitHub.
All sizes
| Model | Base | Weights |
|---|---|---|
| OneJev-0.8B | Qwen3.5-0.8B | 2.2 GB |
| OneJev-4B | Qwen3.5-4B | 10.4 GB |
| OneJev-9B | Qwen3.5-9B | 18.8 GB |
| OneJev-27B | Qwen3.8-27B | 54.7 GB |
| OneJev-27B-FP8 | OneJev-27B in 8-bit | 30.4 GB |
All sizes are in the OneJev collection.
Citation
@misc{onejev2026,
title = {{OneJev}: A Multimodal System One Decision Model},
author = {{OmniJev Team}},
year = {2026},
howpublished = {\url{https://github.com/OmniJev/OneJev}}
}
License
Apache 2.0
- Downloads last month
- 22