Instructions to use Valen-Team/Valen-0.8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Valen-Team/Valen-0.8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Valen-Team/Valen-0.8B", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Valen-Team/Valen-0.8B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Valen-Team/Valen-0.8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Valen-Team/Valen-0.8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Valen-Team/Valen-0.8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Valen-Team/Valen-0.8B
- SGLang
How to use Valen-Team/Valen-0.8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Valen-Team/Valen-0.8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Valen-Team/Valen-0.8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Valen-Team/Valen-0.8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Valen-Team/Valen-0.8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Valen-Team/Valen-0.8B with Docker Model Runner:
docker model run hf.co/Valen-Team/Valen-0.8B
Valen-0.8B
Qwen3.5 with a two-layer MLP-Mixer decision head. Supports text, images and videos, and returns Choice, Noul and Score decisions. Multiple questions can share one state encoding with execution="shared_state".
This revision contains the final full-parameter SFT model. Training uses 1,225,000 records (1,626,911 QA): the previous 1,195k mixture plus JevBench 30k. Two stages: 67 head-warmup steps on a 10% subset, then 325 joint steps covering all records once on 32 GPUs. Joint SFT updates the entire language backbone, vision backbone, visual merger and Mixer. No RL or LoRA is used in this revision.
Inference
Install Python 3.10+, PyTorch, Transformers 5.4.0, torchvision, Pillow and av. Flash Attention 2 is optional with a compatible CUDA build. The model contains its tokenizer, processor and custom inference code; a separate base-model download is unnecessary.
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained(
"Valen-Team/Valen-0.8B", trust_remote_code=True,
dtype="auto", attn_implementation="sdpa",
).to("cuda").eval()
torch.set_float32_matmul_precision("highest")
torch.backends.cudnn.allow_tf32 = False
print(model.predict({
"state": "A cat is on the sofa.",
"questions": {
"animal": {"type": "choice", "instructions": "Which animal is present?",
"criteria": {"cat": "A cat", "dog": "A dog"}},
"on_sofa": {"type": "noul", "instructions": "The cat is on the sofa."},
},
}, execution="shared_state"))
Use attn_implementation="flash_attention_2" for Flash Attention, or execution="question" for independent questions. Video defaults to 16 frames. Image/video request examples and training instructions are in the Valen repository.
dtype="auto" preserves the trained FP32 parameters. The backbone runs under BF16 autocast and the Mixer stays FP32, matching native checkpoint evaluation. Loading all weights as BF16 rounds the trained parameters and can change outputs.
Evaluation
Final native checkpoint, shared-state inference, the same fixed evaluation sets used for the earlier release:
| Benchmark | Accuracy (%) |
|---|---|
| eval_v1 / VisualDecisionBench Image (2,000 QA) | 75.94 |
| General (10,000 QA) | 87.16 |
| Video / VisualDecisionBench Video (7,143 QA) | 79.03 |
| Public JevBench (231 tasks) | 77.92 |
Accuracy includes hard labels only; soft-label Score questions contribute probability metrics. JevBench reports accuracy over all 231 tasks. Export validation checks trained tensor preservation, reloading, real text/image/video inputs, all three question types, and both execution paths; it is separate from benchmark evaluation. See validation.json and export_manifest.json for reproducibility details.
Earlier LoRA release
The original LoRA SFT weights, including unmerged/, remain available at revision lora-sft-1225k. This is an earlier trained model. It is not an adapter representation of this full-SFT revision.
from huggingface_hub import snapshot_download
folder = snapshot_download("Valen-Team/Valen-0.8B", revision="lora-sft-1225k", allow_patterns="unmerged/*")
model = AutoModel.from_pretrained(
f"{folder}/unmerged", trust_remote_code=True,
dtype=torch.bfloat16, attn_implementation="sdpa", local_files_only=False,
).to("cuda").eval()
The LoRA loader downloads its original pinned Qwen base. To reuse a local original base, pass base_model_path="/path/to/Qwen3.5-0.8B".
- Downloads last month
- 25