Valen-0.8B

Qwen3.5 with a two-layer MLP-Mixer decision head. Supports text, images and videos, and returns Choice, Noul and Score decisions. Multiple questions can share one state encoding with execution="shared_state".

This revision contains the final full-parameter SFT model. Training uses 1,225,000 records (1,626,911 QA): the previous 1,195k mixture plus JevBench 30k. Two stages: 67 head-warmup steps on a 10% subset, then 325 joint steps covering all records once on 32 GPUs. Joint SFT updates the entire language backbone, vision backbone, visual merger and Mixer. No RL or LoRA is used in this revision.

Inference

Install Python 3.10+, PyTorch, Transformers 5.4.0, torchvision, Pillow and av. Flash Attention 2 is optional with a compatible CUDA build. The model contains its tokenizer, processor and custom inference code; a separate base-model download is unnecessary.

import torch
from transformers import AutoModel

model = AutoModel.from_pretrained(
    "Valen-Team/Valen-0.8B", trust_remote_code=True,
    dtype="auto", attn_implementation="sdpa",
).to("cuda").eval()
torch.set_float32_matmul_precision("highest")
torch.backends.cudnn.allow_tf32 = False
print(model.predict({
    "state": "A cat is on the sofa.",
    "questions": {
        "animal": {"type": "choice", "instructions": "Which animal is present?",
                   "criteria": {"cat": "A cat", "dog": "A dog"}},
        "on_sofa": {"type": "noul", "instructions": "The cat is on the sofa."},
    },
}, execution="shared_state"))

Use attn_implementation="flash_attention_2" for Flash Attention, or execution="question" for independent questions. Video defaults to 16 frames. Image/video request examples and training instructions are in the Valen repository.

dtype="auto" preserves the trained FP32 parameters. The backbone runs under BF16 autocast and the Mixer stays FP32, matching native checkpoint evaluation. Loading all weights as BF16 rounds the trained parameters and can change outputs.

Evaluation

Final native checkpoint, shared-state inference, the same fixed evaluation sets used for the earlier release:

Benchmark Accuracy (%)
eval_v1 / VisualDecisionBench Image (2,000 QA) 75.94
General (10,000 QA) 87.16
Video / VisualDecisionBench Video (7,143 QA) 79.03
Public JevBench (231 tasks) 77.92

Accuracy includes hard labels only; soft-label Score questions contribute probability metrics. JevBench reports accuracy over all 231 tasks. Export validation checks trained tensor preservation, reloading, real text/image/video inputs, all three question types, and both execution paths; it is separate from benchmark evaluation. See validation.json and export_manifest.json for reproducibility details.

Earlier LoRA release

The original LoRA SFT weights, including unmerged/, remain available at revision lora-sft-1225k. This is an earlier trained model. It is not an adapter representation of this full-SFT revision.

from huggingface_hub import snapshot_download
folder = snapshot_download("Valen-Team/Valen-0.8B", revision="lora-sft-1225k", allow_patterns="unmerged/*")
model = AutoModel.from_pretrained(
    f"{folder}/unmerged", trust_remote_code=True,
    dtype=torch.bfloat16, attn_implementation="sdpa", local_files_only=False,
).to("cuda").eval()

The LoRA loader downloads its original pinned Qwen base. To reuse a local original base, pass base_model_path="/path/to/Qwen3.5-0.8B".

Downloads last month
25
Safetensors
Model size
0.9B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Valen-Team/Valen-0.8B

Finetuned
(472)
this model