VisAlloc — LookAway 9B

This repository contains the completed step 390 checkpoint from the LookAway-Qwen3.5-9B-b16-765d1f5 training run, based on Qwen/Qwen3.5-9B and the LookAway codebase.

The FSDP actor checkpoint was merged into a complete Hugging Face model. The release contains 9,409,813,744 parameters in BF16, tokenizer, image processor, generation configuration, and the checkpoint's chat template. The base-model revision is c202236235762e1c871ad0ccb60c8ee5ba337b9a. The training code was based on commit 765d1f5 with local runtime and memory adjustments. The checkpoint completed 390 training steps and one epoch.

Usage

The evaluation environment used Transformers 5.5.0, PyTorch 2.10.0, and vLLM 0.18.0. Keep the supplied processor and chat template when loading this checkpoint.

import torch
from PIL import Image
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration

model_id = "pianzhikuang/VisAlloc"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
    model_id, dtype=torch.bfloat16, device_map="auto"
).eval()

image = Image.open("image.jpg").convert("RGB")
messages = [{"role": "user", "content": [
    {"type": "image", "image": image},
    {"type": "text", "text": "Describe this image."},
]}]
inputs = processor.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    return_dict=True, return_tensors="pt",
).to(model.device)
with torch.inference_mode():
    output = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(processor.decode(
    output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True
))

Evaluation

The following results were measured on this checkpoint on October 9, 2026. Inference used full images, BF16, greedy decoding (temperature=0), up to 4,096 generated tokens, and a 32,768-token engine context. The checkpoint's chat template was used without an enable_thinking override.

Judging used strict matching for unambiguous answers and GPT-6 Astra for the remaining semantic decisions. All 16,074 inference requests completed; 16,073 samples have valid ground-truth labels and were scored.

Benchmark Metric Score (%) Scored samples
V* Accuracy 83.25 191
ZoomBench Valid-subset accuracy 58.29 844 / 845
HR-4K Accuracy 74.625 800
HR-8K Accuracy 72.25 800
MMVP Per-question accuracy 78.00 300
CV-Bench Mean of 2D and 3D accuracy 74.77 2,638
MMStar Accuracy 62.20 1,500
POPE Accuracy 89.03 9,000
  • ZoomBench contains one question with only A/B choices but a ground-truth label of D. That sample remains unscored; 58.29% is 492/844, not a full-dataset score. The exception is recorded in the evaluation manifest.
  • MMVP paired accuracy is 57.33% (86/150 pairs).
  • POPE F1 is 88.78%. Evaluation includes all three 3,000-question subsets.
  • HR-4K is exactly 597/800 = 74.625%; the CSV uses Python's two-decimal formatting.

See results.json, results.csv, and the evaluation manifest for exact scores, pinned dataset revisions, generation settings, and integrity hashes. Results depend on this prompting and judging protocol.

Release integrity

model.safetensors is the exact weight file used in the evaluation:

d7977253c2b469499e0b9f9e509d9cf60c0ae16d112b119cb2ebce072e59b0ef

The published top-level config.json dtype is corrected to bfloat16 to match all saved tensors and the evaluated inference dtype. The tensor weights and chat template are unchanged.

The base model is released under Apache-2.0; its license is included in LICENSE. This repository contains the model artifacts and aggregate evaluation metadata; training images and benchmark images are separate resources.

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pianzhikuang/VisAlloc

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(988)
this model