How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("image-text-to-text", model="AvaXiao/ReToken-InternVL3.5-8B", trust_remote_code=True)
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    },
]
pipe(text=messages)
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("AvaXiao/ReToken-InternVL3.5-8B", trust_remote_code=True, device_map="auto")
Quick Links

ReToken: One Token to Improve Vision–Language Models for Visual Retrieval

This repository contains the model from the paper ReToken: One Token to Improve Vision-Language Models for Visual Retrieval.

ReToken is a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache.

Code: https://github.com/avaxiao/ReToken

Downloads last month
16
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for AvaXiao/ReToken-InternVL3.5-8B