Spaces:
Running on Zero
Running on Zero
A newer version of the Gradio SDK is available: 6.22.0
metadata
title: Demo of Encoder-Free VLM Trained for $100
emoji: 🖼️
colorFrom: indigo
colorTo: pink
sdk: gradio
sdk_version: 5.49.1
app_file: app.py
pinned: false
license: apache-2.0
short_description: Ask questions about images and get answers instantly
Encoder-Free VLM Demo
An encoder-free vision-language model: there is no pretrained vision
encoder. A learned patch embedder splits a 512×512 image into 32×32 patches and
projects them directly into the hidden space of
Qwen3-1.7B. The projected patch
embeddings are spliced into <|image|> placeholder positions and the decoder
generates from inputs_embeds.
This Space serves the best checkpoint trained on the DenseFusion-1M + ShareGPT4V data mix.
Upload a single image, ask a question, and the model streams a response. The demo is single-image and stateless — each turn is a fresh conversation.
Inspired by the SmolVLM-256M-Demo Space.