Image-Text-to-Text
Transformers
Safetensors
Turkish
English
internvl_chat
feature-extraction
computer-vision
multimodal
e-commerce
catalog-moderation
vision-language-model
product-understanding
conversational
custom_code
Instructions to use Trendyol/Trendyol-Vision-Flash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Trendyol/Trendyol-Vision-Flash with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Trendyol/Trendyol-Vision-Flash", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Trendyol/Trendyol-Vision-Flash", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Trendyol/Trendyol-Vision-Flash with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Trendyol/Trendyol-Vision-Flash" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Trendyol/Trendyol-Vision-Flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Trendyol/Trendyol-Vision-Flash
- SGLang
How to use Trendyol/Trendyol-Vision-Flash with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Trendyol/Trendyol-Vision-Flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Trendyol/Trendyol-Vision-Flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Trendyol/Trendyol-Vision-Flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Trendyol/Trendyol-Vision-Flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Trendyol/Trendyol-Vision-Flash with Docker Model Runner:
docker model run hf.co/Trendyol/Trendyol-Vision-Flash
File size: 15,126 Bytes
3520a83 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 | ---
license: cc-by-4.0
pipeline_tag: image-text-to-text
language:
- tr
- en
base_model:
- OpenGVLab/InternVL3_5-1B-Instruct
model_type: vision-language-model
tags:
- computer-vision
- multimodal
- e-commerce
- catalog-moderation
- vision-language-model
- product-understanding
library_name: transformers
---
# Trendyol-Vision-Flash
_Trendyol-Vision-Flash is a fine-tuned **vision-language model (VLM)** built on [OpenGVLab/InternVL3_5-1B-Instruct](https://huggingface.co/OpenGVLab/InternVL3_5-1B-Instruct) for Trendyol catalog quality and moderation workflows. It is the lightweight Flash-tier VLM in the Trendyol-Vision family — built for production serving under high traffic, with low-latency inference on a single GPU for brand detection, product similarity, attribute extraction, title generation, content moderation, product captioning, and other **e-commerce catalog operations**._
**Trendyol-Vision-Flash vs Master:** Flash is the production-tier VLM for day-to-day, high-traffic catalog workloads on a single GPU. [Trendyol-Vision-Master](https://huggingface.co/Trendyol/Trendyol-Vision-Master) is the larger expert VLM for critical or harder cases (e.g. category detection) and retains stronger general capabilities.
## Model Details
- **Architecture**: InternVL3.5-1B (InternViT-300M + Qwen3-0.6B LLM backbone)
- **Base model**: [OpenGVLab/InternVL3_5-1B-Instruct](https://huggingface.co/OpenGVLab/InternVL3_5-1B-Instruct)
- **Training**: Full SFT on Trendyol catalog multimodal data
- **Languages**: Turkish (primary), English (secondary)
- **Modalities**: Image + text (up to 16 images per prompt)
- **Serving**: Single GPU
## Intended Use
- Detect product brands from images with category context.
- Determine whether two product listings represent the same SKU.
- Extract structured attributes, titles, and captions from product images.
- Moderate unsafe or policy-violating product content.
- Run low-latency catalog enrichment on a single GPU under high-traffic production load.
- Support research and evaluation use cases in e-commerce catalog operations.
**Not intended for** Master-tier category detection, general open-domain chat, medical/legal advice, surveillance, or any use described under Ethical Considerations. Optimized for production catalog workflows rather than general-purpose assistance. For category detection at scale, use [Trendyol/Trendyol-Vision-Master](https://huggingface.co/Trendyol/Trendyol-Vision-Master).
## Supported Use Cases
Specialized for **e-commerce catalog operations** (moderation and enrichment). The tasks below are the primary, production-validated workloads; related catalog workflows can be prompted similarly, with best results on these patterns.
- **Brand detection** — Infer the brand from product images with optional category context.
- **Product similarity** — Decide whether two listings (images and titles) refer to the same product.
- **Attribute extraction** — Extract structured product attributes from image and title/description.
- **Title generation** — Produce a clean catalog title from the product image and a reference title.
- **Product caption** — Generate a grounded product description from the image and optional metadata.
- **Content safety classification** — Classify product image and title for catalog content safety (`0` = Forbidden, `1` = Fantasy, `2` = Safe). **Fantasy** means content that may be published but should be treated as adult (+18).
- **Per-pack quantity extraction** — Extract pack quantity and unit from image and long title.
See **Task Prompts** below for per-task prompt templates.
---
## Quickstart
Install dependencies:
```bash
pip install "transformers==4.56.2" accelerate sentencepiece torch torchvision pillow requests
```
Load the model, preprocess images from URLs, and run inference with the native InternVL `model.chat()` API:
```python
import torch
import requests
from PIL import Image
from io import BytesIO
from torchvision import transforms as T
from torchvision.transforms.functional import InterpolationMode
from transformers import AutoModel, AutoTokenizer
MODEL_ID = "Trendyol/Trendyol-Vision-Flash"
IMAGENET_MEAN = (0.485, 0.456, 0.406)
IMAGENET_STD = (0.229, 0.224, 0.225)
model = AutoModel.from_pretrained(
MODEL_ID,
trust_remote_code=True,
dtype=torch.bfloat16,
low_cpu_mem_usage=True,
use_flash_attn=False,
).eval().cuda()
tokenizer = AutoTokenizer.from_pretrained(
MODEL_ID,
trust_remote_code=True,
use_fast=False,
)
def build_transform(input_size=448):
return T.Compose([
T.Lambda(lambda img: img.convert("RGB") if img.mode != "RGB" else img),
T.Resize((input_size, input_size), interpolation=InterpolationMode.BICUBIC),
T.ToTensor(),
T.Normalize(mean=IMAGENET_MEAN, std=IMAGENET_STD),
])
def load_image_from_url(url, input_size=448):
response = requests.get(url, timeout=30)
response.raise_for_status()
image = Image.open(BytesIO(response.content)).convert("RGB")
return build_transform(input_size)(image).unsqueeze(0)
urls = [
"https://cdn.dsmcdn.com/mnresize/620/920/ty1571/prod/QC/20240924/23/a77c7933-c626-3753-9cad-fbb7388bff45/1_org_zoom.jpg",
"https://cdn.dsmcdn.com/mnresize/620/920/ty1569/prod/QC/20240924/23/356395c3-c71c-37c5-88c4-5ef97899b3d8/1_org_zoom.jpg",
"https://cdn.dsmcdn.com/mnresize/620/920/ty1571/prod/QC/20240924/23/112305b3-ba5d-3a28-8fd3-b1857749afca/1_org_zoom.jpg",
]
pixel_values = torch.cat([load_image_from_url(url) for url in urls], dim=0)
pixel_values = pixel_values.to(dtype=torch.bfloat16, device="cuda")
question = """<image>
Görsellerden, Eldiven kategorisinde yer alan ürünün markasını çıkar.
Sadece verilen görsellerde doğrulanabilen bilgilere dayan.
Emin olmadığında "Unknown" şeklinde cevap ver.
Sadece marka adını döndür."""
generation_config = {"max_new_tokens": 32, "do_sample": False}
response = model.chat(tokenizer, pixel_values, question, generation_config)
print(response) # Adidas
```
**Image ordering:** concatenate `pixel_values` in the same order as `<image>` placeholders in the prompt. For a single `<image>` token with multiple product photos, pass all images in one batch (default behavior).
> **GPU note:** Trendyol-Vision-Flash is designed for **single-GPU** inference. No tensor parallelism required.
---
## Serving (vLLM)
Trendyol-Vision-Flash can be served with vLLM as an OpenAI-compatible API on a **single GPU**.
```bash
vllm serve Trendyol/Trendyol-Vision-Flash \
--trust-remote-code \
--max-model-len 16384 \
--structured-outputs-config.backend xgrammar \
--served-model-name Trendyol-Vision-Flash \
--interleave-mm-strings \
```
When calling the vLLM OpenAI API, use `image_url` content blocks:
```bash
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer EMPTY" \
-d '{
"model": "Trendyol-Vision-Flash",
"temperature": 0.0,
"max_tokens": 32,
"messages": [
{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "https://cdn.dsmcdn.com/mnresize/620/920/ty1571/prod/QC/20240924/23/a77c7933-c626-3753-9cad-fbb7388bff45/1_org_zoom.jpg"}},
{"type": "image_url", "image_url": {"url": "https://cdn.dsmcdn.com/mnresize/620/920/ty1569/prod/QC/20240924/23/356395c3-c71c-37c5-88c4-5ef97899b3d8/1_org_zoom.jpg"}},
{"type": "image_url", "image_url": {"url": "https://cdn.dsmcdn.com/mnresize/620/920/ty1571/prod/QC/20240924/23/112305b3-ba5d-3a28-8fd3-b1857749afca/1_org_zoom.jpg"}},
{"type": "text", "text": "Görsellerden, Eldiven kategorisinde yer alan ürünün markasını çıkar.\nSadece verilen görsellerde doğrulanabilen bilgilere dayan.\nEmin olmadığında \"Unknown\" şeklinde cevap ver.\nSadece marka adını döndür."}
]
}
]
}'
```
---
## Task Prompts
Pass any prompt below to `model.chat()` using the same image-loading flow from Quickstart. Replace `{placeholders}` with your data. Use one `<image>` per image group; for tasks with multiple image slots, pass `num_patches_list` (e.g. `[1, 1]` for before/after).
### 1. Brand Detection
**Images:** product photos (1–16)
**Output:** brand name or `Unknown`
```
<image>
Görsellerden, {category} kategorisinde yer alan ürünün markasını çıkar.
Sadece verilen görsellerde doğrulanabilen bilgilere dayan.
Emin olmadığında "Unknown" şeklinde cevap ver.
Sadece marka adını döndür.
```
### 2. Product Similarity
**Images:** before listing images, then after listing images
**Output:** `1` (same SKU) or `0` (different)
```
E-ticaret kataloğundaki ürün benzerliği konusunda uzmansınız.
İki ürünün önceki ve sonraki başlık/görsellerini karşılaştırarak aynı ürün olup olmadığını belirle.
Ambalaj, arka plan veya model farkları tek başına fark sayılmaz; renk, boyut, miktar veya varyant farkları fark sayılır.
Sadece "1" veya "0" döndür.
Önceki Ürün Başlığı:
{before_title}
Önceki Ürün Resimleri:
<image>
Sonraki Ürün Başlığı:
{after_title}
Sonraki Ürün Resimleri:
<image>
```
### 3. Attribute Extraction
**Images:** 1 product image
**Output:** JSON object
```
<image>
Bu görseldeki, başlığı '{title}' ve açıklaması '{description}' olan ürünün {attribute_list} bilgilerini json formatında çıkarır mısın?
```
Example:
```
<image>
Bu görseldeki, başlığı 'Helen Bar Sandalyesi-mavi-9519q0119' ve açıklaması 'Maksimum kargolanma süresi: Sipariş tarihinden sonraki 7. gün Ortalama montaj süresi: 1 dakika Genişlik: 62cm Derinlik: 52cm Yükseklik: 96cm Oturak Yüksekliği: 44cm Ürün Ağırlığı: 10kg Kullanılan malzeme: Metal Kullanılan sünger: Yüksek yoğunluklu gri sünger Kullanılan kol: Metal Kollu Kullanılan ayak: Metal Ayaklı + Kromajlı Garanti süresi - Menşei: 24 Ay - Yerli Sipariş bazlı üretim-tedarik yapıldığından sipariş iptali yapılamamaktadır Anlaşmalı olunan ambar ve kargolarla bina kapısında teslimat yapılmaktadır' olan ürünün Garanti Süresi, Materyal, Model, Sandalye Kumaşı, Sandalye Sayısı, Tema / Stil bilgilerini json formatında çıkarır mısın?
```
Image URL: `https://cdn.dsmcdn.com/ty1325/product/media/images/prod/QC/20240522/17/f63a7316-5998-3c7d-b45f-93ff8e9c34a2/1_org_zoom.jpg`
Expected output:
```json
{
"Garanti Süresi": "2 Yıl",
"Materyal": "Metal",
"Model": "Bar Sandalyesi",
"Sandalye Kumaşı": "çıkarılamadı",
"Sandalye Sayısı": "1",
"Tema / Stil": "Modern"
}
```
### 4. Title Generation
**Images:** 1 product image
**Output:** plain-text title
```
<image>
Ürün fotoğrafı ile '{reference_title}' bilgisini karşılaştırıp, Trendyol katalog moderasyon kurallarına göre yanıltıcı ifadelerden kaçınarak net bir başlık üret; çıktıyı düz metin ver.
```
### 5. Product Caption
**Images:** 1 product image
**Output:** English product description
```
<image>
Without speculating about details you cannot see, describe this product based on the image and the information provided.
Product title: {title}
Brand: {brand}
First decide which object is the product, review OCR for brand/model/title clues, then analyze colors, shape, material, pattern, and other grounded details.
```
### 6. Content Safety Classification
**Images:** 1 product image
**Labels:** `0` (Forbidden) = not publishable, `1` (Fantasy) = publishable but adult (+18) content, `2` (Safe) = publishable without restriction.
```
<image>
Ürün Başlığı: {title}
Bu ürün görselini ve başlığını inceleyerek moderasyon sınıflandırması yap. Sonucu 0 (Forbidden), 1 (Fantasy) veya 2 (Safe) olarak ver.
```
### 7. Per-Pack Quantity Extraction
**Images:** 1 product image
**Output:** `{"amount": <number|null>, "unit": "KG"|"L"|"PIECE"|null}`
```
<image>
Extract total product quantity from the title and image. Title is the primary source; use the image only when the title is missing or ambiguous.
Return only JSON: {"amount": <number|null>, "unit": "KG"|"L"|"PIECE"|null}
Convert g→KG, ml/cc→L, counts→PIECE. Ignore model numbers, storage, wattage, and dimensions. If unclear, return {"amount": null, "unit": null}.
Product Title: {title}
```
---
## Limitations
- **Domain specificity**: Optimized for Trendyol e-commerce product images and Turkish catalog text; may not generalize to other domains.
- **Not for Master-tier category detection**: Category detection at taxonomy scale is handled by [Trendyol/Trendyol-Vision-Master](https://huggingface.co/Trendyol/Trendyol-Vision-Master).
- **Multi-image tasks**: Product similarity and brand detection may use multiple images; ensure all images are included in the prompt.
- **Language bias**: Turkish prompts and outputs are primary; English performance varies by task.
- **Not a general assistant**: Fine-tuned for structured catalog tasks in production settings; open-domain chat is out of scope, and quality is strongest on the production-validated tasks above.
## Ethical Considerations
- Designed for catalog quality and moderation workflows.
- Ensure compliance with data protection regulations when processing product and user-generated content.
- Monitor for biased decisions across product categories and brands.
- Not intended for surveillance, misinformation, or harmful content generation.
- Safety and moderation labels reflect training for catalog workflows and may not match every jurisdiction or platform policy; human review is recommended for high-impact decisions.
## Citation
```bibtex
@misc{trendyol-vision-flash,
title={Trendyol-Vision-Flash: Lightweight Catalog Quality Vision-Language Model (VLM)},
author={Trendyol Data Science Team},
year={2026},
howpublished={\url{https://huggingface.co/Trendyol/Trendyol-Vision-Flash}}
}
```
```bibtex
@article{zhu2025internvl3_5,
title={InternVL3.5: Advancing Open-Source Multimodal Models},
author={Zhu, Weiyun and others},
year={2025},
url={https://huggingface.co/OpenGVLab/InternVL3_5-1B-Instruct}
}
```
## Model Card Authors
- Trendyol Data Science Team
## License
This model is licensed under the [Creative Commons Attribution 4.0 International License (CC BY 4.0)](https://creativecommons.org/licenses/by/4.0/).
You are free to share and adapt the model for any purpose, even commercially, as long as you give appropriate credit and indicate if changes were made.
This release is a fine-tune of [OpenGVLab/InternVL3_5-1B-Instruct](https://huggingface.co/OpenGVLab/InternVL3_5-1B-Instruct), which is licensed under [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0) (including its Qwen3 language-model component). Redistribution of this derivative continues to satisfy Apache-2.0 notice and attribution requirements for the base model; retain the Apache-2.0 license text and any NOTICE attributions shipped with InternVL3.5 when redistributing.
For the full CC BY 4.0 license text, see: https://creativecommons.org/licenses/by/4.0/legalcode
---
_Released by the Trendyol Data Science Team._
|