Instructions to use tencent/Youtu-Parsing with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use tencent/Youtu-Parsing with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("image-text-to-text", model="tencent/Youtu-Parsing", trust_remote_code=True)
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    },
]
pipe(text=messages)

# Load model directly
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("tencent/Youtu-Parsing", trust_remote_code=True, dtype="auto")

Notebooks
Google Colab
Kaggle
Local Apps

vLLM

How to use tencent/Youtu-Parsing with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "tencent/Youtu-Parsing"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "tencent/Youtu-Parsing",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'

Use Docker

docker model run hf.co/tencent/Youtu-Parsing

SGLang

How to use tencent/Youtu-Parsing with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "tencent/Youtu-Parsing" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "tencent/Youtu-Parsing",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "tencent/Youtu-Parsing" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "tencent/Youtu-Parsing",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'

Docker Model Runner
How to use tencent/Youtu-Parsing with Docker Model Runner:
```
docker model run hf.co/tencent/Youtu-Parsing
```

Improve model card: add transformers library, pipeline tag, and paper link

by nielsr HF Staff - opened Jan 29

base: refs/heads/main

←

from: refs/pr/2

Discussion Files changed

+76

-35

Files changed (1) hide show

README.md +76 -35

README.md CHANGED Viewed

@@ -1,12 +1,14 @@
 ---
 license: other
 license_name: youtu-parsing
 license_link: https://huggingface.co/tencent/Youtu-Parsing/blob/main/LICENSE.txt
-pipeline_tag: image-text-to-text
-base_model:
-  - tencent/Youtu-LLM-2B
 base_model_relation: finetune
 ---
 <div align="center">
 # <img src="assets/youtu-parsing-logo.png" alt="Youtu-Parsing Logo" height="100px">
@@ -22,7 +24,9 @@ base_model_relation: finetune
 ## 🎯 Introduction
-**Youtu-Parsing** is a specialized document parsing model built upon the open-source Youtu-LLM 2B foundation. By extending the capabilities of the base model with a prompt-guided framework and NaViT-style dynamic visual encoder, Youtu-Parsing offers enhanced parsing capabilities for diverse document elements including text, tables, formulas, and charts. The model incorporates an efficient parallel decoding mechanism that significantly accelerates inference, making it practical for real-world document analysis applications. We share Youtu-Parsing with the community to facilitate research and development in document understanding.
 ## ✨ Key Features
@@ -62,33 +66,70 @@ base_model_relation: finetune
 <a id="quickstart"></a>
 ## 🚀 Quick Start
-### Install packages
-```bash
-conda create -n youtu_parsing python=3.10
-conda activate youtu_parsing
-pip install git+https://github.com/TencentCloudADP/youtu-parsing.git#subdirectory=youtu_hf_parser
-# install the flash-attn2
-# For CUDA 12.x + PyTorch 2.6 + Python 3.10 + Linux x86_64:
-pip install https://github.com/Dao-AILab/flash-attention/releases/download/v2.7.4.post1/flash_attn-2.7.4.post1+cu12torch2.6cxx11abiFALSE-cp310-cp310-linux_x86_64.whl
-# Alternative: Install from PyPI
-pip install flash-attn==2.7.0
 ```
-### Usage with transformers
 ```python
-from youtu_hf_parser import YoutuOCRParserHF
-# Initialize the parser
-parser = YoutuOCRParserHF(
-    model_path=model_path,
-    enable_angle_correct=True,  # Set to False to disable angle correction
-    angle_correct_model_path=angle_correct_model_path
 )
-# Parse an image
-parser.parse_file(input_path=image_path, output_dir=output_dir)
 ```
 ## 🎨 Visualization
@@ -138,16 +179,6 @@ We would like to thank [Youtu-LLM](https://github.com/TencentCloudADP/youtu-tip/
 If you find our work useful in your research, please consider citing the following paper:
 ```
-@article{youtu-parsing,
-  title={Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding},
-  author={Tencent Youtu Lab},
-  year={2026},
-  eprint={},
-  archivePrefix={},
-  primaryClass={},
-  url={},
-}
 @article{youtu-vl,
   title={Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision},
   author={Tencent Youtu Lab},
@@ -158,6 +189,16 @@ If you find our work useful in your research, please consider citing the followi
   url={https://arxiv.org/abs/2601.19798},
 }
 @article{youtu-llm,
   title={Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models},
   author={Tencent Youtu Lab},
@@ -167,4 +208,4 @@ If you find our work useful in your research, please consider citing the followi
   primaryClass={cs.CL},
   url={https://arxiv.org/abs/2512.24618},
 }
-```

 ---
+base_model:
+- tencent/Youtu-LLM-2B
 license: other
 license_name: youtu-parsing
 license_link: https://huggingface.co/tencent/Youtu-Parsing/blob/main/LICENSE.txt
+pipeline_tag: image-segmentation
+library_name: transformers
 base_model_relation: finetune
 ---
 <div align="center">
 # <img src="assets/youtu-parsing-logo.png" alt="Youtu-Parsing Logo" height="100px">
 ## 🎯 Introduction
+**Youtu-Parsing** is a specialized document parsing model built upon the open-source Youtu-LLM 2B foundation, as presented in the paper [Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision](https://huggingface.co/papers/2601.19798).
+By extending the capabilities of the base model with a prompt-guided framework and NaViT-style dynamic visual encoder, Youtu-Parsing offers enhanced parsing capabilities for diverse document elements including text, tables, formulas, and charts. The model incorporates an efficient parallel decoding mechanism that significantly accelerates inference, making it practical for real-world document analysis applications. We share Youtu-Parsing with the community to facilitate research and development in document understanding.
 ## ✨ Key Features
 <a id="quickstart"></a>
 ## 🚀 Quick Start
+### Installation
+Ensure your Python environment has the `transformers` library installed:
+```bash
+pip install "transformers>=4.56.0,<=4.57.1" torch accelerate pillow torchvision opencv-python-headless
 ```
+### Usage with Transformers
+You can interact with the model using the `transformers` library:
 ```python
+from transformers import AutoProcessor, AutoModelForCausalLM
+import torch
+model = AutoModelForCausalLM.from_pretrained(
+    "tencent/Youtu-VL-4B-Instruct",
+    attn_implementation="flash_attention_2",
+    torch_dtype="auto",
+    device_map="cuda",
+    trust_remote_code=True
+).eval()
+processor = AutoProcessor.from_pretrained(
+    "tencent/Youtu-VL-4B-Instruct",
+    use_fast=True,
+    trust_remote_code=True
+)
+img_path = "path/to/your/image.png"
+messages = [
+    {
+        "role": "user",
+        "content": [
+            {"type": "image", "image": img_path},
+            {"type": "text",  "text": "Describe the image"},
+        ],
+    }
+]
+inputs = processor.apply_chat_template(
+    messages,
+    tokenize=True,
+    add_generation_prompt=True,
+    return_dict=True,
+    return_tensors="pt"
+).to(model.device)
+generated_ids = model.generate(
+    **inputs,
+    temperature=0.1,
+    top_p=0.001,
+    repetition_penalty=1.05,
+    do_sample=True,
+    max_new_tokens=32768,
 )
+generated_ids_trimmed = [
+    out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
+]
+outputs = processor.batch_decode(
+    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
+)
+print(f"Youtu-VL output:
+{outputs[0]}")
 ```
 ## 🎨 Visualization
 If you find our work useful in your research, please consider citing the following paper:
 ```
 @article{youtu-vl,
   title={Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision},
   author={Tencent Youtu Lab},
   url={https://arxiv.org/abs/2601.19798},
 }
+@article{youtu-parsing,
+  title={Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding},
+  author={Tencent Youtu Lab},
+  year={2026},
+  eprint={},
+  archivePrefix={},
+  primaryClass={},
+  url={},
+}
 @article{youtu-llm,
   title={Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models},
   author={Tencent Youtu Lab},
   primaryClass={cs.CL},
   url={https://arxiv.org/abs/2512.24618},
 }
+```