license: apache-2.0 language: en base_model: zai-org/GLM-OCR tags:
- ocr
- document-understanding
- vision-language
- glm-ocr
- synthos
- z.ai pipeline_tag: image-to-text library_name: transformers
Cizi: Enhanced OCR Fine-Tuned on Z.AI GLM-OCR
Cizi is a specialized fine-tune of Z.AI's GLM-OCR, optimized for [insert your specific niche, e.g., historical documents / medical forms / low-resource languages].
Built on top of GLM-OCR's lightweight 0.9B parameter GLM-V encoder-decoder architecture with Multi-Token Prediction (MTP), Synthos inherits state-of-the-art base OCR capabilities while being further aligned for [your specific use case] through targeted fine-tuning [[2]][[3]].
π Key Features
- Ultra-Lightweight: Only ~0.9B parameters, enabling fast local inference on consumer hardware and edge devices [[1]].
- Multi-Token Prediction: Leverages MTP for faster decoding and improved contextual coherence in dense text regions [[2]].
- [Your Niche] Specialization: Fine-tuned on [X]K domain-specific samples to outperform base GLM-OCR on [specific benchmark/task].
- Structured Markdown Output: Faithfully preserves tables, headers, and reading order in complex layouts.
- End-to-End Architecture: No separate detection/recognition stages β single forward pass from pixels to structured text [[3]].
π Quick Start
from transformers import AutoModelForVision2Seq, AutoProcessor
import torch
from PIL import Image
model_id = "your-org/synthos"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForVision2Seq.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True
)
image = Image.open("document.png")
messages = [
{"role": "user", "content": [
{"type": "image"},
{"type": "text", "text": "Transcribe this document faithfully in Markdown format."}
]}
]
inputs = processor.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
output_ids = model.generate(inputs, max_new_tokens=2048, do_sample=False)
response = processor.batch_decode(output_ids, skip_special_tokens=True)[0]
print(response)
Model tree for TensorVizion/Cizi-OCR
Base model
zai-org/GLM-OCR