Instructions to use aeddix-labs/aeddix-alpine-ocr with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use aeddix-labs/aeddix-alpine-ocr with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="aeddix-labs/aeddix-alpine-ocr") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("aeddix-labs/aeddix-alpine-ocr") model = AutoModelForMultimodalLM.from_pretrained("aeddix-labs/aeddix-alpine-ocr", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use aeddix-labs/aeddix-alpine-ocr with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "aeddix-labs/aeddix-alpine-ocr" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aeddix-labs/aeddix-alpine-ocr", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/aeddix-labs/aeddix-alpine-ocr
- SGLang
How to use aeddix-labs/aeddix-alpine-ocr with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "aeddix-labs/aeddix-alpine-ocr" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aeddix-labs/aeddix-alpine-ocr", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "aeddix-labs/aeddix-alpine-ocr" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aeddix-labs/aeddix-alpine-ocr", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use aeddix-labs/aeddix-alpine-ocr with Docker Model Runner:
docker model run hf.co/aeddix-labs/aeddix-alpine-ocr
Aeddix Alpine OCR (1.2B)
Aeddix Alpine OCR turns a page image into Markdown: text in reading order, tables as HTML, charts as Markdown data tables, formulas as LaTeX, and inline styling (bold, italic, superscript, subscript) where the page has it. It is a fine-tune of MinerU2.5-Pro-2605-1.2B by OpenDataLab and keeps its architecture, prompts and two-step pipeline: the model first finds the page's blocks and their order, then reads each block.
This is a research release. It is free to use under Apache-2.0, but it was tuned against one benchmark and has the limitations listed below.
Results on ParseBench
ParseBench scores document parsers on about 2,000 real enterprise pages in five dimensions.
All rows below were measured by us with the same scorer, ParseBench main at commit afb36bd (2026-09-29), on all 2,078 pages, with the page server in inference/ (vLLM 0.28.0, mineru-vl-utils 2.0.5, 150 dpi, one NVIDIA L4).
| Model | Tables | Charts | Content faithfulness | Semantic formatting | Visual grounding | Overall |
|---|---|---|---|---|---|---|
| MinerU2.5-Pro-2605-1.2B (base) | 77.30 | 59.13 | 87.86 | 49.59 | 72.83 | 69.34 |
| Aeddix Alpine OCR | 77.12 | 65.28 | 88.27 | 65.74 | 72.78 | 73.84 |
- Overall is the unweighted mean of the five headline metrics:
grits_trm_composite, chartrule_pass_rate,content_faithfulness,semantic_formattingandlayout_element_rule_pass_rate. - Paired on the same documents, the fine-tune adds +4.38 overall to the base model (document bootstrap 95% interval +3.60 to +5.17), almost all of it in charts (+6.15, standard error 1.27) and semantic formatting (+15.99, standard error 1.45). That comparison is before the two Markdown rules described below, which add about +0.1 more.
- Tables, content and grounding are level with the base model within one standard error.
- Each number is one full pass; repeated passes of the same stack matched to four decimals.
Comparing with the ParseBench leaderboard.
The published leaderboard row for the base model (72.78) was scored in June 2026 by the maintainers' own harness, which credited formatting and grounding more generously than the public scorer does today.
On the current public scorer the same base model scores 69.34, so read our 73.84 against 69.34, not against 72.78.
Rows on the leaderboard carry the scorer of their date; ours is afb36bd.
How to run it
The page server in this repository's inference/ folder serves the model behind ParseBench's mineru2605pro page API and is what produced the numbers above.
hf download aeddix-labs/aeddix-alpine-ocr --revision v0.1 --include "inference/*" --local-dir alpine-ocr
pip install ./alpine-ocr/inference # vLLM 0.28.0, mineru-vl-utils 2.0.5
alpine-ocr-server --model aeddix-labs/aeddix-alpine-ocr --revision v0.1 --port 8765
import base64, requests
page = base64.b64encode(open("page.png", "rb").read()).decode() # PDFs: render each page, 150 dpi
result = requests.post("http://127.0.0.1:8765/predict", json={"image_base64": page}).json()
print(result["markdown"]) # the page as Markdown
print(result["blocks"][:3]) # typed blocks with boxes normalised to [0, 1]
To reproduce the ParseBench number: REVISION=v0.1 bash inference/scripts/reproduce_parsebench.sh (about 55 minutes on an L4).
The weights also work with mineru-vl-utils directly, as for the base model:
from PIL import Image
from vllm import LLM
from mineru_vl_utils import MinerUClient, MinerULogitsProcessor
from mineru_vl_utils.post_process import json2md
llm = LLM(model="aeddix-labs/aeddix-alpine-ocr", revision="v0.1", logits_processors=[MinerULogitsProcessor])
client = MinerUClient(backend="vllm-engine", vllm_llm=llm, image_analysis=True, enable_table_formula_eq_wrap=True)
print(json2md(client.two_step_extract(Image.open("page.png"))))
That path skips the server's three additions: nested-chart analysis, and the two Markdown rules.
Use image_analysis=True, or charts come back empty.
What the server adds
| Addition | What it does | Effect on ParseBench |
|---|---|---|
| Nested charts | mineru-vl-utils 2.0.5 skips any chart or image lying inside a multi-chart figure (image_block); the server analyses each panel like a standalone chart |
base model charts 59.13 to 64.47 (+5.34) |
| Heading levels | a numbered title takes its number's depth (2.1 is ###); other titles are ranked by font size on the page; MinerU itself gives every heading the same label |
formatting +0.16 (standard error 0.11) |
| Photo descriptions out | a block MinerU classes as a natural photo carries a description the model wrote ("Portrait of a man in a suit"), not page text; it stays out of the Markdown, and its block and box stay | content +0.42 (standard error 0.11) |
None of them adds text or styling the model did not read.
--md-rules "" and --nested-charts 0 switch them off.
How it was trained
Three steps, each from the previous one.
- Base:
opendatalab/MinerU2.5-Pro-2605-1.2Bat commitbff20d4. - Supervised fine-tune: 20,000 samples (19,600 train, 400 validation), one epoch, full fine-tune of the language model and the vision-language projector with the vision tower frozen (524.8M trainable parameters). AdamW, learning rate 6e-6 with cosine decay and 10% warm-up, effective batch 16, bf16 with fp32 master weights, MinerU's own prompts and image preprocessing, loss on answer tokens only. The recipe follows the alignment stage of jina-ocr-v1 (arXiv 2609.03181): its learning rate, trainable set and data cleaning (degeneration-loop filter, duplicate images, samples under 64 tokens except charts and formatting).
- GRPO: two epochs of 100 steps on 1,600 prompts (783 charts, 611 formatting, 206 tables), LoRA rank 64 on the language model, merged into the weights afterwards. DAPO loss (token-level, clip 0.2 and 0.28, no KL), 8 samples per prompt at temperature 1.0, and a ReMax baseline (advantage = sampled reward minus the greedy answer's reward). Rewards are built from ParseBench's own metric code applied to the training items' known answers, never to ParseBench pages: chart data-point rules, the GriTS table metric and styled-span F-beta, each multiplied by a structure-validity term and a repetition penalty. Each epoch was kept only if no ParseBench dimension fell more than one paired standard error below the starting checkpoint.
Training data
Every training item was checked against all 2,078 ParseBench pages and removed on any overlap: a 256-bit perceptual page hash, the source file name, any shared 13-token sequence with the page text or ground truth, and for Chinese and Japanese text any shared 8-character sequence. That removed 310 of 28,600 supervised candidates and 254 of 77,350 GRPO-pool candidates, most of them real sentence overlaps with public financial filings.
| Dataset | Licence | Used for | How |
|---|---|---|---|
ibm-granite/ChartNet, core_permissive subset only |
CDLA-Permissive-2.0 | 4,200 SFT charts, 686 GRPO prompts | chart image; its data CSV as the Markdown table |
| docling-project/SynthChartNet | CDLA-Permissive-2.0 | 2,000 SFT charts, 97 GRPO prompts | chart image; its data table |
| docling-project/SynthTabNet_OTSL | CDLA-Permissive-1.0 (IBM/SynthTabNet) | 2,500 SFT tables, 64 GRPO prompts | table image; its structure label |
| docling-project/FinTabNet_OTSL | CDLA-Permissive-1.0 (IBM FinTabNet) | 142 GRPO prompts | table image (annual-report tables); reward against its structure label |
| docling-project/DocLayNet-v1.2 | CDLA-Permissive-1.0 | 5,199 layout pages, 2,700 formatting crops, 901 text crops | human layout boxes; bold and italic from the PDF's font names |
| lightonai/LightOnOCR-mix-0126 | Common Crawl and Digital Corpora terms | sentences for 2,500 SFT renders and 546 GRPO prompts | plain prose sentences only, rendered by us with styling |
| HuggingFaceFW/fineweb-2, Japanese and Chinese | ODC-By 1.0 | sentences for 65 GRPO prompts | rendered by us with Noto Sans CJK |
| The base model's own readings | Apache-2.0 (the base model) | targets of the DocLayNet crops, reading order of the layout pages | self-distillation; no other model wrote any label |
- Rendered samples use DejaVu, Liberation and Noto Sans CJK fonts; the images are ours.
- The page images in DocLayNet and FinTabNet are public documents whose copyright stays with their publishers; the datasets' licences cover their use as training data.
- Not used, because of their licences: datasets labelled by closed models (olmOCR-mix, GPT-4.1 labels; olmOCR-synthmix, Claude labels), ChartNet's other subsets (Mistral Research Licence), and every non-commercial dataset.
Limitations
- Charts are the weak spot, and one training attempt made them worse. A larger 65,000-sample fine-tune with 26,000 more ChartNet charts lowered chart reading by 2.25 points (standard error 1.03) against this model's SFT step: synthetic plotting-library charts did not transfer to the report charts ParseBench uses. That checkpoint was not released. What fixed the chart scores was not that data: the server's nested-chart analysis (+5.34 on the base model), the 20,000-sample fine-tune (+4.57 on the base, about half of it because the model stopped wrapping single charts in multi-chart containers) and GRPO with a chart-rule reward (+0.43 more, within noise). 85 of the 568 chart pages (15%) still score zero (the base model: 154, or 111 with the nested-chart analysis).
- Tables did not improve (77.12 against the base model's 77.30, within noise). Leading parsers score 85 or more on ParseBench tables.
- Visual grounding did not move (72.78 against 72.83).
The model learned part of DocLayNet's layout taxonomy and now uses MinerU's
ref_text,aside_textandcodelabels far less often (on ParseBench's pages 76, 147 and 0 blocks, against the base model's 459, 379 and 13). ParseBench maps them to Text, so its score is unaffected, but other users of the block types may notice. - GRPO's own effect is within noise (+0.17 overall, 95% interval −0.20 to +0.52). Most of the gain over the base model comes from the supervised step and the server.
- Underline, strikethrough and highlight are still rarely written. The gains in formatting are bold, italic, superscript and subscript.
- Tuned for one benchmark. Data was weighted toward ParseBench's weak dimensions, and checkpoints were selected on ParseBench scores; expect smaller gains on other document types, scans, handwriting and languages other than English. MinerU2.5-Pro's own OmniDocBench results were not re-measured for this model.
- It inherits the base model's limits: one page at a time, no cross-page table merging, and the page image must be rendered first.
Intended use
Research and evaluation of document parsing, and as a starting point for further fine-tuning. It can be used in products under Apache-2.0, but check its output on your own documents first; we measured it only on ParseBench.
Licence and attribution
- Weights and code in this repository: Apache-2.0 (see
LICENSEandNOTICE). - Base model: MinerU2.5-Pro-2605-1.2B by OpenDataLab, Apache-2.0; please cite it:
- Training data: see the table above; the CDLA-Permissive datasets are by IBM, FineWeb-2 by Hugging Face (ODC-By 1.0).
@misc{wang2026mineru25propushinglimitsdatacentric,
title={MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale},
author={Bin, Wang and Tianyao, He and Linke, Ouyang and Fan, Wu and Zhiyuan, Zhao and Tao, Chu and Yuan, Qu and Zhenjiang, Jin and Weijun, Zeng and Ziyang, Miao and Bangrui, Xu and Junbo, Niu and others},
year={2026},
eprint={2604.04771},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.04771},
}
- Downloads last month
- -
Model tree for aeddix-labs/aeddix-alpine-ocr
Base model
opendatalab/MinerU2.5-Pro-2605-1.2BDatasets used to train aeddix-labs/aeddix-alpine-ocr
ibm-granite/ChartNet
docling-project/DocLayNet-v1.2
Paper for aeddix-labs/aeddix-alpine-ocr
Evaluation results
- llamaindex/ParseBench leaderboard
- Mean View evaluation resultssource
Pipeline name: aeddix_alpine_ocr_vllm (the mineru2605pro provider API, served by inference/ in this repository at v0.1). ParseBench main afb36bd (2026-09-29), all 2,078 pages, 0 inference failures, vLLM 0.28.0, mineru-vl-utils 2.0.5, 150 dpi, 32 concurrent pages, one NVIDIA L4. Markdown rules headings,photos and the nested-chart analysis on. Unweighted mean of the five headline metrics. Same pipeline and scorer on the base model MinerU2.5-Pro-2605-1.2B: 69.34.73.84 * - Table View evaluation resultssource
aeddix_alpine_ocr_vllm; grits_trm_composite, 503 documents; base on the same pipeline 77.3077.12 * - Chart View evaluation resultssource
aeddix_alpine_ocr_vllm; rule_pass_rate, 568 documents; base on the same pipeline without the nested-chart analysis 59.13, with it 64.4765.28 *