Aeddix Alpine OCR (1.2B)

Aeddix Alpine OCR turns a page image into Markdown: text in reading order, tables as HTML, charts as Markdown data tables, formulas as LaTeX, and inline styling (bold, italic, superscript, subscript) where the page has it. It is a fine-tune of MinerU2.5-Pro-2605-1.2B by OpenDataLab and keeps its architecture, prompts and two-step pipeline: the model first finds the page's blocks and their order, then reads each block.

This is a research release. It is free to use under Apache-2.0, but it was tuned against one benchmark and has the limitations listed below.

Results on ParseBench

ParseBench scores document parsers on about 2,000 real enterprise pages in five dimensions. All rows below were measured by us with the same scorer, ParseBench main at commit afb36bd (2026-09-29), on all 2,078 pages, with the page server in inference/ (vLLM 0.28.0, mineru-vl-utils 2.0.5, 150 dpi, one NVIDIA L4).

Model Tables Charts Content faithfulness Semantic formatting Visual grounding Overall
MinerU2.5-Pro-2605-1.2B (base) 77.30 59.13 87.86 49.59 72.83 69.34
Aeddix Alpine OCR 77.12 65.28 88.27 65.74 72.78 73.84
  • Overall is the unweighted mean of the five headline metrics: grits_trm_composite, chart rule_pass_rate, content_faithfulness, semantic_formatting and layout_element_rule_pass_rate.
  • Paired on the same documents, the fine-tune adds +4.38 overall to the base model (document bootstrap 95% interval +3.60 to +5.17), almost all of it in charts (+6.15, standard error 1.27) and semantic formatting (+15.99, standard error 1.45). That comparison is before the two Markdown rules described below, which add about +0.1 more.
  • Tables, content and grounding are level with the base model within one standard error.
  • Each number is one full pass; repeated passes of the same stack matched to four decimals.

Comparing with the ParseBench leaderboard. The published leaderboard row for the base model (72.78) was scored in June 2026 by the maintainers' own harness, which credited formatting and grounding more generously than the public scorer does today. On the current public scorer the same base model scores 69.34, so read our 73.84 against 69.34, not against 72.78. Rows on the leaderboard carry the scorer of their date; ours is afb36bd.

How to run it

The page server in this repository's inference/ folder serves the model behind ParseBench's mineru2605pro page API and is what produced the numbers above.

hf download aeddix-labs/aeddix-alpine-ocr --revision v0.1 --include "inference/*" --local-dir alpine-ocr
pip install ./alpine-ocr/inference          # vLLM 0.28.0, mineru-vl-utils 2.0.5
alpine-ocr-server --model aeddix-labs/aeddix-alpine-ocr --revision v0.1 --port 8765
import base64, requests

page = base64.b64encode(open("page.png", "rb").read()).decode()   # PDFs: render each page, 150 dpi
result = requests.post("http://127.0.0.1:8765/predict", json={"image_base64": page}).json()
print(result["markdown"])          # the page as Markdown
print(result["blocks"][:3])        # typed blocks with boxes normalised to [0, 1]

To reproduce the ParseBench number: REVISION=v0.1 bash inference/scripts/reproduce_parsebench.sh (about 55 minutes on an L4).

The weights also work with mineru-vl-utils directly, as for the base model:

from PIL import Image
from vllm import LLM
from mineru_vl_utils import MinerUClient, MinerULogitsProcessor
from mineru_vl_utils.post_process import json2md

llm = LLM(model="aeddix-labs/aeddix-alpine-ocr", revision="v0.1", logits_processors=[MinerULogitsProcessor])
client = MinerUClient(backend="vllm-engine", vllm_llm=llm, image_analysis=True, enable_table_formula_eq_wrap=True)
print(json2md(client.two_step_extract(Image.open("page.png"))))

That path skips the server's three additions: nested-chart analysis, and the two Markdown rules. Use image_analysis=True, or charts come back empty.

What the server adds

Addition What it does Effect on ParseBench
Nested charts mineru-vl-utils 2.0.5 skips any chart or image lying inside a multi-chart figure (image_block); the server analyses each panel like a standalone chart base model charts 59.13 to 64.47 (+5.34)
Heading levels a numbered title takes its number's depth (2.1 is ###); other titles are ranked by font size on the page; MinerU itself gives every heading the same label formatting +0.16 (standard error 0.11)
Photo descriptions out a block MinerU classes as a natural photo carries a description the model wrote ("Portrait of a man in a suit"), not page text; it stays out of the Markdown, and its block and box stay content +0.42 (standard error 0.11)

None of them adds text or styling the model did not read. --md-rules "" and --nested-charts 0 switch them off.

How it was trained

Three steps, each from the previous one.

  1. Base: opendatalab/MinerU2.5-Pro-2605-1.2B at commit bff20d4.
  2. Supervised fine-tune: 20,000 samples (19,600 train, 400 validation), one epoch, full fine-tune of the language model and the vision-language projector with the vision tower frozen (524.8M trainable parameters). AdamW, learning rate 6e-6 with cosine decay and 10% warm-up, effective batch 16, bf16 with fp32 master weights, MinerU's own prompts and image preprocessing, loss on answer tokens only. The recipe follows the alignment stage of jina-ocr-v1 (arXiv 2609.03181): its learning rate, trainable set and data cleaning (degeneration-loop filter, duplicate images, samples under 64 tokens except charts and formatting).
  3. GRPO: two epochs of 100 steps on 1,600 prompts (783 charts, 611 formatting, 206 tables), LoRA rank 64 on the language model, merged into the weights afterwards. DAPO loss (token-level, clip 0.2 and 0.28, no KL), 8 samples per prompt at temperature 1.0, and a ReMax baseline (advantage = sampled reward minus the greedy answer's reward). Rewards are built from ParseBench's own metric code applied to the training items' known answers, never to ParseBench pages: chart data-point rules, the GriTS table metric and styled-span F-beta, each multiplied by a structure-validity term and a repetition penalty. Each epoch was kept only if no ParseBench dimension fell more than one paired standard error below the starting checkpoint.

Training data

Every training item was checked against all 2,078 ParseBench pages and removed on any overlap: a 256-bit perceptual page hash, the source file name, any shared 13-token sequence with the page text or ground truth, and for Chinese and Japanese text any shared 8-character sequence. That removed 310 of 28,600 supervised candidates and 254 of 77,350 GRPO-pool candidates, most of them real sentence overlaps with public financial filings.

Dataset Licence Used for How
ibm-granite/ChartNet, core_permissive subset only CDLA-Permissive-2.0 4,200 SFT charts, 686 GRPO prompts chart image; its data CSV as the Markdown table
docling-project/SynthChartNet CDLA-Permissive-2.0 2,000 SFT charts, 97 GRPO prompts chart image; its data table
docling-project/SynthTabNet_OTSL CDLA-Permissive-1.0 (IBM/SynthTabNet) 2,500 SFT tables, 64 GRPO prompts table image; its structure label
docling-project/FinTabNet_OTSL CDLA-Permissive-1.0 (IBM FinTabNet) 142 GRPO prompts table image (annual-report tables); reward against its structure label
docling-project/DocLayNet-v1.2 CDLA-Permissive-1.0 5,199 layout pages, 2,700 formatting crops, 901 text crops human layout boxes; bold and italic from the PDF's font names
lightonai/LightOnOCR-mix-0126 Common Crawl and Digital Corpora terms sentences for 2,500 SFT renders and 546 GRPO prompts plain prose sentences only, rendered by us with styling
HuggingFaceFW/fineweb-2, Japanese and Chinese ODC-By 1.0 sentences for 65 GRPO prompts rendered by us with Noto Sans CJK
The base model's own readings Apache-2.0 (the base model) targets of the DocLayNet crops, reading order of the layout pages self-distillation; no other model wrote any label
  • Rendered samples use DejaVu, Liberation and Noto Sans CJK fonts; the images are ours.
  • The page images in DocLayNet and FinTabNet are public documents whose copyright stays with their publishers; the datasets' licences cover their use as training data.
  • Not used, because of their licences: datasets labelled by closed models (olmOCR-mix, GPT-4.1 labels; olmOCR-synthmix, Claude labels), ChartNet's other subsets (Mistral Research Licence), and every non-commercial dataset.

Limitations

  • Charts are the weak spot, and one training attempt made them worse. A larger 65,000-sample fine-tune with 26,000 more ChartNet charts lowered chart reading by 2.25 points (standard error 1.03) against this model's SFT step: synthetic plotting-library charts did not transfer to the report charts ParseBench uses. That checkpoint was not released. What fixed the chart scores was not that data: the server's nested-chart analysis (+5.34 on the base model), the 20,000-sample fine-tune (+4.57 on the base, about half of it because the model stopped wrapping single charts in multi-chart containers) and GRPO with a chart-rule reward (+0.43 more, within noise). 85 of the 568 chart pages (15%) still score zero (the base model: 154, or 111 with the nested-chart analysis).
  • Tables did not improve (77.12 against the base model's 77.30, within noise). Leading parsers score 85 or more on ParseBench tables.
  • Visual grounding did not move (72.78 against 72.83). The model learned part of DocLayNet's layout taxonomy and now uses MinerU's ref_text, aside_text and code labels far less often (on ParseBench's pages 76, 147 and 0 blocks, against the base model's 459, 379 and 13). ParseBench maps them to Text, so its score is unaffected, but other users of the block types may notice.
  • GRPO's own effect is within noise (+0.17 overall, 95% interval −0.20 to +0.52). Most of the gain over the base model comes from the supervised step and the server.
  • Underline, strikethrough and highlight are still rarely written. The gains in formatting are bold, italic, superscript and subscript.
  • Tuned for one benchmark. Data was weighted toward ParseBench's weak dimensions, and checkpoints were selected on ParseBench scores; expect smaller gains on other document types, scans, handwriting and languages other than English. MinerU2.5-Pro's own OmniDocBench results were not re-measured for this model.
  • It inherits the base model's limits: one page at a time, no cross-page table merging, and the page image must be rendered first.

Intended use

Research and evaluation of document parsing, and as a starting point for further fine-tuning. It can be used in products under Apache-2.0, but check its output on your own documents first; we measured it only on ParseBench.

Licence and attribution

  • Weights and code in this repository: Apache-2.0 (see LICENSE and NOTICE).
  • Base model: MinerU2.5-Pro-2605-1.2B by OpenDataLab, Apache-2.0; please cite it:
  • Training data: see the table above; the CDLA-Permissive datasets are by IBM, FineWeb-2 by Hugging Face (ODC-By 1.0).
@misc{wang2026mineru25propushinglimitsdatacentric,
      title={MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale},
      author={Bin, Wang and Tianyao, He and Linke, Ouyang and Fan, Wu and Zhiyuan, Zhao and Tao, Chu and Yuan, Qu and Zhenjiang, Jin and Weijun, Zeng and Ziyang, Miao and Bangrui, Xu and Junbo, Niu and others},
      year={2026},
      eprint={2604.04771},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2604.04771},
}
Downloads last month
-
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aeddix-labs/aeddix-alpine-ocr

Finetuned
(2)
this model

Datasets used to train aeddix-labs/aeddix-alpine-ocr

Paper for aeddix-labs/aeddix-alpine-ocr

Evaluation results

  • llamaindex/ParseBench leaderboard
  • Mean View evaluation results
    source
    Pipeline name: aeddix_alpine_ocr_vllm (the mineru2605pro provider API, served by inference/ in this repository at v0.1). ParseBench main afb36bd (2026-09-29), all 2,078 pages, 0 inference failures, vLLM 0.28.0, mineru-vl-utils 2.0.5, 150 dpi, 32 concurrent pages, one NVIDIA L4. Markdown rules headings,photos and the nested-chart analysis on. Unweighted mean of the five headline metrics. Same pipeline and scorer on the base model MinerU2.5-Pro-2605-1.2B: 69.34.
    73.84 *
  • Table View evaluation results
    source
    aeddix_alpine_ocr_vllm; grits_trm_composite, 503 documents; base on the same pipeline 77.30
    77.12 *
  • Chart View evaluation results
    source
    aeddix_alpine_ocr_vllm; rule_pass_rate, 568 documents; base on the same pipeline without the nested-chart analysis 59.13, with it 64.47
    65.28 *