Papers
arxiv:2609.03595

How Far Can Synthetic Data Take Thai OCR?

Published on Sep 3
· Submitted by
Kunat Pipatanakul
on Sep 14
Authors:

Abstract

A Thai OCR model trained solely on synthetic documents achieves strong real-world performance by isolating key transfer factors like typography, layout, and handwriting glyphs.

We investigate what makes synthetic OCR supervision transfer to real Thai documents and use the resulting insights to build Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages. Synthetic data provide exact labels at scale, but "realism" conflates source domain, page context, typography, spatial structure, and glyph variation. We disentangle these factors with a controlled document-reconstruction pipeline and evaluate each variant under page- and crop-level training on printed and handwritten Thai documents. Non-text context has little consistent effect, whereas typeface diversity, two-dimensional structure, and real handwriting glyphs improve transfer; moreover, source-domain matching depends on training granularity, with in-domain reconstruction approaching real printed supervision under page-level training (1.82% versus 1.31% median character error rate) but underperforming out-of-domain reconstruction under crop-level training (15.59% versus 5.52%). Guided by these findings, we adapt the 0.9B-parameter PaddleOCR-VL-1.6 into Wayu-Paxa-OCR-Zero using 45,723 synthetic pages: relative to its base checkpoint, it reduces median character error rate from 6.64% to 1.24% on printed pages and from 74.87% to 20.55% on handwriting and outperforms Typhoon OCR v1 7B on all five evaluation sets, showing that synthetic-only training can be competitive.

Community

Hi
Kunat Pipatanakul,

Congratulations on this project. I read your paper and found it both practical and very interesting. It is clear that a great deal of effort went into preparing the training data and developing the model.

I’m currently working part-time as an AI Engineer in Bangkok. At my company, we frequently translate Thai PDF documents into English using frontier models. I have also experimented with building OCR solutions, but I have found it challenging to achieve both fast and reliable performance for Thai documents.

After reading the model card for your 0.9B-parameter model, I’m very excited to try it in our workflow. I would love to connect, share any results or progress from testing it, and potentially contribute if there is an opportunity.

Thank you again for making this work available!

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.03595
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.03595 in a dataset README.md to link it from this page.

Spaces citing this paper 2

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.