File size: 9,328 Bytes
f340984
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
---
pipeline_tag: image-text-to-text
license: mit
tags:
  - rocm
  - amd
  - rdna4
  - gfx1201
  - ocr
  - runtime
---

# Unlimited-OCR RDNA 4

Single-GPU [Baidu Unlimited-OCR](https://huggingface.co/baidu/Unlimited-OCR) inference for AMD RDNA 4
(`gfx1200` / `gfx1201`) using native ROCm PyTorch, verified end to end on `gfx1201`.

This project does **not** redistribute or rename Baidu's weights. It provides a reproducible runtime around the original
BF16 checkpoint: an isolated AMD wheel bootstrap, pinned model preparation, audited custom-code patching, single-GPU
selection, local-only model resolution during inference, atomic output, bounded PDF handling, diagnostics, and
validation evidence.

> Status: beta. Verified on 31 July 2026 with an AMD Radeon RX 9070 XT (`gfx1201`). This is not an AMD or Baidu project.

## Verified configuration

| Component | Verified value |
|---|---|
| GPU | AMD Radeon RX 9070 XT, 16 GB |
| GPU allocation | One isolated GPU |
| OS | Ubuntu 24.04.4 |
| Kernel | 6.17.0-40-generic |
| Host ROCm | 7.2.1 |
| PyTorch | 2.9.1 + ROCm 7.2.1 |
| Transformers | 4.57.1 |
| Matplotlib | 3.10.8 |
| Model revision | `07dea832e22aefee32ad281d4b80551282e1c168` |
| Weight precision | Original BF16; no quantization |
| Peak allocated VRAM | 8.25 GiB on the 1600×2000 validation page |

The initial validation recovered headings, reading order, a structured table, euro amounts, a formula, and an exact
checksum. Repeated greedy generation was byte-identical after deterministic tag cleanup. See
[docs/VALIDATION.md](docs/VALIDATION.md) for the evidence and its limits.

## Quick start

The bootstrap downloads checksum-pinned AMD ROCm wheels into your XDG cache, installs every other runtime/build
dependency from `requirements/bootstrap.lock` with hashes required, builds this checkout without build isolation, and
creates a repository-local virtual environment. It does not use `sudo`, alter `/opt/rocm`, or touch the system Python.

```bash
git clone https://huggingface.co/dougvk/Unlimited-OCR-RDNA4
cd Unlimited-OCR-RDNA4

./scripts/bootstrap-rocm.sh
.venv/bin/unlimited-ocr-rdna4 prepare
.venv/bin/unlimited-ocr-rdna4 run --input page.png
```

The model download is approximately 6.78 GB. The verified AMD wheel set is approximately 1.9 GB before installation.

PDFs are rendered and parsed page by page:

```bash
.venv/bin/unlimited-ocr-rdna4 run \
  --input document.pdf \
  --output document.md \
  --max-pages 10 \
  --dpi 200
```

## Commands

### `prepare`

Downloads the pinned original checkpoint, fully hashes the 6.67 GB safetensors file and every behavior-defining model,
configuration, tokenizer, and index file, preserves Baidu's original model source, and applies a narrowly scoped audited
patch.

```bash
unlimited-ocr-rdna4 prepare
unlimited-ocr-rdna4 prepare --dry-run --json
unlimited-ocr-rdna4 prepare --model-dir /data/models/unlimited-ocr
```

Preparation is serialized and idempotent. It uses a unique partial directory, validates the exact manifest and file set,
and refuses unknown, extra, symlinked, or mismatched files instead of executing them.

### `doctor`

Reports the active PyTorch/HIP stack, visible devices, RDNA 4 architecture, BF16 capability, and model readiness.

```bash
unlimited-ocr-rdna4 doctor
unlimited-ocr-rdna4 doctor --device 0 --require-model
unlimited-ocr-rdna4 doctor --json
```

### `run`

Parses one image or PDF and publishes untrusted model output atomically. The content may include Markdown and raw HTML;
do not render it in a privileged origin without sanitization.

```bash
unlimited-ocr-rdna4 run --input scan.png
unlimited-ocr-rdna4 run --input scan.png --output scan.md --device 0
unlimited-ocr-rdna4 run --input report.pdf --start-page 21 --max-pages 10
```

Important flags:

| Flag | Default | Meaning |
|---|---:|---|
| `--device` | `$UNLIMITED_OCR_DEVICE` or `0` | ROCm ordinal or stable ROCr UUID |
| `--mode` | `gundam` | `gundam` for detailed single-page parsing; `base` for lower-detail input |
| `--max-length` | `4096` | Total input + output sequence limit; maximum `32768` |
| `--dpi` | `200` | PDF rendering resolution |
| `--max-pages` | `20` | Per-run PDF safety cap |
| `--max-page-pixels` | `60000000` | Per-page rendered-pixel cap |
| `--max-total-pixels` | `400000000` | Aggregate rendered-pixel cap |
| `--max-page-rendered-mib` | `512` | Per-page rendered temporary-byte cap |
| `--max-rendered-mib` | `2048` | Aggregate rendered temporary-byte cap |
| `--force` | off | Replace an existing output file |
| `--json` | off | Stable machine-readable summary on stdout |
| `--quiet` | off | Suppress progress diagnostics |

OCR content is written to the output file. Progress and warnings go to stderr. Human or JSON summaries go to stdout.

## GPU selection

PyTorch uses the CUDA-compatible API name on ROCm. `--device` sets `ROCR_VISIBLE_DEVICES` before PyTorch is imported, so
the process sees one logical `cuda:0` backed by the selected AMD card.

For ordinary one-GPU systems:

```bash
unlimited-ocr-rdna4 run --input page.png --device 0
```

For multi-GPU systems, a stable ROCr UUID avoids dependence on enumeration order:

```bash
unlimited-ocr-rdna4 run --input page.png --device GPU-0123456789abcdef
```

Run `unlimited-ocr-rdna4 doctor` before choosing a device. Do not assume another application's GPU numbering matches
ROCm's ordinal numbering.

## What the RDNA 4 adaptation changes

- Pins the Baidu checkpoint and verifies the complete expected local model tree before custom code loads.
- Uses AMD's production ROCm 7.2.1 PyTorch, torchvision, and Triton wheels for Python 3.12.
- Isolates one `gfx1200`/`gfx1201` GPU before importing PyTorch.
- Replaces unsafe `eval()` calls in optional model-output geometry parsing with `ast.literal_eval()`.
- Makes one internal mask transfer follow the active tensor device.
- Supplies the missing all-ones attention mask and pad token for single-sequence generation.
- Loads only the prepared local model during inference and enables Hugging Face offline mode.
- Avoids vLLM, SGLang, quantization, tensor parallelism, and unverified custom serving kernels.
- Removes only complete, exact layout sentinel pairs without rewriting recognized Unicode or TeX.
- Detects obvious terminal repetition and warns without silently rewriting recognition content.

The model weights and mathematical operators are unchanged. PyTorch's ROCm backend provides the RDNA 4 kernels.

## Configuration

Precedence is flags, then environment variables, then XDG defaults.

| Environment variable | Purpose |
|---|---|
| `UNLIMITED_OCR_DEVICE` | Default ROCm ordinal or UUID |
| `UNLIMITED_OCR_MODEL_DIR` | Prepared model directory |
| `XDG_DATA_HOME` | Default model storage root |
| `XDG_CACHE_HOME` | Wheel, Hugging Face, and temporary-work cache root |
| `NO_COLOR` | Accepted implicitly; the CLI currently emits no color |

This project contains no analytics or telemetry code.

## Safety and limitations

- The default 4096 sequence limit includes visual-prefill tokens. Dense pages may require a larger value.
- PDF mode uses permissively licensed PDFium bindings and renders one bounded page at a time. A failed run leaves no
  published partial document; retry with `--start-page` to resume manually.
- Generative OCR can omit or hallucinate content. Verify consequential documents against the source.
- Rotated text and repetitive pages are known upstream weak spots.
- Only the configuration above has completed the full repository acceptance test. Newer ROCm/PyTorch versions may work,
  but `doctor` reports them as unverified until measured.
- The hardware guard accepts `gfx1200`, but reports `hardware_verified=false`; only `gfx1201` has completed this GPU gate.
- The bootstrap currently supports Linux x86_64, Python 3.12, and host ROCm 7.2.1.

See [SECURITY.md](SECURITY.md) before processing untrusted documents.

## Development

CPU-only tests do not install PyTorch. The commands below are convenient for development; the release bootstrap is the
hash-locked installation path.

```bash
python3.12 -m venv .venv-dev
.venv-dev/bin/pip install -e '.[dev]'
.venv-dev/bin/ruff check .
.venv-dev/bin/ruff format --check .
.venv-dev/bin/pytest
.venv-dev/bin/python -m build
.venv-dev/bin/twine check dist/*
```

The GPU smoke gate is intentionally separate. It uses tracked, hash-checked fixtures; parses structured JSON; runs the
image twice in fresh processes; and checks image/PDF output hashes, structure, stack identity, revision, VRAM, and process
release. On the validated host, run it in a root-created private network namespace:

```bash
VALIDATION_NETWORK_ISOLATED=1 \
UNLIMITED_OCR_DEVICE=GPU-0123456789abcdef \
AMD_SMI_GPU=3 \
./scripts/validate-smoke.sh
```

## Credits and license

- Model and model code: [Baidu Unlimited-OCR](https://github.com/baidu/Unlimited-OCR), MIT licensed.
- ROCm evaluation and batching research: [AIwork4me/Unlimited-OCR-ROCm](https://github.com/AIwork4me/Unlimited-OCR-ROCm).
- AMD wheel source and compatibility guidance: [ROCm documentation](https://rocm.docs.amd.com/).
- PDF rendering: [pypdfium2](https://pypi.org/project/pypdfium2/) and PDFium, under permissive licenses.

The runtime is MIT licensed. Baidu's original copyright and license are preserved; see
[THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md).