Spaces:
Running on Zero
Running on Zero
electblake commited on
Commit ·
508a5ff
1
Parent(s): fed3258
not references atm
Browse files- .agents/references/gradio/2026-07-20/README.md +0 -5
- .agents/references/gradio/2026-07-20/REFERENCE_INDEX.md +0 -7
- .agents/references/gradio/2026-07-20/REFERENCE_SUMMARY.md +0 -8
- .agents/references/gradio/2026-07-20/SOURCE_URLS.md +0 -3
- .agents/references/nuextract3/2026-07-20/README.md +0 -5
- .agents/references/nuextract3/2026-07-20/REFERENCE_INDEX.md +0 -22
- .agents/references/nuextract3/2026-07-20/REFERENCE_SUMMARY.md +0 -20
- .agents/references/nuextract3/2026-07-20/SOURCE_URLS.md +0 -10
- .agents/references/nuextract3/2026-07-20/fixtures/user-drop-001.txt +0 -942
- .agents/references/nuextract3/2026-07-20/huggingface.co/numind/NuExtract3.md +0 -1010
- .agents/references/nuextract3/2026-07-20/huggingface.co/numind/NuExtract3/blob/main/README.md +0 -1001
.agents/references/gradio/2026-07-20/README.md
DELETED
|
@@ -1,5 +0,0 @@
|
|
| 1 |
-
# gradio
|
| 2 |
-
|
| 3 |
-
- Topic: `gradio`
|
| 4 |
-
- Capture Date: `2026-07-20`
|
| 5 |
-
- Scope: TODO
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
.agents/references/gradio/2026-07-20/REFERENCE_INDEX.md
DELETED
|
@@ -1,7 +0,0 @@
|
|
| 1 |
-
# Reference Index
|
| 2 |
-
|
| 3 |
-
- `README.md`: bundle scope
|
| 4 |
-
- `SOURCE_URLS.md`: verified sources
|
| 5 |
-
- `REFERENCE_SUMMARY.md`: findings and version notes
|
| 6 |
-
- `fixtures/`: user-provided artifacts
|
| 7 |
-
- `<domain>/`: mirrored markdown references by domain
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
.agents/references/gradio/2026-07-20/REFERENCE_SUMMARY.md
DELETED
|
@@ -1,8 +0,0 @@
|
|
| 1 |
-
# Reference Summary
|
| 2 |
-
|
| 3 |
-
- Reference Date: `2026-07-20`
|
| 4 |
-
- Official Version: TODO
|
| 5 |
-
- Key Findings:
|
| 6 |
-
- TODO
|
| 7 |
-
- Compatibility Notes:
|
| 8 |
-
- TODO
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
.agents/references/gradio/2026-07-20/SOURCE_URLS.md
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
# Source URLs
|
| 2 |
-
|
| 3 |
-
- TODO: add every official URL used in this bundle
|
|
|
|
|
|
|
|
|
|
|
|
.agents/references/nuextract3/2026-07-20/README.md
DELETED
|
@@ -1,5 +0,0 @@
|
|
| 1 |
-
# nuextract3
|
| 2 |
-
|
| 3 |
-
- Topic: `nuextract3`
|
| 4 |
-
- Capture Date: `2026-07-20`
|
| 5 |
-
- Scope: TODO
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
.agents/references/nuextract3/2026-07-20/REFERENCE_INDEX.md
DELETED
|
@@ -1,22 +0,0 @@
|
|
| 1 |
-
# Reference Index
|
| 2 |
-
|
| 3 |
-
- Bundle Path: `S:\Spaces\Data-Extraction\NuMarkApp\.agents\references\nuextract3\2026-07-20`
|
| 4 |
-
- Mirrored References: `1`
|
| 5 |
-
- Fixtures: `1`
|
| 6 |
-
|
| 7 |
-
- `README.md`: bundle scope
|
| 8 |
-
- `SOURCE_URLS.md`: verified sources
|
| 9 |
-
- `REFERENCE_SUMMARY.md`: findings and version notes
|
| 10 |
-
- `fixtures/`: user-provided artifacts
|
| 11 |
-
- `<domain>/`: mirrored markdown references by domain
|
| 12 |
-
|
| 13 |
-
## References
|
| 14 |
-
|
| 15 |
-
### huggingface.co
|
| 16 |
-
|
| 17 |
-
- `huggingface.co/numind/NuExtract3.md`: `https://huggingface.co/numind/NuExtract3`
|
| 18 |
-
|
| 19 |
-
## Fixtures
|
| 20 |
-
|
| 21 |
-
- `fixtures/user-drop-001.txt`
|
| 22 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
.agents/references/nuextract3/2026-07-20/REFERENCE_SUMMARY.md
DELETED
|
@@ -1,20 +0,0 @@
|
|
| 1 |
-
# NuExtract3 Reference Summary
|
| 2 |
-
|
| 3 |
-
- Reference date: 2026-07-20
|
| 4 |
-
- Official model: `numind/NuExtract3`
|
| 5 |
-
- Library: Transformers
|
| 6 |
-
- Base model: Qwen/Qwen3.5-4B
|
| 7 |
-
- License: Apache-2.0
|
| 8 |
-
|
| 9 |
-
NuExtract3 is a 4B vision-language reasoning model for structured extraction and document-to-Markdown conversion. It accepts text, images, or both.
|
| 10 |
-
|
| 11 |
-
For direct Transformers inference, the official README loads `AutoProcessor` and `AutoModelForImageTextToText`, then passes `template`, `mode`, and `enable_thinking` through `processor.apply_chat_template`. Non-thinking structured extraction uses `enable_thinking=False`; the documented starting temperature is 0.2.
|
| 12 |
-
|
| 13 |
-
Runtime-specific note: this project uses Transformers 5.14.1, Python 3.12, BF16, automatic device placement, and FlashAttention 2.
|
| 14 |
-
|
| 15 |
-
- Reference Date: `2026-07-20`
|
| 16 |
-
- Official Version: TODO
|
| 17 |
-
- Key Findings:
|
| 18 |
-
- TODO
|
| 19 |
-
- Compatibility Notes:
|
| 20 |
-
- TODO
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
.agents/references/nuextract3/2026-07-20/SOURCE_URLS.md
DELETED
|
@@ -1,10 +0,0 @@
|
|
| 1 |
-
# Source URLs
|
| 2 |
-
|
| 3 |
-
- Bundle: `2026-07-20`
|
| 4 |
-
- Generated: `2026-07-20T17:29:26.022745+00:00`
|
| 5 |
-
- Mirrored References: `1`
|
| 6 |
-
|
| 7 |
-
## huggingface.co
|
| 8 |
-
|
| 9 |
-
- `https://huggingface.co/numind/NuExtract3` -> `huggingface.co/numind/NuExtract3.md`
|
| 10 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
.agents/references/nuextract3/2026-07-20/fixtures/user-drop-001.txt
DELETED
|
@@ -1,942 +0,0 @@
|
|
| 1 |
-
---
|
| 2 |
-
license: apache-2.0
|
| 3 |
-
license_link: https://huggingface.co/numind/NuExtract3/blob/main/LICENSE
|
| 4 |
-
library_name: transformers
|
| 5 |
-
pipeline_tag: image-to-text
|
| 6 |
-
tags:
|
| 7 |
-
- image-text-to-text
|
| 8 |
-
- transformers
|
| 9 |
-
- safetensors
|
| 10 |
-
- qwen3_5
|
| 11 |
-
- vision-language
|
| 12 |
-
- vlm
|
| 13 |
-
- document-understanding
|
| 14 |
-
- structured-extraction
|
| 15 |
-
- information-extraction
|
| 16 |
-
- ocr
|
| 17 |
-
- document-to-markdown
|
| 18 |
-
- markdown
|
| 19 |
-
- rag
|
| 20 |
-
- reasoning
|
| 21 |
-
- multilingual
|
| 22 |
-
- conversational
|
| 23 |
-
base_model:
|
| 24 |
-
- Qwen/Qwen3.5-4B
|
| 25 |
-
model_name: NuExtract3
|
| 26 |
-
---
|
| 27 |
-
|
| 28 |
-
<p align="center">
|
| 29 |
-
<a href="https://nuextract.ai/">
|
| 30 |
-
<img src="header.svg" width="900px"/>
|
| 31 |
-
</a>
|
| 32 |
-
</p>
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
<p align="center">
|
| 36 |
-
🖥️ <a href="https://nuextract.ai/">API / Platform</a> |
|
| 37 |
-
📑 <a href="https://numind.ai/blog">Blog</a> |
|
| 38 |
-
🗣️ <a href="https://discord.gg/3tsEtJNCDe">Discord</a> |
|
| 39 |
-
🛠️ <a href="https://github.com/numindai/nuextract">GitHub</a>
|
| 40 |
-
</p>
|
| 41 |
-
|
| 42 |
-
**NuExtract3** is a unified **4B** vision-language reasoning model for document understanding.
|
| 43 |
-
|
| 44 |
-
It combines strong **structured information extraction** with high-quality **image-to-Markdown** conversion, making it suitable for extraction pipelines, OCR, and RAG preprocessing for all types of documents such as scans, receipts, forms, invoices, contracts or tables.
|
| 45 |
-
|
| 46 |
-
Try it out in [the 🤗 space!](https://huggingface.co/spaces/numind/NuExtract-3-4B)
|
| 47 |
-
|
| 48 |
-
## Overview
|
| 49 |
-
|
| 50 |
-
- **Structured extraction**: input (text/images) + JSON template + instructions --> JSON output
|
| 51 |
-
- **Markdown conversion**: input (text/images) --> Markdown
|
| 52 |
-
- **Multimodal inputs**: text, images, or text + images.
|
| 53 |
-
- **Multilingual** documents.
|
| 54 |
-
- **Reasoning** and non-reasoning inference modes.
|
| 55 |
-
- **Template generation** for structured extraction from natural language or input document.
|
| 56 |
-
|
| 57 |
-
# Benchmark results
|
| 58 |
-
|
| 59 |
-
## Structured Extraction
|
| 60 |
-
|
| 61 |
-
We benchmarked NuExtract on NuMind's internal structured benchmark, measuring model's performances on ~600 documents of diverse types including invoices, movie posters or floor plans. These documents and their ground-truth cover diverse use-cases testing model visual understanding, OCR, reasoning and ability to handle long input and output contexts.
|
| 62 |
-
We plan to open-source this benchmark in the coming weeks, along with a extensive leaderboard including most popular open-weight and closed-sourced APIs and a Python library allowing to easily measure model performances on structured extraction.
|
| 63 |
-
|
| 64 |
-
<img src="st.svg" width="1000"/>
|
| 65 |
-
|
| 66 |
-
To measure a pair of predicted and ground-truth JSONs, we represent both as trees which we align based on node names, compute metric scores for aligned leaves and report the average of these scores. `string` and `verbatim-string` leaves are evaluated with indel distance (i.e. Levenshtein without replacement), while all others are evaluated with exact-match.
|
| 67 |
-
Models were evaluated using vllm, with a temperature of 0.25 and a maximum of 65000 output token (for both thinking and answer), which largely exceeds 22000 which is the number of tokens of the largest ground truth output.
|
| 68 |
-
|
| 69 |
-
<figure>
|
| 70 |
-
|
| 71 |
-
|Model name |Average score|Num. failed⁽¹⁾|Avg. num tokens thinking|Avg. num tokens answer|
|
| 72 |
-
|--------------------|-------------|-----------|------------------------|----------------------|
|
| 73 |
-
|NuExtract3.4_4B-RL |**0.651 ± 0.019**|27 |2036 |1856 |
|
| 74 |
-
|gemma-4-E4B-it |0.538 ± 0.023|31 |3005 |1287 |
|
| 75 |
-
|Qwen3.5-9B |0.479 ± 0.030|170 |22409 |1257 |
|
| 76 |
-
|Qwen3.5-4B |0.417 ± 0.031|229 |27177 |1201 |
|
| 77 |
-
|GLM-4.6V-Flash |0.435 ± 0.026|153 |2989 |1357 |
|
| 78 |
-
|Nemotron-3-Nano-Omni|0.387 ± 0.028|204 |25827 |522 |
|
| 79 |
-
|Ministral-3-3B |0.240 ± 0.022|344 |27586 |362 |
|
| 80 |
-
|
| 81 |
-
<figcaption>
|
| 82 |
-
<small>
|
| 83 |
-
(1) number of model outputs that were not JSON deserializable, either directly or by removing leading and trailing backticks.<br>
|
| 84 |
-
95% confidence intervals computed using a nonparametric bootstrap over scores distributions.
|
| 85 |
-
</small>
|
| 86 |
-
</figcaption>
|
| 87 |
-
</figure>
|
| 88 |
-
|
| 89 |
-
The benchmark include samples containing multiple images resulting in large input context, and some with ground-truth containing large numbers of items to extract resulting in large outputs. We found that the reasoning of small models significantly negatively impact their performances. The reason is that many models ended up falling in repetition loops, hitting the output tokens limit and resulting in failed requests.
|
| 90 |
-
|
| 91 |
-
## Document to Markdown
|
| 92 |
-
|
| 93 |
-
NuExtract can also convert document images into clean Markdown. Output will be Markdown for text (headers etc), HTML for tables, LaTeX for math and ```<figure data-type="image" data-id="img_n"><img src="/NM-dev/model_card-A/resolve/main/img_n.png" alt="Detail description of the images"/> ```
|
| 94 |
-
|
| 95 |
-
Modern, format-agnostic benchmarks for complex document understanding are limited, so we explored a new evaluation approach.
|
| 96 |
-
We selected 100 documents with challenging layouts and tables, asked each model to convert them into a structured representation, then used Gemini 3 Flash to compare model outputs against the source document and choose the most accurate result.
|
| 97 |
-
The rankings aligned with human votes, suggesting this is a promising method for evaluating document-to-Markdown capabilities. More details will be shared in an upcoming technical report.
|
| 98 |
-
Here are some results:
|
| 99 |
-
|
| 100 |
-
<img src="ocr_preferences.svg" width="1000"/>
|
| 101 |
-
|
| 102 |
-
### Using "Markdown-to-structured"
|
| 103 |
-
|
| 104 |
-
To add other evaluate references, we used our structured extraction benchmark to evaluate models in a two-step fashion: convert the benchmark inputs to Markdown, then use Qwen3.6 27B to perform the structured extraction task on them. Intuitively, it allows to evaluate how models achieve to keep the input document content and layout: good models will allow the "structured extractor" model to perform better scores.
|
| 105 |
-
|
| 106 |
-
<img src="md2st.svg" width="1000"/>
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
# Using NuExtract
|
| 110 |
-
|
| 111 |
-
## Structured extraction
|
| 112 |
-
|
| 113 |
-
Structured extraction takes as inputs:
|
| 114 |
-
|
| 115 |
-
1. An input document, which can be text, image, or both;
|
| 116 |
-
2. A JSON template describing the information to extract;
|
| 117 |
-
3. (Optional) Instructions, allowing to specify expected output formats or values, to provide with the `instructions` chat template kwarg;
|
| 118 |
-
4. (Optional) In-Context Learning (ICL) examples.
|
| 119 |
-
|
| 120 |
-
### Input JSON template
|
| 121 |
-
|
| 122 |
-
NuExtract uses a input JSON template whose structure is identical to the output JSON. Its leaf values are specify the **types** of the output JSON leaves. For examples:
|
| 123 |
-
|
| 124 |
-
```json
|
| 125 |
-
{
|
| 126 |
-
"invoice_number": "verbatim-string",
|
| 127 |
-
"invoice_date": "date",
|
| 128 |
-
"total_amount": "number",
|
| 129 |
-
"currency": "currency",
|
| 130 |
-
"line_items": [
|
| 131 |
-
{
|
| 132 |
-
"description": "verbatim-string",
|
| 133 |
-
"item_type": ["electronics", "clothing", "vehicle", "furniture", "other"],
|
| 134 |
-
"quantity": "integer",
|
| 135 |
-
"unit_price": "number",
|
| 136 |
-
"total": "number"
|
| 137 |
-
}
|
| 138 |
-
]
|
| 139 |
-
}
|
| 140 |
-
```
|
| 141 |
-
|
| 142 |
-
Supported template types include:
|
| 143 |
-
|
| 144 |
-
- `verbatim-string`: extract text exactly as it appears in the document;
|
| 145 |
-
- `string`: generic string field, allowing abstraction or light paraphrasing;
|
| 146 |
-
- `integer`: whole number;
|
| 147 |
-
- `number`: integer or decimal number;
|
| 148 |
-
- `date-time`: ISO-8601 date, time or date-time;
|
| 149 |
-
- Other specific types such as `data`, `time`, `country`, `currency`, `email` and so on.
|
| 150 |
-
[**For more details, read the complete types specifications and examples**](TYPES.md)
|
| 151 |
-
|
| 152 |
-
Template constructors:
|
| 153 |
-
|
| 154 |
-
- Arrays, for example `["string"]`;
|
| 155 |
-
- Enums, for example `["yes", "no", "maybe"]`;
|
| 156 |
-
- Multi-enums (multiple possible values), for example `[["A", "B", "C"]]`.
|
| 157 |
-
|
| 158 |
-
If the model does not find relevant information for a field, it returns `null` or `[]`.
|
| 159 |
-
|
| 160 |
-
### Converting JSON schema / Pydantic models to NuExtract template
|
| 161 |
-
|
| 162 |
-
Our Python SDK (`pip install numind`) offers a method to convert JSON schemas to NuExtract templates:
|
| 163 |
-
|
| 164 |
-
```Python
|
| 165 |
-
from typing import Literal
|
| 166 |
-
|
| 167 |
-
from pydantic import Field, BaseModel
|
| 168 |
-
from numind.nuextract_utils import convert_json_schema_to_nuextract_template
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
class HotelBooking(BaseModel):
|
| 172 |
-
city: str
|
| 173 |
-
check_in_date: str = Field(description="date")
|
| 174 |
-
check_out_date: str = Field(description="date")
|
| 175 |
-
number_of_guests: int
|
| 176 |
-
room_type: Literal["single", "double", "suite"]
|
| 177 |
-
|
| 178 |
-
|
| 179 |
-
template, dropped_branches = convert_json_schema_to_nuextract_template(
|
| 180 |
-
HotelBooking.model_json_schema()
|
| 181 |
-
)
|
| 182 |
-
|
| 183 |
-
# {'check_in_date': 'date', 'check_out_date': 'date', 'city': 'string', 'number_of_guests': 'integer', 'room_type': ['single', 'double', 'suite']}
|
| 184 |
-
```
|
| 185 |
-
|
| 186 |
-
## Document-to-Markdown
|
| 187 |
-
|
| 188 |
-
NuExtract can also convert document images into clean Markdown. Output will be markdown for text (headers etc), html for tables, latex for mat and ```<figure data-type="image" data-id="img_n"><img src="/NM-dev/model_card-A/resolve/main/img_n.png" alt="Detail description of the images"/> ```
|
| 189 |
-
|
| 190 |
-
Markdown example:
|
| 191 |
-
|
| 192 |
-
```markdown
|
| 193 |
-
<figure data-type="image" data-id="img_1">
|
| 194 |
-
<img src="img_1.png" alt="Logo of Mobilier 2000 with contact information: Tél.: (418) 275-4232, 1654, boul. Marcotte, Roberval (Qc) G8H 2P2"/>
|
| 195 |
-
</figure>
|
| 196 |
-
|
| 197 |
-
# COMMANDE
|
| 198 |
-
**NUMÉRO 72259**
|
| 199 |
-
|
| 200 |
-
1
|
| 201 |
-
|
| 202 |
-
**Vendu à**
|
| 203 |
-
TREMBLAY ERIC
|
| 204 |
-
ERIC TREMBLAY
|
| 205 |
-
348 BOUL. DE L'ANSE
|
| 206 |
-
ROBERVAL
|
| 207 |
-
G8H 1Y9
|
| 208 |
-
|
| 209 |
-
**Livré à**
|
| 210 |
-
TREMBLAY ERIC
|
| 211 |
-
ERIC TREMBLAY
|
| 212 |
-
348 BOUL. DE L'ANSE
|
| 213 |
-
ROBERVAL
|
| 214 |
-
G8H 1Y9
|
| 215 |
-
|
| 216 |
-
<table>
|
| 217 |
-
<thead>
|
| 218 |
-
<tr>
|
| 219 |
-
<th># CLIENT</th>
|
| 220 |
-
<th>EXPÉDITEUR</th>
|
| 221 |
-
<th>TERME DE CRÉDIT</th>
|
| 222 |
-
<th>DATE</th>
|
| 223 |
-
</tr>
|
| 224 |
-
</thead>
|
| 225 |
-
<tbody>
|
| 226 |
-
<tr>
|
| 227 |
-
<td>2753133</td>
|
| 228 |
-
<td>Notre camion</td>
|
| 229 |
-
<td>à la livraison</td>
|
| 230 |
-
<td>22/06/2023</td>
|
| 231 |
-
</tr>
|
| 232 |
-
</tbody>
|
| 233 |
-
</table>
|
| 234 |
-
|
| 235 |
-
<table>
|
| 236 |
-
<thead>
|
| 237 |
-
<tr>
|
| 238 |
-
<th>NOM DU VENDEUR</th>
|
| 239 |
-
<th>VOTRE ÉCONOMIE !</th>
|
| 240 |
-
<th># COMMANDE</th>
|
| 241 |
-
</tr>
|
| 242 |
-
</thead>
|
| 243 |
-
<tbody>
|
| 244 |
-
<tr>
|
| 245 |
-
<td>Éric</td>
|
| 246 |
-
<td>0.00</td>
|
| 247 |
-
<td></td>
|
| 248 |
-
</tr>
|
| 249 |
-
</tbody>
|
| 250 |
-
</table>
|
| 251 |
-
```
|
| 252 |
-
|
| 253 |
-
---
|
| 254 |
-
|
| 255 |
-
## Reasoning and non-reasoning modes
|
| 256 |
-
|
| 257 |
-
NuExtract supports both reasoning and non-reasoning inference.
|
| 258 |
-
|
| 259 |
-
### Non-thinking mode
|
| 260 |
-
|
| 261 |
-
Use this for fast and deterministic extraction or Markdown conversion.
|
| 262 |
-
|
| 263 |
-
```python
|
| 264 |
-
enable_thinking = False
|
| 265 |
-
temperature = 0.2
|
| 266 |
-
```
|
| 267 |
-
|
| 268 |
-
### Thinking mode
|
| 269 |
-
|
| 270 |
-
Use this for difficult documents, complex layouts, ambiguous fields, or cases where the document structure requires additional reasoning.
|
| 271 |
-
|
| 272 |
-
```python
|
| 273 |
-
enable_thinking = True
|
| 274 |
-
temperature = 0.6
|
| 275 |
-
```
|
| 276 |
-
|
| 277 |
-
For production extraction workloads, we recommend starting with **non-reasoning mode** and enabling reasoning only for difficult examples.
|
| 278 |
-
|
| 279 |
-
|
| 280 |
-
---
|
| 281 |
-
|
| 282 |
-
## vLLM deployment
|
| 283 |
-
|
| 284 |
-
NuExtract can be served with vLLM using an OpenAI-compatible API.
|
| 285 |
-
|
| 286 |
-
```bash
|
| 287 |
-
vllm serve numind/NuExtract3 \
|
| 288 |
-
--trust-remote-code \
|
| 289 |
-
--limit-mm-per-prompt '{"image": 99, "video": 0}' \
|
| 290 |
-
--chat-template-content-format openai \
|
| 291 |
-
--generation-config vllm \
|
| 292 |
-
--max-model-len 131072 \
|
| 293 |
-
--speculative-config '{"method": "qwen3_next_mtp", "num_speculative_tokens": 2}'
|
| 294 |
-
```
|
| 295 |
-
|
| 296 |
-
|
| 297 |
-
### Multi Token Prediction
|
| 298 |
-
<details>
|
| 299 |
-
The deployment commands above enable Multi Token Prediction (MTP) through vLLM speculative decoding:
|
| 300 |
-
|
| 301 |
-
```bash
|
| 302 |
-
--speculative-config '{"method": "qwen3_next_mtp", "num_speculative_tokens": 2}'
|
| 303 |
-
```
|
| 304 |
-
|
| 305 |
-
MTP can improve decoding throughput without changing the OpenAI-compatible request payload. You can tune `num_speculative_tokens` for your hardware and workload, or remove `--speculative-config` if your vLLM version or environment does not support this speculative decoding method.
|
| 306 |
-
|
| 307 |
-
If you encounter memory issues, reduce the maximum model length and the maximum number of images:
|
| 308 |
-
|
| 309 |
-
```bash
|
| 310 |
-
vllm serve numind/NuExtract-3 \
|
| 311 |
-
--trust-remote-code \
|
| 312 |
-
--limit-mm-per-prompt '{"image": 6, "video": 0}' \
|
| 313 |
-
--chat-template-content-format openai \
|
| 314 |
-
--generation-config vllm \
|
| 315 |
-
--max-model-len 16384 \
|
| 316 |
-
--speculative-config '{"method": "qwen3_next_mtp", "num_speculative_tokens": 2}'
|
| 317 |
-
```
|
| 318 |
-
</details>
|
| 319 |
-
|
| 320 |
-
## vLLM inference: structured extraction: text
|
| 321 |
-
```python
|
| 322 |
-
import json
|
| 323 |
-
from openai import OpenAI
|
| 324 |
-
|
| 325 |
-
client = OpenAI(
|
| 326 |
-
api_key="EMPTY",
|
| 327 |
-
base_url="http://localhost:8000/v1",
|
| 328 |
-
)
|
| 329 |
-
|
| 330 |
-
template = {
|
| 331 |
-
"store": "verbatim-string",
|
| 332 |
-
"date": "date-time",
|
| 333 |
-
"total": "number",
|
| 334 |
-
"currency": ["USD", "EUR", "GBP", "JPY", "Other"],
|
| 335 |
-
"items": [
|
| 336 |
-
{
|
| 337 |
-
"name": "verbatim-string",
|
| 338 |
-
"price": "number"
|
| 339 |
-
}
|
| 340 |
-
]
|
| 341 |
-
}
|
| 342 |
-
|
| 343 |
-
response = client.chat.completions.create(
|
| 344 |
-
model="numind/NuExtract3",
|
| 345 |
-
temperature=0.2,
|
| 346 |
-
messages=[
|
| 347 |
-
{
|
| 348 |
-
"role": "user",
|
| 349 |
-
"content": [
|
| 350 |
-
{
|
| 351 |
-
"type": "text",
|
| 352 |
-
"text": "Yesterday I bought apples and coffee at Trader Joe's for a total of $12.40."
|
| 353 |
-
}
|
| 354 |
-
],
|
| 355 |
-
}
|
| 356 |
-
],
|
| 357 |
-
extra_body={
|
| 358 |
-
"chat_template_kwargs": {
|
| 359 |
-
"template": json.dumps(template),
|
| 360 |
-
"instructions": "Specify the time for the `date` entry only if it is present, otherwise only output the date component.",
|
| 361 |
-
"enable_thinking": False
|
| 362 |
-
}
|
| 363 |
-
}
|
| 364 |
-
)
|
| 365 |
-
|
| 366 |
-
print(response.choices[0].message.content)
|
| 367 |
-
```
|
| 368 |
-
|
| 369 |
-
Example output:
|
| 370 |
-
|
| 371 |
-
```json
|
| 372 |
-
{
|
| 373 |
-
"store": "Trader Joe's",
|
| 374 |
-
"date": null,
|
| 375 |
-
"total": 12.40,
|
| 376 |
-
"currency": "USD",
|
| 377 |
-
"items": [
|
| 378 |
-
{
|
| 379 |
-
"name": "apples",
|
| 380 |
-
"price": null
|
| 381 |
-
},
|
| 382 |
-
{
|
| 383 |
-
"name": "coffee",
|
| 384 |
-
"price": null
|
| 385 |
-
}
|
| 386 |
-
]
|
| 387 |
-
}
|
| 388 |
-
```
|
| 389 |
-
|
| 390 |
-
---
|
| 391 |
-
|
| 392 |
-
## vLLM inference: structured extraction: image
|
| 393 |
-
|
| 394 |
-
```python
|
| 395 |
-
import json
|
| 396 |
-
import base64
|
| 397 |
-
from openai import OpenAI
|
| 398 |
-
|
| 399 |
-
client = OpenAI(
|
| 400 |
-
api_key="EMPTY",
|
| 401 |
-
base_url="http://localhost:8000/v1",
|
| 402 |
-
)
|
| 403 |
-
|
| 404 |
-
def encode_image(image_path):
|
| 405 |
-
with open(image_path, "rb") as image_file:
|
| 406 |
-
return base64.b64encode(image_file.read()).decode("utf-8")
|
| 407 |
-
|
| 408 |
-
image_base64 = encode_image("receipt.png")
|
| 409 |
-
data_url = f"data:image/png;base64,{image_base64}"
|
| 410 |
-
|
| 411 |
-
template = {
|
| 412 |
-
"store": "verbatim-string",
|
| 413 |
-
"date": "date-time",
|
| 414 |
-
"total": "number",
|
| 415 |
-
"payment_method": "verbatim-string"
|
| 416 |
-
}
|
| 417 |
-
|
| 418 |
-
response = client.chat.completions.create(
|
| 419 |
-
model="numind/NuExtract3",
|
| 420 |
-
temperature=0.2,
|
| 421 |
-
messages=[
|
| 422 |
-
{
|
| 423 |
-
"role": "user",
|
| 424 |
-
"content": [
|
| 425 |
-
{
|
| 426 |
-
"type": "image_url",
|
| 427 |
-
"image_url": {"url": data_url}
|
| 428 |
-
}
|
| 429 |
-
],
|
| 430 |
-
}
|
| 431 |
-
],
|
| 432 |
-
extra_body={
|
| 433 |
-
"chat_template_kwargs": {
|
| 434 |
-
"template": json.dumps(template, indent=4),
|
| 435 |
-
"enable_thinking": False
|
| 436 |
-
}
|
| 437 |
-
}
|
| 438 |
-
)
|
| 439 |
-
|
| 440 |
-
print(response.choices[0].message.content)
|
| 441 |
-
```
|
| 442 |
-
|
| 443 |
-
Example output:
|
| 444 |
-
|
| 445 |
-
```json
|
| 446 |
-
{
|
| 447 |
-
"store": "Trader Joe's",
|
| 448 |
-
"date": "2025-04-12",
|
| 449 |
-
"total": 42.85,
|
| 450 |
-
"payment_method": "Visa"
|
| 451 |
-
}
|
| 452 |
-
```
|
| 453 |
-
|
| 454 |
-
### Multiple page PDF
|
| 455 |
-
<details>
|
| 456 |
-
You can render a PDF to one PNG image per page with PyMuPDF, then pass the images to vLLM in page order.
|
| 457 |
-
|
| 458 |
-
```python
|
| 459 |
-
import base64
|
| 460 |
-
import json
|
| 461 |
-
|
| 462 |
-
import fitz # pip install pymupdf
|
| 463 |
-
from openai import OpenAI
|
| 464 |
-
|
| 465 |
-
client = OpenAI(
|
| 466 |
-
api_key="EMPTY",
|
| 467 |
-
base_url="http://localhost:8000/v1",
|
| 468 |
-
)
|
| 469 |
-
|
| 470 |
-
def pdf_to_png_data_urls(pdf_path, dpi=170):
|
| 471 |
-
data_urls = []
|
| 472 |
-
|
| 473 |
-
with fitz.open(pdf_path) as doc:
|
| 474 |
-
for page in doc:
|
| 475 |
-
pix = page.get_pixmap(dpi=dpi, alpha=False)
|
| 476 |
-
png_bytes = pix.tobytes("png")
|
| 477 |
-
png_base64 = base64.b64encode(png_bytes).decode("utf-8")
|
| 478 |
-
data_urls.append(f"data:image/png;base64,{png_base64}")
|
| 479 |
-
|
| 480 |
-
return data_urls
|
| 481 |
-
|
| 482 |
-
data_urls = pdf_to_png_data_urls("invoice.pdf", dpi=170)
|
| 483 |
-
|
| 484 |
-
template = {
|
| 485 |
-
"invoice_number": "verbatim-string",
|
| 486 |
-
"invoice_date": "date",
|
| 487 |
-
"total": "number",
|
| 488 |
-
"currency": "currency",
|
| 489 |
-
"line_items": [
|
| 490 |
-
{
|
| 491 |
-
"description": "verbatim-string",
|
| 492 |
-
"quantity": "number",
|
| 493 |
-
"unit_price": "number",
|
| 494 |
-
"total": "number"
|
| 495 |
-
}
|
| 496 |
-
]
|
| 497 |
-
}
|
| 498 |
-
|
| 499 |
-
response = client.chat.completions.create(
|
| 500 |
-
model="numind/NuExtract3",
|
| 501 |
-
temperature=0.2,
|
| 502 |
-
messages=[
|
| 503 |
-
{
|
| 504 |
-
"role": "user",
|
| 505 |
-
"content": [
|
| 506 |
-
{
|
| 507 |
-
"type": "image_url",
|
| 508 |
-
"image_url": {"url": data_url}
|
| 509 |
-
}
|
| 510 |
-
for data_url in data_urls
|
| 511 |
-
],
|
| 512 |
-
}
|
| 513 |
-
],
|
| 514 |
-
extra_body={
|
| 515 |
-
"chat_template_kwargs": {
|
| 516 |
-
"template": json.dumps(template, indent=4),
|
| 517 |
-
"enable_thinking": False
|
| 518 |
-
}
|
| 519 |
-
}
|
| 520 |
-
)
|
| 521 |
-
|
| 522 |
-
print(response.choices[0].message.content)
|
| 523 |
-
```
|
| 524 |
-
</details>
|
| 525 |
-
|
| 526 |
-
|
| 527 |
-
|
| 528 |
-
## vLLM inference: document-to-Markdown
|
| 529 |
-
|
| 530 |
-
For Markdown OCR, use `mode="markdown"` or `mode="content"` without a template.
|
| 531 |
-
|
| 532 |
-
```python
|
| 533 |
-
import base64
|
| 534 |
-
from openai import OpenAI
|
| 535 |
-
|
| 536 |
-
client = OpenAI(
|
| 537 |
-
api_key="EMPTY",
|
| 538 |
-
base_url="http://localhost:8000/v1",
|
| 539 |
-
)
|
| 540 |
-
|
| 541 |
-
def encode_image(image_path):
|
| 542 |
-
with open(image_path, "rb") as image_file:
|
| 543 |
-
return base64.b64encode(image_file.read()).decode("utf-8")
|
| 544 |
-
|
| 545 |
-
image_base64 = encode_image("document.png")
|
| 546 |
-
data_url = f"data:image/png;base64,{image_base64}"
|
| 547 |
-
|
| 548 |
-
response = client.chat.completions.create(
|
| 549 |
-
model="numind/NuExtract3",
|
| 550 |
-
temperature=1,
|
| 551 |
-
messages=[
|
| 552 |
-
{
|
| 553 |
-
"role": "user",
|
| 554 |
-
"content": [
|
| 555 |
-
{
|
| 556 |
-
"type": "image_url",
|
| 557 |
-
"image_url": {"url": data_url}
|
| 558 |
-
}
|
| 559 |
-
],
|
| 560 |
-
}
|
| 561 |
-
],
|
| 562 |
-
extra_body={
|
| 563 |
-
"chat_template_kwargs": {
|
| 564 |
-
"mode": "markdown",
|
| 565 |
-
"enable_thinking": False
|
| 566 |
-
}
|
| 567 |
-
}
|
| 568 |
-
)
|
| 569 |
-
|
| 570 |
-
print(response.choices[0].message.content)
|
| 571 |
-
```
|
| 572 |
-
|
| 573 |
-
---
|
| 574 |
-
|
| 575 |
-
## vLLM inference: reasoning mode
|
| 576 |
-
<details>
|
| 577 |
-
Reasoning can be enabled for harder structured extraction or Markdown tasks.
|
| 578 |
-
|
| 579 |
-
```python
|
| 580 |
-
response = client.chat.completions.create(
|
| 581 |
-
model="numind/NuExtract3",
|
| 582 |
-
temperature=1,
|
| 583 |
-
messages=[
|
| 584 |
-
{
|
| 585 |
-
"role": "user",
|
| 586 |
-
"content": [
|
| 587 |
-
{
|
| 588 |
-
"type": "image_url",
|
| 589 |
-
"image_url": {"url": data_url}
|
| 590 |
-
}
|
| 591 |
-
],
|
| 592 |
-
}
|
| 593 |
-
],
|
| 594 |
-
extra_body={
|
| 595 |
-
"chat_template_kwargs": {
|
| 596 |
-
"mode": "markdown",
|
| 597 |
-
"enable_thinking": True
|
| 598 |
-
}
|
| 599 |
-
}
|
| 600 |
-
)
|
| 601 |
-
|
| 602 |
-
result = response.choices[0].message.content
|
| 603 |
-
|
| 604 |
-
if "</think>" in result:
|
| 605 |
-
reasoning, answer = [part.strip() for part in result.split("</think>")]
|
| 606 |
-
else:
|
| 607 |
-
reasoning, answer = None, result
|
| 608 |
-
|
| 609 |
-
print(answer)
|
| 610 |
-
```
|
| 611 |
-
</details>
|
| 612 |
-
|
| 613 |
-
|
| 614 |
-
## In-context examples for extraction
|
| 615 |
-
<details>
|
| 616 |
-
NuExtract supports in-context examples for structured extraction.
|
| 617 |
-
|
| 618 |
-
Examples are especially useful when the desired formatting is ambiguous or when the schema requires task-specific conventions. Examples can be provided by using `developer` messages, for which all items of the contents except the last one are the input, and the last one is the expected output.
|
| 619 |
-
|
| 620 |
-
```python
|
| 621 |
-
import json
|
| 622 |
-
from openai import OpenAI
|
| 623 |
-
|
| 624 |
-
client = OpenAI(
|
| 625 |
-
api_key="EMPTY",
|
| 626 |
-
base_url="http://localhost:8000/v1",
|
| 627 |
-
)
|
| 628 |
-
|
| 629 |
-
template = {
|
| 630 |
-
"names": ["string"]
|
| 631 |
-
}
|
| 632 |
-
|
| 633 |
-
response = client.chat.completions.create(
|
| 634 |
-
model="numind/NuExtract3",
|
| 635 |
-
temperature=0.2,
|
| 636 |
-
messages=[
|
| 637 |
-
{
|
| 638 |
-
"role": "developer",
|
| 639 |
-
"content": [
|
| 640 |
-
{
|
| 641 |
-
"type": "text",
|
| 642 |
-
"text": "Stephen is the manager at Susan's store.",
|
| 643 |
-
},
|
| 644 |
-
{
|
| 645 |
-
"type": "text",
|
| 646 |
-
"text": "{\"names\": [\"-STEPHEN-\", \"-SUSAN-\"]}",
|
| 647 |
-
}
|
| 648 |
-
],
|
| 649 |
-
},
|
| 650 |
-
{
|
| 651 |
-
"role": "user",
|
| 652 |
-
"content": [
|
| 653 |
-
{
|
| 654 |
-
"type": "text",
|
| 655 |
-
"text": "John went to the restaurant with Mary. James went to the cinema."
|
| 656 |
-
}
|
| 657 |
-
],
|
| 658 |
-
}
|
| 659 |
-
],
|
| 660 |
-
extra_body={
|
| 661 |
-
"chat_template_kwargs": {
|
| 662 |
-
"template": json.dumps(template, indent=4),
|
| 663 |
-
"enable_thinking": False
|
| 664 |
-
}
|
| 665 |
-
}
|
| 666 |
-
)
|
| 667 |
-
|
| 668 |
-
print(response.choices[0].message.content)
|
| 669 |
-
```
|
| 670 |
-
|
| 671 |
-
Example output:
|
| 672 |
-
|
| 673 |
-
```json
|
| 674 |
-
{
|
| 675 |
-
"names": ["-JOHN-", "-MARY-", "-JAMES-"]
|
| 676 |
-
}
|
| 677 |
-
```
|
| 678 |
-
</details>
|
| 679 |
-
|
| 680 |
-
|
| 681 |
-
## vLLM inference: template generation
|
| 682 |
-
|
| 683 |
-
NuExtract can generate an extraction template from a natural language description.
|
| 684 |
-
|
| 685 |
-
```python
|
| 686 |
-
from openai import OpenAI
|
| 687 |
-
|
| 688 |
-
client = OpenAI(
|
| 689 |
-
api_key="EMPTY",
|
| 690 |
-
base_url="http://localhost:8000/v1",
|
| 691 |
-
)
|
| 692 |
-
|
| 693 |
-
response = client.chat.completions.create(
|
| 694 |
-
model="numind/NuExtract3",
|
| 695 |
-
temperature=0.2,
|
| 696 |
-
messages=[
|
| 697 |
-
{
|
| 698 |
-
"role": "user",
|
| 699 |
-
"content": [
|
| 700 |
-
{
|
| 701 |
-
"type": "text",
|
| 702 |
-
"text": "I want to extract the key details from a rental contract."
|
| 703 |
-
}
|
| 704 |
-
],
|
| 705 |
-
}
|
| 706 |
-
],
|
| 707 |
-
extra_body={
|
| 708 |
-
"chat_template_kwargs": {
|
| 709 |
-
"mode": "template-generation"
|
| 710 |
-
}
|
| 711 |
-
}
|
| 712 |
-
)
|
| 713 |
-
|
| 714 |
-
print(response.choices[0].message.content)
|
| 715 |
-
```
|
| 716 |
-
|
| 717 |
-
Example output:
|
| 718 |
-
|
| 719 |
-
```json
|
| 720 |
-
{
|
| 721 |
-
"contract_title": "verbatim-string",
|
| 722 |
-
"landlord": "verbatim-string",
|
| 723 |
-
"tenant": "verbatim-string",
|
| 724 |
-
"property_address": "verbatim-string",
|
| 725 |
-
"start_date": "date-time",
|
| 726 |
-
"end_date": "date-time",
|
| 727 |
-
"monthly_rent": "number",
|
| 728 |
-
"currency": "verbatim-string",
|
| 729 |
-
"deposit": "number",
|
| 730 |
-
"signatories": ["verbatim-string"]
|
| 731 |
-
}
|
| 732 |
-
```
|
| 733 |
-
|
| 734 |
-
## Curl examples
|
| 735 |
-
<details>
|
| 736 |
-
|
| 737 |
-
The following examples assume that vLLM is running locally on port 8000. They use `jq` to build valid JSON request bodies without manually escaping the image data or template string.
|
| 738 |
-
|
| 739 |
-
### Single image structured extraction
|
| 740 |
-
|
| 741 |
-
```bash
|
| 742 |
-
API_KEY="EMPTY"
|
| 743 |
-
IMAGE_BASE64_FILE=$(mktemp)
|
| 744 |
-
REQUEST_BODY_FILE=$(mktemp)
|
| 745 |
-
|
| 746 |
-
base64 < receipt.png | tr -d '\n' > "$IMAGE_BASE64_FILE"
|
| 747 |
-
|
| 748 |
-
TEMPLATE=$(cat <<'JSON'
|
| 749 |
-
{
|
| 750 |
-
"store": "verbatim-string",
|
| 751 |
-
"date": "date-time",
|
| 752 |
-
"total": "number",
|
| 753 |
-
"payment_method": "verbatim-string"
|
| 754 |
-
}
|
| 755 |
-
JSON
|
| 756 |
-
)
|
| 757 |
-
|
| 758 |
-
jq -n \
|
| 759 |
-
--rawfile image_base64 "$IMAGE_BASE64_FILE" \
|
| 760 |
-
--arg template "$TEMPLATE" \
|
| 761 |
-
'{
|
| 762 |
-
model: "numind/NuExtract3",
|
| 763 |
-
temperature: 0.6,
|
| 764 |
-
messages: [
|
| 765 |
-
{
|
| 766 |
-
role: "user",
|
| 767 |
-
content: [
|
| 768 |
-
{
|
| 769 |
-
type: "image_url",
|
| 770 |
-
image_url: {url: ("data:image/png;base64," + $image_base64)}
|
| 771 |
-
}
|
| 772 |
-
]
|
| 773 |
-
}
|
| 774 |
-
],
|
| 775 |
-
chat_template_kwargs: {
|
| 776 |
-
template: $template,
|
| 777 |
-
enable_thinking: false
|
| 778 |
-
}
|
| 779 |
-
}' > "$REQUEST_BODY_FILE"
|
| 780 |
-
|
| 781 |
-
curl http://localhost:8000/v1/chat/completions \
|
| 782 |
-
-H "Content-Type: application/json" \
|
| 783 |
-
-H "Authorization: Bearer $API_KEY" \
|
| 784 |
-
--data-binary "@$REQUEST_BODY_FILE"
|
| 785 |
-
|
| 786 |
-
rm "$IMAGE_BASE64_FILE" "$REQUEST_BODY_FILE"
|
| 787 |
-
```
|
| 788 |
-
|
| 789 |
-
### Single image content extraction
|
| 790 |
-
|
| 791 |
-
```bash
|
| 792 |
-
API_KEY="EMPTY"
|
| 793 |
-
IMAGE_BASE64_FILE=$(mktemp)
|
| 794 |
-
REQUEST_BODY_FILE=$(mktemp)
|
| 795 |
-
|
| 796 |
-
base64 < document.png | tr -d '\n' > "$IMAGE_BASE64_FILE"
|
| 797 |
-
|
| 798 |
-
jq -n \
|
| 799 |
-
--rawfile image_base64 "$IMAGE_BASE64_FILE" \
|
| 800 |
-
'{
|
| 801 |
-
model: "numind/NuExtract3",
|
| 802 |
-
temperature: 0.6,
|
| 803 |
-
messages: [
|
| 804 |
-
{
|
| 805 |
-
role: "user",
|
| 806 |
-
content: [
|
| 807 |
-
{
|
| 808 |
-
type: "image_url",
|
| 809 |
-
image_url: {url: ("data:image/png;base64," + $image_base64)}
|
| 810 |
-
}
|
| 811 |
-
]
|
| 812 |
-
}
|
| 813 |
-
],
|
| 814 |
-
chat_template_kwargs: {
|
| 815 |
-
mode: "content",
|
| 816 |
-
enable_thinking: false
|
| 817 |
-
}
|
| 818 |
-
}' > "$REQUEST_BODY_FILE"
|
| 819 |
-
|
| 820 |
-
curl http://localhost:8000/v1/chat/completions \
|
| 821 |
-
-H "Content-Type: application/json" \
|
| 822 |
-
-H "Authorization: Bearer $API_KEY" \
|
| 823 |
-
--data-binary "@$REQUEST_BODY_FILE"
|
| 824 |
-
|
| 825 |
-
rm "$IMAGE_BASE64_FILE" "$REQUEST_BODY_FILE"
|
| 826 |
-
```
|
| 827 |
-
</details>
|
| 828 |
-
|
| 829 |
-
|
| 830 |
-
## Transformers example
|
| 831 |
-
<details>
|
| 832 |
-
You can also run NuExtract directly with `transformers`. The same `template`, `mode`, and `enable_thinking` options are passed to `processor.apply_chat_template`.
|
| 833 |
-
|
| 834 |
-
```python
|
| 835 |
-
import json
|
| 836 |
-
|
| 837 |
-
import torch
|
| 838 |
-
from PIL import Image
|
| 839 |
-
from transformers import AutoModelForImageTextToText, AutoProcessor
|
| 840 |
-
|
| 841 |
-
model_id = "numind/NuExtract3"
|
| 842 |
-
|
| 843 |
-
processor = AutoProcessor.from_pretrained(
|
| 844 |
-
model_id,
|
| 845 |
-
trust_remote_code=True,
|
| 846 |
-
)
|
| 847 |
-
model = AutoModelForImageTextToText.from_pretrained(
|
| 848 |
-
model_id,
|
| 849 |
-
dtype=torch.bfloat16,
|
| 850 |
-
device_map="auto",
|
| 851 |
-
trust_remote_code=True,
|
| 852 |
-
).eval()
|
| 853 |
-
|
| 854 |
-
def run_nuextract(messages, **chat_template_kwargs):
|
| 855 |
-
inputs = processor.apply_chat_template(
|
| 856 |
-
messages,
|
| 857 |
-
add_generation_prompt=True,
|
| 858 |
-
tokenize=True,
|
| 859 |
-
return_dict=True,
|
| 860 |
-
return_tensors="pt",
|
| 861 |
-
**chat_template_kwargs,
|
| 862 |
-
).to(model.device)
|
| 863 |
-
|
| 864 |
-
with torch.inference_mode():
|
| 865 |
-
generated_ids = model.generate(
|
| 866 |
-
**inputs,
|
| 867 |
-
max_new_tokens=4096,
|
| 868 |
-
do_sample=False,
|
| 869 |
-
)
|
| 870 |
-
|
| 871 |
-
generated_ids = generated_ids[:, inputs.input_ids.shape[1]:]
|
| 872 |
-
return processor.batch_decode(
|
| 873 |
-
generated_ids,
|
| 874 |
-
skip_special_tokens=True,
|
| 875 |
-
clean_up_tokenization_spaces=False,
|
| 876 |
-
)[0].strip()
|
| 877 |
-
|
| 878 |
-
# Single image structured extraction
|
| 879 |
-
receipt_image = Image.open("receipt.png").convert("RGB")
|
| 880 |
-
receipt_messages = [
|
| 881 |
-
{
|
| 882 |
-
"role": "user",
|
| 883 |
-
"content": [
|
| 884 |
-
{
|
| 885 |
-
"type": "image",
|
| 886 |
-
"image": receipt_image,
|
| 887 |
-
}
|
| 888 |
-
],
|
| 889 |
-
}
|
| 890 |
-
]
|
| 891 |
-
|
| 892 |
-
template = {
|
| 893 |
-
"store": "verbatim-string",
|
| 894 |
-
"date": "date-time",
|
| 895 |
-
"total": "number",
|
| 896 |
-
"payment_method": "verbatim-string"
|
| 897 |
-
}
|
| 898 |
-
|
| 899 |
-
structured_output = run_nuextract(
|
| 900 |
-
receipt_messages,
|
| 901 |
-
template=json.dumps(template, indent=4),
|
| 902 |
-
enable_thinking=False,
|
| 903 |
-
)
|
| 904 |
-
print(structured_output)
|
| 905 |
-
|
| 906 |
-
# Single image content extraction
|
| 907 |
-
document_image = Image.open("document.png").convert("RGB")
|
| 908 |
-
document_messages = [
|
| 909 |
-
{
|
| 910 |
-
"role": "user",
|
| 911 |
-
"content": [
|
| 912 |
-
{
|
| 913 |
-
"type": "image",
|
| 914 |
-
"image": document_image,
|
| 915 |
-
}
|
| 916 |
-
],
|
| 917 |
-
}
|
| 918 |
-
]
|
| 919 |
-
|
| 920 |
-
content_output = run_nuextract(
|
| 921 |
-
document_messages,
|
| 922 |
-
mode="content",
|
| 923 |
-
enable_thinking=False,
|
| 924 |
-
)
|
| 925 |
-
print(content_output)
|
| 926 |
-
```
|
| 927 |
-
</details>
|
| 928 |
-
|
| 929 |
-
Special thanks to the Lambda.ai team for the compute that made this project a success.
|
| 930 |
-
|
| 931 |
-
## Citation
|
| 932 |
-
|
| 933 |
-
If you use NuExtract, please cite NuMind and link to the model page.
|
| 934 |
-
|
| 935 |
-
```bibtex
|
| 936 |
-
@misc{nuextract3,
|
| 937 |
-
title = {NuExtract3},
|
| 938 |
-
author = {NuMind},
|
| 939 |
-
year = {2026},
|
| 940 |
-
url = {https://nuextract.ai/}
|
| 941 |
-
}
|
| 942 |
-
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
.agents/references/nuextract3/2026-07-20/huggingface.co/numind/NuExtract3.md
DELETED
|
@@ -1,1010 +0,0 @@
|
|
| 1 |
-
<!--
|
| 2 |
-
source: https://huggingface.co/numind/NuExtract3
|
| 3 |
-
retrieved: 2026-07-20T17:29:25.403047+00:00
|
| 4 |
-
final_url: https://huggingface.co/numind/NuExtract3
|
| 5 |
-
content_type: text/html
|
| 6 |
-
-->
|
| 7 |
-
|
| 8 |
-
# numind/NuExtract3 · Hugging Face
|
| 9 |
-
|
| 10 |
-
[Hugging Face](https://huggingface.co/)
|
| 11 |
-
|
| 12 |
-
#
|
| 13 |
-
|
| 14 |
-
[https://huggingface.co/numind](https://huggingface.co/numind)
|
| 15 |
-
|
| 16 |
-
[numind](https://huggingface.co/numind)
|
| 17 |
-
|
| 18 |
-
/
|
| 19 |
-
|
| 20 |
-
[NuExtract3](https://huggingface.co/numind/NuExtract3)
|
| 21 |
-
|
| 22 |
-
like 295
|
| 23 |
-
|
| 24 |
-
Follow
|
| 25 |
-
|
| 26 |
-
NuMind 649
|
| 27 |
-
|
| 28 |
-
[Image-to-Text](https://huggingface.co/models?pipeline_tag=image-to-text)
|
| 29 |
-
|
| 30 |
-
[Transformers](https://huggingface.co/models?library=transformers)
|
| 31 |
-
|
| 32 |
-
[Safetensors](https://huggingface.co/models?library=safetensors)
|
| 33 |
-
|
| 34 |
-
[qwen3_5](https://huggingface.co/models?other=qwen3_5)
|
| 35 |
-
|
| 36 |
-
[image-text-to-text](https://huggingface.co/models?other=image-text-to-text)
|
| 37 |
-
|
| 38 |
-
[vision-language](https://huggingface.co/models?other=vision-language)
|
| 39 |
-
|
| 40 |
-
[vlm](https://huggingface.co/models?other=vlm)
|
| 41 |
-
|
| 42 |
-
[document-understanding](https://huggingface.co/models?other=document-understanding)
|
| 43 |
-
|
| 44 |
-
[structured-extraction](https://huggingface.co/models?other=structured-extraction)
|
| 45 |
-
|
| 46 |
-
[information-extraction](https://huggingface.co/models?other=information-extraction)
|
| 47 |
-
|
| 48 |
-
[ocr](https://huggingface.co/models?other=ocr)
|
| 49 |
-
|
| 50 |
-
[document-to-markdown](https://huggingface.co/models?other=document-to-markdown)
|
| 51 |
-
|
| 52 |
-
[markdown](https://huggingface.co/models?other=markdown)
|
| 53 |
-
|
| 54 |
-
[rag](https://huggingface.co/models?other=rag)
|
| 55 |
-
|
| 56 |
-
[reasoning](https://huggingface.co/models?other=reasoning)
|
| 57 |
-
|
| 58 |
-
[multilingual](https://huggingface.co/models?other=multilingual)
|
| 59 |
-
|
| 60 |
-
[conversational](https://huggingface.co/models?other=conversational)
|
| 61 |
-
|
| 62 |
-
[Eval Results](https://huggingface.co/models?other=eval-results)
|
| 63 |
-
|
| 64 |
-
License: apache-2.0
|
| 65 |
-
|
| 66 |
-
[Model card](https://huggingface.co/numind/NuExtract3)
|
| 67 |
-
|
| 68 |
-
[Files Files and versions xet](https://huggingface.co/numind/NuExtract3/tree/main)
|
| 69 |
-
|
| 70 |
-
[Community 2](https://huggingface.co/numind/NuExtract3/discussions)
|
| 71 |
-
|
| 72 |
-
Deploy
|
| 73 |
-
|
| 74 |
-
Copy to bucket new
|
| 75 |
-
|
| 76 |
-
Use this model
|
| 77 |
-
|
| 78 |
-
[https://nuextract.ai/](https://nuextract.ai/)
|
| 79 |
-
|
| 80 |
-
🖥️ [API / Platform](https://nuextract.ai/) | 📑 [Blog](https://numind.ai/blog) | 🗣️ [Discord](https://discord.gg/3tsEtJNCDe) | 🛠️ [GitHub](https://github.com/numindai/nuextract)
|
| 81 |
-
|
| 82 |
-
NuExtract3 is a unified 4B vision-language reasoning model for document understanding.
|
| 83 |
-
|
| 84 |
-
It combines strong structured information extraction with high-quality image-to-Markdown conversion, making it suitable for extraction pipelines, OCR, and RAG preprocessing for all types of documents such as scans, receipts, forms, invoices, contracts or tables.
|
| 85 |
-
|
| 86 |
-
Try it out in [the 🤗 space!](https://huggingface.co/spaces/numind/NuExtract-3-4B)
|
| 87 |
-
|
| 88 |
-
## [https://huggingface.co/numind/NuExtract3#overview](https://huggingface.co/numind/NuExtract3#overview) Overview
|
| 89 |
-
|
| 90 |
-
- Structured extraction : input (text/images) + JSON template + instructions --> JSON output
|
| 91 |
-
|
| 92 |
-
- Markdown conversion : input (text/images) --> Markdown
|
| 93 |
-
|
| 94 |
-
- Multimodal inputs : text, images, or text + images.
|
| 95 |
-
|
| 96 |
-
- Multilingual documents.
|
| 97 |
-
|
| 98 |
-
- Reasoning and non-reasoning inference modes.
|
| 99 |
-
|
| 100 |
-
- Template generation for structured extraction from natural language or input document.
|
| 101 |
-
|
| 102 |
-
# [https://huggingface.co/numind/NuExtract3#benchmark-results](https://huggingface.co/numind/NuExtract3#benchmark-results) Benchmark results
|
| 103 |
-
|
| 104 |
-
## [https://huggingface.co/numind/NuExtract3#structured-extraction](https://huggingface.co/numind/NuExtract3#structured-extraction) Structured Extraction
|
| 105 |
-
|
| 106 |
-
We benchmarked NuExtract on NuMind's internal structured benchmark, measuring model's performances on ~600 documents of diverse types including invoices, movie posters or floor plans. These documents and their ground-truth cover diverse use-cases testing model visual understanding, OCR, reasoning and ability to handle long input and output contexts. We plan to open-source this benchmark in the coming weeks, along with a extensive leaderboard including most popular open-weight and closed-sourced APIs and a Python library allowing to easily measure model performances on structured extraction.
|
| 107 |
-
|
| 108 |
-
To measure a pair of predicted and ground-truth JSONs, we represent both as trees which we align based on node names, compute metric scores for aligned leaves and report the average of these scores. `string` and `verbatim-string` leaves are evaluated with indel distance (i.e. Levenshtein without replacement), while all others are evaluated with exact-match. Models were evaluated using vllm, with a temperature of 0.25 and a maximum of 65000 output token (for both thinking and answer), which largely exceeds 22000 which is the number of tokens of the largest ground truth output.
|
| 109 |
-
|
| 110 |
-
Model name Average score Num. failed⁽¹⁾ Avg. num tokens thinking Avg. num tokens answer NuExtract3.4_4B-RL 0.651 ± 0.019 27 2036 1856 gemma-4-E4B-it 0.538 ± 0.023 31 3005 1287 Qwen3.5-9B 0.479 ± 0.030 170 22409 1257 Qwen3.5-4B 0.417 ± 0.031 229 27177 1201 GLM-4.6V-Flash 0.435 ± 0.026 153 2989 1357 Nemotron-3-Nano-Omni 0.387 ± 0.028 204 25827 522 Ministral-3-3B 0.240 ± 0.022 344 27586 362
|
| 111 |
-
|
| 112 |
-
(1) number of model outputs that were not JSON deserializable, either directly or by removing leading and trailing backticks.
|
| 113 |
-
95% confidence intervals computed using a nonparametric bootstrap over scores distributions.
|
| 114 |
-
|
| 115 |
-
The benchmark include samples containing multiple images resulting in large input context, and some with ground-truth containing large numbers of items to extract resulting in large outputs. We found that the reasoning of small models significantly negatively impact their performances. The reason is that many models ended up falling in repetition loops, hitting the output tokens limit and resulting in failed requests.
|
| 116 |
-
|
| 117 |
-
## [https://huggingface.co/numind/NuExtract3#document-to-markdown](https://huggingface.co/numind/NuExtract3#document-to-markdown) Document to Markdown
|
| 118 |
-
|
| 119 |
-
NuExtract can also convert document images into clean Markdown. Output will be Markdown for text (headers etc), HTML for tables, LaTeX for math and `<figure data-type="image" data-id="img_n"><img src="/NM-dev/model_card-A/resolve/main/img_n.png" alt="Detail description of the images"/> `
|
| 120 |
-
|
| 121 |
-
Modern, format-agnostic benchmarks for complex document understanding are limited, so we explored a new evaluation approach. We selected 100 documents with challenging layouts and tables, asked each model to convert them into a structured representation, then used Gemini 3 Flash to compare model outputs against the source document and choose the most accurate result. The rankings aligned with human votes, suggesting this is a promising method for evaluating document-to-Markdown capabilities. More details will be shared in an upcoming technical report. Here are some results:
|
| 122 |
-
|
| 123 |
-
### [https://huggingface.co/numind/NuExtract3#using-markdown-to-structured](https://huggingface.co/numind/NuExtract3#using-markdown-to-structured) Using "Markdown-to-structured"
|
| 124 |
-
|
| 125 |
-
To add other evaluate references, we used our structured extraction benchmark to evaluate models in a two-step fashion: convert the benchmark inputs to Markdown, then use Qwen3.6 27B to perform the structured extraction task on them. Intuitively, it allows to evaluate how models achieve to keep the input document content and layout: good models will allow the "structured extractor" model to perform better scores.
|
| 126 |
-
|
| 127 |
-
# [https://huggingface.co/numind/NuExtract3#using-nuextract](https://huggingface.co/numind/NuExtract3#using-nuextract) Using NuExtract
|
| 128 |
-
|
| 129 |
-
## [https://huggingface.co/numind/NuExtract3#structured-extraction-1](https://huggingface.co/numind/NuExtract3#structured-extraction-1) Structured extraction
|
| 130 |
-
|
| 131 |
-
Structured extraction takes as inputs:
|
| 132 |
-
|
| 133 |
-
1. An input document, which can be text, image, or both;
|
| 134 |
-
|
| 135 |
-
2. A JSON template describing the information to extract;
|
| 136 |
-
|
| 137 |
-
3. (Optional) Instructions, allowing to specify expected output formats or values, to provide with the `instructions` chat template kwarg;
|
| 138 |
-
|
| 139 |
-
4. (Optional) In-Context Learning (ICL) examples.
|
| 140 |
-
|
| 141 |
-
### [https://huggingface.co/numind/NuExtract3#input-json-template](https://huggingface.co/numind/NuExtract3#input-json-template) Input JSON template
|
| 142 |
-
|
| 143 |
-
NuExtract uses a input JSON template whose structure is identical to the output JSON. Its leaf values are specify the types of the output JSON leaves. For examples:
|
| 144 |
-
|
| 145 |
-
```
|
| 146 |
-
{
|
| 147 |
-
"invoice_number": "verbatim-string",
|
| 148 |
-
"invoice_date": "date",
|
| 149 |
-
"total_amount": "number",
|
| 150 |
-
"currency": "currency",
|
| 151 |
-
"line_items": [
|
| 152 |
-
{
|
| 153 |
-
"description": "verbatim-string",
|
| 154 |
-
"item_type": ["electronics", "clothing", "vehicle", "furniture", "other"],
|
| 155 |
-
"quantity": "integer",
|
| 156 |
-
"unit_price": "number",
|
| 157 |
-
"total": "number"
|
| 158 |
-
}
|
| 159 |
-
]
|
| 160 |
-
}
|
| 161 |
-
```
|
| 162 |
-
|
| 163 |
-
Supported template types include:
|
| 164 |
-
|
| 165 |
-
- `verbatim-string`: extract text exactly as it appears in the document;
|
| 166 |
-
|
| 167 |
-
- `string`: generic string field, allowing abstraction or light paraphrasing;
|
| 168 |
-
|
| 169 |
-
- `integer`: whole number;
|
| 170 |
-
|
| 171 |
-
- `number`: integer or decimal number;
|
| 172 |
-
|
| 173 |
-
- `date-time`: ISO-8601 date, time or date-time;
|
| 174 |
-
|
| 175 |
-
- Other specific types such as `data`, `time`, `country`, `currency`, `email` and so on. [For more details, read the complete types specifications and examples](https://huggingface.co/numind/NuExtract3/blob/main/TYPES.md)
|
| 176 |
-
|
| 177 |
-
Template constructors:
|
| 178 |
-
|
| 179 |
-
- Arrays, for example `["string"]`;
|
| 180 |
-
|
| 181 |
-
- Enums, for example `["yes", "no", "maybe"]`;
|
| 182 |
-
|
| 183 |
-
- Multi-enums (multiple possible values), for example `[["A", "B", "C"]]`.
|
| 184 |
-
|
| 185 |
-
If the model does not find relevant information for a field, it returns `null` or `[]`.
|
| 186 |
-
|
| 187 |
-
### [https://huggingface.co/numind/NuExtract3#converting-json-schema--pydantic-models-to-nuextract-template](https://huggingface.co/numind/NuExtract3#converting-json-schema--pydantic-models-to-nuextract-template) Converting JSON schema / Pydantic models to NuExtract template
|
| 188 |
-
|
| 189 |
-
Our Python SDK (`pip install numind`) offers a method to convert JSON schemas to NuExtract templates:
|
| 190 |
-
|
| 191 |
-
```
|
| 192 |
-
from typing import Literal
|
| 193 |
-
|
| 194 |
-
from pydantic import Field, BaseModel
|
| 195 |
-
from numind.nuextract_utils import convert_json_schema_to_nuextract_template
|
| 196 |
-
|
| 197 |
-
class HotelBooking(BaseModel):
|
| 198 |
-
city: str
|
| 199 |
-
check_in_date: str = Field(description="date")
|
| 200 |
-
check_out_date: str = Field(description="date")
|
| 201 |
-
number_of_guests: int
|
| 202 |
-
room_type: Literal["single", "double", "suite"]
|
| 203 |
-
|
| 204 |
-
template, dropped_branches = convert_json_schema_to_nuextract_template(
|
| 205 |
-
HotelBooking.model_json_schema()
|
| 206 |
-
)
|
| 207 |
-
|
| 208 |
-
# {'check_in_date': 'date', 'check_out_date': 'date', 'city': 'string', 'number_of_guests': 'integer', 'room_type': ['single', 'double', 'suite']}
|
| 209 |
-
```
|
| 210 |
-
|
| 211 |
-
## [https://huggingface.co/numind/NuExtract3#document-to-markdown-1](https://huggingface.co/numind/NuExtract3#document-to-markdown-1) Document-to-Markdown
|
| 212 |
-
|
| 213 |
-
NuExtract can also convert document images into clean Markdown. Output will be markdown for text (headers etc), html for tables, latex for mat and `<figure data-type="image" data-id="img_n"><img src="/NM-dev/model_card-A/resolve/main/img_n.png" alt="Detail description of the images"/> `
|
| 214 |
-
|
| 215 |
-
Markdown example:
|
| 216 |
-
|
| 217 |
-
```
|
| 218 |
-
<figure data-type="image" data-id="img_1">
|
| 219 |
-
<img src="/numind/NuExtract3/resolve/main/img_1.png" alt="Logo of Mobilier 2000 with contact information: Tél.: (418) 275-4232, 1654, boul. Marcotte, Roberval (Qc) G8H 2P2"/>
|
| 220 |
-
</figure>
|
| 221 |
-
|
| 222 |
-
# COMMANDE
|
| 223 |
-
**NUMÉRO 72259**
|
| 224 |
-
|
| 225 |
-
1
|
| 226 |
-
|
| 227 |
-
**Vendu à**
|
| 228 |
-
TREMBLAY ERIC
|
| 229 |
-
ERIC TREMBLAY
|
| 230 |
-
348 BOUL. DE L'ANSE
|
| 231 |
-
ROBERVAL
|
| 232 |
-
G8H 1Y9
|
| 233 |
-
|
| 234 |
-
**Livré à**
|
| 235 |
-
TREMBLAY ERIC
|
| 236 |
-
ERIC TREMBLAY
|
| 237 |
-
348 BOUL. DE L'ANSE
|
| 238 |
-
ROBERVAL
|
| 239 |
-
G8H 1Y9
|
| 240 |
-
|
| 241 |
-
<table>
|
| 242 |
-
<thead>
|
| 243 |
-
<tr>
|
| 244 |
-
<th># CLIENT</th>
|
| 245 |
-
<th>EXPÉDITEUR</th>
|
| 246 |
-
<th>TERME DE CRÉDIT</th>
|
| 247 |
-
<th>DATE</th>
|
| 248 |
-
</tr>
|
| 249 |
-
</thead>
|
| 250 |
-
<tbody>
|
| 251 |
-
<tr>
|
| 252 |
-
<td>2753133</td>
|
| 253 |
-
<td>Notre camion</td>
|
| 254 |
-
<td>à la livraison</td>
|
| 255 |
-
<td>22/06/2023</td>
|
| 256 |
-
</tr>
|
| 257 |
-
</tbody>
|
| 258 |
-
</table>
|
| 259 |
-
|
| 260 |
-
<table>
|
| 261 |
-
<thead>
|
| 262 |
-
<tr>
|
| 263 |
-
<th>NOM DU VENDEUR</th>
|
| 264 |
-
<th>VOTRE ÉCONOMIE !</th>
|
| 265 |
-
<th># COMMANDE</th>
|
| 266 |
-
</tr>
|
| 267 |
-
</thead>
|
| 268 |
-
<tbody>
|
| 269 |
-
<tr>
|
| 270 |
-
<td>Éric</td>
|
| 271 |
-
<td>0.00</td>
|
| 272 |
-
<td></td>
|
| 273 |
-
</tr>
|
| 274 |
-
</tbody>
|
| 275 |
-
</table>
|
| 276 |
-
```
|
| 277 |
-
|
| 278 |
-
## [https://huggingface.co/numind/NuExtract3#reasoning-and-non-reasoning-modes](https://huggingface.co/numind/NuExtract3#reasoning-and-non-reasoning-modes) Reasoning and non-reasoning modes
|
| 279 |
-
|
| 280 |
-
NuExtract supports both reasoning and non-reasoning inference.
|
| 281 |
-
|
| 282 |
-
### [https://huggingface.co/numind/NuExtract3#non-thinking-mode](https://huggingface.co/numind/NuExtract3#non-thinking-mode) Non-thinking mode
|
| 283 |
-
|
| 284 |
-
Use this for fast and deterministic extraction or Markdown conversion.
|
| 285 |
-
|
| 286 |
-
```
|
| 287 |
-
enable_thinking = False
|
| 288 |
-
temperature = 0.2
|
| 289 |
-
```
|
| 290 |
-
|
| 291 |
-
### [https://huggingface.co/numind/NuExtract3#thinking-mode](https://huggingface.co/numind/NuExtract3#thinking-mode) Thinking mode
|
| 292 |
-
|
| 293 |
-
Use this for difficult documents, complex layouts, ambiguous fields, or cases where the document structure requires additional reasoning.
|
| 294 |
-
|
| 295 |
-
```
|
| 296 |
-
enable_thinking = True
|
| 297 |
-
temperature = 0.6
|
| 298 |
-
```
|
| 299 |
-
|
| 300 |
-
For production extraction workloads, we recommend starting with non-reasoning mode and enabling reasoning only for difficult examples.
|
| 301 |
-
|
| 302 |
-
## [https://huggingface.co/numind/NuExtract3#vllm-deployment](https://huggingface.co/numind/NuExtract3#vllm-deployment) vLLM deployment
|
| 303 |
-
|
| 304 |
-
NuExtract can be served with vLLM using an OpenAI-compatible API.
|
| 305 |
-
|
| 306 |
-
```
|
| 307 |
-
vllm serve numind/NuExtract3 \
|
| 308 |
-
--trust-remote-code \
|
| 309 |
-
--limit-mm-per-prompt '{"image": 99, "video": 0}' \
|
| 310 |
-
--chat-template-content-format openai \
|
| 311 |
-
--generation-config vllm \
|
| 312 |
-
--max-model-len 131072 \
|
| 313 |
-
--speculative-config '{"method": "qwen3_next_mtp", "num_speculative_tokens": 2}'
|
| 314 |
-
```
|
| 315 |
-
|
| 316 |
-
### [https://huggingface.co/numind/NuExtract3#multi-token-prediction](https://huggingface.co/numind/NuExtract3#multi-token-prediction) Multi Token Prediction
|
| 317 |
-
|
| 318 |
-
The deployment commands above enable Multi Token Prediction (MTP) through vLLM speculative decoding:
|
| 319 |
-
|
| 320 |
-
```
|
| 321 |
-
--speculative-config '{"method": "qwen3_next_mtp", "num_speculative_tokens": 2}'
|
| 322 |
-
```
|
| 323 |
-
|
| 324 |
-
MTP can improve decoding throughput without changing the OpenAI-compatible request payload. You can tune `num_speculative_tokens` for your hardware and workload, or remove `--speculative-config` if your vLLM version or environment does not support this speculative decoding method.
|
| 325 |
-
|
| 326 |
-
If you encounter memory issues, reduce the maximum model length and the maximum number of images:
|
| 327 |
-
|
| 328 |
-
```
|
| 329 |
-
vllm serve numind/NuExtract-3 \
|
| 330 |
-
--trust-remote-code \
|
| 331 |
-
--limit-mm-per-prompt '{"image": 6, "video": 0}' \
|
| 332 |
-
--chat-template-content-format openai \
|
| 333 |
-
--generation-config vllm \
|
| 334 |
-
--max-model-len 16384 \
|
| 335 |
-
--speculative-config '{"method": "qwen3_next_mtp", "num_speculative_tokens": 2}'
|
| 336 |
-
```
|
| 337 |
-
|
| 338 |
-
## [https://huggingface.co/numind/NuExtract3#vllm-inference-structured-extraction-text](https://huggingface.co/numind/NuExtract3#vllm-inference-structured-extraction-text) vLLM inference: structured extraction: text
|
| 339 |
-
|
| 340 |
-
```
|
| 341 |
-
import json
|
| 342 |
-
from openai import OpenAI
|
| 343 |
-
|
| 344 |
-
client = OpenAI(
|
| 345 |
-
api_key="EMPTY",
|
| 346 |
-
base_url="http://localhost:8000/v1",
|
| 347 |
-
)
|
| 348 |
-
|
| 349 |
-
template = {
|
| 350 |
-
"store": "verbatim-string",
|
| 351 |
-
"date": "date-time",
|
| 352 |
-
"total": "number",
|
| 353 |
-
"currency": ["USD", "EUR", "GBP", "JPY", "Other"],
|
| 354 |
-
"items": [
|
| 355 |
-
{
|
| 356 |
-
"name": "verbatim-string",
|
| 357 |
-
"price": "number"
|
| 358 |
-
}
|
| 359 |
-
]
|
| 360 |
-
}
|
| 361 |
-
|
| 362 |
-
response = client.chat.completions.create(
|
| 363 |
-
model="numind/NuExtract3",
|
| 364 |
-
temperature=0.2,
|
| 365 |
-
messages=[
|
| 366 |
-
{
|
| 367 |
-
"role": "user",
|
| 368 |
-
"content": [
|
| 369 |
-
{
|
| 370 |
-
"type": "text",
|
| 371 |
-
"text": "Yesterday I bought apples and coffee at Trader Joe's for a total of $12.40."
|
| 372 |
-
}
|
| 373 |
-
],
|
| 374 |
-
}
|
| 375 |
-
],
|
| 376 |
-
extra_body={
|
| 377 |
-
"chat_template_kwargs": {
|
| 378 |
-
"template": json.dumps(template),
|
| 379 |
-
"instructions": "Specify the time for the `date` entry only if it is present, otherwise only output the date component.",
|
| 380 |
-
"enable_thinking": False
|
| 381 |
-
}
|
| 382 |
-
}
|
| 383 |
-
)
|
| 384 |
-
|
| 385 |
-
print(response.choices[0].message.content)
|
| 386 |
-
```
|
| 387 |
-
|
| 388 |
-
Example output:
|
| 389 |
-
|
| 390 |
-
```
|
| 391 |
-
{
|
| 392 |
-
"store": "Trader Joe's",
|
| 393 |
-
"date": null,
|
| 394 |
-
"total": 12.40,
|
| 395 |
-
"currency": "USD",
|
| 396 |
-
"items": [
|
| 397 |
-
{
|
| 398 |
-
"name": "apples",
|
| 399 |
-
"price": null
|
| 400 |
-
},
|
| 401 |
-
{
|
| 402 |
-
"name": "coffee",
|
| 403 |
-
"price": null
|
| 404 |
-
}
|
| 405 |
-
]
|
| 406 |
-
}
|
| 407 |
-
```
|
| 408 |
-
|
| 409 |
-
## [https://huggingface.co/numind/NuExtract3#vllm-inference-structured-extraction-image](https://huggingface.co/numind/NuExtract3#vllm-inference-structured-extraction-image) vLLM inference: structured extraction: image
|
| 410 |
-
|
| 411 |
-
```
|
| 412 |
-
import json
|
| 413 |
-
import base64
|
| 414 |
-
from openai import OpenAI
|
| 415 |
-
|
| 416 |
-
client = OpenAI(
|
| 417 |
-
api_key="EMPTY",
|
| 418 |
-
base_url="http://localhost:8000/v1",
|
| 419 |
-
)
|
| 420 |
-
|
| 421 |
-
def encode_image(image_path):
|
| 422 |
-
with open(image_path, "rb") as image_file:
|
| 423 |
-
return base64.b64encode(image_file.read()).decode("utf-8")
|
| 424 |
-
|
| 425 |
-
image_base64 = encode_image("receipt.png")
|
| 426 |
-
data_url = f"data:image/png;base64,{image_base64}"
|
| 427 |
-
|
| 428 |
-
template = {
|
| 429 |
-
"store": "verbatim-string",
|
| 430 |
-
"date": "date-time",
|
| 431 |
-
"total": "number",
|
| 432 |
-
"payment_method": "verbatim-string"
|
| 433 |
-
}
|
| 434 |
-
|
| 435 |
-
response = client.chat.completions.create(
|
| 436 |
-
model="numind/NuExtract3",
|
| 437 |
-
temperature=0.2,
|
| 438 |
-
messages=[
|
| 439 |
-
{
|
| 440 |
-
"role": "user",
|
| 441 |
-
"content": [
|
| 442 |
-
{
|
| 443 |
-
"type": "image_url",
|
| 444 |
-
"image_url": {"url": data_url}
|
| 445 |
-
}
|
| 446 |
-
],
|
| 447 |
-
}
|
| 448 |
-
],
|
| 449 |
-
extra_body={
|
| 450 |
-
"chat_template_kwargs": {
|
| 451 |
-
"template": json.dumps(template, indent=4),
|
| 452 |
-
"enable_thinking": False
|
| 453 |
-
}
|
| 454 |
-
}
|
| 455 |
-
)
|
| 456 |
-
|
| 457 |
-
print(response.choices[0].message.content)
|
| 458 |
-
```
|
| 459 |
-
|
| 460 |
-
Example output:
|
| 461 |
-
|
| 462 |
-
```
|
| 463 |
-
{
|
| 464 |
-
"store": "Trader Joe's",
|
| 465 |
-
"date": "2025-04-12",
|
| 466 |
-
"total": 42.85,
|
| 467 |
-
"payment_method": "Visa"
|
| 468 |
-
}
|
| 469 |
-
```
|
| 470 |
-
|
| 471 |
-
### [https://huggingface.co/numind/NuExtract3#multiple-page-pdf](https://huggingface.co/numind/NuExtract3#multiple-page-pdf) Multiple page PDF
|
| 472 |
-
|
| 473 |
-
You can render a PDF to one PNG image per page with PyMuPDF, then pass the images to vLLM in page order.
|
| 474 |
-
|
| 475 |
-
```
|
| 476 |
-
import base64
|
| 477 |
-
import json
|
| 478 |
-
|
| 479 |
-
import fitz # pip install pymupdf
|
| 480 |
-
from openai import OpenAI
|
| 481 |
-
|
| 482 |
-
client = OpenAI(
|
| 483 |
-
api_key="EMPTY",
|
| 484 |
-
base_url="http://localhost:8000/v1",
|
| 485 |
-
)
|
| 486 |
-
|
| 487 |
-
def pdf_to_png_data_urls(pdf_path, dpi=170):
|
| 488 |
-
data_urls = []
|
| 489 |
-
|
| 490 |
-
with fitz.open(pdf_path) as doc:
|
| 491 |
-
for page in doc:
|
| 492 |
-
pix = page.get_pixmap(dpi=dpi, alpha=False)
|
| 493 |
-
png_bytes = pix.tobytes("png")
|
| 494 |
-
png_base64 = base64.b64encode(png_bytes).decode("utf-8")
|
| 495 |
-
data_urls.append(f"data:image/png;base64,{png_base64}")
|
| 496 |
-
|
| 497 |
-
return data_urls
|
| 498 |
-
|
| 499 |
-
data_urls = pdf_to_png_data_urls("invoice.pdf", dpi=170)
|
| 500 |
-
|
| 501 |
-
template = {
|
| 502 |
-
"invoice_number": "verbatim-string",
|
| 503 |
-
"invoice_date": "date",
|
| 504 |
-
"total": "number",
|
| 505 |
-
"currency": "currency",
|
| 506 |
-
"line_items": [
|
| 507 |
-
{
|
| 508 |
-
"description": "verbatim-string",
|
| 509 |
-
"quantity": "number",
|
| 510 |
-
"unit_price": "number",
|
| 511 |
-
"total": "number"
|
| 512 |
-
}
|
| 513 |
-
]
|
| 514 |
-
}
|
| 515 |
-
|
| 516 |
-
response = client.chat.completions.create(
|
| 517 |
-
model="numind/NuExtract3",
|
| 518 |
-
temperature=0.2,
|
| 519 |
-
messages=[
|
| 520 |
-
{
|
| 521 |
-
"role": "user",
|
| 522 |
-
"content": [
|
| 523 |
-
{
|
| 524 |
-
"type": "image_url",
|
| 525 |
-
"image_url": {"url": data_url}
|
| 526 |
-
}
|
| 527 |
-
for data_url in data_urls
|
| 528 |
-
],
|
| 529 |
-
}
|
| 530 |
-
],
|
| 531 |
-
extra_body={
|
| 532 |
-
"chat_template_kwargs": {
|
| 533 |
-
"template": json.dumps(template, indent=4),
|
| 534 |
-
"enable_thinking": False
|
| 535 |
-
}
|
| 536 |
-
}
|
| 537 |
-
)
|
| 538 |
-
|
| 539 |
-
print(response.choices[0].message.content)
|
| 540 |
-
```
|
| 541 |
-
|
| 542 |
-
## [https://huggingface.co/numind/NuExtract3#vllm-inference-document-to-markdown](https://huggingface.co/numind/NuExtract3#vllm-inference-document-to-markdown) vLLM inference: document-to-Markdown
|
| 543 |
-
|
| 544 |
-
For Markdown OCR, use `mode="markdown"` or `mode="content"` without a template.
|
| 545 |
-
|
| 546 |
-
```
|
| 547 |
-
import base64
|
| 548 |
-
from openai import OpenAI
|
| 549 |
-
|
| 550 |
-
client = OpenAI(
|
| 551 |
-
api_key="EMPTY",
|
| 552 |
-
base_url="http://localhost:8000/v1",
|
| 553 |
-
)
|
| 554 |
-
|
| 555 |
-
def encode_image(image_path):
|
| 556 |
-
with open(image_path, "rb") as image_file:
|
| 557 |
-
return base64.b64encode(image_file.read()).decode("utf-8")
|
| 558 |
-
|
| 559 |
-
image_base64 = encode_image("document.png")
|
| 560 |
-
data_url = f"data:image/png;base64,{image_base64}"
|
| 561 |
-
|
| 562 |
-
response = client.chat.completions.create(
|
| 563 |
-
model="numind/NuExtract3",
|
| 564 |
-
temperature=1,
|
| 565 |
-
messages=[
|
| 566 |
-
{
|
| 567 |
-
"role": "user",
|
| 568 |
-
"content": [
|
| 569 |
-
{
|
| 570 |
-
"type": "image_url",
|
| 571 |
-
"image_url": {"url": data_url}
|
| 572 |
-
}
|
| 573 |
-
],
|
| 574 |
-
}
|
| 575 |
-
],
|
| 576 |
-
extra_body={
|
| 577 |
-
"chat_template_kwargs": {
|
| 578 |
-
"mode": "markdown",
|
| 579 |
-
"enable_thinking": False
|
| 580 |
-
}
|
| 581 |
-
}
|
| 582 |
-
)
|
| 583 |
-
|
| 584 |
-
print(response.choices[0].message.content)
|
| 585 |
-
```
|
| 586 |
-
|
| 587 |
-
## [https://huggingface.co/numind/NuExtract3#vllm-inference-reasoning-mode](https://huggingface.co/numind/NuExtract3#vllm-inference-reasoning-mode) vLLM inference: reasoning mode
|
| 588 |
-
|
| 589 |
-
Reasoning can be enabled for harder structured extraction or Markdown tasks.
|
| 590 |
-
|
| 591 |
-
```
|
| 592 |
-
response = client.chat.completions.create(
|
| 593 |
-
model="numind/NuExtract3",
|
| 594 |
-
temperature=1,
|
| 595 |
-
messages=[
|
| 596 |
-
{
|
| 597 |
-
"role": "user",
|
| 598 |
-
"content": [
|
| 599 |
-
{
|
| 600 |
-
"type": "image_url",
|
| 601 |
-
"image_url": {"url": data_url}
|
| 602 |
-
}
|
| 603 |
-
],
|
| 604 |
-
}
|
| 605 |
-
],
|
| 606 |
-
extra_body={
|
| 607 |
-
"chat_template_kwargs": {
|
| 608 |
-
"mode": "markdown",
|
| 609 |
-
"enable_thinking": True
|
| 610 |
-
}
|
| 611 |
-
}
|
| 612 |
-
)
|
| 613 |
-
|
| 614 |
-
result = response.choices[0].message.content
|
| 615 |
-
|
| 616 |
-
if "</think>" in result:
|
| 617 |
-
reasoning, answer = [part.strip() for part in result.split("</think>")]
|
| 618 |
-
else:
|
| 619 |
-
reasoning, answer = None, result
|
| 620 |
-
|
| 621 |
-
print(answer)
|
| 622 |
-
```
|
| 623 |
-
|
| 624 |
-
## [https://huggingface.co/numind/NuExtract3#in-context-examples-for-extraction](https://huggingface.co/numind/NuExtract3#in-context-examples-for-extraction) In-context examples for extraction
|
| 625 |
-
|
| 626 |
-
NuExtract supports in-context examples for structured extraction.
|
| 627 |
-
|
| 628 |
-
Examples are especially useful when the desired formatting is ambiguous or when the schema requires task-specific conventions. Examples can be provided by using `developer` messages, for which all items of the contents except the last one are the input, and the last one is the expected output.
|
| 629 |
-
|
| 630 |
-
```
|
| 631 |
-
import json
|
| 632 |
-
from openai import OpenAI
|
| 633 |
-
|
| 634 |
-
client = OpenAI(
|
| 635 |
-
api_key="EMPTY",
|
| 636 |
-
base_url="http://localhost:8000/v1",
|
| 637 |
-
)
|
| 638 |
-
|
| 639 |
-
template = {
|
| 640 |
-
"names": ["string"]
|
| 641 |
-
}
|
| 642 |
-
|
| 643 |
-
response = client.chat.completions.create(
|
| 644 |
-
model="numind/NuExtract3",
|
| 645 |
-
temperature=0.2,
|
| 646 |
-
messages=[
|
| 647 |
-
{
|
| 648 |
-
"role": "developer",
|
| 649 |
-
"content": [
|
| 650 |
-
{
|
| 651 |
-
"type": "text",
|
| 652 |
-
"text": "Stephen is the manager at Susan's store.",
|
| 653 |
-
},
|
| 654 |
-
{
|
| 655 |
-
"type": "text",
|
| 656 |
-
"text": "{\"names\": [\"-STEPHEN-\", \"-SUSAN-\"]}",
|
| 657 |
-
}
|
| 658 |
-
],
|
| 659 |
-
},
|
| 660 |
-
{
|
| 661 |
-
"role": "user",
|
| 662 |
-
"content": [
|
| 663 |
-
{
|
| 664 |
-
"type": "text",
|
| 665 |
-
"text": "John went to the restaurant with Mary. James went to the cinema."
|
| 666 |
-
}
|
| 667 |
-
],
|
| 668 |
-
}
|
| 669 |
-
],
|
| 670 |
-
extra_body={
|
| 671 |
-
"chat_template_kwargs": {
|
| 672 |
-
"template": json.dumps(template, indent=4),
|
| 673 |
-
"enable_thinking": False
|
| 674 |
-
}
|
| 675 |
-
}
|
| 676 |
-
)
|
| 677 |
-
|
| 678 |
-
print(response.choices[0].message.content)
|
| 679 |
-
```
|
| 680 |
-
|
| 681 |
-
Example output:
|
| 682 |
-
|
| 683 |
-
```
|
| 684 |
-
{
|
| 685 |
-
"names": ["-JOHN-", "-MARY-", "-JAMES-"]
|
| 686 |
-
}
|
| 687 |
-
```
|
| 688 |
-
|
| 689 |
-
## [https://huggingface.co/numind/NuExtract3#vllm-inference-template-generation](https://huggingface.co/numind/NuExtract3#vllm-inference-template-generation) vLLM inference: template generation
|
| 690 |
-
|
| 691 |
-
NuExtract can generate an extraction template from a natural language description.
|
| 692 |
-
|
| 693 |
-
```
|
| 694 |
-
from openai import OpenAI
|
| 695 |
-
|
| 696 |
-
client = OpenAI(
|
| 697 |
-
api_key="EMPTY",
|
| 698 |
-
base_url="http://localhost:8000/v1",
|
| 699 |
-
)
|
| 700 |
-
|
| 701 |
-
response = client.chat.completions.create(
|
| 702 |
-
model="numind/NuExtract3",
|
| 703 |
-
temperature=0.2,
|
| 704 |
-
messages=[
|
| 705 |
-
{
|
| 706 |
-
"role": "user",
|
| 707 |
-
"content": [
|
| 708 |
-
{
|
| 709 |
-
"type": "text",
|
| 710 |
-
"text": "I want to extract the key details from a rental contract."
|
| 711 |
-
}
|
| 712 |
-
],
|
| 713 |
-
}
|
| 714 |
-
],
|
| 715 |
-
extra_body={
|
| 716 |
-
"chat_template_kwargs": {
|
| 717 |
-
"mode": "template-generation"
|
| 718 |
-
}
|
| 719 |
-
}
|
| 720 |
-
)
|
| 721 |
-
|
| 722 |
-
print(response.choices[0].message.content)
|
| 723 |
-
```
|
| 724 |
-
|
| 725 |
-
Example output:
|
| 726 |
-
|
| 727 |
-
```
|
| 728 |
-
{
|
| 729 |
-
"contract_title": "verbatim-string",
|
| 730 |
-
"landlord": "verbatim-string",
|
| 731 |
-
"tenant": "verbatim-string",
|
| 732 |
-
"property_address": "verbatim-string",
|
| 733 |
-
"start_date": "date-time",
|
| 734 |
-
"end_date": "date-time",
|
| 735 |
-
"monthly_rent": "number",
|
| 736 |
-
"currency": "verbatim-string",
|
| 737 |
-
"deposit": "number",
|
| 738 |
-
"signatories": ["verbatim-string"]
|
| 739 |
-
}
|
| 740 |
-
```
|
| 741 |
-
|
| 742 |
-
## [https://huggingface.co/numind/NuExtract3#curl-examples](https://huggingface.co/numind/NuExtract3#curl-examples) Curl examples
|
| 743 |
-
|
| 744 |
-
The following examples assume that vLLM is running locally on port 8000. They use `jq` to build valid JSON request bodies without manually escaping the image data or template string.
|
| 745 |
-
|
| 746 |
-
### [https://huggingface.co/numind/NuExtract3#single-image-structured-extraction](https://huggingface.co/numind/NuExtract3#single-image-structured-extraction) Single image structured extraction
|
| 747 |
-
|
| 748 |
-
```
|
| 749 |
-
API_KEY="EMPTY"
|
| 750 |
-
IMAGE_BASE64_FILE=$(mktemp)
|
| 751 |
-
REQUEST_BODY_FILE=$(mktemp)
|
| 752 |
-
|
| 753 |
-
base64 < receipt.png | tr -d '\n' > "$IMAGE_BASE64_FILE"
|
| 754 |
-
|
| 755 |
-
TEMPLATE=$(cat <<'JSON'
|
| 756 |
-
{
|
| 757 |
-
"store": "verbatim-string",
|
| 758 |
-
"date": "date-time",
|
| 759 |
-
"total": "number",
|
| 760 |
-
"payment_method": "verbatim-string"
|
| 761 |
-
}
|
| 762 |
-
JSON
|
| 763 |
-
)
|
| 764 |
-
|
| 765 |
-
jq -n \
|
| 766 |
-
--rawfile image_base64 "$IMAGE_BASE64_FILE" \
|
| 767 |
-
--arg template "$TEMPLATE" \
|
| 768 |
-
'{
|
| 769 |
-
model: "numind/NuExtract3",
|
| 770 |
-
temperature: 0.6,
|
| 771 |
-
messages: [
|
| 772 |
-
{
|
| 773 |
-
role: "user",
|
| 774 |
-
content: [
|
| 775 |
-
{
|
| 776 |
-
type: "image_url",
|
| 777 |
-
image_url: {url: ("data:image/png;base64," + $image_base64)}
|
| 778 |
-
}
|
| 779 |
-
]
|
| 780 |
-
}
|
| 781 |
-
],
|
| 782 |
-
chat_template_kwargs: {
|
| 783 |
-
template: $template,
|
| 784 |
-
enable_thinking: false
|
| 785 |
-
}
|
| 786 |
-
}' > "$REQUEST_BODY_FILE"
|
| 787 |
-
|
| 788 |
-
curl http://localhost:8000/v1/chat/completions \
|
| 789 |
-
-H "Content-Type: application/json" \
|
| 790 |
-
-H "Authorization: Bearer $API_KEY" \
|
| 791 |
-
--data-binary "@$REQUEST_BODY_FILE"
|
| 792 |
-
|
| 793 |
-
rm "$IMAGE_BASE64_FILE" "$REQUEST_BODY_FILE"
|
| 794 |
-
```
|
| 795 |
-
|
| 796 |
-
### [https://huggingface.co/numind/NuExtract3#single-image-content-extraction](https://huggingface.co/numind/NuExtract3#single-image-content-extraction) Single image content extraction
|
| 797 |
-
|
| 798 |
-
```
|
| 799 |
-
API_KEY="EMPTY"
|
| 800 |
-
IMAGE_BASE64_FILE=$(mktemp)
|
| 801 |
-
REQUEST_BODY_FILE=$(mktemp)
|
| 802 |
-
|
| 803 |
-
base64 < document.png | tr -d '\n' > "$IMAGE_BASE64_FILE"
|
| 804 |
-
|
| 805 |
-
jq -n \
|
| 806 |
-
--rawfile image_base64 "$IMAGE_BASE64_FILE" \
|
| 807 |
-
'{
|
| 808 |
-
model: "numind/NuExtract3",
|
| 809 |
-
temperature: 0.6,
|
| 810 |
-
messages: [
|
| 811 |
-
{
|
| 812 |
-
role: "user",
|
| 813 |
-
content: [
|
| 814 |
-
{
|
| 815 |
-
type: "image_url",
|
| 816 |
-
image_url: {url: ("data:image/png;base64," + $image_base64)}
|
| 817 |
-
}
|
| 818 |
-
]
|
| 819 |
-
}
|
| 820 |
-
],
|
| 821 |
-
chat_template_kwargs: {
|
| 822 |
-
mode: "content",
|
| 823 |
-
enable_thinking: false
|
| 824 |
-
}
|
| 825 |
-
}' > "$REQUEST_BODY_FILE"
|
| 826 |
-
|
| 827 |
-
curl http://localhost:8000/v1/chat/completions \
|
| 828 |
-
-H "Content-Type: application/json" \
|
| 829 |
-
-H "Authorization: Bearer $API_KEY" \
|
| 830 |
-
--data-binary "@$REQUEST_BODY_FILE"
|
| 831 |
-
|
| 832 |
-
rm "$IMAGE_BASE64_FILE" "$REQUEST_BODY_FILE"
|
| 833 |
-
```
|
| 834 |
-
|
| 835 |
-
## [https://huggingface.co/numind/NuExtract3#transformers-example](https://huggingface.co/numind/NuExtract3#transformers-example) Transformers example
|
| 836 |
-
|
| 837 |
-
You can also run NuExtract directly with `transformers`. The same `template`, `mode`, and `enable_thinking` options are passed to `processor.apply_chat_template`.
|
| 838 |
-
|
| 839 |
-
```
|
| 840 |
-
import json
|
| 841 |
-
|
| 842 |
-
import torch
|
| 843 |
-
from PIL import Image
|
| 844 |
-
from transformers import AutoModelForImageTextToText, AutoProcessor
|
| 845 |
-
|
| 846 |
-
model_id = "numind/NuExtract3"
|
| 847 |
-
|
| 848 |
-
processor = AutoProcessor.from_pretrained(
|
| 849 |
-
model_id,
|
| 850 |
-
trust_remote_code=True,
|
| 851 |
-
)
|
| 852 |
-
model = AutoModelForImageTextToText.from_pretrained(
|
| 853 |
-
model_id,
|
| 854 |
-
dtype=torch.bfloat16,
|
| 855 |
-
device_map="auto",
|
| 856 |
-
trust_remote_code=True,
|
| 857 |
-
).eval()
|
| 858 |
-
|
| 859 |
-
def run_nuextract(messages, **chat_template_kwargs):
|
| 860 |
-
inputs = processor.apply_chat_template(
|
| 861 |
-
messages,
|
| 862 |
-
add_generation_prompt=True,
|
| 863 |
-
tokenize=True,
|
| 864 |
-
return_dict=True,
|
| 865 |
-
return_tensors="pt",
|
| 866 |
-
**chat_template_kwargs,
|
| 867 |
-
).to(model.device)
|
| 868 |
-
|
| 869 |
-
with torch.inference_mode():
|
| 870 |
-
generated_ids = model.generate(
|
| 871 |
-
**inputs,
|
| 872 |
-
max_new_tokens=4096,
|
| 873 |
-
do_sample=False,
|
| 874 |
-
)
|
| 875 |
-
|
| 876 |
-
generated_ids = generated_ids[:, inputs.input_ids.shape[1]:]
|
| 877 |
-
return processor.batch_decode(
|
| 878 |
-
generated_ids,
|
| 879 |
-
skip_special_tokens=True,
|
| 880 |
-
clean_up_tokenization_spaces=False,
|
| 881 |
-
)[0].strip()
|
| 882 |
-
|
| 883 |
-
# Single image structured extraction
|
| 884 |
-
receipt_image = Image.open("receipt.png").convert("RGB")
|
| 885 |
-
receipt_messages = [
|
| 886 |
-
{
|
| 887 |
-
"role": "user",
|
| 888 |
-
"content": [
|
| 889 |
-
{
|
| 890 |
-
"type": "image",
|
| 891 |
-
"image": receipt_image,
|
| 892 |
-
}
|
| 893 |
-
],
|
| 894 |
-
}
|
| 895 |
-
]
|
| 896 |
-
|
| 897 |
-
template = {
|
| 898 |
-
"store": "verbatim-string",
|
| 899 |
-
"date": "date-time",
|
| 900 |
-
"total": "number",
|
| 901 |
-
"payment_method": "verbatim-string"
|
| 902 |
-
}
|
| 903 |
-
|
| 904 |
-
structured_output = run_nuextract(
|
| 905 |
-
receipt_messages,
|
| 906 |
-
template=json.dumps(template, indent=4),
|
| 907 |
-
enable_thinking=False,
|
| 908 |
-
)
|
| 909 |
-
print(structured_output)
|
| 910 |
-
|
| 911 |
-
# Single image content extraction
|
| 912 |
-
document_image = Image.open("document.png").convert("RGB")
|
| 913 |
-
document_messages = [
|
| 914 |
-
{
|
| 915 |
-
"role": "user",
|
| 916 |
-
"content": [
|
| 917 |
-
{
|
| 918 |
-
"type": "image",
|
| 919 |
-
"image": document_image,
|
| 920 |
-
}
|
| 921 |
-
],
|
| 922 |
-
}
|
| 923 |
-
]
|
| 924 |
-
|
| 925 |
-
content_output = run_nuextract(
|
| 926 |
-
document_messages,
|
| 927 |
-
mode="content",
|
| 928 |
-
enable_thinking=False,
|
| 929 |
-
)
|
| 930 |
-
print(content_output)
|
| 931 |
-
```
|
| 932 |
-
|
| 933 |
-
Special thanks to the Lambda.ai team for the compute that made this project a success.
|
| 934 |
-
|
| 935 |
-
## [https://huggingface.co/numind/NuExtract3#citation](https://huggingface.co/numind/NuExtract3#citation) Citation
|
| 936 |
-
|
| 937 |
-
If you use NuExtract, please cite NuMind and link to the model page.
|
| 938 |
-
|
| 939 |
-
```
|
| 940 |
-
@misc{nuextract3,
|
| 941 |
-
title = {NuExtract3},
|
| 942 |
-
author = {NuMind},
|
| 943 |
-
year = {2026},
|
| 944 |
-
url = {https://nuextract.ai/}
|
| 945 |
-
}
|
| 946 |
-
```
|
| 947 |
-
|
| 948 |
-
Downloads last month 1,236,855
|
| 949 |
-
|
| 950 |
-
Safetensors[https://huggingface.co/docs/safetensors](https://huggingface.co/docs/safetensors)
|
| 951 |
-
|
| 952 |
-
Model size
|
| 953 |
-
|
| 954 |
-
5B params
|
| 955 |
-
|
| 956 |
-
Tensor type
|
| 957 |
-
|
| 958 |
-
BF16
|
| 959 |
-
|
| 960 |
-
·
|
| 961 |
-
|
| 962 |
-
Chat template
|
| 963 |
-
|
| 964 |
-
Files info
|
| 965 |
-
|
| 966 |
-
Inference Providers[NEW](https://huggingface.co/docs/inference-providers)
|
| 967 |
-
|
| 968 |
-
[Image-to-Text](https://huggingface.co/tasks/image-to-text)
|
| 969 |
-
|
| 970 |
-
This model isn't deployed by any Inference Provider.[🙋 3 Ask for provider support](https://huggingface.co/spaces/huggingface/InferenceSupport/discussions/10487)
|
| 971 |
-
|
| 972 |
-
## Model tree for numind/NuExtract3[https://huggingface.co/docs/hub/model-cards#specifying-a-base-model](https://huggingface.co/docs/hub/model-cards#specifying-a-base-model)
|
| 973 |
-
|
| 974 |
-
Base model
|
| 975 |
-
|
| 976 |
-
[Qwen/Qwen3.5-4B-Base](https://huggingface.co/Qwen/Qwen3.5-4B-Base)
|
| 977 |
-
|
| 978 |
-
Finetuned
|
| 979 |
-
|
| 980 |
-
[Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)
|
| 981 |
-
|
| 982 |
-
Finetuned
|
| 983 |
-
|
| 984 |
-
([392](https://huggingface.co/models?other=base_model:finetune:Qwen/Qwen3.5-4B) )
|
| 985 |
-
|
| 986 |
-
this model
|
| 987 |
-
|
| 988 |
-
Finetunes
|
| 989 |
-
|
| 990 |
-
[4 models](https://huggingface.co/models?other=base_model:finetune:numind/NuExtract3)
|
| 991 |
-
|
| 992 |
-
Quantizations
|
| 993 |
-
|
| 994 |
-
[11 models](https://huggingface.co/models?other=base_model:quantized:numind/NuExtract3)
|
| 995 |
-
|
| 996 |
-
## Spaces using numind/NuExtract3 9
|
| 997 |
-
|
| 998 |
-
## Collection including numind/NuExtract3
|
| 999 |
-
|
| 1000 |
-
####
|
| 1001 |
-
|
| 1002 |
-
[NuExtract3 Collection 12 items • Updated May 26 • 15](https://huggingface.co/collections/numind/nuextract3)
|
| 1003 |
-
|
| 1004 |
-
## Evaluation results [https://huggingface.co/docs/hub/eval-results](https://huggingface.co/docs/hub/eval-results)
|
| 1005 |
-
|
| 1006 |
-
- [allenai/olmOCR-bench](https://huggingface.co/datasets/allenai/olmOCR-bench) · Old Scans[View evaluation results](https://huggingface.co/numind/NuExtract3/blob/main/.eval_results/olmocrbench.yaml)
|
| 1007 |
-
|
| 1008 |
-
[source](https://github.com/davanstrien/ocr-bench/blob/99f7550c/experiments/olmocr-bench-oldscans/BENCHMARKING.md)[leaderboard](https://huggingface.co/datasets/allenai/olmOCR-bench?eval_result=numind/NuExtract3&leaderboard_task_id=old_scans)
|
| 1009 |
-
|
| 1010 |
-
37.8 *
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
.agents/references/nuextract3/2026-07-20/huggingface.co/numind/NuExtract3/blob/main/README.md
DELETED
|
@@ -1,1001 +0,0 @@
|
|
| 1 |
-
<!--
|
| 2 |
-
source: https://huggingface.co/numind/NuExtract3/blob/main/README.md
|
| 3 |
-
retrieved: 2026-07-20T17:29:25.828402+00:00
|
| 4 |
-
final_url: https://huggingface.co/numind/NuExtract3/blob/main/README.md
|
| 5 |
-
content_type: text/html
|
| 6 |
-
-->
|
| 7 |
-
|
| 8 |
-
# README.md · numind/NuExtract3 at main
|
| 9 |
-
|
| 10 |
-
[Hugging Face](https://huggingface.co/)
|
| 11 |
-
|
| 12 |
-
#
|
| 13 |
-
|
| 14 |
-
[https://huggingface.co/numind](https://huggingface.co/numind)
|
| 15 |
-
|
| 16 |
-
[numind](https://huggingface.co/numind)
|
| 17 |
-
|
| 18 |
-
/
|
| 19 |
-
|
| 20 |
-
[NuExtract3](https://huggingface.co/numind/NuExtract3)
|
| 21 |
-
|
| 22 |
-
like 295
|
| 23 |
-
|
| 24 |
-
Follow
|
| 25 |
-
|
| 26 |
-
NuMind 649
|
| 27 |
-
|
| 28 |
-
[Image-to-Text](https://huggingface.co/models?pipeline_tag=image-to-text)
|
| 29 |
-
|
| 30 |
-
[Transformers](https://huggingface.co/models?library=transformers)
|
| 31 |
-
|
| 32 |
-
[Safetensors](https://huggingface.co/models?library=safetensors)
|
| 33 |
-
|
| 34 |
-
[qwen3_5](https://huggingface.co/models?other=qwen3_5)
|
| 35 |
-
|
| 36 |
-
[image-text-to-text](https://huggingface.co/models?other=image-text-to-text)
|
| 37 |
-
|
| 38 |
-
[vision-language](https://huggingface.co/models?other=vision-language)
|
| 39 |
-
|
| 40 |
-
[vlm](https://huggingface.co/models?other=vlm)
|
| 41 |
-
|
| 42 |
-
[document-understanding](https://huggingface.co/models?other=document-understanding)
|
| 43 |
-
|
| 44 |
-
[structured-extraction](https://huggingface.co/models?other=structured-extraction)
|
| 45 |
-
|
| 46 |
-
[information-extraction](https://huggingface.co/models?other=information-extraction)
|
| 47 |
-
|
| 48 |
-
[ocr](https://huggingface.co/models?other=ocr)
|
| 49 |
-
|
| 50 |
-
[document-to-markdown](https://huggingface.co/models?other=document-to-markdown)
|
| 51 |
-
|
| 52 |
-
[markdown](https://huggingface.co/models?other=markdown)
|
| 53 |
-
|
| 54 |
-
[rag](https://huggingface.co/models?other=rag)
|
| 55 |
-
|
| 56 |
-
[reasoning](https://huggingface.co/models?other=reasoning)
|
| 57 |
-
|
| 58 |
-
[multilingual](https://huggingface.co/models?other=multilingual)
|
| 59 |
-
|
| 60 |
-
[conversational](https://huggingface.co/models?other=conversational)
|
| 61 |
-
|
| 62 |
-
[Eval Results](https://huggingface.co/models?other=eval-results)
|
| 63 |
-
|
| 64 |
-
License: apache-2.0
|
| 65 |
-
|
| 66 |
-
[Model card](https://huggingface.co/numind/NuExtract3)
|
| 67 |
-
|
| 68 |
-
[Files Files and versions xet](https://huggingface.co/numind/NuExtract3/tree/main)
|
| 69 |
-
|
| 70 |
-
[Community 2](https://huggingface.co/numind/NuExtract3/discussions)
|
| 71 |
-
|
| 72 |
-
Deploy
|
| 73 |
-
|
| 74 |
-
Copy to bucket new
|
| 75 |
-
|
| 76 |
-
Use this model
|
| 77 |
-
|
| 78 |
-
main
|
| 79 |
-
|
| 80 |
-
[NuExtract3](https://huggingface.co/numind/NuExtract3/tree/main) / README.md
|
| 81 |
-
|
| 82 |
-
[NathanFradet](https://huggingface.co/NathanFradet)
|
| 83 |
-
|
| 84 |
-
Update README.md
|
| 85 |
-
|
| 86 |
-
[acaf70e](https://huggingface.co/numind/NuExtract3/commit/acaf70ecff9c3dbbfcbae651b82b66a0d8dbd0c6) verified about 2 months ago
|
| 87 |
-
|
| 88 |
-
[preview](https://huggingface.co/numind/NuExtract3/blob/main/README.md)[code](https://huggingface.co/numind/NuExtract3/blob/main/README.md?code=true)
|
| 89 |
-
|
| 90 |
-
|
|
| 91 |
-
|
| 92 |
-
[Raw](https://huggingface.co/numind/NuExtract3/raw/main/README.md)
|
| 93 |
-
|
| 94 |
-
Download with hf CLI
|
| 95 |
-
|
| 96 |
-
Copy download link
|
| 97 |
-
|
| 98 |
-
[History](https://huggingface.co/numind/NuExtract3/commits/main/README.md)[Blame](https://huggingface.co/numind/NuExtract3/blame/main/README.md)[Contribute](https://huggingface.co/numind/NuExtract3/edit/main/README.md)[Delete](https://huggingface.co/numind/NuExtract3/delete/main/README.md)
|
| 99 |
-
|
| 100 |
-
Safe
|
| 101 |
-
|
| 102 |
-
25.5 kB
|
| 103 |
-
|
| 104 |
-
[metadata](https://huggingface.co/docs/hub/model-cards#model-card-metadata)
|
| 105 |
-
|
| 106 |
-
```
|
| 107 |
-
license: apache-2.0
|
| 108 |
-
license_link: https://huggingface.co/numind/NuExtract3/blob/main/LICENSE
|
| 109 |
-
library_name: transformers
|
| 110 |
-
pipeline_tag: image-to-text
|
| 111 |
-
tags:
|
| 112 |
-
- image-text-to-text
|
| 113 |
-
- transformers
|
| 114 |
-
- safetensors
|
| 115 |
-
- qwen3_5
|
| 116 |
-
- vision-language
|
| 117 |
-
- vlm
|
| 118 |
-
- document-understanding
|
| 119 |
-
- structured-extraction
|
| 120 |
-
- information-extraction
|
| 121 |
-
- ocr
|
| 122 |
-
- document-to-markdown
|
| 123 |
-
- markdown
|
| 124 |
-
- rag
|
| 125 |
-
- reasoning
|
| 126 |
-
- multilingual
|
| 127 |
-
- conversational
|
| 128 |
-
base_model:
|
| 129 |
-
- Qwen/Qwen3.5-4B
|
| 130 |
-
model_name: NuExtract3
|
| 131 |
-
```
|
| 132 |
-
|
| 133 |
-
[https://nuextract.ai/](https://nuextract.ai/)
|
| 134 |
-
|
| 135 |
-
🖥️ [API / Platform](https://nuextract.ai/) | 📑 [Blog](https://numind.ai/blog) | 🗣️ [Discord](https://discord.gg/3tsEtJNCDe) | 🛠️ [GitHub](https://github.com/numindai/nuextract)
|
| 136 |
-
|
| 137 |
-
NuExtract3 is a unified 4B vision-language reasoning model for document understanding.
|
| 138 |
-
|
| 139 |
-
It combines strong structured information extraction with high-quality image-to-Markdown conversion, making it suitable for extraction pipelines, OCR, and RAG preprocessing for all types of documents such as scans, receipts, forms, invoices, contracts or tables.
|
| 140 |
-
|
| 141 |
-
Try it out in [the 🤗 space!](https://huggingface.co/spaces/numind/NuExtract-3-4B)
|
| 142 |
-
|
| 143 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#overview](https://huggingface.co/numind/NuExtract3/blob/main/README.md#overview) Overview
|
| 144 |
-
|
| 145 |
-
- Structured extraction : input (text/images) + JSON template + instructions --> JSON output
|
| 146 |
-
|
| 147 |
-
- Markdown conversion : input (text/images) --> Markdown
|
| 148 |
-
|
| 149 |
-
- Multimodal inputs : text, images, or text + images.
|
| 150 |
-
|
| 151 |
-
- Multilingual documents.
|
| 152 |
-
|
| 153 |
-
- Reasoning and non-reasoning inference modes.
|
| 154 |
-
|
| 155 |
-
- Template generation for structured extraction from natural language or input document.
|
| 156 |
-
|
| 157 |
-
# [https://huggingface.co/numind/NuExtract3/blob/main/README.md#benchmark-results](https://huggingface.co/numind/NuExtract3/blob/main/README.md#benchmark-results) Benchmark results
|
| 158 |
-
|
| 159 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#structured-extraction](https://huggingface.co/numind/NuExtract3/blob/main/README.md#structured-extraction) Structured Extraction
|
| 160 |
-
|
| 161 |
-
We benchmarked NuExtract on NuMind's internal structured benchmark, measuring model's performances on ~600 documents of diverse types including invoices, movie posters or floor plans. These documents and their ground-truth cover diverse use-cases testing model visual understanding, OCR, reasoning and ability to handle long input and output contexts. We plan to open-source this benchmark in the coming weeks, along with a extensive leaderboard including most popular open-weight and closed-sourced APIs and a Python library allowing to easily measure model performances on structured extraction.
|
| 162 |
-
|
| 163 |
-
To measure a pair of predicted and ground-truth JSONs, we represent both as trees which we align based on node names, compute metric scores for aligned leaves and report the average of these scores. `string` and `verbatim-string` leaves are evaluated with indel distance (i.e. Levenshtein without replacement), while all others are evaluated with exact-match. Models were evaluated using vllm, with a temperature of 0.25 and a maximum of 65000 output token (for both thinking and answer), which largely exceeds 22000 which is the number of tokens of the largest ground truth output.
|
| 164 |
-
|
| 165 |
-
Model name Average score Num. failed⁽¹⁾ Avg. num tokens thinking Avg. num tokens answer NuExtract3.4_4B-RL 0.651 ± 0.019 27 2036 1856 gemma-4-E4B-it 0.538 ± 0.023 31 3005 1287 Qwen3.5-9B 0.479 ± 0.030 170 22409 1257 Qwen3.5-4B 0.417 ± 0.031 229 27177 1201 GLM-4.6V-Flash 0.435 ± 0.026 153 2989 1357 Nemotron-3-Nano-Omni 0.387 ± 0.028 204 25827 522 Ministral-3-3B 0.240 ± 0.022 344 27586 362
|
| 166 |
-
|
| 167 |
-
(1) number of model outputs that were not JSON deserializable, either directly or by removing leading and trailing backticks.
|
| 168 |
-
95% confidence intervals computed using a nonparametric bootstrap over scores distributions.
|
| 169 |
-
|
| 170 |
-
The benchmark include samples containing multiple images resulting in large input context, and some with ground-truth containing large numbers of items to extract resulting in large outputs. We found that the reasoning of small models significantly negatively impact their performances. The reason is that many models ended up falling in repetition loops, hitting the output tokens limit and resulting in failed requests.
|
| 171 |
-
|
| 172 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#document-to-markdown](https://huggingface.co/numind/NuExtract3/blob/main/README.md#document-to-markdown) Document to Markdown
|
| 173 |
-
|
| 174 |
-
NuExtract can also convert document images into clean Markdown. Output will be Markdown for text (headers etc), HTML for tables, LaTeX for math and `<figure data-type="image" data-id="img_n"><img src="/NM-dev/model_card-A/resolve/main/img_n.png" alt="Detail description of the images"/> `
|
| 175 |
-
|
| 176 |
-
Modern, format-agnostic benchmarks for complex document understanding are limited, so we explored a new evaluation approach. We selected 100 documents with challenging layouts and tables, asked each model to convert them into a structured representation, then used Gemini 3 Flash to compare model outputs against the source document and choose the most accurate result. The rankings aligned with human votes, suggesting this is a promising method for evaluating document-to-Markdown capabilities. More details will be shared in an upcoming technical report. Here are some results:
|
| 177 |
-
|
| 178 |
-
### [https://huggingface.co/numind/NuExtract3/blob/main/README.md#using-markdown-to-structured](https://huggingface.co/numind/NuExtract3/blob/main/README.md#using-markdown-to-structured) Using "Markdown-to-structured"
|
| 179 |
-
|
| 180 |
-
To add other evaluate references, we used our structured extraction benchmark to evaluate models in a two-step fashion: convert the benchmark inputs to Markdown, then use Qwen3.6 27B to perform the structured extraction task on them. Intuitively, it allows to evaluate how models achieve to keep the input document content and layout: good models will allow the "structured extractor" model to perform better scores.
|
| 181 |
-
|
| 182 |
-
# [https://huggingface.co/numind/NuExtract3/blob/main/README.md#using-nuextract](https://huggingface.co/numind/NuExtract3/blob/main/README.md#using-nuextract) Using NuExtract
|
| 183 |
-
|
| 184 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#structured-extraction-1](https://huggingface.co/numind/NuExtract3/blob/main/README.md#structured-extraction-1) Structured extraction
|
| 185 |
-
|
| 186 |
-
Structured extraction takes as inputs:
|
| 187 |
-
|
| 188 |
-
1. An input document, which can be text, image, or both;
|
| 189 |
-
|
| 190 |
-
2. A JSON template describing the information to extract;
|
| 191 |
-
|
| 192 |
-
3. (Optional) Instructions, allowing to specify expected output formats or values, to provide with the `instructions` chat template kwarg;
|
| 193 |
-
|
| 194 |
-
4. (Optional) In-Context Learning (ICL) examples.
|
| 195 |
-
|
| 196 |
-
### [https://huggingface.co/numind/NuExtract3/blob/main/README.md#input-json-template](https://huggingface.co/numind/NuExtract3/blob/main/README.md#input-json-template) Input JSON template
|
| 197 |
-
|
| 198 |
-
NuExtract uses a input JSON template whose structure is identical to the output JSON. Its leaf values are specify the types of the output JSON leaves. For examples:
|
| 199 |
-
|
| 200 |
-
```
|
| 201 |
-
{
|
| 202 |
-
"invoice_number": "verbatim-string",
|
| 203 |
-
"invoice_date": "date",
|
| 204 |
-
"total_amount": "number",
|
| 205 |
-
"currency": "currency",
|
| 206 |
-
"line_items": [
|
| 207 |
-
{
|
| 208 |
-
"description": "verbatim-string",
|
| 209 |
-
"item_type": ["electronics", "clothing", "vehicle", "furniture", "other"],
|
| 210 |
-
"quantity": "integer",
|
| 211 |
-
"unit_price": "number",
|
| 212 |
-
"total": "number"
|
| 213 |
-
}
|
| 214 |
-
]
|
| 215 |
-
}
|
| 216 |
-
```
|
| 217 |
-
|
| 218 |
-
Supported template types include:
|
| 219 |
-
|
| 220 |
-
- `verbatim-string`: extract text exactly as it appears in the document;
|
| 221 |
-
|
| 222 |
-
- `string`: generic string field, allowing abstraction or light paraphrasing;
|
| 223 |
-
|
| 224 |
-
- `integer`: whole number;
|
| 225 |
-
|
| 226 |
-
- `number`: integer or decimal number;
|
| 227 |
-
|
| 228 |
-
- `date-time`: ISO-8601 date, time or date-time;
|
| 229 |
-
|
| 230 |
-
- Other specific types such as `data`, `time`, `country`, `currency`, `email` and so on. [For more details, read the complete types specifications and examples](https://huggingface.co/numind/NuExtract3/blob/main/TYPES.md)
|
| 231 |
-
|
| 232 |
-
Template constructors:
|
| 233 |
-
|
| 234 |
-
- Arrays, for example `["string"]`;
|
| 235 |
-
|
| 236 |
-
- Enums, for example `["yes", "no", "maybe"]`;
|
| 237 |
-
|
| 238 |
-
- Multi-enums (multiple possible values), for example `[["A", "B", "C"]]`.
|
| 239 |
-
|
| 240 |
-
If the model does not find relevant information for a field, it returns `null` or `[]`.
|
| 241 |
-
|
| 242 |
-
### [https://huggingface.co/numind/NuExtract3/blob/main/README.md#converting-json-schema--pydantic-models-to-nuextract-template](https://huggingface.co/numind/NuExtract3/blob/main/README.md#converting-json-schema--pydantic-models-to-nuextract-template) Converting JSON schema / Pydantic models to NuExtract template
|
| 243 |
-
|
| 244 |
-
Our Python SDK (`pip install numind`) offers a method to convert JSON schemas to NuExtract templates:
|
| 245 |
-
|
| 246 |
-
```
|
| 247 |
-
from typing import Literal
|
| 248 |
-
|
| 249 |
-
from pydantic import Field, BaseModel
|
| 250 |
-
from numind.nuextract_utils import convert_json_schema_to_nuextract_template
|
| 251 |
-
|
| 252 |
-
class HotelBooking(BaseModel):
|
| 253 |
-
city: str
|
| 254 |
-
check_in_date: str = Field(description="date")
|
| 255 |
-
check_out_date: str = Field(description="date")
|
| 256 |
-
number_of_guests: int
|
| 257 |
-
room_type: Literal["single", "double", "suite"]
|
| 258 |
-
|
| 259 |
-
template, dropped_branches = convert_json_schema_to_nuextract_template(
|
| 260 |
-
HotelBooking.model_json_schema()
|
| 261 |
-
)
|
| 262 |
-
|
| 263 |
-
# {'check_in_date': 'date', 'check_out_date': 'date', 'city': 'string', 'number_of_guests': 'integer', 'room_type': ['single', 'double', 'suite']}
|
| 264 |
-
```
|
| 265 |
-
|
| 266 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#document-to-markdown-1](https://huggingface.co/numind/NuExtract3/blob/main/README.md#document-to-markdown-1) Document-to-Markdown
|
| 267 |
-
|
| 268 |
-
NuExtract can also convert document images into clean Markdown. Output will be markdown for text (headers etc), html for tables, latex for mat and `<figure data-type="image" data-id="img_n"><img src="/NM-dev/model_card-A/resolve/main/img_n.png" alt="Detail description of the images"/> `
|
| 269 |
-
|
| 270 |
-
Markdown example:
|
| 271 |
-
|
| 272 |
-
```
|
| 273 |
-
<figure data-type="image" data-id="img_1">
|
| 274 |
-
<img src="/numind/NuExtract3/resolve/main/img_1.png" alt="Logo of Mobilier 2000 with contact information: Tél.: (418) 275-4232, 1654, boul. Marcotte, Roberval (Qc) G8H 2P2"/>
|
| 275 |
-
</figure>
|
| 276 |
-
|
| 277 |
-
# COMMANDE
|
| 278 |
-
**NUMÉRO 72259**
|
| 279 |
-
|
| 280 |
-
1
|
| 281 |
-
|
| 282 |
-
**Vendu à**
|
| 283 |
-
TREMBLAY ERIC
|
| 284 |
-
ERIC TREMBLAY
|
| 285 |
-
348 BOUL. DE L'ANSE
|
| 286 |
-
ROBERVAL
|
| 287 |
-
G8H 1Y9
|
| 288 |
-
|
| 289 |
-
**Livré à**
|
| 290 |
-
TREMBLAY ERIC
|
| 291 |
-
ERIC TREMBLAY
|
| 292 |
-
348 BOUL. DE L'ANSE
|
| 293 |
-
ROBERVAL
|
| 294 |
-
G8H 1Y9
|
| 295 |
-
|
| 296 |
-
<table>
|
| 297 |
-
<thead>
|
| 298 |
-
<tr>
|
| 299 |
-
<th># CLIENT</th>
|
| 300 |
-
<th>EXPÉDITEUR</th>
|
| 301 |
-
<th>TERME DE CRÉDIT</th>
|
| 302 |
-
<th>DATE</th>
|
| 303 |
-
</tr>
|
| 304 |
-
</thead>
|
| 305 |
-
<tbody>
|
| 306 |
-
<tr>
|
| 307 |
-
<td>2753133</td>
|
| 308 |
-
<td>Notre camion</td>
|
| 309 |
-
<td>à la livraison</td>
|
| 310 |
-
<td>22/06/2023</td>
|
| 311 |
-
</tr>
|
| 312 |
-
</tbody>
|
| 313 |
-
</table>
|
| 314 |
-
|
| 315 |
-
<table>
|
| 316 |
-
<thead>
|
| 317 |
-
<tr>
|
| 318 |
-
<th>NOM DU VENDEUR</th>
|
| 319 |
-
<th>VOTRE ÉCONOMIE !</th>
|
| 320 |
-
<th># COMMANDE</th>
|
| 321 |
-
</tr>
|
| 322 |
-
</thead>
|
| 323 |
-
<tbody>
|
| 324 |
-
<tr>
|
| 325 |
-
<td>Éric</td>
|
| 326 |
-
<td>0.00</td>
|
| 327 |
-
<td></td>
|
| 328 |
-
</tr>
|
| 329 |
-
</tbody>
|
| 330 |
-
</table>
|
| 331 |
-
```
|
| 332 |
-
|
| 333 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#reasoning-and-non-reasoning-modes](https://huggingface.co/numind/NuExtract3/blob/main/README.md#reasoning-and-non-reasoning-modes) Reasoning and non-reasoning modes
|
| 334 |
-
|
| 335 |
-
NuExtract supports both reasoning and non-reasoning inference.
|
| 336 |
-
|
| 337 |
-
### [https://huggingface.co/numind/NuExtract3/blob/main/README.md#non-thinking-mode](https://huggingface.co/numind/NuExtract3/blob/main/README.md#non-thinking-mode) Non-thinking mode
|
| 338 |
-
|
| 339 |
-
Use this for fast and deterministic extraction or Markdown conversion.
|
| 340 |
-
|
| 341 |
-
```
|
| 342 |
-
enable_thinking = False
|
| 343 |
-
temperature = 0.2
|
| 344 |
-
```
|
| 345 |
-
|
| 346 |
-
### [https://huggingface.co/numind/NuExtract3/blob/main/README.md#thinking-mode](https://huggingface.co/numind/NuExtract3/blob/main/README.md#thinking-mode) Thinking mode
|
| 347 |
-
|
| 348 |
-
Use this for difficult documents, complex layouts, ambiguous fields, or cases where the document structure requires additional reasoning.
|
| 349 |
-
|
| 350 |
-
```
|
| 351 |
-
enable_thinking = True
|
| 352 |
-
temperature = 0.6
|
| 353 |
-
```
|
| 354 |
-
|
| 355 |
-
For production extraction workloads, we recommend starting with non-reasoning mode and enabling reasoning only for difficult examples.
|
| 356 |
-
|
| 357 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#vllm-deployment](https://huggingface.co/numind/NuExtract3/blob/main/README.md#vllm-deployment) vLLM deployment
|
| 358 |
-
|
| 359 |
-
NuExtract can be served with vLLM using an OpenAI-compatible API.
|
| 360 |
-
|
| 361 |
-
```
|
| 362 |
-
vllm serve numind/NuExtract3 \
|
| 363 |
-
--trust-remote-code \
|
| 364 |
-
--limit-mm-per-prompt '{"image": 99, "video": 0}' \
|
| 365 |
-
--chat-template-content-format openai \
|
| 366 |
-
--generation-config vllm \
|
| 367 |
-
--max-model-len 131072 \
|
| 368 |
-
--speculative-config '{"method": "qwen3_next_mtp", "num_speculative_tokens": 2}'
|
| 369 |
-
```
|
| 370 |
-
|
| 371 |
-
### [https://huggingface.co/numind/NuExtract3/blob/main/README.md#multi-token-prediction](https://huggingface.co/numind/NuExtract3/blob/main/README.md#multi-token-prediction) Multi Token Prediction
|
| 372 |
-
|
| 373 |
-
The deployment commands above enable Multi Token Prediction (MTP) through vLLM speculative decoding:
|
| 374 |
-
|
| 375 |
-
```
|
| 376 |
-
--speculative-config '{"method": "qwen3_next_mtp", "num_speculative_tokens": 2}'
|
| 377 |
-
```
|
| 378 |
-
|
| 379 |
-
MTP can improve decoding throughput without changing the OpenAI-compatible request payload. You can tune `num_speculative_tokens` for your hardware and workload, or remove `--speculative-config` if your vLLM version or environment does not support this speculative decoding method.
|
| 380 |
-
|
| 381 |
-
If you encounter memory issues, reduce the maximum model length and the maximum number of images:
|
| 382 |
-
|
| 383 |
-
```
|
| 384 |
-
vllm serve numind/NuExtract-3 \
|
| 385 |
-
--trust-remote-code \
|
| 386 |
-
--limit-mm-per-prompt '{"image": 6, "video": 0}' \
|
| 387 |
-
--chat-template-content-format openai \
|
| 388 |
-
--generation-config vllm \
|
| 389 |
-
--max-model-len 16384 \
|
| 390 |
-
--speculative-config '{"method": "qwen3_next_mtp", "num_speculative_tokens": 2}'
|
| 391 |
-
```
|
| 392 |
-
|
| 393 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#vllm-inference-structured-extraction-text](https://huggingface.co/numind/NuExtract3/blob/main/README.md#vllm-inference-structured-extraction-text) vLLM inference: structured extraction: text
|
| 394 |
-
|
| 395 |
-
```
|
| 396 |
-
import json
|
| 397 |
-
from openai import OpenAI
|
| 398 |
-
|
| 399 |
-
client = OpenAI(
|
| 400 |
-
api_key="EMPTY",
|
| 401 |
-
base_url="http://localhost:8000/v1",
|
| 402 |
-
)
|
| 403 |
-
|
| 404 |
-
template = {
|
| 405 |
-
"store": "verbatim-string",
|
| 406 |
-
"date": "date-time",
|
| 407 |
-
"total": "number",
|
| 408 |
-
"currency": ["USD", "EUR", "GBP", "JPY", "Other"],
|
| 409 |
-
"items": [
|
| 410 |
-
{
|
| 411 |
-
"name": "verbatim-string",
|
| 412 |
-
"price": "number"
|
| 413 |
-
}
|
| 414 |
-
]
|
| 415 |
-
}
|
| 416 |
-
|
| 417 |
-
response = client.chat.completions.create(
|
| 418 |
-
model="numind/NuExtract3",
|
| 419 |
-
temperature=0.2,
|
| 420 |
-
messages=[
|
| 421 |
-
{
|
| 422 |
-
"role": "user",
|
| 423 |
-
"content": [
|
| 424 |
-
{
|
| 425 |
-
"type": "text",
|
| 426 |
-
"text": "Yesterday I bought apples and coffee at Trader Joe's for a total of $12.40."
|
| 427 |
-
}
|
| 428 |
-
],
|
| 429 |
-
}
|
| 430 |
-
],
|
| 431 |
-
extra_body={
|
| 432 |
-
"chat_template_kwargs": {
|
| 433 |
-
"template": json.dumps(template),
|
| 434 |
-
"instructions": "Specify the time for the `date` entry only if it is present, otherwise only output the date component.",
|
| 435 |
-
"enable_thinking": False
|
| 436 |
-
}
|
| 437 |
-
}
|
| 438 |
-
)
|
| 439 |
-
|
| 440 |
-
print(response.choices[0].message.content)
|
| 441 |
-
```
|
| 442 |
-
|
| 443 |
-
Example output:
|
| 444 |
-
|
| 445 |
-
```
|
| 446 |
-
{
|
| 447 |
-
"store": "Trader Joe's",
|
| 448 |
-
"date": null,
|
| 449 |
-
"total": 12.40,
|
| 450 |
-
"currency": "USD",
|
| 451 |
-
"items": [
|
| 452 |
-
{
|
| 453 |
-
"name": "apples",
|
| 454 |
-
"price": null
|
| 455 |
-
},
|
| 456 |
-
{
|
| 457 |
-
"name": "coffee",
|
| 458 |
-
"price": null
|
| 459 |
-
}
|
| 460 |
-
]
|
| 461 |
-
}
|
| 462 |
-
```
|
| 463 |
-
|
| 464 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#vllm-inference-structured-extraction-image](https://huggingface.co/numind/NuExtract3/blob/main/README.md#vllm-inference-structured-extraction-image) vLLM inference: structured extraction: image
|
| 465 |
-
|
| 466 |
-
```
|
| 467 |
-
import json
|
| 468 |
-
import base64
|
| 469 |
-
from openai import OpenAI
|
| 470 |
-
|
| 471 |
-
client = OpenAI(
|
| 472 |
-
api_key="EMPTY",
|
| 473 |
-
base_url="http://localhost:8000/v1",
|
| 474 |
-
)
|
| 475 |
-
|
| 476 |
-
def encode_image(image_path):
|
| 477 |
-
with open(image_path, "rb") as image_file:
|
| 478 |
-
return base64.b64encode(image_file.read()).decode("utf-8")
|
| 479 |
-
|
| 480 |
-
image_base64 = encode_image("receipt.png")
|
| 481 |
-
data_url = f"data:image/png;base64,{image_base64}"
|
| 482 |
-
|
| 483 |
-
template = {
|
| 484 |
-
"store": "verbatim-string",
|
| 485 |
-
"date": "date-time",
|
| 486 |
-
"total": "number",
|
| 487 |
-
"payment_method": "verbatim-string"
|
| 488 |
-
}
|
| 489 |
-
|
| 490 |
-
response = client.chat.completions.create(
|
| 491 |
-
model="numind/NuExtract3",
|
| 492 |
-
temperature=0.2,
|
| 493 |
-
messages=[
|
| 494 |
-
{
|
| 495 |
-
"role": "user",
|
| 496 |
-
"content": [
|
| 497 |
-
{
|
| 498 |
-
"type": "image_url",
|
| 499 |
-
"image_url": {"url": data_url}
|
| 500 |
-
}
|
| 501 |
-
],
|
| 502 |
-
}
|
| 503 |
-
],
|
| 504 |
-
extra_body={
|
| 505 |
-
"chat_template_kwargs": {
|
| 506 |
-
"template": json.dumps(template, indent=4),
|
| 507 |
-
"enable_thinking": False
|
| 508 |
-
}
|
| 509 |
-
}
|
| 510 |
-
)
|
| 511 |
-
|
| 512 |
-
print(response.choices[0].message.content)
|
| 513 |
-
```
|
| 514 |
-
|
| 515 |
-
Example output:
|
| 516 |
-
|
| 517 |
-
```
|
| 518 |
-
{
|
| 519 |
-
"store": "Trader Joe's",
|
| 520 |
-
"date": "2025-04-12",
|
| 521 |
-
"total": 42.85,
|
| 522 |
-
"payment_method": "Visa"
|
| 523 |
-
}
|
| 524 |
-
```
|
| 525 |
-
|
| 526 |
-
### [https://huggingface.co/numind/NuExtract3/blob/main/README.md#multiple-page-pdf](https://huggingface.co/numind/NuExtract3/blob/main/README.md#multiple-page-pdf) Multiple page PDF
|
| 527 |
-
|
| 528 |
-
You can render a PDF to one PNG image per page with PyMuPDF, then pass the images to vLLM in page order.
|
| 529 |
-
|
| 530 |
-
```
|
| 531 |
-
import base64
|
| 532 |
-
import json
|
| 533 |
-
|
| 534 |
-
import fitz # pip install pymupdf
|
| 535 |
-
from openai import OpenAI
|
| 536 |
-
|
| 537 |
-
client = OpenAI(
|
| 538 |
-
api_key="EMPTY",
|
| 539 |
-
base_url="http://localhost:8000/v1",
|
| 540 |
-
)
|
| 541 |
-
|
| 542 |
-
def pdf_to_png_data_urls(pdf_path, dpi=170):
|
| 543 |
-
data_urls = []
|
| 544 |
-
|
| 545 |
-
with fitz.open(pdf_path) as doc:
|
| 546 |
-
for page in doc:
|
| 547 |
-
pix = page.get_pixmap(dpi=dpi, alpha=False)
|
| 548 |
-
png_bytes = pix.tobytes("png")
|
| 549 |
-
png_base64 = base64.b64encode(png_bytes).decode("utf-8")
|
| 550 |
-
data_urls.append(f"data:image/png;base64,{png_base64}")
|
| 551 |
-
|
| 552 |
-
return data_urls
|
| 553 |
-
|
| 554 |
-
data_urls = pdf_to_png_data_urls("invoice.pdf", dpi=170)
|
| 555 |
-
|
| 556 |
-
template = {
|
| 557 |
-
"invoice_number": "verbatim-string",
|
| 558 |
-
"invoice_date": "date",
|
| 559 |
-
"total": "number",
|
| 560 |
-
"currency": "currency",
|
| 561 |
-
"line_items": [
|
| 562 |
-
{
|
| 563 |
-
"description": "verbatim-string",
|
| 564 |
-
"quantity": "number",
|
| 565 |
-
"unit_price": "number",
|
| 566 |
-
"total": "number"
|
| 567 |
-
}
|
| 568 |
-
]
|
| 569 |
-
}
|
| 570 |
-
|
| 571 |
-
response = client.chat.completions.create(
|
| 572 |
-
model="numind/NuExtract3",
|
| 573 |
-
temperature=0.2,
|
| 574 |
-
messages=[
|
| 575 |
-
{
|
| 576 |
-
"role": "user",
|
| 577 |
-
"content": [
|
| 578 |
-
{
|
| 579 |
-
"type": "image_url",
|
| 580 |
-
"image_url": {"url": data_url}
|
| 581 |
-
}
|
| 582 |
-
for data_url in data_urls
|
| 583 |
-
],
|
| 584 |
-
}
|
| 585 |
-
],
|
| 586 |
-
extra_body={
|
| 587 |
-
"chat_template_kwargs": {
|
| 588 |
-
"template": json.dumps(template, indent=4),
|
| 589 |
-
"enable_thinking": False
|
| 590 |
-
}
|
| 591 |
-
}
|
| 592 |
-
)
|
| 593 |
-
|
| 594 |
-
print(response.choices[0].message.content)
|
| 595 |
-
```
|
| 596 |
-
|
| 597 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#vllm-inference-document-to-markdown](https://huggingface.co/numind/NuExtract3/blob/main/README.md#vllm-inference-document-to-markdown) vLLM inference: document-to-Markdown
|
| 598 |
-
|
| 599 |
-
For Markdown OCR, use `mode="markdown"` or `mode="content"` without a template.
|
| 600 |
-
|
| 601 |
-
```
|
| 602 |
-
import base64
|
| 603 |
-
from openai import OpenAI
|
| 604 |
-
|
| 605 |
-
client = OpenAI(
|
| 606 |
-
api_key="EMPTY",
|
| 607 |
-
base_url="http://localhost:8000/v1",
|
| 608 |
-
)
|
| 609 |
-
|
| 610 |
-
def encode_image(image_path):
|
| 611 |
-
with open(image_path, "rb") as image_file:
|
| 612 |
-
return base64.b64encode(image_file.read()).decode("utf-8")
|
| 613 |
-
|
| 614 |
-
image_base64 = encode_image("document.png")
|
| 615 |
-
data_url = f"data:image/png;base64,{image_base64}"
|
| 616 |
-
|
| 617 |
-
response = client.chat.completions.create(
|
| 618 |
-
model="numind/NuExtract3",
|
| 619 |
-
temperature=1,
|
| 620 |
-
messages=[
|
| 621 |
-
{
|
| 622 |
-
"role": "user",
|
| 623 |
-
"content": [
|
| 624 |
-
{
|
| 625 |
-
"type": "image_url",
|
| 626 |
-
"image_url": {"url": data_url}
|
| 627 |
-
}
|
| 628 |
-
],
|
| 629 |
-
}
|
| 630 |
-
],
|
| 631 |
-
extra_body={
|
| 632 |
-
"chat_template_kwargs": {
|
| 633 |
-
"mode": "markdown",
|
| 634 |
-
"enable_thinking": False
|
| 635 |
-
}
|
| 636 |
-
}
|
| 637 |
-
)
|
| 638 |
-
|
| 639 |
-
print(response.choices[0].message.content)
|
| 640 |
-
```
|
| 641 |
-
|
| 642 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#vllm-inference-reasoning-mode](https://huggingface.co/numind/NuExtract3/blob/main/README.md#vllm-inference-reasoning-mode) vLLM inference: reasoning mode
|
| 643 |
-
|
| 644 |
-
Reasoning can be enabled for harder structured extraction or Markdown tasks.
|
| 645 |
-
|
| 646 |
-
```
|
| 647 |
-
response = client.chat.completions.create(
|
| 648 |
-
model="numind/NuExtract3",
|
| 649 |
-
temperature=1,
|
| 650 |
-
messages=[
|
| 651 |
-
{
|
| 652 |
-
"role": "user",
|
| 653 |
-
"content": [
|
| 654 |
-
{
|
| 655 |
-
"type": "image_url",
|
| 656 |
-
"image_url": {"url": data_url}
|
| 657 |
-
}
|
| 658 |
-
],
|
| 659 |
-
}
|
| 660 |
-
],
|
| 661 |
-
extra_body={
|
| 662 |
-
"chat_template_kwargs": {
|
| 663 |
-
"mode": "markdown",
|
| 664 |
-
"enable_thinking": True
|
| 665 |
-
}
|
| 666 |
-
}
|
| 667 |
-
)
|
| 668 |
-
|
| 669 |
-
result = response.choices[0].message.content
|
| 670 |
-
|
| 671 |
-
if "</think>" in result:
|
| 672 |
-
reasoning, answer = [part.strip() for part in result.split("</think>")]
|
| 673 |
-
else:
|
| 674 |
-
reasoning, answer = None, result
|
| 675 |
-
|
| 676 |
-
print(answer)
|
| 677 |
-
```
|
| 678 |
-
|
| 679 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#in-context-examples-for-extraction](https://huggingface.co/numind/NuExtract3/blob/main/README.md#in-context-examples-for-extraction) In-context examples for extraction
|
| 680 |
-
|
| 681 |
-
NuExtract supports in-context examples for structured extraction.
|
| 682 |
-
|
| 683 |
-
Examples are especially useful when the desired formatting is ambiguous or when the schema requires task-specific conventions. Examples can be provided by using `developer` messages, for which all items of the contents except the last one are the input, and the last one is the expected output.
|
| 684 |
-
|
| 685 |
-
```
|
| 686 |
-
import json
|
| 687 |
-
from openai import OpenAI
|
| 688 |
-
|
| 689 |
-
client = OpenAI(
|
| 690 |
-
api_key="EMPTY",
|
| 691 |
-
base_url="http://localhost:8000/v1",
|
| 692 |
-
)
|
| 693 |
-
|
| 694 |
-
template = {
|
| 695 |
-
"names": ["string"]
|
| 696 |
-
}
|
| 697 |
-
|
| 698 |
-
response = client.chat.completions.create(
|
| 699 |
-
model="numind/NuExtract3",
|
| 700 |
-
temperature=0.2,
|
| 701 |
-
messages=[
|
| 702 |
-
{
|
| 703 |
-
"role": "developer",
|
| 704 |
-
"content": [
|
| 705 |
-
{
|
| 706 |
-
"type": "text",
|
| 707 |
-
"text": "Stephen is the manager at Susan's store.",
|
| 708 |
-
},
|
| 709 |
-
{
|
| 710 |
-
"type": "text",
|
| 711 |
-
"text": "{\"names\": [\"-STEPHEN-\", \"-SUSAN-\"]}",
|
| 712 |
-
}
|
| 713 |
-
],
|
| 714 |
-
},
|
| 715 |
-
{
|
| 716 |
-
"role": "user",
|
| 717 |
-
"content": [
|
| 718 |
-
{
|
| 719 |
-
"type": "text",
|
| 720 |
-
"text": "John went to the restaurant with Mary. James went to the cinema."
|
| 721 |
-
}
|
| 722 |
-
],
|
| 723 |
-
}
|
| 724 |
-
],
|
| 725 |
-
extra_body={
|
| 726 |
-
"chat_template_kwargs": {
|
| 727 |
-
"template": json.dumps(template, indent=4),
|
| 728 |
-
"enable_thinking": False
|
| 729 |
-
}
|
| 730 |
-
}
|
| 731 |
-
)
|
| 732 |
-
|
| 733 |
-
print(response.choices[0].message.content)
|
| 734 |
-
```
|
| 735 |
-
|
| 736 |
-
Example output:
|
| 737 |
-
|
| 738 |
-
```
|
| 739 |
-
{
|
| 740 |
-
"names": ["-JOHN-", "-MARY-", "-JAMES-"]
|
| 741 |
-
}
|
| 742 |
-
```
|
| 743 |
-
|
| 744 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#vllm-inference-template-generation](https://huggingface.co/numind/NuExtract3/blob/main/README.md#vllm-inference-template-generation) vLLM inference: template generation
|
| 745 |
-
|
| 746 |
-
NuExtract can generate an extraction template from a natural language description.
|
| 747 |
-
|
| 748 |
-
```
|
| 749 |
-
from openai import OpenAI
|
| 750 |
-
|
| 751 |
-
client = OpenAI(
|
| 752 |
-
api_key="EMPTY",
|
| 753 |
-
base_url="http://localhost:8000/v1",
|
| 754 |
-
)
|
| 755 |
-
|
| 756 |
-
response = client.chat.completions.create(
|
| 757 |
-
model="numind/NuExtract3",
|
| 758 |
-
temperature=0.2,
|
| 759 |
-
messages=[
|
| 760 |
-
{
|
| 761 |
-
"role": "user",
|
| 762 |
-
"content": [
|
| 763 |
-
{
|
| 764 |
-
"type": "text",
|
| 765 |
-
"text": "I want to extract the key details from a rental contract."
|
| 766 |
-
}
|
| 767 |
-
],
|
| 768 |
-
}
|
| 769 |
-
],
|
| 770 |
-
extra_body={
|
| 771 |
-
"chat_template_kwargs": {
|
| 772 |
-
"mode": "template-generation"
|
| 773 |
-
}
|
| 774 |
-
}
|
| 775 |
-
)
|
| 776 |
-
|
| 777 |
-
print(response.choices[0].message.content)
|
| 778 |
-
```
|
| 779 |
-
|
| 780 |
-
Example output:
|
| 781 |
-
|
| 782 |
-
```
|
| 783 |
-
{
|
| 784 |
-
"contract_title": "verbatim-string",
|
| 785 |
-
"landlord": "verbatim-string",
|
| 786 |
-
"tenant": "verbatim-string",
|
| 787 |
-
"property_address": "verbatim-string",
|
| 788 |
-
"start_date": "date-time",
|
| 789 |
-
"end_date": "date-time",
|
| 790 |
-
"monthly_rent": "number",
|
| 791 |
-
"currency": "verbatim-string",
|
| 792 |
-
"deposit": "number",
|
| 793 |
-
"signatories": ["verbatim-string"]
|
| 794 |
-
}
|
| 795 |
-
```
|
| 796 |
-
|
| 797 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#curl-examples](https://huggingface.co/numind/NuExtract3/blob/main/README.md#curl-examples) Curl examples
|
| 798 |
-
|
| 799 |
-
The following examples assume that vLLM is running locally on port 8000. They use `jq` to build valid JSON request bodies without manually escaping the image data or template string.
|
| 800 |
-
|
| 801 |
-
### [https://huggingface.co/numind/NuExtract3/blob/main/README.md#single-image-structured-extraction](https://huggingface.co/numind/NuExtract3/blob/main/README.md#single-image-structured-extraction) Single image structured extraction
|
| 802 |
-
|
| 803 |
-
```
|
| 804 |
-
API_KEY="EMPTY"
|
| 805 |
-
IMAGE_BASE64_FILE=$(mktemp)
|
| 806 |
-
REQUEST_BODY_FILE=$(mktemp)
|
| 807 |
-
|
| 808 |
-
base64 < receipt.png | tr -d '\n' > "$IMAGE_BASE64_FILE"
|
| 809 |
-
|
| 810 |
-
TEMPLATE=$(cat <<'JSON'
|
| 811 |
-
{
|
| 812 |
-
"store": "verbatim-string",
|
| 813 |
-
"date": "date-time",
|
| 814 |
-
"total": "number",
|
| 815 |
-
"payment_method": "verbatim-string"
|
| 816 |
-
}
|
| 817 |
-
JSON
|
| 818 |
-
)
|
| 819 |
-
|
| 820 |
-
jq -n \
|
| 821 |
-
--rawfile image_base64 "$IMAGE_BASE64_FILE" \
|
| 822 |
-
--arg template "$TEMPLATE" \
|
| 823 |
-
'{
|
| 824 |
-
model: "numind/NuExtract3",
|
| 825 |
-
temperature: 0.6,
|
| 826 |
-
messages: [
|
| 827 |
-
{
|
| 828 |
-
role: "user",
|
| 829 |
-
content: [
|
| 830 |
-
{
|
| 831 |
-
type: "image_url",
|
| 832 |
-
image_url: {url: ("data:image/png;base64," + $image_base64)}
|
| 833 |
-
}
|
| 834 |
-
]
|
| 835 |
-
}
|
| 836 |
-
],
|
| 837 |
-
chat_template_kwargs: {
|
| 838 |
-
template: $template,
|
| 839 |
-
enable_thinking: false
|
| 840 |
-
}
|
| 841 |
-
}' > "$REQUEST_BODY_FILE"
|
| 842 |
-
|
| 843 |
-
curl http://localhost:8000/v1/chat/completions \
|
| 844 |
-
-H "Content-Type: application/json" \
|
| 845 |
-
-H "Authorization: Bearer $API_KEY" \
|
| 846 |
-
--data-binary "@$REQUEST_BODY_FILE"
|
| 847 |
-
|
| 848 |
-
rm "$IMAGE_BASE64_FILE" "$REQUEST_BODY_FILE"
|
| 849 |
-
```
|
| 850 |
-
|
| 851 |
-
### [https://huggingface.co/numind/NuExtract3/blob/main/README.md#single-image-content-extraction](https://huggingface.co/numind/NuExtract3/blob/main/README.md#single-image-content-extraction) Single image content extraction
|
| 852 |
-
|
| 853 |
-
```
|
| 854 |
-
API_KEY="EMPTY"
|
| 855 |
-
IMAGE_BASE64_FILE=$(mktemp)
|
| 856 |
-
REQUEST_BODY_FILE=$(mktemp)
|
| 857 |
-
|
| 858 |
-
base64 < document.png | tr -d '\n' > "$IMAGE_BASE64_FILE"
|
| 859 |
-
|
| 860 |
-
jq -n \
|
| 861 |
-
--rawfile image_base64 "$IMAGE_BASE64_FILE" \
|
| 862 |
-
'{
|
| 863 |
-
model: "numind/NuExtract3",
|
| 864 |
-
temperature: 0.6,
|
| 865 |
-
messages: [
|
| 866 |
-
{
|
| 867 |
-
role: "user",
|
| 868 |
-
content: [
|
| 869 |
-
{
|
| 870 |
-
type: "image_url",
|
| 871 |
-
image_url: {url: ("data:image/png;base64," + $image_base64)}
|
| 872 |
-
}
|
| 873 |
-
]
|
| 874 |
-
}
|
| 875 |
-
],
|
| 876 |
-
chat_template_kwargs: {
|
| 877 |
-
mode: "content",
|
| 878 |
-
enable_thinking: false
|
| 879 |
-
}
|
| 880 |
-
}' > "$REQUEST_BODY_FILE"
|
| 881 |
-
|
| 882 |
-
curl http://localhost:8000/v1/chat/completions \
|
| 883 |
-
-H "Content-Type: application/json" \
|
| 884 |
-
-H "Authorization: Bearer $API_KEY" \
|
| 885 |
-
--data-binary "@$REQUEST_BODY_FILE"
|
| 886 |
-
|
| 887 |
-
rm "$IMAGE_BASE64_FILE" "$REQUEST_BODY_FILE"
|
| 888 |
-
```
|
| 889 |
-
|
| 890 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#transformers-example](https://huggingface.co/numind/NuExtract3/blob/main/README.md#transformers-example) Transformers example
|
| 891 |
-
|
| 892 |
-
You can also run NuExtract directly with `transformers`. The same `template`, `mode`, and `enable_thinking` options are passed to `processor.apply_chat_template`.
|
| 893 |
-
|
| 894 |
-
```
|
| 895 |
-
import json
|
| 896 |
-
|
| 897 |
-
import torch
|
| 898 |
-
from PIL import Image
|
| 899 |
-
from transformers import AutoModelForImageTextToText, AutoProcessor
|
| 900 |
-
|
| 901 |
-
model_id = "numind/NuExtract3"
|
| 902 |
-
|
| 903 |
-
processor = AutoProcessor.from_pretrained(
|
| 904 |
-
model_id,
|
| 905 |
-
trust_remote_code=True,
|
| 906 |
-
)
|
| 907 |
-
model = AutoModelForImageTextToText.from_pretrained(
|
| 908 |
-
model_id,
|
| 909 |
-
dtype=torch.bfloat16,
|
| 910 |
-
device_map="auto",
|
| 911 |
-
trust_remote_code=True,
|
| 912 |
-
).eval()
|
| 913 |
-
|
| 914 |
-
def run_nuextract(messages, **chat_template_kwargs):
|
| 915 |
-
inputs = processor.apply_chat_template(
|
| 916 |
-
messages,
|
| 917 |
-
add_generation_prompt=True,
|
| 918 |
-
tokenize=True,
|
| 919 |
-
return_dict=True,
|
| 920 |
-
return_tensors="pt",
|
| 921 |
-
**chat_template_kwargs,
|
| 922 |
-
).to(model.device)
|
| 923 |
-
|
| 924 |
-
with torch.inference_mode():
|
| 925 |
-
generated_ids = model.generate(
|
| 926 |
-
**inputs,
|
| 927 |
-
max_new_tokens=4096,
|
| 928 |
-
do_sample=False,
|
| 929 |
-
)
|
| 930 |
-
|
| 931 |
-
generated_ids = generated_ids[:, inputs.input_ids.shape[1]:]
|
| 932 |
-
return processor.batch_decode(
|
| 933 |
-
generated_ids,
|
| 934 |
-
skip_special_tokens=True,
|
| 935 |
-
clean_up_tokenization_spaces=False,
|
| 936 |
-
)[0].strip()
|
| 937 |
-
|
| 938 |
-
# Single image structured extraction
|
| 939 |
-
receipt_image = Image.open("receipt.png").convert("RGB")
|
| 940 |
-
receipt_messages = [
|
| 941 |
-
{
|
| 942 |
-
"role": "user",
|
| 943 |
-
"content": [
|
| 944 |
-
{
|
| 945 |
-
"type": "image",
|
| 946 |
-
"image": receipt_image,
|
| 947 |
-
}
|
| 948 |
-
],
|
| 949 |
-
}
|
| 950 |
-
]
|
| 951 |
-
|
| 952 |
-
template = {
|
| 953 |
-
"store": "verbatim-string",
|
| 954 |
-
"date": "date-time",
|
| 955 |
-
"total": "number",
|
| 956 |
-
"payment_method": "verbatim-string"
|
| 957 |
-
}
|
| 958 |
-
|
| 959 |
-
structured_output = run_nuextract(
|
| 960 |
-
receipt_messages,
|
| 961 |
-
template=json.dumps(template, indent=4),
|
| 962 |
-
enable_thinking=False,
|
| 963 |
-
)
|
| 964 |
-
print(structured_output)
|
| 965 |
-
|
| 966 |
-
# Single image content extraction
|
| 967 |
-
document_image = Image.open("document.png").convert("RGB")
|
| 968 |
-
document_messages = [
|
| 969 |
-
{
|
| 970 |
-
"role": "user",
|
| 971 |
-
"content": [
|
| 972 |
-
{
|
| 973 |
-
"type": "image",
|
| 974 |
-
"image": document_image,
|
| 975 |
-
}
|
| 976 |
-
],
|
| 977 |
-
}
|
| 978 |
-
]
|
| 979 |
-
|
| 980 |
-
content_output = run_nuextract(
|
| 981 |
-
document_messages,
|
| 982 |
-
mode="content",
|
| 983 |
-
enable_thinking=False,
|
| 984 |
-
)
|
| 985 |
-
print(content_output)
|
| 986 |
-
```
|
| 987 |
-
|
| 988 |
-
Special thanks to the Lambda.ai team for the compute that made this project a success.
|
| 989 |
-
|
| 990 |
-
## [https://huggingface.co/numind/NuExtract3/blob/main/README.md#citation](https://huggingface.co/numind/NuExtract3/blob/main/README.md#citation) Citation
|
| 991 |
-
|
| 992 |
-
If you use NuExtract, please cite NuMind and link to the model page.
|
| 993 |
-
|
| 994 |
-
```
|
| 995 |
-
@misc{nuextract3,
|
| 996 |
-
title = {NuExtract3},
|
| 997 |
-
author = {NuMind},
|
| 998 |
-
year = {2026},
|
| 999 |
-
url = {https://nuextract.ai/}
|
| 1000 |
-
}
|
| 1001 |
-
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|