Spaces:
Sleeping
Sleeping
| title: Invoice & Receipt Extractor | |
| emoji: π§Ύ | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: gradio | |
| sdk_version: 5.9.1 | |
| app_file: app.py | |
| pinned: false | |
| # Invoice & Receipt Extractor (Gemma 3 4B, ZeroGPU) | |
| Upload an invoice/receipt image, get structured JSON back. | |
| ## Why 4B and not 1B? | |
| `unsloth/gemma-3-1b-it-qat` is **text-only** β Gemma 3's SigLIP vision encoder was | |
| only added to the 4B, 12B, and 27B sizes. This Space uses | |
| `unsloth/gemma-3-4b-it-qat` instead, which is still small/fast enough to run | |
| comfortably on a ZeroGPU slot. | |
| ## Deploying | |
| 1. Create a new Space on huggingface.co: **SDK = Gradio**. | |
| 2. You need a **HF PRO** subscription (or an Enterprise Hub org) for the | |
| **ZeroGPU** hardware option to appear β select it in the Space's Settings. | |
| 3. Push these three files (`app.py`, `requirements.txt`, this `README.md`) to | |
| the Space repo (via `git push`, the web UI, or `huggingface_hub`'s | |
| `upload_folder`). | |
| 4. Accept the Gemma license on the | |
| [google/gemma-3-4b-it](https://huggingface.co/google/gemma-3-4b-it) model | |
| page with the same account/org the Space runs as (unsloth's repo mirrors | |
| the same gated license). | |
| 5. First build will take a few minutes to download the model weights. | |
| ## Calling it via API | |
| Gradio auto-generates an API for any Space with an `api_name` set (`extract` | |
| here). Two easy ways to call it: | |
| **Python, via `gradio_client`:** | |
| ```python | |
| from gradio_client import Client, handle_file | |
| client = Client("your-username/invoice-extractor-space") | |
| result = client.predict( | |
| image=handle_file("receipt.jpg"), | |
| high_accuracy=False, # set True only for long/narrow thermal-paper receipts | |
| api_name="/extract", | |
| ) | |
| print(result) # JSON string β see "Output schema & tax/item edge cases" below | |
| ``` | |
| ## Image size handling | |
| `app.py` downscales any incoming image to a 1568px long edge before it reaches | |
| the model β Gemma's vision encoder resizes everything to a fixed 896x896 | |
| square internally anyway (256 tokens/pass), so nothing above that buys extra | |
| accuracy. This guard protects against 12MP phone photos costing you upload | |
| time and Space RAM for zero benefit. | |
| If you're calling from a mobile client or somewhere bandwidth is tight, it's | |
| still worth compressing/resizing client-side before upload (e.g. JPEG quality | |
| ~85, long edge ~1600px) β the server-side guard only kicks in after the full | |
| file has already been transferred. | |
| The `high_accuracy` checkbox turns on Gemma's `pan_and_scan`, which crops the | |
| image into extra tiles instead of squashing it into one square β useful for | |
| long thermal-paper receipts where line items near the top/bottom would | |
| otherwise get compressed away. It costs roughly +256 tokens per extra tile, | |
| so leave it off unless you're seeing missed line items. | |
| **Raw HTTP**, if you'd rather not add the `gradio_client` dependency β click | |
| "Use via API" at the bottom of your Space's page once it's live; it gives you | |
| the exact `POST` endpoint and payload shape for your Space (Gradio's queueing | |
| API requires a submit + poll call pair, which `gradio_client` handles for | |
| you β that's the easier route for most use cases). | |
| If the Space is private, pass `hf_token="hf_..."` to `Client(...)`. | |
| ## Output schema & tax/item edge cases | |
| `app.py`'s prompt asks for a schema built to handle the messy realities of | |
| real receipts, not just the happy path: | |
| - **Multiple tax lines** (GST+PST, state+county, a VAT rate table) go in a | |
| `taxes: [{label, rate_percent, amount}]` array; `tax` stays as a | |
| convenience sum. A receipt with one combined tax figure just uses `tax` | |
| and leaves `taxes` empty. | |
| - **`tax_inclusive`** flags whether listed prices already include tax | |
| (common outside the US) β matters if you're recomputing anything downstream. | |
| - **Discounts and refunds** are captured as a positive `discount` amount | |
| (meant to be subtracted) at the document level, and as a *negative* | |
| `amount` on the specific line item if it's an itemized discount/return. | |
| - **`service_charge` vs `tip`** are kept separate β a mandatory service fee | |
| isn't the same as a voluntary gratuity, and conflating them breaks | |
| downstream accounting. | |
| - **Non-receipt images** (or unreadable ones) get `document_type: "unknown"`, | |
| everything else `null`, and a reason in `notes` β instead of the model | |
| forcing garbage into the schema. | |
| - **`validation`** is appended after parsing, not part of what the model | |
| generates: it recomputes `subtotal` from `line_items`, sums `taxes`, adds | |
| `service_charge`/`tip`, subtracts `discount`, and compares against the | |
| stated `total` (Β±$0.02 tolerance). Use `matches_stated_total: false` as a | |
| signal to flag a document for manual review β it usually means either the | |
| model misread a digit or the receipt itself doesn't add up. | |
| - If generation hits the token limit before finishing, `validation.warning` | |
| says so β treat line items near the end of the list as unverified rather | |
| than silently trusting a cut-off list. | |
| This is inherently a **known-limitations tool, not a guarantee**: a 4B model | |
| can still misread a smudged digit that happens to produce internally | |
| consistent (but wrong) totals β `validation` only catches *inconsistency*, | |
| not every possible misread. It also currently assumes **one document per | |
| image** (a photo containing two separate receipts side by side isn't | |
| handled) and doesn't support multi-page PDFs in a single call. | |
| ## Notes / tuning | |
| - `@spaces.GPU(duration=90)` caps each call at 90s of GPU time β bumped up | |
| from 60s since `MAX_NEW_TOKENS` was raised to 1536 to fit long itemized | |
| receipts. Raise further if you still see timeouts on dense invoices. | |
| - `MAX_NEW_TOKENS = 1536` in `app.py` β the ceiling on how long a generated | |
| JSON response can be. Grocery-length receipts (40+ line items) can get | |
| close to this; the truncation check (`validation.warning`) tells you when | |
| a response likely got cut off so you know to raise it further. | |
| - `do_sample=False` (greedy decoding) is used for repeatability; extraction | |
| tasks don't benefit from sampling. | |
| - JSON parsing has two fallback layers before giving up: stripping stray | |
| prose/fences around the `{...}`, then repairing trailing commas β the two | |
| most common small-model output quirks. | |
| - The prompt in `app.py` defines the JSON schema. Edit it directly if you | |
| need extra fields (e.g. `po_number`, `tax_id`). | |
| - For stricter production reliability, consider validating the model's JSON | |
| output against a `pydantic` schema and retrying once on failure/mismatch | |
| before falling back to returning the raw output for manual handling. | |