File size: 6,564 Bytes
736f22a
77f7821
 
 
 
736f22a
2636347
736f22a
 
 
 
77f7821
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
04a481e
77f7821
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
04a481e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
77f7821
 
04a481e
 
 
 
 
 
 
77f7821
 
04a481e
 
 
77f7821
 
04a481e
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
---
title: Invoice & Receipt Extractor
emoji: 🧾
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 5.9.1
app_file: app.py
pinned: false
---

# Invoice & Receipt Extractor (Gemma 3 4B, ZeroGPU)

Upload an invoice/receipt image, get structured JSON back.

## Why 4B and not 1B?

`unsloth/gemma-3-1b-it-qat` is **text-only** β€” Gemma 3's SigLIP vision encoder was
only added to the 4B, 12B, and 27B sizes. This Space uses
`unsloth/gemma-3-4b-it-qat` instead, which is still small/fast enough to run
comfortably on a ZeroGPU slot.

## Deploying

1. Create a new Space on huggingface.co: **SDK = Gradio**.
2. You need a **HF PRO** subscription (or an Enterprise Hub org) for the
   **ZeroGPU** hardware option to appear β€” select it in the Space's Settings.
3. Push these three files (`app.py`, `requirements.txt`, this `README.md`) to
   the Space repo (via `git push`, the web UI, or `huggingface_hub`'s
   `upload_folder`).
4. Accept the Gemma license on the
   [google/gemma-3-4b-it](https://huggingface.co/google/gemma-3-4b-it) model
   page with the same account/org the Space runs as (unsloth's repo mirrors
   the same gated license).
5. First build will take a few minutes to download the model weights.

## Calling it via API

Gradio auto-generates an API for any Space with an `api_name` set (`extract`
here). Two easy ways to call it:

**Python, via `gradio_client`:**

```python
from gradio_client import Client, handle_file

client = Client("your-username/invoice-extractor-space")
result = client.predict(
    image=handle_file("receipt.jpg"),
    high_accuracy=False,  # set True only for long/narrow thermal-paper receipts
    api_name="/extract",
)
print(result)  # JSON string β€” see "Output schema & tax/item edge cases" below
```

## Image size handling

`app.py` downscales any incoming image to a 1568px long edge before it reaches
the model β€” Gemma's vision encoder resizes everything to a fixed 896x896
square internally anyway (256 tokens/pass), so nothing above that buys extra
accuracy. This guard protects against 12MP phone photos costing you upload
time and Space RAM for zero benefit.

If you're calling from a mobile client or somewhere bandwidth is tight, it's
still worth compressing/resizing client-side before upload (e.g. JPEG quality
~85, long edge ~1600px) β€” the server-side guard only kicks in after the full
file has already been transferred.

The `high_accuracy` checkbox turns on Gemma's `pan_and_scan`, which crops the
image into extra tiles instead of squashing it into one square β€” useful for
long thermal-paper receipts where line items near the top/bottom would
otherwise get compressed away. It costs roughly +256 tokens per extra tile,
so leave it off unless you're seeing missed line items.

**Raw HTTP**, if you'd rather not add the `gradio_client` dependency β€” click
"Use via API" at the bottom of your Space's page once it's live; it gives you
the exact `POST` endpoint and payload shape for your Space (Gradio's queueing
API requires a submit + poll call pair, which `gradio_client` handles for
you β€” that's the easier route for most use cases).

If the Space is private, pass `hf_token="hf_..."` to `Client(...)`.

## Output schema & tax/item edge cases

`app.py`'s prompt asks for a schema built to handle the messy realities of
real receipts, not just the happy path:

- **Multiple tax lines** (GST+PST, state+county, a VAT rate table) go in a
  `taxes: [{label, rate_percent, amount}]` array; `tax` stays as a
  convenience sum. A receipt with one combined tax figure just uses `tax`
  and leaves `taxes` empty.
- **`tax_inclusive`** flags whether listed prices already include tax
  (common outside the US) β€” matters if you're recomputing anything downstream.
- **Discounts and refunds** are captured as a positive `discount` amount
  (meant to be subtracted) at the document level, and as a *negative*
  `amount` on the specific line item if it's an itemized discount/return.
- **`service_charge` vs `tip`** are kept separate β€” a mandatory service fee
  isn't the same as a voluntary gratuity, and conflating them breaks
  downstream accounting.
- **Non-receipt images** (or unreadable ones) get `document_type: "unknown"`,
  everything else `null`, and a reason in `notes` β€” instead of the model
  forcing garbage into the schema.
- **`validation`** is appended after parsing, not part of what the model
  generates: it recomputes `subtotal` from `line_items`, sums `taxes`, adds
  `service_charge`/`tip`, subtracts `discount`, and compares against the
  stated `total` (Β±$0.02 tolerance). Use `matches_stated_total: false` as a
  signal to flag a document for manual review β€” it usually means either the
  model misread a digit or the receipt itself doesn't add up.
- If generation hits the token limit before finishing, `validation.warning`
  says so β€” treat line items near the end of the list as unverified rather
  than silently trusting a cut-off list.

This is inherently a **known-limitations tool, not a guarantee**: a 4B model
can still misread a smudged digit that happens to produce internally
consistent (but wrong) totals β€” `validation` only catches *inconsistency*,
not every possible misread. It also currently assumes **one document per
image** (a photo containing two separate receipts side by side isn't
handled) and doesn't support multi-page PDFs in a single call.

## Notes / tuning

- `@spaces.GPU(duration=90)` caps each call at 90s of GPU time β€” bumped up
  from 60s since `MAX_NEW_TOKENS` was raised to 1536 to fit long itemized
  receipts. Raise further if you still see timeouts on dense invoices.
- `MAX_NEW_TOKENS = 1536` in `app.py` β€” the ceiling on how long a generated
  JSON response can be. Grocery-length receipts (40+ line items) can get
  close to this; the truncation check (`validation.warning`) tells you when
  a response likely got cut off so you know to raise it further.
- `do_sample=False` (greedy decoding) is used for repeatability; extraction
  tasks don't benefit from sampling.
- JSON parsing has two fallback layers before giving up: stripping stray
  prose/fences around the `{...}`, then repairing trailing commas β€” the two
  most common small-model output quirks.
- The prompt in `app.py` defines the JSON schema. Edit it directly if you
  need extra fields (e.g. `po_number`, `tax_id`).
- For stricter production reliability, consider validating the model's JSON
  output against a `pydantic` schema and retrying once on failure/mismatch
  before falling back to returning the raw output for manual handling.