File size: 3,047 Bytes
1254239
ee7e5b7
 
 
 
 
1254239
 
ee7e5b7
1254239
ee7e5b7
 
 
 
1254239
ee7e5b7
1254239
ee7e5b7
 
1254239
ee7e5b7
 
 
 
 
1254239
ee7e5b7
1254239
ee7e5b7
1254239
ee7e5b7
 
1254239
ee7e5b7
 
 
 
 
 
 
 
1254239
ee7e5b7
 
1254239
ee7e5b7
1254239
ee7e5b7
 
 
 
1254239
ee7e5b7
 
 
 
 
1254239
ee7e5b7
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
---
license: gemma
base_model: unsloth/gemma-2-2b-it
library_name: peft
tags: [hebrew, lora, qlora, information-extraction, israeli-tax, text2receipt]
language: [he]
---

# Text2Receipt — Parser (LoRA adapter)

LoRA adapter over `unsloth/gemma-2-2b-it` that extracts a structured `parse` from a messy free-text **Hebrew**
income note. Part of the [Text2Receipt](https://huggingface.co/datasets/yonilev/Text2Receipt) project. The model
predicts **only** the linguistically-present fields (`client_name`, `client_is_business`, `items`); all fiscal
arithmetic (VAT, totals, allocation number, serials) is handled deterministically by `complete()`.

## Recommendation encoder (bake-off)

Three multilingual encoders were scored on Recall@k for same-category retrieval; the winner powers the
"similar past receipts" feature in the Space.

| model                                                       |   params_M |   dim |   encode_sec |   recall@1 |   recall@3 |   recall@5 |
|:------------------------------------------------------------|-----------:|------:|-------------:|-----------:|-----------:|-----------:|
| intfloat/multilingual-e5-small                              |      117.7 |   384 |         15   |     0.9675 |     0.934  |     0.9114 |
| sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 |      117.7 |   384 |         17.4 |     0.909  |     0.872  |     0.8484 |
| sentence-transformers/distiluse-base-multilingual-cased-v2  |      134.7 |   512 |         19.5 |     0.8765 |     0.8338 |     0.8083 |

**Winner:** `intfloat/multilingual-e5-small`

## Fine-tune vs baseline

Parse-extraction quality, baseline (zero-shot) vs LoRA fine-tune, across the in-distribution test, the
disjoint-vocabulary OOD test, and the hand-written human test:

| model                | split      |   valid_json |   exact_match |   field_f1 |
|:---------------------|:-----------|-------------:|--------------:|-----------:|
| baseline (zero-shot) | iid_test   |         0.85 |         0.23  |     0.6093 |
| baseline (zero-shot) | ood_test   |         0.9  |         0.17  |     0.6233 |
| baseline (zero-shot) | human_test |         0.8  |         0.08  |     0.5066 |
| fine-tuned (LoRA)    | iid_test   |         1    |         0.915 |     0.9735 |
| fine-tuned (LoRA)    | ood_test   |         1    |         0.325 |     0.8382 |
| fine-tuned (LoRA)    | human_test |         1    |         0.08  |     0.73   |

![bake-off](assets/embedding_bakeoff.png)
![fine-tune](assets/finetune_comparison.png)

## Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch

base = "unsloth/gemma-2-2b-it"
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, device_map="auto", torch_dtype=torch.float16)
model = PeftModel.from_pretrained(model, "yonilev/Text2Receipt-parser")
```

Adapter + embeddings artifacts (`receipts_store.parquet`, `receipts_embeddings.npy`,
`embeddings_manifest.json`) are in this repo for the application to consume.