File size: 10,989 Bytes
a463f96 2edb151 54349ae 177c99a 54349ae 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 177c99a 2edb151 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 | ---
license: apache-2.0
pretty_name: keys-Auto Receipts Studio
tags:
- receipts
- gemma-4
- vllm
- iphone
- autonomous-os
---
# keys-Auto Receipts Studio (iPhone / may add Autonomous Lamp Skill)
**v1.0 alpha**
iPhone (or desktop) photo → local **Gemma 4 12B-it** vision extract → SQLite. Optional [Autonomous OS](https://github.com/autonomous-ai/autonomous-os) Lamp skill (camera on the robot, 12B on the GPU box — the Lamp’s 6 GB cannot hold 12B).
GitHub: [`drowzeys/keys-Auto-Receipts-Studio`](https://github.com/drowzeys/keys-Auto-Receipts-Studio)
Hugging Face: [`drowzeys/keys-Auto-Receipts-Studio`](https://huggingface.co/drowzeys/keys-Auto-Receipts-Studio)
## Pathways (pick a recipe)
Gemma 4 12B-it **never** runs on the Lamp (6 GB) and typically **not** on a Windows/Mac laptop. The GPU box is the brain. PC / Mac / Linux run the inbox + Review UI. iPhone is a Safari capture client on the same Wi‑Fi.
| Pathway | One-shot recipe | Gemma 4 12B | iPhone page |
|---|---|---|---|
| **A. Linux NVIDIA GPU box** (Spark / Omen) | `bash oneshot.sh` | started here (`:8080`, util **0.15**, FP8, never above **0.85**) | `http://<this-box-lan-ip>:7860/phone` |
| **B. Linux, no NVIDIA** | `bash oneshot.sh` then edit `.env` | skip local vLLM; point `.env` at pathway A | same, this machine’s LAN IP |
| **C. Windows PC** | `oneshot.bat` then `scripts\install-launcher.bat` | on pathway A (or WSL2) | `http://<this-pc-lan-ip>:7860/phone` |
| **D. macOS** | venv + `bash scripts/install-launcher.sh` | on pathway A | `http://<this-mac-lan-ip>:7860/phone` |
| **E. iPhone** | Safari only, after A/B/C/D UI is up | n/a (uploads to the UI box) | `/phone` — **Take photo** or **Choose file** |
| **F. Autonomous Lamp** | `skills/keys-receipt-scanner/` | on pathway A | n/a (HAL camera → Studio `/api/inbox`) |
```bash
git clone https://github.com/drowzeys/keys-Auto-Receipts-Studio.git
cd keys-Auto-Receipts-Studio
```
Desktop icon after first UI start: Linux/macOS `bash scripts/install-launcher.sh` · Windows `scripts\install-launcher.bat`. Step-by-step for each letter is [below](#a-linux-nvidia-gpu-box-gemma--ui--iphone).
## Screenshots (v1.0 alpha)
iPhone Safari — take a photo, then completion notice:
<p>
<img src="docs/screenshots/05-iphone-home.jpg" alt="iPhone: Take photo / Choose file" width="280"/>
<img src="docs/screenshots/06-iphone-upload.jpg" alt="iPhone: upload in progress" width="280"/>
<img src="docs/screenshots/07-iphone-complete.jpg" alt="iPhone: Completed Whole Foods Market groceries #5" width="280"/>
</p>
Desktop Review queue (click a row to edit kind / category / vendor / date / total, or Delete):

Inbox drop zone, catalog SKUs, and backend health:
<p>
<img src="docs/screenshots/01-inbox.png" alt="Inbox" width="420"/>
<img src="docs/screenshots/03-catalog.png" alt="Catalog" width="420"/>
<img src="docs/screenshots/04-settings.png" alt="Settings backends" width="420"/>
</p>
## Promo (v1.0 alpha)
[~50s story cut](docs/promo/Keys-Auto-Receipts-Studio-promo.mp4) — messy dorm → iPhone (lay the receipt flat, then photograph) → charcoal Autonomous Lamp scans the last receipt → trash → clean room. Bed: Beethoven Symphony No. 5, I. Allegro con brio (public-domain recording).
Also attached on the [v1.0-alpha-promo release](https://github.com/drowzeys/keys-Auto-Receipts-Studio/releases/tag/v1.0-alpha-promo) (plays in the browser).
## Pathway details
One-click / `app.launch` binds the UI on the LAN (`RECEIPT_UI_SHARE_LAN=true`) and opens Review in the browser. Phone URL is printed in the terminal.
Desktop icons (do **not** copy `scripts/start-ui.bat` itself to the Desktop):
| OS | Command | Opens |
|---|---|---|
| Linux | `bash scripts/install-launcher.sh` | Desktop **Receipt Studio.desktop** |
| macOS | `bash scripts/install-launcher.sh` | Desktop **Receipt Studio.command** (first time: right-click → **Open**) |
| Windows | `scripts\install-launcher.bat` | Desktop **Receipt Studio.bat** |
### A. Linux NVIDIA GPU box (Gemma + UI + iPhone)
```bash
bash oneshot.sh
```
Idempotent. Installs `.venv`, fetches **google/gemma-4-12B-it** to `~/models-gemma4-12b-it` if missing (`hf auth login` once), starts `scripts/serve-gemma.sh` if `:8080` is down, starts the UI on **0.0.0.0:7860**.
When it prints READY:
| | |
|---|---|
| Review | http://127.0.0.1:7860 |
| iPhone Safari (same Wi‑Fi) | http://<this-box-lan-ip>:7860/phone |
| Gemma OpenAI-compat | http://127.0.0.1:8080/v1 |
vLLM flags (also used by `scripts/start-ui.sh` on a CUDA box if Gemma is down):
- `--gpu-memory-utilization **0.15**` (~18.3 GiB of 121.7 GiB; **never above 0.85**)
- **FP8** (BF16 weights ~23 GB cannot fit in that pool)
- `--max-model-len **8192**`
If Gemma is already running, oneshot / launch leave it alone. To apply the 0.15 cap: stop that `vllm` process, then `bash oneshot.sh` again.
Gemma-only (UI already running): `bash scripts/serve-gemma.sh`
### B. Linux without NVIDIA (UI + iPhone only)
```bash
bash oneshot.sh
```
If `vllm` is not on PATH and nothing answers on `:8080`, oneshot **skips** weights and vLLM (it does **not** die) and still starts the UI. Then set `.env` at the GPU box from pathway A:
```
RECEIPT_LLM_BASE_URL=http://<spark-lan-ip>:8080/v1
RECEIPT_EMBED_BASE_URL=http://<spark-lan-ip>:8080/v1
RECEIPT_LLM_MODEL=google/gemma-4-12B-it
RECEIPT_EMBED_MODEL=google/gemma-4-12B-it
RECEIPT_EMBED_DIM=3840
```
Restart the UI (`./scripts/start-ui.sh` or the desktop icon).
### C. Windows PC (UI + iPhone hotspot)
12B is **not** started by the Windows scripts. Serve Gemma on pathway A (or WSL2 + `bash oneshot.sh`).
1. Install [Python 3.12](https://www.python.org/downloads/) — check **Add python.exe to PATH**.
2. In the cloned repo, double-click **`oneshot.bat`** (creates `.venv`, opens Review).
3. Edit `.env` as in pathway B (Spark LAN IP, port **8080**).
4. Optional desktop icon: `scripts\install-launcher.bat` → **Desktop\Receipt Studio.bat**.
5. iPhone: `http://<this-windows-lan-ip>:7860/phone` in **Safari**.
Do not copy `scripts\start-ui.bat` to the Desktop; it must run from the repo (or from the wrapper `install-launcher.bat` writes).
If this PC **is** NVIDIA + `vllm` on PATH (Git Bash / WSL): point `.env` at `127.0.0.1:8080` and `bash scripts/serve-gemma.sh`.
### D. macOS (UI + iPhone)
Apple Silicon does not load 12B with this vLLM-NVIDIA recipe. Use pathway A for Gemma.
```bash
python3 -m venv .venv
.venv/bin/pip install -e ".[dev]"
cp .env.example .env
# edit RECEIPT_LLM_BASE_URL / RECEIPT_EMBED_BASE_URL like pathway B
bash scripts/install-launcher.sh
```
Double-click **Desktop → Receipt Studio.command** (first time: right-click → **Open**). Same as `./scripts/start-ui.sh`.
### E. iPhone (Safari)
Works against **whichever machine runs the UI** (A, B, C, or D), same Wi‑Fi. Chrome/Firefox on iOS are not the supported path.
1. Open `http://<ui-machine-lan-ip>:7860/phone`
2. **Take photo** (camera) or **Choose file** (Photos / Files, including HEIC)
3. Wait for **Completed: vendor · total · category · #id**
4. Confirm / edit / delete on the desktop **Review** tab
The phone page POSTs `/api/inbox` and polls `/api/jobs/{id}`. The GPU box (or the UI machine talking to Gemma) does extract + embed. Inbox drop and Syncthing into `inbox/` use the same pipeline (files idle 30s unless the phone API processes immediately).
### F. Autonomous Lamp skill
Repo path: **`skills/keys-receipt-scanner/`** (`SKILL.md`, `skill.json` `capabilities: ["vision"]`, `scripts/scan.py`).
```
skills/keys-receipt-scanner/
SKILL.md
skill.json
scripts/scan.py
references/hardware.md
```
Install (no reboot):
```bash
make push-skill SKILL=./skills/keys-receipt-scanner TARGET=pi@lamp-xxxx.local
```
On the Lamp (camera only — **do not** load 12B):
```bash
export RECEIPT_GPU_HOST=<spark-lan-ip>
export RECEIPT_STUDIO_URL=http://<spark-lan-ip>:7860
export RECEIPT_CAMERA_URL=http://127.0.0.1:5001
```
Say **“scan this receipt”**. `scripts/scan.py` snapshots HAL (`:5001/camera/snapshot`, 1280 px), then POSTs to Receipt Studio `/api/inbox` on the GPU box. Fallback: `RECEIPT_LLM_BASE_URL=http://<spark-lan-ip>:8080/v1` (direct vLLM) or `RECEIPT_HERMES_BASE_URL` (same OpenAI-compat URL Hermes registered). On the GPU box, `python scripts/register-hermes-gemma.py` adds `custom_providers` `gemma4-12b-it`.
Intern has no camera; this skill will not install there.
To ship in a fork of [autonomous-os](https://github.com/autonomous-ai/autonomous-os): copy `skills/keys-receipt-scanner/` into `skills/`, run `python skills/skill-creator/scripts/quick_validate.py`, `make skills-catalog`, open the PR.
---
## Does Gemma 4 12B Unified fit on the Lamp?
**No.** Lamp is 8-core ARM64 with **6 GB RAM**. Gemma 4 12B Unified is a 12B dense omni model (`Gemma4UnifiedForConditionalGeneration`, hidden size **3840**).
| Piece | Lamp (6 GB) | GPU box |
|---|---|---|
| `skills/keys-receipt-scanner/` | yes | yes |
| HAL snapshot | yes | n/a |
| SQLite + HTTP client | yes | yes |
| **Gemma 4 12B Unified weights** | **no** | yes (vLLM util **0.15**) |
## Gemma 4 12B Unified (omni) on the GPU box
One OpenAI-compat server: vision `POST /v1/chat/completions` **and** `POST /v1/embeddings` (dim **3840**). Do not raise util above **0.85**.
```bash
bash scripts/serve-gemma.sh
# util 0.15, FP8, max-model-len 8192
```
Context at util 0.15 (after ~12.5 GB FP8 weights):
| Estimate | Tokens |
|---|---|
| Conservative (48 layers full attn) | **~12k** |
| Hybrid (8 full + 40× sliding-1024) | **~65k** |
| Model native window | 262,144 (not at 0.15) |
A receipt scan uses ~1–2k tokens. Raise `RECEIPT_VLLM_MAX_MODEL_LEN` only after the boot log shows a large enough `GPU KV cache size`.
If `/v1/embeddings` 404s, either serve a pooling convert on another port (`RECEIPT_EMBED_BASE_URL`) or `RECEIPT_EMBED_BACKEND=nvidia` + Nemotron-3-Embed-1B at dim **2048**. **Never mix 3840 and 2048 in one DB.**
Fallback vision (not omni embed): Qwen3.8-27B ADay777 at `:8078` (`RECEIPT_LLM_BACKEND=nvidia`, `RECEIPT_LLM_MODEL=qwen38-nvfp4`). Lightning is **text-only** — never send images.
## Manual UI (any OS, after venv)
```bash
python3 -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
cp .env.example .env
# set RECEIPT_LLM_* like pathway B unless Gemma is local
RECEIPT_UI_SHARE_LAN=true python -m app.cli ui
```
CLI:
```bash
python -m app.cli scan --image path/to/receipt.jpg
python -m app.cli query --category groceries
pytest
```
Syncthing: camera/share folder → `inbox/` (idle **30s** then process).
## What you still run yourself
- Pathway A (or equivalent) so Gemma is reachable on the LAN
- iPhone on the **same Wi‑Fi**, Safari
- Lamp: pair in the Autonomous app, `make push-skill`, set `RECEIPT_STUDIO_URL`
- Optional: Syncthing
No model weights in this repo. No PyInstaller in this release.
|