pageparse.ai / README.md
Varun2007's picture
initial clean deployment commit with compilers
8c3e275
|
Raw
History Blame Contribute Delete
15.5 kB
---
title: PageParse
emoji: πŸ“„
colorFrom: blue
colorTo: green
sdk: docker
app_port: 7860
---
<!--
SPDX-FileCopyrightText: 2026 Team Centurions
SPDX-License-Identifier: AGPL-3.0-or-later
-->
# PageParse
**Any handwritten page β†’ clean structured data. On CPU. Fully offline.**
PageParse reads a scanned or photographed page of handwritten notes entirely on
your own device and turns it into structured, schema-validated records in a local
database β€” no GPU, no CUDA, no cloud. Point it at a handwritten page; it cleans the
image, recognises the ink on CPU, maps the messy text to a defined JSON schema, and
saves it to SQLite. It keeps working with the Wi-Fi switched off.
> Built for **The CPU-First Hackathon** β€” _"Build AI that runs anywhere."_
<p>
<img alt="CPU-first" src="https://img.shields.io/badge/inference-CPU--only-blue">
<img alt="Offline" src="https://img.shields.io/badge/network-offline--first-green">
<img alt="License" src="https://img.shields.io/badge/license-AGPL--3.0-orange">
<img alt="Python" src="https://img.shields.io/badge/python-3.11-yellow">
</p>
---
## Table of Contents
- [The Idea](#the-idea)
- [Why It Matters](#why-it-matters)
- [How It Works](#how-it-works)
- [Architecture](#architecture)
- [Model & Runtime (CPU-first)](#model--runtime-cpu-first)
- [Output Schema](#output-schema)
- [Data Model (SQLite)](#data-model-sqlite)
- [Tech Stack](#tech-stack)
- [Project Structure](#project-structure)
- [Getting Started](#getting-started)
- [Usage](#usage)
- [Offline-First by Design](#offline-first-by-design)
- [Quality Gates (CI/CD)](#quality-gates-cicd)
- [Judging Criteria Mapping](#judging-criteria-mapping)
- [Roadmap](#roadmap)
- [Team & Work Division](#team--work-division)
- [Contributing](#contributing)
- [License](#license)
---
## The Idea
People still write on paper β€” to-do lists, planners, field surveys, recipe cards,
clinic notes. That ink almost never makes it into a system you can search, sort,
or query. **PageParse bridges paper and database, entirely on-device.**
Drop in a scan or photo of a handwritten page. PageParse:
1. **reads** the handwriting on CPU,
2. **maps** it to a defined, relational schema using a small local language model, and
3. **stores** clean structured rows in SQLite you can query and export β€”
all while completely air-gapped.
The reference build is anchored to one note type β€” a **handwritten to-do / planner
page** β€” because the schema is tight and objectively measurable. The same engine
generalises to survey forms, recipe cards, and meeting notes with only a new schema
and grammar.
---
## Why It Matters
GPUs are scarce, costly, and online. Most computing isn't. PageParse is a
demonstration that **the CPU is enough**, and that a genuinely useful app keeps
working when the network doesn't β€” in clinics, fields, classrooms, and anywhere
connectivity is unreliable.
It also tackles the hackathon's core mission head-on: handwriting is about as
**unstructured** as input gets, and PageParse maps it to a strict, **structured**
schema. Handwriting recognition is a hard, real problem β€” and that difficulty is
exactly what makes the result worth scoring.
---
## How It Works
PageParse follows the required **Ingestion β†’ Processing β†’ Transformation β†’ Storage**
workflow, with every inference step on CPU and offline.
| Stage | What happens | Runs on |
| ----- | ------------ | ------- |
| **1. Ingestion** | Load a handwritten page (JPG / PNG / PDF). | local |
| **2. Processing** | OpenCV cleanup: grayscale, deskew, adaptive threshold, denoise. | CPU |
| **3. Recognition** | On-device OCR converts ink β†’ raw text (handwriting tier + printed fallback). | CPU |
| **4. Transformation** | A small language model maps raw text β†’ JSON, grammar-constrained to the schema. | CPU |
| **5. Storage** | Validated rows persisted to SQLite; optional local semantic search. | local |
---
## Architecture
```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ PageParse β”‚
β”‚ (100% on-device, CPU) β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
scan / photo OpenCV ONNX / Tesseract llama.cpp + GBNF SQLite
JPG Β· PNG Β· PDF β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
───────────────▢│ Ingestion│──────▢│ Preprocessing │──▢│ OCR (on CPU) │──▢│ Transform│──▢│ Storage β”‚
│ │ clean │ deskew·binarize│ │ TrOCR / Surya /│ │ SLM→JSON │ │ + vec │
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ denoiseβ””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ Tesseract β”‚ β”‚ validatedβ”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ query Β· search β”‚
β”‚ export Β· review β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```
---
## Model & Runtime (CPU-first)
Declared explicitly per the rules. **No GPU or CUDA anywhere in the inference path.**
All weights are fetched once and cached locally, so runtime needs no network.
| Stage | Model | Runtime | Precision |
| ------------------- | --------------------------------------- | ---------------------- | --------- |
| Handwriting OCR | TrOCR-small (handwritten) / Surya OCR 2 | ONNX Runtime (CPU EP) | INT8 |
| Printed fallback | Tesseract LSTM | Tesseract (CPU) | β€” |
| Text β†’ schema (SLM) | Qwen2.5-1.5B-Instruct | llama.cpp | Q4_K_M |
| Semantic search | all-MiniLM-L6-v2 | ONNX Runtime (CPU EP) | INT8 |
Schema validity is **guaranteed** by **llama.cpp GBNF grammar-constrained
decoding**: the model is structurally prevented from emitting anything that is not
valid against the PageParse schema, then re-validated with Pydantic.
---
## Output Schema
The structured target that unstructured handwriting is mapped to:
```json
{
"source_page": "string (filename)",
"captured_date": "YYYY-MM-DD",
"tasks": [
{
"task": "string",
"due_date": "YYYY-MM-DD | null",
"priority": "high | medium | low",
"category": "string | null",
"status": "todo | done",
"ocr_confidence": 0.0
}
]
}
```
---
## Data Model (SQLite)
```sql
CREATE TABLE pages (
id INTEGER PRIMARY KEY,
filename TEXT,
captured_date TEXT,
raw_ocr_text TEXT,
created_at TEXT
);
CREATE TABLE tasks (
id INTEGER PRIMARY KEY,
page_id INTEGER REFERENCES pages(id),
task TEXT NOT NULL,
due_date TEXT,
priority TEXT CHECK (priority IN ('high', 'medium', 'low')),
category TEXT,
status TEXT DEFAULT 'todo',
ocr_confidence REAL
);
```
---
## Tech Stack
**Language & core**
- Python 3.11
**Computer vision / OCR**
- OpenCV β€” image preprocessing
- ONNX Runtime (CPU EP) β€” TrOCR / Surya handwriting OCR
- Tesseract β€” printed-text fallback
**AI / structuring**
- llama.cpp (`llama-cpp-python`) β€” Qwen2.5-1.5B Q4_K_M GGUF inference
- GBNF grammar β€” guaranteed-valid JSON
- Pydantic β€” schema validation
**Storage & retrieval**
- SQLite β€” structured persistence
- sqlite-vec + all-MiniLM-L6-v2 (ONNX) β€” optional semantic search
**Interface**
- Typer β€” CLI (primary)
- FastAPI + PWA β€” web interface
**Observability**
- psutil β€” live CPU & memory telemetry
**Tooling & quality**
- Ruff, mypy, Bandit, pip-audit, detect-secrets, pytest, gitlint, yamllint,
markdownlint, REUSE
- pre-commit, GitLab CI (self-hosted Docker runner)
**License**
- AGPL-3.0-or-later
---
## Project Structure
```
pageparse/
β”œβ”€β”€ src/
β”‚ └── pageparse/
β”‚ β”œβ”€β”€ __init__.py
β”‚ β”œβ”€β”€ ingest.py # load image / PDF input
β”‚ β”œβ”€β”€ preprocess.py # OpenCV cleanup pipeline
β”‚ β”œβ”€β”€ ocr/
β”‚ β”‚ β”œβ”€β”€ __init__.py
β”‚ β”‚ β”œβ”€β”€ handwriting.py # TrOCR / Surya (ONNX)
β”‚ β”‚ └── printed.py # Tesseract fallback
β”‚ β”œβ”€β”€ extract.py # SLM + GBNF β†’ JSON
β”‚ β”œβ”€β”€ schema.py # Pydantic models
β”‚ β”œβ”€β”€ store.py # SQLite persistence
β”‚ β”œβ”€β”€ search.py # optional semantic search
β”‚ β”œβ”€β”€ telemetry.py # psutil CPU/RAM panel
β”‚ β”œβ”€β”€ config.py # paths, model settings, airgap flag
β”‚ β”œβ”€β”€ cli.py # Typer CLI entry point
β”‚ └── web.py # FastAPI + PWA
β”œβ”€β”€ grammars/
β”‚ └── task.gbnf # schema-constraining grammar
β”œβ”€β”€ models/ # bundled GGUF / ONNX weights (gitignored)
β”œβ”€β”€ scripts/
β”‚ β”œβ”€β”€ fetch_models.py # one-time online model download
β”‚ └── init_db.py # initialise SQLite
β”œβ”€β”€ samples/ # real handwritten demo pages
β”œβ”€β”€ tests/
β”‚ β”œβ”€β”€ test_preprocess.py
β”‚ β”œβ”€β”€ test_extract.py
β”‚ β”œβ”€β”€ test_schema.py
β”‚ └── test_store.py
β”œβ”€β”€ web/
β”‚ β”œβ”€β”€ static/ # PWA assets, service worker
β”‚ └── templates/
β”œβ”€β”€ docs/
β”‚ └── architecture.md
β”œβ”€β”€ .gitlab-ci.yml
β”œβ”€β”€ .pre-commit-config.yaml
β”œβ”€β”€ .gitignore
β”œβ”€β”€ .secrets.baseline
β”œβ”€β”€ pyproject.toml
β”œβ”€β”€ CONTRIBUTING.md
β”œβ”€β”€ CHANGELOG.md
β”œβ”€β”€ LICENSE # AGPL-3.0
└── README.md
```
---
## Getting Started
### Prerequisites
- Python 3.11+
- Tesseract OCR β€” `apt install tesseract-ocr` / `brew install tesseract`
- ~3 GB free disk for bundled models
### Installation
```bash
# Clone
git clone https://code.swecha.org/centurions/pageparse.git
cd pageparse
# Environment
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# Install (with dev tooling)
pip install -e ".[dev]"
# Fetch models ONCE while online β€” cached locally thereafter
python scripts/fetch_models.py
# Initialise the local database
python scripts/init_db.py
```
After `fetch_models.py` runs once, **disconnect from the network entirely** β€”
everything below works offline.
---
## Usage
### CLI
```bash
# Process a single handwritten page
pageparse process samples/todo_page_01.jpg
# Process a whole folder
pageparse process samples/ --batch
# Force air-gapped mode (any outbound call fails loudly)
pageparse process samples/todo_page_01.jpg --airgap
# Query and search
pageparse list --priority high
pageparse search "supplier"
# Export
pageparse export --format json > out.json
```
### Web (PWA)
```bash
pageparse serve # http://localhost:8000
```
Upload a page, watch the **queued β†’ processing β†’ saved** status, and view the
structured result. Installable as a PWA for offline use.
---
## Offline-First by Design
The core feature works with **no cloud calls**, and the demo proves it:
- **Air-gap toggle** β€” a UI switch and `--airgap` flag assert zero network access;
any accidental outbound call fails loudly rather than silently.
- **Status badge** β€” every page shows `queued β†’ processing β†’ saved`, making caching
and graceful handling observable.
- **Graceful degradation** β€” low-confidence OCR lines are flagged for review
instead of crashing; pages queue and process as the CPU frees up.
- **Live resource panel** β€” a psutil readout of CPU % and RAM during inference, so
the footprint is measured, not guessed.
**Demo procedure:** enable airplane mode / pull the Ethernet, then run a full
ingestion β†’ extraction β†’ storage cycle on a real handwritten page.
---
## Quality Gates (CI/CD)
All checks run locally via **pre-commit** and in CI on a **self-hosted GitLab
runner** (Docker executor). These are real checks β€” no stub jobs, no `exit 0`.
| # | Check | Tool |
| -- | ------------------- | ---------------- |
| 1 | Lint | Ruff |
| 2 | Format | Ruff format |
| 3 | Type-check | mypy |
| 4 | Security (SAST) | Bandit |
| 5 | Dependency CVEs | pip-audit |
| 6 | Secret scan | detect-secrets |
| 7 | Unit tests | pytest |
| 8 | Semantic commits | gitlint |
| 9 | CI YAML lint | yamllint |
| 10 | Docs lint | markdownlint |
| 11 | License compliance | REUSE |
| 12 | Smoke build | app boots headless |
```bash
pre-commit run --all-files
pytest
```
Conventional Commits are enforced (`feat:`, `fix:`, `chore:`, …).
---
## Judging Criteria Mapping
| Criterion | How PageParse addresses it |
| ---------------------- | ------------------------------------------------------------ |
| **Model performance** | OCR confidence + extraction accuracy reported on real pages. |
| **Resource efficiency**| Live psutil CPU/RAM panel during inference. |
| **Offline resiliency** | Air-gap toggle; full cycle demoed with the network off. |
| **Schema alignment** | GBNF grammar guarantees valid JSON every run. |
| **Graceful handling** | Queued processing + flagged low-confidence lines + caching. |
---
## Roadmap
- [ ] **MVP:** scanned to-do page β†’ validated JSON β†’ SQLite (offline, CPU)
- [ ] Printed-text fallback path
- [ ] Web PWA with air-gap toggle and status badge
- [ ] Live CPU/RAM telemetry panel
- [ ] Semantic search over extracted notes (`sqlite-vec`)
- [ ] Additional note types: survey forms, recipe cards, meeting notes
- [ ] Indic-script handwriting support
---
## Contributing
See [CONTRIBUTING.md](CONTRIBUTING.md). Branch from `main`, keep commits
Conventional, ensure `pre-commit run --all-files` and `pytest` pass, then open a
merge request. All contributions are licensed under AGPL-3.0-or-later.
---
## License
Licensed under the **GNU Affero General Public License v3.0 or later**
(AGPL-3.0-or-later). See [LICENSE](LICENSE).
AGPL is chosen deliberately: as strong copyleft, it closes the network/SaaS
loophole that plain GPL leaves open, ensuring anyone who runs a modified PageParse
as a network service must also share their source.