--- title: HTR ยท VLM as Annotator emoji: ๐Ÿ“œ colorFrom: indigo colorTo: purple sdk: docker app_port: 7860 pinned: false license: mit short_description: HTR via VLM and MoE post-correction --- # HTR VLM + MoE Post-correction Custom HuggingFace-style web app for Handwritten Text Recognition by Vision-Language Models (via OpenRouter) and Mixture-of-Experts post-correction. Logic : *LLM as annotator* โ€” the model proposes a transcription, the human validates / corrects, the validated examples feed an in-context-learning (ICL) pool that is re-injected as few-shot for subsequent pages. ## Two modes ### 1. OCR - Upload one or more page images (manuscripts, prints, โ€ฆ). - Select an active page, run the OCR with a single VLM of your choice (any OpenRouter model slug). - Edit the result inline. After "Validate & score", edited words are highlighted in light green and CER / WER are computed (corrected = reference). - "Add to ICL" stores `{image, corrected_text}` in the in-memory pool. Future OCR calls inject up to *N* recent examples as few-shot, so the VLM learns the hand / orthography on the fly. - A quick-add tool also lets you paste a validated transcription directly without going through the editor. ### 2. Post-correction (MoE) - Image + OCR text (paste your own, or come from mode 1). - Two expert VLMs are called **in parallel** with the same prompt. - If they agree โ†’ consensus result. Otherwise โ†’ a third "judge" VLM arbitrates. - Defaults : `google/gemini-3-flash-preview`, `qwen/qwen3.5-flash-02-23`, judge `mistralai/mistral-large-2512` (all editable from Settings). ## Customisation - Language / script field and free-text guidelines are merged into the prompts. - Any model slug from OpenRouter can be entered manually. - Advanced : custom system / user prompt overrides for the OCR call. - All settings are reachable from the header menus (no hidden CLI flags). ## Run locally ```bash conda activate research cd Webapp pip install -r requirements.txt uvicorn app:app --reload --port 7860 ``` Then open . Open the *API key* modal, paste your OpenRouter key (or export `OPENROUTER_API_KEY` in the environment as fallback), click *Test*, then start uploading images. ## Run in Docker / HF Spaces ```bash docker build -t htr-vlm . docker run -p 7860:7860 -e OPENROUTER_API_KEY=sk-or-v1-โ€ฆ htr-vlm ``` Pushing the directory to a HuggingFace Space (Docker SDK) deploys it as-is. ## Notes - State is **in-memory and single-process** (intended single-user demo / annotator workflow). Use the *Export JSONL* button on the ICL pool to persist before reloading. - Images are normalised to JPEG, โ‰ค 2200 px on the longest side, sent as base64 data-URLs (OpenRouter multimodal request shape). - No telemetry, no persistence on disk except what you explicitly download. ## Architecture - `app.py` โ€” FastAPI routes, in-memory session. - `provider.py` โ€” async OpenRouter client (multimodal). - `prompts.py` + `data/prompts/*.txt` โ€” system / user templates and ICL pool. - `moe.py` โ€” two-experts + judge pipeline. - `io_utils.py` โ€” image normalisation, word-level diff, CER / WER (jiwer-backed). - `schemas.py` โ€” Pydantic request shapes. - `static/` โ€” Alpine.js + Tailwind single-page UI.