Spaces:
Runtime error
Runtime error
| # CLAUDE.md | |
| This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. | |
| ## Workflow gate (READ FIRST) | |
| Before changing anything, read **`SPEC.md`** β it defines the mandatory process: | |
| sanity tests (`tests/test_sanity.py`) must pass before any commit (enforced by | |
| the git pre-commit hook installed via `bash scripts/install-hooks.sh`), and the | |
| **How To Guide** tab in `ui/app.py` must be updated on any user-facing UI change. | |
| ## Project | |
| **G.U.I.D.E.** (Grievance Utility for Information Extraction, Drafting and Enrichment) β a privacy-first consumer complaint assistant for Indian consumers. Users describe their complaint in plain language; the system redacts PII locally, drafts a formal letter, and routes to the correct Indian regulatory authority. | |
| ## Startup | |
| ### 1. Configure environment | |
| Create a `.env` file in the project root with the following keys (all three are required): | |
| ```bash | |
| cp .env.example .env | |
| ``` | |
| Then edit `.env` and fill in: | |
| ``` | |
| ANTHROPIC_API_KEY=sk-ant-... # Claude API β powers the managed agent | |
| HF_TOKEN=hf_... # HuggingFace token β needed to download model checkpoints from sarav95/guide-models | |
| LANGCHAIN_API_KEY=ls__... # LangSmith API key β used for tracing / observability | |
| ``` | |
| > The `.env` file is git-ignored and must never be committed. | |
| ### 2. Install dependencies | |
| ```bash | |
| pip install -r requirements.txt | |
| python -m spacy download en_core_web_lg | |
| ``` | |
| ### 3. Run | |
| ```bash | |
| # Run everything (auto-downloads models from HF if missing, then starts both servers) | |
| python start.py | |
| # Access | |
| # API docs β http://localhost:8000/docs | |
| # UI β http://localhost:7860 | |
| ``` | |
| ### `start.py` flags | |
| | Flag | Effect | | |
| |------|--------| | |
| | _(none)_ | Auto-download missing checkpoints from HF, auto-train any still missing, then serve | | |
| | `--download-models` | Force download all checkpoints from `sarav95/guide-models` on HuggingFace | | |
| | `--no-train` | Skip training; start servers immediately (models must exist) | | |
| | `--train` | Force re-train all three models even if checkpoints exist | | |
| | `--train-only` | Train then exit β no servers started | | |
| | `--cfpb_csv PATH` | Override default CFPB CSV path (`data/raw/complaints.csv`) | | |
| `--no-train` and `--train` are mutually exclusive. `--train-only` can be combined with `--train`. | |
| Model checkpoints are stored in `models/` (git-ignored). They are hosted on HuggingFace at `sarav95/guide-models` and auto-downloaded on first run when the directory is missing. | |
| ### Train individual models | |
| ```bash | |
| # EvidenceNER β synthetic data generated in-memory, no download required | |
| python -m src.ner.train --output_dir models/evidence_ner | |
| # NextActionPredictor β synthetic data, <30 s on CPU | |
| python -m src.next_action.train --output_path models/next_action/model.pt | |
| # DomainClassifier β needs CFPB CSV first (one-time Kaggle download) | |
| python -m src.classifier.train \ | |
| --cfpb_csv data/raw/complaints.csv \ | |
| --output_dir models/domain_classifier | |
| ``` | |
| Without a DomainClassifier checkpoint the system falls back to keyword heuristics (confidence=0.0 sentinel). Without a NextActionPredictor checkpoint it falls back to `DOMAIN_ACTION_PRIORS` rule mapping. Both fallbacks keep the pipeline fully functional. | |
| ## Architecture | |
| ### Layer 1 β Privacy (runs first, locally, on every message) | |
| `src/privacy/redactor.py` β Microsoft Presidio (`presidio-analyzer` + `presidio-anonymizer`) with spaCy `en_core_web_lg`. Detects and replaces PII with `<ENTITY_TYPE>` placeholders **before** any text leaves the process. Fails open: on any error the original text is returned unchanged so the pipeline is never blocked. | |
| Entity types redacted: `PERSON`, `PHONE_NUMBER`, `EMAIL_ADDRESS`, `CREDIT_CARD`, `IBAN_CODE`, `US_BANK_NUMBER`, `IN_AADHAAR`, `IN_PAN`, `IN_VEHICLE_REGISTRATION`. | |
| ### Layer 2 β Deep Learning Models | |
| | Module | Architecture | Checkpoint | Fallback | | |
| |--------|-------------|-----------|---------| | |
| | `src/classifier/` | DistilBERT + linear head, 6 classes | `models/domain_classifier/` | keyword heuristics | | |
| | `src/ner/` | DistilBERT token classifier, BIO tags, 6 entity types | `models/evidence_ner/` | _(none β required)_ | | |
| | `src/next_action/` | MLP 12β64β64β6 | `models/next_action/model.pt` | `DOMAIN_ACTION_PRIORS` | | |
| | `src/document_processor/` | Tesseract OCR + ViT patch-variance scoring | HuggingFace `google/vit-base-patch16-224` | OCR-only | | |
| **6 domain classes:** `ecommerce` Β· `telecom` Β· `banking` Β· `cibil` Β· `insurance` Β· `general` | |
| **6 NER entity types:** `ORG` Β· `AMOUNT` Β· `DATE` Β· `REF_ID` Β· `ACCOUNT` Β· `PERSON` | |
| **6 escalation actions:** `company_support` Β· `nch` Β· `trai` Β· `rbi_ombudsman` Β· `irdai` Β· `legal` | |
| Document processing path: PDF (pdfplumber) β EvidenceNER; image (pytesseract + Pillow greyscale/threshold/deskew) β EvidenceNER + DocumentViT. Both paths merge entities by `(text.lower(), label)` key, keeping the higher-confidence span. | |
| ### Layer 3 β Claude Managed Agent | |
| `src/agent/` β one `GUIDEAgent` instance per session; model is env-driven via `GUIDE_MODEL` (default `claude-sonnet-4-6`). Agent loop: stream response β on `stop_reason == "tool_use"` dispatch all tool_use blocks via `execute_tool()` β append `tool_result` messages β repeat until `end_turn`. Loop capped at 12 rounds. | |
| **7 tools:** `classify_domain` Β· `extract_entities` Β· `process_document` Β· `draft_complaint` Β· `recommend_action` Β· `store_memory` Β· `get_memory` | |
| `draft_complaint` is handled internally by Claude (returns `{"status": "proceed"}` from Python); the actual letter is generated as text in the next assistant turn. | |
| **HITL gate (mandatory):** Claude must present a numbered entity summary and receive `[USER CONFIRMED]: {...}` (injected by `POST /api/session/{id}/validate-entities`) before calling `draft_complaint()`. This is enforced in the system prompt (Rule 4 + Rule 5) and in `GUIDEAgent.confirm_entities()`. | |
| ### Layer 4 β FastAPI + Gradio | |
| **FastAPI** (`src/api/`) β singleton startup order in `lifespan()`: Presidio β DomainClassifier β EvidenceNER β NextActionPredictor β (DocumentProcessor lazy). All blocking model calls and CMA calls use `run_in_threadpool`. | |
| **Gradio** (`ui/app.py`) β tabs: Chat (π privacy badge + side-by-side redaction reveal showing original-vs-sent text) Β· Verify Entities (HITL editable fields + Confirm button) Β· Complaint Draft (copy + .txt/PDF download via `reportlab`) Β· Escalation Guide Β· **Privacy Audit** (timestamped outbound/local event log with verified "0 raw identifiers transmitted" guarantee) Β· About. | |
| **Privacy reveal & audit:** `POST /message` returns `original_text`, `redacted_text`, and `redactions` (per-span `RedactionSpan`s) so the UI renders a side-by-side reveal β these are returned ONLY to the same user's browser, never forwarded externally. Each `/message` and `/upload` appends an entry to `Session.audit`; `GET /session/{id}/audit` serves the trail. The audit's `leak_check` re-verifies that no original PII value survives into the transmitted text. | |
| ## Critical Files | |
| | File | Role | | |
| |------|------| | |
| | `start.py` | Single entry point β training + server launch + `[API]`/`[UI]` log prefixing | | |
| | `src/privacy/redactor.py` | `PIIRedactor` singleton; `init_redactor()` / `get_redactor()` | | |
| | `src/agent/prompts.py` | `SYSTEM_PROMPT` β all 7 CMA rules including HITL gate | | |
| | `src/agent/tools.py` | `TOOL_DEFINITIONS` (JSON Schema) + `execute_tool()` dispatcher | | |
| | `src/agent/memory.py` | `SessionMemory` β per-session key-value store | | |
| | `src/agent/agent.py` | `GUIDEAgent` β streaming agent loop, `send_message()`, `confirm_entities()`, `add_document()` | | |
| | `src/api/main.py` | FastAPI app, `lifespan()` startup order, `_component_status` dict | | |
| | `src/api/routes.py` | All HTTP endpoints; Presidio redaction wired into `POST /message` | | |
| | `src/api/sessions.py` | Session registry (`session_id β Session(agent, history)`) | | |
| | `src/next_action/priors.py` | `DOMAIN_ACTION_PRIORS` + `ACTION_METADATA` (URLs) β training source of truth | | |
| | `ui/app.py` | Gradio Blocks app (`build_app()`), all event handlers, API client helpers | | |
| | `.env` | `ANTHROPIC_API_KEY`, `HF_TOKEN`, `LANGCHAIN_API_KEY` β never committed (in `.gitignore`) | | |
| ## Key Invariants | |
| - **PII never reaches Anthropic.** Presidio runs in-process before every `POST /message`. The Gradio UI sends raw text to FastAPI; FastAPI redacts before calling `GUIDEAgent.send_message()`. The agent only ever sees `<ENTITY_TYPE>` placeholders. | |
| - **HITL gate is mandatory before drafting.** `draft_complaint()` must not be called until `[USER CONFIRMED]` arrives. This is encoded in both the system prompt (Rules 4-5) and the route handler separation (`/message` vs `/validate-entities`). | |
| - **All model inference is blocking I/O.** Every call to a DL model or the Anthropic API uses `run_in_threadpool` to avoid blocking the FastAPI async event loop. | |
| - **Fallbacks keep the pipeline alive.** DomainClassifier β keyword fallback. NextActionPredictor β `DOMAIN_ACTION_PRIORS`. Both degrade gracefully with no code changes required. | |
| - **`add_document()` queues, not processes.** The document path is held in `_pending_documents` and prepended as `[Document uploaded: <path>]` on the next `send_message()` call, triggering Rule 3 of the system prompt. | |
| ## Model Checkpoint Paths | |
| ``` | |
| models/ | |
| evidence_ner/ β HuggingFace format (config.json + pytorch_model.bin) | |
| domain_classifier/ β HuggingFace format | |
| next_action/ | |
| model.pt β PyTorch state_dict + metadata dict | |
| ``` | |
| `start.py` detects checkpoints by checking for `config.json` (HuggingFace) or file existence (`.pt`). The `--train` flag bypasses these checks and retrains unconditionally. | |
| ## Environment Variables | |
| | Variable | Default | Purpose | | |
| |----------|---------|---------| | |
| | `ANTHROPIC_API_KEY` | _(required)_ | Anthropic API key β powers the Claude managed agent | | |
| | `HF_TOKEN` | _(required)_ | HuggingFace token β authenticates `snapshot_download` from `sarav95/guide-models` | | |
| | `LANGCHAIN_API_KEY` | _(required)_ | LangSmith API key β tracing and observability | | |
| | `GUIDE_MODEL` | `claude-sonnet-4-6` | Claude model for the CMA. Use an Anthropic ID (e.g. `claude-opus-4-8`, `claude-haiku-4-5-20251001`) or a gateway/Bedrock ID when `LITELLM_PROXY_URL` is set | | |
| | `LITELLM_PROXY_URL` | _(unset)_ | If set, route Anthropic calls through a LiteLLM gateway instead of the direct API | | |
| | `API_PORT` | `8000` | FastAPI port | | |
| | `GRADIO_PORT` | `7860` | Gradio port | | |
| | `LOG_LEVEL` | `info` | Uvicorn log level | | |