medishield / ARCHITECTURE.md
sriny2131's picture
deploy: sync from local repo
d86db02 verified
|
Raw
History Blame Contribute Delete
19.2 kB

MediShield AI Document Classifier β€” Architecture & Design

Complete technical architecture documentation for the insurance document classification system.


πŸ“ System Architecture Overview

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                         Browser / Client                                 β”‚
β”‚              Drag & Drop UI Β· frontend/index.html                        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚ POST /classify (multipart/form-data)
                                 β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚              FastAPI Server Β· src/api.py Β· Port 8000                     β”‚
β”‚                                                                          β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚   β”‚              asyncio.gather β†’ Concurrent Processing             β”‚  β”‚
β”‚   β”‚                                                                  β”‚  β”‚
β”‚   β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”‚  β”‚
β”‚   β”‚  β”‚     Orchestrator: src/classifier.py                      β”‚    β”‚  β”‚
β”‚   β”‚  β”‚                                                          β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  bill_*   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β”‚ Stage 1      │──────────▢│  doc_type = "bill"   β”‚   β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β”‚ Rules Engine β”‚           β”‚  method   = "rules"  β”‚   β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β”‚              β”‚           β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜                                       β”‚    β”‚  β”‚
β”‚   β”‚  β”‚         β”‚ others                                        β”‚    β”‚  β”‚
β”‚   β”‚  β”‚         β–Ό                                               β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  KYC kw   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β”‚ Stage 2      │──────────▢│  doc_type = "kyc"    β”‚   β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β”‚ KYC OCR      β”‚           β”‚  method   = "ocr"    β”‚   β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β”‚ (easyocr)    β”‚           β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜                                       β”‚    β”‚  β”‚
β”‚   β”‚  β”‚         β”‚ no KYC match                                  β”‚    β”‚  β”‚
β”‚   β”‚  β”‚         β–Ό                                               β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”           β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β”‚ Stage 3      │──────────▢│  doc_type = "image"  β”‚   β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β”‚ Gemini LLM   β”‚           β”‚  sub_type = category β”‚   β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β”‚              β”‚           β”‚  method   = "llm"    β”‚   β”‚    β”‚  β”‚
β”‚   β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜           β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚    β”‚  β”‚
β”‚   β”‚  β”‚                                                          β”‚    β”‚  β”‚
β”‚   β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β”‚  β”‚
β”‚   β”‚                                                                  β”‚  β”‚
β”‚   β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”‚  β”‚
β”‚   β”‚  β”‚           Monitoring: src/monitoring.py                β”‚    β”‚  β”‚
β”‚   β”‚  β”‚   @traceable spans Β· token counts Β· latency metrics     β”‚    β”‚  β”‚
β”‚   β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β”‚  β”‚
β”‚   β”‚                                                                  β”‚  β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚                                                                          β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚         Response JSON: {                                       β”‚   β”‚
β”‚  β”‚           filename, doc_type, sub_type, method,               β”‚   β”‚
β”‚  β”‚           latency_ms, tokens_used, confidence_score            β”‚   β”‚
β”‚  β”‚         }                                                       β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”˜   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”˜
                                                                    β”‚
                                                                    β–Ό
                                                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                                    β”‚   Frontend UI        β”‚
                                                    β”‚ Results table with   β”‚
                                                    β”‚ color-coded badges   β”‚
                                                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ”„ Three-Stage Classification Pipeline

Stage 1: Rules Engine (src/rules_engine.py)

Trigger: Filename analysis
Method: Regex pattern matching
Cost: $0
Speed: <1ms
Success Rate: ~70% of documents

def rules_engine(filename: str, image_data: bytes) -> Optional[Dict]:
    """
    Match documents by filename patterns.
    Fast, zero-cost, high-precision.
    """
    # Example patterns
    if filename.startswith("bill_"):
        return {
            "doc_type": "bill",
            "method": "rules",
            "confidence": 1.0,
            "latency_ms": 0.5
        }
    return None

Patterns matched:

  • bill_* β†’ doc_type = "bill"
  • invoice_* β†’ doc_type = "invoice"
  • receipt_* β†’ doc_type = "receipt"

When to use: Fast, immediate classification when metadata is reliable.


Stage 2: KYC OCR Detection (src/kyc_detector.py)

Trigger: Documents unmatched at Stage 1
Method: easyOCR text extraction + keyword matching
Cost: $0.001 per document
Speed: 1-2 seconds
Success Rate: ~20% of remaining documents (
6% of total)

def kyc_detector(image_data: bytes) -> Optional[Dict]:
    """
    Detect KYC documents (Aadhaar, PAN, Passport) via OCR.
    Medium speed, cheap, reliable for document types.
    """
    # Use easyOCR to extract text
    ocr_text = easyocr.recognize(image_data)
    
    # Match against KYC keywords
    kyc_keywords = ["aadhaar", "pan", "passport", "voter id", "driving license"]
    
    if any(keyword in ocr_text.lower() for keyword in kyc_keywords):
        return {
            "doc_type": "kyc",
            "sub_type": detect_kyc_subtype(ocr_text),
            "method": "ocr",
            "confidence": 0.95,
            "latency_ms": 1500
        }
    return None

Document types detected:

  • Aadhaar β†’ sub_type = "aadhaar"
  • PAN β†’ sub_type = "pan"
  • Passport β†’ sub_type = "passport"
  • Voter ID β†’ sub_type = "voter_id"
  • Driving License β†’ sub_type = "driving_license"

When to use: When OCR is fast enough and keyword matching is reliable.


Stage 3: Gemini LLM Classification (src/llm_classifier.py)

Trigger: All documents unmatched at Stages 1 & 2
Method: Gemini API (gemma-4-31b-it)
Cost: ~$0.01 per document
Speed: 2-4 seconds
Success Rate: ~10% of documents (only complex cases)

def llm_classifier(image_data: bytes, ocr_text: str) -> Dict:
    """
    Full AI classification for complex/ambiguous documents.
    Slow, expensive, but handles edge cases.
    """
    prompt = f"""
    Classify this insurance document.
    
    OCR extracted text (may be partial/noisy):
    {ocr_text}
    
    Image provided for visual analysis.
    
    Respond with JSON:
    {{
      "doc_type": "image|letter|form|other",
      "sub_type": "prescription|lab_report|claim_form|...",
      "confidence": 0.0-1.0,
      "reasoning": "brief explanation"
    }}
    """
    
    response = gemini_api.generate(
        image=image_data,
        text=prompt
    )
    
    return {
        "doc_type": response.doc_type,
        "sub_type": response.sub_type,
        "method": "llm",
        "confidence": response.confidence,
        "latency_ms": elapsed_time,
        "tokens_used": response.usage.total_tokens
    }

Document types classified:

  • Prescriptions
  • Lab reports
  • Claim forms
  • Medical letters
  • X-ray reports
  • Test certificates
  • Insurance documents (various)

When to use: Complex, ambiguous cases where Rules + OCR aren't sufficient.


⚑ Async Concurrency Architecture

Request Flow

FastAPI Handler (src/api.py)
    β”‚
    β”œβ”€ Read multipart form data
    β”œβ”€ Extract file list: [file1.pdf, file2.pdf, ...]
    β”‚
    β–Ό
asyncio.gather([
    executor.submit(classify_document, file1),
    executor.submit(classify_document, file2),
    ...
])
    β”‚
    └─ Runs all files in PARALLEL thread pool
       Each file: Stage1 β†’ Stage2 β†’ Stage3 (cascading, any can short-circuit)
    β”‚
    β–Ό
Collect results, emit LangSmith traces
    β”‚
    β–Ό
Return JSON response

Performance Characteristics

Scenario A: All Stage 1 matches (bills with bill_ prefix)

  • Input: 10 bills
  • Processing: 10 Γ— <1ms = ~10ms total
  • Response time: ~100ms (overhead)

Scenario B: Mixed (70% rules, 20% OCR, 10% LLM)

  • Input: 100 documents
  • Stage 1: 70 docs Γ— <1ms = ~0.07s
  • Stage 2: 20 docs Γ— 1.5s = ~30s parallel (1 thread per doc)
  • Stage 3: 10 docs Γ— 3s = ~30s parallel (1 thread per doc)
  • Total: ~30s (OCR & LLM run in parallel, limited by slowest)

Optimization: With N worker threads and M documents:

  • If M ≀ N: perfect parallelism, time β‰ˆ max(latencies)
  • If M > N: queued, time β‰ˆ sum(latencies) / N

πŸ“Š Monitoring & Observability

LangSmith Tracing (@traceable decorator)

Each stage emits structured spans to LangSmith:

from langsmith import traceable

@traceable(name="rules_engine")
def rules_engine(filename: str):
    # Automatic tracing:
    # - Execution time
    # - Input/output
    # - Errors
    pass

@traceable(name="kyc_ocr")
def kyc_detector(image_data: bytes):
    # Token usage tracked automatically
    # Latency metrics collected
    pass

@traceable(name="gemini_llm")
def llm_classifier(image_data: bytes):
    # Detailed LLM call tracing
    # Token usage: input + output
    # Model name, parameters, latency
    pass

Metrics Collected

Per-request metrics:

  • Request ID
  • Number of documents
  • Document types (histogram)
  • Total latency (ms)
  • Breakdown by stage
  • Token usage (for LLM stage)
  • Classification confidence scores

Real-time dashboards:

  • Azure Monitor: logs, alerts, performance
  • LangSmith: traces, token usage, latency percentiles

πŸš€ Deployment Architecture

Azure Container Apps Stack

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Azure Container Apps (ACA)             β”‚
β”‚  - 0.5 vCPU, 2GB RAM per replica        β”‚
β”‚  - Min 1 replica, Max 3 (auto-scale)    β”‚
β”‚  - HTTPS endpoint                       β”‚
β”‚  - Auto-scaling on CPU/memory           β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                 β”‚
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β”‚            β”‚            β”‚
    β–Ό            β–Ό            β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚Replica1β”‚  β”‚Replica2β”‚  β”‚Replica3β”‚
β”‚(running)  β”‚(standby)  β”‚(standby)
β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜
    β”‚
    └─────────────────┬─────────────────┐
                      β”‚                 β”‚
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”
    β”‚  Azure Monitor         β”‚  β”‚  LangSmith    β”‚
    β”‚  - Logs                β”‚  β”‚  - Traces     β”‚
    β”‚  - Performance         β”‚  β”‚  - Token usageβ”‚
    β”‚  - Alerts              β”‚  β”‚  - Latency    β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Container Image

Dockerfile:

FROM python:3.12-slim

WORKDIR /app

# Install dependencies
COPY requirements.txt .
RUN pip install -r requirements.txt

# Copy source
COPY src/ ./src/
COPY frontend/ ./frontend/

# Expose port
EXPOSE 8000

# Run with uvicorn
CMD ["uvicorn", "src.api:app", "--host", "0.0.0.0", "--port", "8000"]

Image stored in: Azure Container Registry (medishieldacr.azurecr.io)


πŸ”’ Security Architecture

API Security

  • HTTPS only via Azure Container Apps
  • CORS: Configured for frontend origin
  • File validation: Size limits, MIME type checks
  • Input sanitization: Filename validation, size bounds

Data Privacy

  • No persistent storage: Processed files deleted after classification
  • Transient memory: Results live only in response
  • Encrypted in transit: TLS 1.2+
  • Audit logs: All classifications logged to Azure Monitor

Secrets Management

  • Environment variables via Azure Container Apps secrets
  • Gemini API key: Never in code, via env
  • LangSmith key: Via env, read-only service key

πŸ“ˆ Scalability Analysis

Current Setup

  • Max throughput: ~1000 documents/hour (with 1 replica, max 3s per doc)
  • Bottleneck: Gemini API rate limits (~100 requests/minute)
    • With 10% of docs hitting Gemini β†’ ~100 docs/min Γ— 10 = 1000 docs/min
    • Actual: limited by API quotas

Scaling Options

  1. Increase replicas: Auto-scaling to 3 replicas β†’ 3x throughput
  2. Batch processing: Collect documents, process asynchronously β†’ decoupled throughput
  3. Queue-based: Azure Service Bus β†’ robust handling of traffic spikes
  4. Cache results: Store common patterns (frequent filenames) β†’ reduce processing

πŸ§ͺ Testing Strategy

117 Passing Tests

Unit tests (70):

  • Stage 1 rules engine (10 test cases)
  • Stage 2 OCR patterns (15 test cases)
  • Stage 3 LLM prompt formatting (10 test cases)
  • Async orchestration (15 test cases)
  • Monitoring/tracing (10 test cases)
  • Response validation (10 test cases)

Integration tests (30):

  • End-to-end classification pipeline (10 test cases)
  • Multi-file concurrent processing (5 test cases)
  • Error handling & retry logic (5 test cases)
  • API endpoint validation (5 test cases)
  • LangSmith trace verification (5 test cases)

Performance tests (10):

  • Latency benchmarks (Stage 1, 2, 3)
  • Concurrency stress tests
  • Memory usage profiling

Edge cases (7):

  • Empty files
  • Corrupt images
  • Unicode filenames
  • Very large files
  • Timeout handling

πŸ“Š Cost Model

Per-Document Costs

Stage Cost Trigger Frequency
Stage 1 (Rules) $0.00 Filename pattern 70%
Stage 2 (OCR) ~$0.001 easyOCR library 20%
Stage 3 (LLM) ~$0.01 Gemini API call 10%

Expected cost per document: (0.7 Γ— $0) + (0.2 Γ— $0.001) + (0.1 Γ— $0.01) = $0.0013

Annual cost (1M documents): 1,000,000 Γ— $0.0013 = $1,300 (AI costs only)

Infrastructure: Azure Container Apps ~$50-100/month (compute + storage)

Total monthly: ~$200 (AI + compute + monitoring)

vs. Manual Labor

  • Manual operator: ~$3,000/month salary
  • 12 operators: $36,000/month
  • 2 remaining operators: $6,000/month
  • Savings: ~$30,000/month

ROI: Pays for itself in <1 week of savings.


πŸ”„ Continuous Improvement

Metrics to Monitor

  1. Accuracy by stage: Track Stage 3 confidence scores
  2. False negatives: Documents incorrectly classified at Stage 1 or 2
  3. Latency trends: Identify performance regressions
  4. Token usage: Monitor Gemini API efficiency
  5. Cost per document: Optimize stage cascade

Feedback Loop

Production metrics β†’ Identify misclassifications
    ↓
Add new rules β†’ Retrain Stage 1/2
    ↓
Deploy β†’ Test in staging
    ↓
Compare accuracy vs production
    ↓
If better: deploy; else: revert

πŸ“š Related Documentation


Generated: April 26, 2026
Document version: 2.0 (architecture diagrams added)