themis / README.md
vg15o2's picture
Moonley backend (HF Space build)
1d9bd9b
|
Raw
History Blame Contribute Delete
16.2 kB
metadata
title: Moonley API
sdk: docker
app_port: 7860
pinned: false

Moonley

Grounded AI legal research for Indian law. Moonley answers natural-language questions by retrieving the relevant statutes (IPC · BNS · CrPC · BNSS · IEA · BSA) and Supreme Court judgments, fusing them on a single cross-encoder, generating an answer with DeepSeek, and verifying every cited section/case against the retrieved evidence — streaming each reasoning step to the UI over Server-Sent Events.

  • Live app: moonley-pilot.vercel.app (React/Vite SPA)
  • Backend: Hugging Face Spaces (FastAPI/Docker, API only)
  • Production branch: phase1.1 · Repository: vg15o2/themis

Table of contents

  1. Architecture
  2. Tech stack
  3. Repository layout
  4. The retrieval pipeline (in depth)
  5. Data & indices
  6. Backend API contract (SSE)
  7. Frontend
  8. Configuration
  9. Local development
  10. Deployment
  11. Operational notes
  12. Roadmap → V2
  13. Documentation

1. Architecture

flowchart TB
  UI["React SPA (Vercel)<br/>SSE reasoning UI"] -- "POST /ask (SSE)" --> API
  subgraph Space["HF Spaces · Docker · CPU 16GB"]
    API["FastAPI<br/>backend/app.py"] --> R["Unified router<br/>hybrid_rag/unified_legal_Rag.py"]
    R --> SR["Statute pipeline<br/>statute_retrieval.py"]
    R --> JR["Judgment pipeline<br/>llm_retriever.py"]
    R --> VER["Verifier<br/>verifier/verified.py"]
    SR --> CS[("Chroma<br/>indian_statutes")]
    JR --> CJ[("Chroma<br/>sci_judgments_bge_v2")]
    JR --> BM["BM25Okapi (in-memory)"]
    SR --> M["BGE-small + ms-marco CE"]
    JR --> M
  end
  R -- "stream" --> DS["DeepSeek<br/>deepseek-chat"]
  DSset["HF Dataset<br/>vg15o2/themis-judgments"] -. "downloaded at build" .-> CJ

Request lifecycle: route_query decides statute / judgment / hybrid → each enabled pipeline retrieves candidates → unified_rerank re-scores all candidates on one cross-encoder (dedupe + per-source cap) → an intent-specific prompt + evidence is streamed through DeepSeek → check_grounding validates citations → done. Every stage is emitted as a typed SSE event.


2. Tech stack

Layer Choice Detail
Frontend React 18 + Vite 5 SSE via fetch+ReadableStream; react-markdown+remark-gfm; deployed on Vercel (root frontend/)
Backend FastAPI + Uvicorn Single SSE endpoint; lazy in-process pipeline load
Backend host HF Spaces (Docker SDK) 2 vCPU / 16 GB free tier; listens on :7860
Vector store ChromaDB PersistentClient 2 collections, local on-disk
Embeddings BAAI/bge-small-en-v1.5 384-dim, normalize_embeddings=True; query prefix "Represent this sentence for searching relevant passages: "
Reranker cross-encoder/ms-marco-MiniLM-L-6-v2 statute rerank, judgment child/parent rerank, and the unified cross-source rerank
Lexical rank_bm25.BM25Okapi built in-memory over all judgment chunks at startup
LLM DeepSeek deepseek-chat OpenAI-compatible client, base_url=https://api.deepseek.com, stream=True, temperature=0.1, max_tokens=2500

3. Repository layout

backend/
  app.py              FastAPI app · POST /ask (SSE) · citation links · lazy get_pipeline()
  requirements.txt    fastapi, uvicorn, chromadb, sentence-transformers, openai, rank-bm25, huggingface_hub, python-dotenv
hybrid_rag/
  unified_legal_Rag.py  route_query · get_statute_evidence · get_judgment_evidence · unified_rerank · classify_intent · PROMPTS · Evidence
statute corpus/
  statute_retrieval.py  Chroma `indian_statutes` · parse_section_references · direct_lookup · expand_query · semantic_search · rerank_results
  (serving data is fetched from a private, pinned Hugging Face dataset at runtime)
llm_retriever.py        Chroma `sci_judgments_bge_v2` · dense_search · bm25_search · rrf_fusion · rerank_children · fetch_parent_chunks · rerank_parents · retrieve
verifier/
  verified.py           check_grounding (Tier-1/1.5) · verify_citations_live (Tier-2, lazy bharat_courts)
  __init__.py
frontend/
  src/App.jsx           turns · ReasoningPanel · StepTimeline · Answer · CopyButton · intent dropdown · stop/new-chat
  src/api.js            streamAsk(query, history, onEvent, signal, intent) — SSE parser
  src/index.css         Harvey-inspired warm-light theme
  public/icon.svg, favicon.ico
Dockerfile              deps → COPY → private runtime artifact download → uvicorn
render.yaml             (legacy) Render blueprint
themis/                  current architecture, scaling and storyline documents

4. The retrieval pipeline (in depth)

4.1 Routing — unified_legal_Rag.route_query(query) -> RouteDecision

explicit INSC citation (judg_rag.extract_citations) AND not statute_signal -> judgment-only
section reference (stat_rag.parse_section_references)  AND not judgment_signal -> statute-only
otherwise -> hybrid (both)

Signal regexes: JUDGMENT_SIGNAL_RE (case/judgment/held/INSC/ratio/…), STATUTE_SIGNAL_RE (section/provision/IPC/BNS/…).

4.2 Statute path — get_statute_evidence(query)statute_retrieval

  1. classify_query{type: direct_lookup|semantic|hybrid, section_refs, acts_mentioned, is_comparative}.
  2. direct_lookup(act, sec) for each parsed (ACT, section) → exact section (score 999.0).
  3. expand_query(query) → DeepSeek rewrites colloquial → statutory language (2–3 variants).
  4. semantic_search(query, top_k=40, act_filter) over Chroma (+ per-expansion and per-act searches, deduped by (act_short, section_number)).
  5. rerank_results(rerank_query, candidates, top_k=15) cross-encoder → wrapped as Evidence.

4.3 Judgment path — get_judgment_evidence(query)llm_retriever.retrieve

Query is enriched with severity context (offence title from the referenced section) + an LLM keyword expansion. retrieve has three modes:

  • citation (1 INSC cite) → exact parent via citation_search.
  • comparison (≥2 cites + comparison words) → each parent.
  • hybriddense_search(top 150) + bm25_search(top 150)rrf_fusion(k=60)deduplicate_children(≤3/case)rerank_children(top 30)rank_cases_from_childrenfetch_parent_chunks(top 20)rerank_parents(top 5).

4.4 Unified rerank — unified_rerank(query, evidence)

Exact hits (score==999) pinned on top; everything else scored on the same ms-marco cross-encoder so statute and judgment candidates are comparable. Then: dedupe by ("s", act, section) / ("j", neutral_citation), per-source cap MAX_SINGLE_TYPE_SHARE=6, final FINAL_EVIDENCE_N=8. The rerank query is severity-enriched so "BNS 103" biases toward murder-class judgments.

4.5 Generation

classify_intent(query) (defaults to LEGAL_RESEARCH) selects one of five system prompts in PROMPTS (research citation-table / case summary / comprehensive study / comparison / story evaluation). All prompts enforce evidence-grounding (no model knowledge beyond evidence). The backend streams stat_rag.llm_client.chat.completions.create(..., stream=True) and emits each delta as a token event. A forced intent from the UI overrides classification.

4.6 Verification — verifier.check_grounding(answer, final_evidence, stat_rag, judg_rag)

Strips markdown bold, extracts cited sections + INSC citations, then:

  • cited section in retrieved evidence → grounded; in full local DB but not retrievedretrieval miss (real law); not in DB → hallucination (flagged).
  • cited case not in evidence → ungrounded citation (flagged).
  • grounded = no hallucinated sections AND no ungrounded citations. Tier-2 (verify_citations_live, live bharat_courts + 30-day cache) is lazily imported and currently deferred.

5. Data & indices

Corpus Chroma collection Count On-disk Provisioning
Statutes (6 acts) indian_statutes 2,353 sections private artifact Downloaded at startup from a pinned private Hugging Face dataset using the Space secret
SCI judgments schema-v5 FAISS + SQLite release-defined private artifact Downloaded at startup from a pinned private Hugging Face dataset using the Space secret
  • Parent-child chunking: parent = full judgment + metadata; children = ~512-token windows (100 overlap). IDs: <citation>__child_NNNN, <citation>__parent.
  • Judgment metadata (per chunk): case_name, neutral_citation, court, date, bench, author_judge, acts, sections, issue, short_summary, full_headnote, outcome, source_url, ….
  • Statute record: {metadata: {act_short, act_name, section_number, title}, retrieval_text}.
  • Neither statute nor judgment serving artifacts are committed to GitHub or baked into the public Space image. start_private_space.py downloads both pinned snapshots with HF_TOKEN, verifies their required entrypoints, removes the token from the API process environment, and then starts FastAPI.

6. Backend API contract (SSE)

POST /asktext/event-stream

// request
{ "query": "string",
  "history": [{"role": "user|assistant", "content": "..."}],
  "intent": "AUTO | LEGAL_RESEARCH | CASE_SUMMARY | COMPREHENSIVE_CASE_STUDY | CASE_COMPARISON | STORY_EVALUATION" }
// events — each emitted as `data: {json}\n\n`
{ "type":"step",     "phase":"planning|retrieval|rerank|answer|verify", "title":"...", "detail":"..." }
{ "type":"evidence", "items":[ /* statute or judgment items, see below */ ] }
{ "type":"token",    "delta":"..." }
{ "type":"verify",   "grounded":true, "hallucinated_sections":[], "retrieval_miss_sections":[], "unverified_citations":[] }
{ "type":"done",     "answer":"...", "intent":"LEGAL_RESEARCH", "route":"...", "citations":[...], "elapsed_seconds":12.3 }
{ "type":"error",    "message":"..." }
// evidence items
{ "kind":"statute",  "act":"BNS", "section":"103", "title":"Punishment for murder.", "score":1.23, "url":"https://indiankanoon.org/search/?formInput=..." }
{ "kind":"judgment", "case":"Sanjay Kumar Sharma v. State of Bihar", "citation":"2026 INSC 223", "title":"<issue/summary>", "score":1.23, "url":"<source_url>" }

GET /health{status, service, pipeline_loaded} · GET / → service info.

Notes: an immediate step:"Warming up" is flushed before the lazy pipeline load so the SSE connection opens promptly; CORS origin via FRONTEND_ORIGIN (default *).


7. Frontend

  • streamAsk(query, history, onEvent, signal, intent) POSTs JSON and parses the SSE frame stream (\n\n-delimited data: lines) off the ReadableStream.
  • Per-turn state {query, steps[], evidence[], answer, citations[], verify, done, intent, elapsed_seconds}.
  • ReasoningPanel — collapsed-by-default "Thinking…" disclosure; expands to the live step timeline. Answer — markdown + top "Copy". Copy answer + sources — appends a formatted Sources: block. StopAbortController.abort() (keeps partial answer). New chat — clears turns + aborts. Intent dropdown — forces the answer style.
  • VITE_API_URL selects the backend (falls back to /api, proxied to localhost:8000 in dev).

8. Configuration

Var Where Purpose
DEEPSEEK_API_KEY backend env / HF secret DeepSeek auth (required; checked before pipeline import)
CLERK_PUBLISHABLE_KEY backend env / HF variable Public Clerk application key returned to both standalone frontends
CLERK_SECRET_KEY backend env / HF secret Clerk backend API credential; never expose in frontend code
CLERK_JWT_KEY backend env / HF secret Optional PEM public key for networkless session-token verification
CLERK_AUTHORIZED_PARTIES backend env / HF variable Exact Moonley and local browser origins allowed by the API
JUDGMENTS_CHROMA_PATH backend env Path to the judgments Chroma dir (Docker sets /app/judgments_data/chroma_bge_v2)
FRONTEND_ORIGIN backend env CORS allowlist (comma-sep or *)
VITE_API_URL frontend build env Backend base URL (set in Vercel)
HF_HOME, SENTENCE_TRANSFORMERS_HOME, TRANSFORMERS_CACHE Docker model caches → /tmp/hf

9. Local development

# Backend (needs private judgment/statute snapshots and DEEPSEEK_API_KEY).
# Production downloads them automatically; new environments use MOONLEY_DATA and
# MOONLEY_STATUTE_CHROMA. Legacy THEMIS_* names remain compatibility fallbacks.
cp .env.example .env            # add DEEPSEEK_API_KEY
pip install -r backend/requirements.txt
uvicorn app:app --app-dir backend --host 0.0.0.0 --port 8000

# Frontend (Vite dev server proxies /api -> http://localhost:8000)
cd frontend && npm install && npm run dev

First /ask is slow (loads both indices + builds BM25 over 28,612 chunks + models), then warm.


10. Deployment

flowchart LR
  GH["GitHub phase1.1"] -->|auto| VC["Vercel (React UI, root=vercel-frontend/)"]
  GH -->|backend-only orphan push| SP["HF Space (FastAPI API)"]
  SP -->|HF_TOKEN + snapshot_download| JD["Private judgment release"]
  SP -->|HF_TOKEN + snapshot_download| SD["Private statute release"]
  • Frontend → Vercel: the phase1.1 production branch builds from Root Directory = vercel-frontend and publishes only the React UI.
  • Backend → HF Space (Docker SDK): run the deploy script. It creates a clean orphan branch, strips both frontend directories + assets/ (HF rejects un-tracked binaries), and injects the HF config frontmatter into README.md — which is kept out of the GitHub README (GitHub renders frontmatter as an ugly table) — then force-pushes to the Space's main:
    git remote add space https://huggingface.co/spaces/<owner>/themis   # once
    bash scripts/deploy-space.sh
    
  • Container startup: download pinned private judgment and statute releases → verify required files → remove HF_TOKEN from the API process environment → start uvicorn on :7860.
  • Set DEEPSEEK_API_KEY and a read-only HF_TOKEN as Space secrets; restart after rotation.

11. Operational notes

  • Cold start: /health is instant; the first /ask pays the full model+index+BM25 load (~1–2 min on free CPU). Free Spaces sleep after ~48 h idle.
  • Memory: both Chroma collections + in-memory BM25 (28,612 chunks) + 2 transformer models fit in 16 GB; this is the ceiling that motivates V2.
  • Citations: parse_section_references matches BNS 103, BNS Section 103, Section 103 of BNS, and Section 439 CrPC (connector optional).
  • Determinism: temperature=0.1 (answers), 0.0 (expansion/classification).

12. Roadmap → V2

V1 is in-process and monolithic (index baked into the image, BM25 in RAM, models in the serving process) — fine for the pilot corpus, but it cannot reach the full Indian corpus. The scaling design re-platforms retrieval behind the same SSE contract: a managed, sharded hybrid search engine (server-side BM25 + dense ANN), GPU embedding/rerank services, Postgres + object store, and a real ingestion ETL with metadata pre-filtering and court×year sharding — so per-query work stays ~constant as the corpus grows.

is kept with the production documentation on phase1.1.


13. Documentation

Deployment

Architecture and scaling — phase1.1