title: Moonley API
sdk: docker
app_port: 7860
pinned: false
Moonley
Grounded AI legal research for Indian law. Moonley answers natural-language questions by retrieving the relevant statutes (IPC · BNS · CrPC · BNSS · IEA · BSA) and Supreme Court judgments, fusing them on a single cross-encoder, generating an answer with DeepSeek, and verifying every cited section/case against the retrieved evidence — streaming each reasoning step to the UI over Server-Sent Events.
- Live app: moonley-pilot.vercel.app (React/Vite SPA)
- Backend: Hugging Face Spaces (FastAPI/Docker, API only)
- Production branch:
phase1.1· Repository:vg15o2/themis
Table of contents
- Architecture
- Tech stack
- Repository layout
- The retrieval pipeline (in depth)
- Data & indices
- Backend API contract (SSE)
- Frontend
- Configuration
- Local development
- Deployment
- Operational notes
- Roadmap → V2
- Documentation
1. Architecture
flowchart TB
UI["React SPA (Vercel)<br/>SSE reasoning UI"] -- "POST /ask (SSE)" --> API
subgraph Space["HF Spaces · Docker · CPU 16GB"]
API["FastAPI<br/>backend/app.py"] --> R["Unified router<br/>hybrid_rag/unified_legal_Rag.py"]
R --> SR["Statute pipeline<br/>statute_retrieval.py"]
R --> JR["Judgment pipeline<br/>llm_retriever.py"]
R --> VER["Verifier<br/>verifier/verified.py"]
SR --> CS[("Chroma<br/>indian_statutes")]
JR --> CJ[("Chroma<br/>sci_judgments_bge_v2")]
JR --> BM["BM25Okapi (in-memory)"]
SR --> M["BGE-small + ms-marco CE"]
JR --> M
end
R -- "stream" --> DS["DeepSeek<br/>deepseek-chat"]
DSset["HF Dataset<br/>vg15o2/themis-judgments"] -. "downloaded at build" .-> CJ
Request lifecycle: route_query decides statute / judgment / hybrid → each enabled
pipeline retrieves candidates → unified_rerank re-scores all candidates on one
cross-encoder (dedupe + per-source cap) → an intent-specific prompt + evidence is streamed
through DeepSeek → check_grounding validates citations → done. Every stage is emitted
as a typed SSE event.
2. Tech stack
| Layer | Choice | Detail |
|---|---|---|
| Frontend | React 18 + Vite 5 | SSE via fetch+ReadableStream; react-markdown+remark-gfm; deployed on Vercel (root frontend/) |
| Backend | FastAPI + Uvicorn | Single SSE endpoint; lazy in-process pipeline load |
| Backend host | HF Spaces (Docker SDK) | 2 vCPU / 16 GB free tier; listens on :7860 |
| Vector store | ChromaDB PersistentClient |
2 collections, local on-disk |
| Embeddings | BAAI/bge-small-en-v1.5 |
384-dim, normalize_embeddings=True; query prefix "Represent this sentence for searching relevant passages: " |
| Reranker | cross-encoder/ms-marco-MiniLM-L-6-v2 |
statute rerank, judgment child/parent rerank, and the unified cross-source rerank |
| Lexical | rank_bm25.BM25Okapi |
built in-memory over all judgment chunks at startup |
| LLM | DeepSeek deepseek-chat |
OpenAI-compatible client, base_url=https://api.deepseek.com, stream=True, temperature=0.1, max_tokens=2500 |
3. Repository layout
backend/
app.py FastAPI app · POST /ask (SSE) · citation links · lazy get_pipeline()
requirements.txt fastapi, uvicorn, chromadb, sentence-transformers, openai, rank-bm25, huggingface_hub, python-dotenv
hybrid_rag/
unified_legal_Rag.py route_query · get_statute_evidence · get_judgment_evidence · unified_rerank · classify_intent · PROMPTS · Evidence
statute corpus/
statute_retrieval.py Chroma `indian_statutes` · parse_section_references · direct_lookup · expand_query · semantic_search · rerank_results
(serving data is fetched from a private, pinned Hugging Face dataset at runtime)
llm_retriever.py Chroma `sci_judgments_bge_v2` · dense_search · bm25_search · rrf_fusion · rerank_children · fetch_parent_chunks · rerank_parents · retrieve
verifier/
verified.py check_grounding (Tier-1/1.5) · verify_citations_live (Tier-2, lazy bharat_courts)
__init__.py
frontend/
src/App.jsx turns · ReasoningPanel · StepTimeline · Answer · CopyButton · intent dropdown · stop/new-chat
src/api.js streamAsk(query, history, onEvent, signal, intent) — SSE parser
src/index.css Harvey-inspired warm-light theme
public/icon.svg, favicon.ico
Dockerfile deps → COPY → private runtime artifact download → uvicorn
render.yaml (legacy) Render blueprint
themis/ current architecture, scaling and storyline documents
4. The retrieval pipeline (in depth)
4.1 Routing — unified_legal_Rag.route_query(query) -> RouteDecision
explicit INSC citation (judg_rag.extract_citations) AND not statute_signal -> judgment-only
section reference (stat_rag.parse_section_references) AND not judgment_signal -> statute-only
otherwise -> hybrid (both)
Signal regexes: JUDGMENT_SIGNAL_RE (case/judgment/held/INSC/ratio/…), STATUTE_SIGNAL_RE
(section/provision/IPC/BNS/…).
4.2 Statute path — get_statute_evidence(query) → statute_retrieval
classify_query→{type: direct_lookup|semantic|hybrid, section_refs, acts_mentioned, is_comparative}.direct_lookup(act, sec)for each parsed(ACT, section)→ exact section (score999.0).expand_query(query)→ DeepSeek rewrites colloquial → statutory language (2–3 variants).semantic_search(query, top_k=40, act_filter)over Chroma (+ per-expansion and per-act searches, deduped by(act_short, section_number)).rerank_results(rerank_query, candidates, top_k=15)cross-encoder → wrapped asEvidence.
4.3 Judgment path — get_judgment_evidence(query) → llm_retriever.retrieve
Query is enriched with severity context (offence title from the referenced section) +
an LLM keyword expansion. retrieve has three modes:
- citation (1 INSC cite) → exact parent via
citation_search. - comparison (≥2 cites + comparison words) → each parent.
- hybrid →
dense_search(top 150)+bm25_search(top 150)→rrf_fusion(k=60)→deduplicate_children(≤3/case)→rerank_children(top 30)→rank_cases_from_children→fetch_parent_chunks(top 20)→rerank_parents(top 5).
4.4 Unified rerank — unified_rerank(query, evidence)
Exact hits (score==999) pinned on top; everything else scored on the same
ms-marco cross-encoder so statute and judgment candidates are comparable. Then:
dedupe by ("s", act, section) / ("j", neutral_citation), per-source cap
MAX_SINGLE_TYPE_SHARE=6, final FINAL_EVIDENCE_N=8. The rerank query is
severity-enriched so "BNS 103" biases toward murder-class judgments.
4.5 Generation
classify_intent(query) (defaults to LEGAL_RESEARCH) selects one of five system prompts in
PROMPTS (research citation-table / case summary / comprehensive study / comparison / story
evaluation). All prompts enforce evidence-grounding (no model knowledge beyond evidence).
The backend streams stat_rag.llm_client.chat.completions.create(..., stream=True) and emits
each delta as a token event. A forced intent from the UI overrides classification.
4.6 Verification — verifier.check_grounding(answer, final_evidence, stat_rag, judg_rag)
Strips markdown bold, extracts cited sections + INSC citations, then:
- cited section in retrieved evidence → grounded; in full local DB but not retrieved → retrieval miss (real law); not in DB → hallucination (flagged).
- cited case not in evidence → ungrounded citation (flagged).
grounded = no hallucinated sections AND no ungrounded citations. Tier-2 (verify_citations_live, livebharat_courts+ 30-day cache) is lazily imported and currently deferred.
5. Data & indices
| Corpus | Chroma collection | Count | On-disk | Provisioning |
|---|---|---|---|---|
| Statutes (6 acts) | indian_statutes |
2,353 sections | private artifact | Downloaded at startup from a pinned private Hugging Face dataset using the Space secret |
| SCI judgments | schema-v5 FAISS + SQLite | release-defined | private artifact | Downloaded at startup from a pinned private Hugging Face dataset using the Space secret |
- Parent-child chunking: parent = full judgment + metadata; children = ~512-token windows
(100 overlap). IDs:
<citation>__child_NNNN,<citation>__parent. - Judgment metadata (per chunk):
case_name, neutral_citation, court, date, bench, author_judge, acts, sections, issue, short_summary, full_headnote, outcome, source_url, …. - Statute record:
{metadata: {act_short, act_name, section_number, title}, retrieval_text}. - Neither statute nor judgment serving artifacts are committed to GitHub or baked into the
public Space image.
start_private_space.pydownloads both pinned snapshots withHF_TOKEN, verifies their required entrypoints, removes the token from the API process environment, and then starts FastAPI.
6. Backend API contract (SSE)
POST /ask → text/event-stream
// request
{ "query": "string",
"history": [{"role": "user|assistant", "content": "..."}],
"intent": "AUTO | LEGAL_RESEARCH | CASE_SUMMARY | COMPREHENSIVE_CASE_STUDY | CASE_COMPARISON | STORY_EVALUATION" }
// events — each emitted as `data: {json}\n\n`
{ "type":"step", "phase":"planning|retrieval|rerank|answer|verify", "title":"...", "detail":"..." }
{ "type":"evidence", "items":[ /* statute or judgment items, see below */ ] }
{ "type":"token", "delta":"..." }
{ "type":"verify", "grounded":true, "hallucinated_sections":[], "retrieval_miss_sections":[], "unverified_citations":[] }
{ "type":"done", "answer":"...", "intent":"LEGAL_RESEARCH", "route":"...", "citations":[...], "elapsed_seconds":12.3 }
{ "type":"error", "message":"..." }
// evidence items
{ "kind":"statute", "act":"BNS", "section":"103", "title":"Punishment for murder.", "score":1.23, "url":"https://indiankanoon.org/search/?formInput=..." }
{ "kind":"judgment", "case":"Sanjay Kumar Sharma v. State of Bihar", "citation":"2026 INSC 223", "title":"<issue/summary>", "score":1.23, "url":"<source_url>" }
GET /health → {status, service, pipeline_loaded} · GET / → service info.
Notes: an immediate step:"Warming up" is flushed before the lazy pipeline load so the
SSE connection opens promptly; CORS origin via FRONTEND_ORIGIN (default *).
7. Frontend
streamAsk(query, history, onEvent, signal, intent)POSTs JSON and parses the SSE frame stream (\n\n-delimiteddata:lines) off theReadableStream.- Per-turn state
{query, steps[], evidence[], answer, citations[], verify, done, intent, elapsed_seconds}. - ReasoningPanel — collapsed-by-default "Thinking…" disclosure; expands to the live
step timeline. Answer — markdown + top "Copy". Copy answer + sources — appends a
formatted
Sources:block. Stop —AbortController.abort()(keeps partial answer). New chat — clears turns + aborts. Intent dropdown — forces the answer style. VITE_API_URLselects the backend (falls back to/api, proxied tolocalhost:8000in dev).
8. Configuration
| Var | Where | Purpose |
|---|---|---|
DEEPSEEK_API_KEY |
backend env / HF secret | DeepSeek auth (required; checked before pipeline import) |
CLERK_PUBLISHABLE_KEY |
backend env / HF variable | Public Clerk application key returned to both standalone frontends |
CLERK_SECRET_KEY |
backend env / HF secret | Clerk backend API credential; never expose in frontend code |
CLERK_JWT_KEY |
backend env / HF secret | Optional PEM public key for networkless session-token verification |
CLERK_AUTHORIZED_PARTIES |
backend env / HF variable | Exact Moonley and local browser origins allowed by the API |
JUDGMENTS_CHROMA_PATH |
backend env | Path to the judgments Chroma dir (Docker sets /app/judgments_data/chroma_bge_v2) |
FRONTEND_ORIGIN |
backend env | CORS allowlist (comma-sep or *) |
VITE_API_URL |
frontend build env | Backend base URL (set in Vercel) |
HF_HOME, SENTENCE_TRANSFORMERS_HOME, TRANSFORMERS_CACHE |
Docker | model caches → /tmp/hf |
9. Local development
# Backend (needs private judgment/statute snapshots and DEEPSEEK_API_KEY).
# Production downloads them automatically; new environments use MOONLEY_DATA and
# MOONLEY_STATUTE_CHROMA. Legacy THEMIS_* names remain compatibility fallbacks.
cp .env.example .env # add DEEPSEEK_API_KEY
pip install -r backend/requirements.txt
uvicorn app:app --app-dir backend --host 0.0.0.0 --port 8000
# Frontend (Vite dev server proxies /api -> http://localhost:8000)
cd frontend && npm install && npm run dev
First /ask is slow (loads both indices + builds BM25 over 28,612 chunks + models), then warm.
10. Deployment
flowchart LR
GH["GitHub phase1.1"] -->|auto| VC["Vercel (React UI, root=vercel-frontend/)"]
GH -->|backend-only orphan push| SP["HF Space (FastAPI API)"]
SP -->|HF_TOKEN + snapshot_download| JD["Private judgment release"]
SP -->|HF_TOKEN + snapshot_download| SD["Private statute release"]
- Frontend → Vercel: the
phase1.1production branch builds from Root Directory =vercel-frontendand publishes only the React UI. - Backend → HF Space (Docker SDK): run the deploy script. It creates a clean orphan
branch, strips both frontend directories +
assets/(HF rejects un-tracked binaries), and injects the HF config frontmatter intoREADME.md— which is kept out of the GitHub README (GitHub renders frontmatter as an ugly table) — then force-pushes to the Space'smain:git remote add space https://huggingface.co/spaces/<owner>/themis # once bash scripts/deploy-space.sh - Container startup: download pinned private judgment and statute releases → verify required
files → remove
HF_TOKENfrom the API process environment → startuvicornon:7860. - Set
DEEPSEEK_API_KEYand a read-onlyHF_TOKENas Space secrets; restart after rotation.
11. Operational notes
- Cold start:
/healthis instant; the first/askpays the full model+index+BM25 load (~1–2 min on free CPU). Free Spaces sleep after ~48 h idle. - Memory: both Chroma collections + in-memory BM25 (28,612 chunks) + 2 transformer models fit in 16 GB; this is the ceiling that motivates V2.
- Citations:
parse_section_referencesmatchesBNS 103,BNS Section 103,Section 103 of BNS, andSection 439 CrPC(connector optional). - Determinism:
temperature=0.1(answers),0.0(expansion/classification).
12. Roadmap → V2
V1 is in-process and monolithic (index baked into the image, BM25 in RAM, models in the
serving process) — fine for the pilot corpus, but it cannot reach the full Indian corpus. The scaling design
re-platforms retrieval behind the same SSE contract: a managed, sharded hybrid search
engine (server-side BM25 + dense ANN), GPU embedding/rerank services, Postgres +
object store, and a real ingestion ETL with metadata pre-filtering and court×year
sharding — so per-query work stays ~constant as the corpus grows.
is kept with the production documentation on phase1.1.