Spaces:
Running
Running
File size: 29,862 Bytes
917730d 5ea3240 34acd3c 917730d 5ea3240 917730d 5ea3240 f0307a2 5ea3240 13784a6 d6ee0a6 f0307a2 b0af996 f79a242 03bed0b b0af996 0e7bbcc e67c48b 0e7bbcc e67c48b 0e7bbcc e67c48b 27716f7 d6ee0a6 27716f7 d6ee0a6 13784a6 b0af996 13784a6 5ea3240 13784a6 b0af996 13784a6 ef2c57c 13784a6 ef2c57c 13784a6 5ea3240 13784a6 5ea3240 27716f7 5ea3240 13784a6 5ea3240 13784a6 5ea3240 13784a6 5ea3240 13784a6 5ea3240 27716f7 5ea3240 13784a6 5ea3240 13784a6 5ea3240 13784a6 5ea3240 13784a6 5ea3240 13784a6 5ea3240 13784a6 5ea3240 27716f7 5ea3240 13784a6 5ea3240 13784a6 5ea3240 13784a6 5ea3240 27716f7 13784a6 d6ee0a6 5ea3240 13784a6 d6ee0a6 5ea3240 13784a6 5ea3240 13784a6 5ea3240 f4b92b8 5ea3240 13784a6 5ea3240 13784a6 5ea3240 13784a6 5ea3240 ef2c57c bb5d2bb b0af996 ef2c57c 5ea3240 13784a6 5ea3240 13784a6 5ea3240 ef2c57c 5ea3240 b0af996 ef2c57c b0af996 ef2c57c bb5d2bb b0af996 5ea3240 ef2c57c 5ea3240 bb5d2bb 5ea3240 27716f7 bb5d2bb 27716f7 5ea3240 13784a6 5ea3240 13784a6 5ea3240 27716f7 d6ee0a6 13784a6 5ea3240 13784a6 5ea3240 13784a6 5ea3240 13784a6 5ea3240 13784a6 5ea3240 ef2c57c 27716f7 ef2c57c f4b92b8 bb5d2bb 27716f7 f4b92b8 ef2c57c 5ea3240 13784a6 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 | ---
title: RAGForge
emoji: 🔎
colorFrom: gray
colorTo: indigo
sdk: docker
app_port: 7860
pinned: false
---
# RAGForge
**RAGForge v1.9 - a production-style, portfolio-ready agentic RAG and analytical synthesis system for Hugging Face Spaces.**
RAGForge combines hybrid document retrieval, source-level/hierarchical retrieval, semantic query planning, corrective RAG, Self-RAG-style verification, Text2SQL and an “Ask-the-Web” research path in one CPU-friendly application. The default LLM is **Google Gemini 3.5 Flash-Lite**; the UI also exposes Gemini 3.1 Flash-Lite and stronger Flash models.
## What is new in v1.9
v1.9 is the scale, context-economics and release-readiness release. It generalizes v1.8's successful fixed top-3 pruning into a corpus-aware policy and bundles the remaining production/evaluation work needed before treating RAGForge as a finished portfolio system. The new local ablations do not add Gemini requests.
- **Adaptive retrieval depth** - candidate depth expands from 6 to 8/10/12 as corpus size grows, while broad tasks retain additional breadth.
- **Dynamic 2-5 chunk context budgets** - focused local lookups choose their generation budget from corpus scale, retrieval confidence, score separation and source ambiguity instead of always keeping three chunks.
- **Focused evidence compression** - a zero-LLM sentence selector trims only the generation copy of focused evidence; source cards and citation evidence retain the original chunks.
- **Three-way context ablation** - Standard/Deep compare full top-k, the v1.8 fixed top-3 policy and the v1.9 adaptive budget. A second ablation measures adaptive context with/without sentence compression and checks labeled answer-signal retention.
- **1x/5x/20x scale stress** - a deterministic distractor harness clones long-document chunks and reuses existing vectors to exercise the real hybrid retriever at roughly demo, medium and ~1000-chunk scale without embedding or Gemini cost.
- **Overview/insight consistency** - the planner now enforces task/strategy structural consistency, while the overview evaluator accepts semantically richer insight synthesis only when route, breadth, citations and no-web behavior remain correct.
- **Prompt and evidence economics** - traces expose retrieval depth, confidence/gap, context budget, compression, prompt/output/total token estimates and cited-source utilization.
- **Pacing-corrected node latency** - evaluation node summaries remove deliberate free-tier pacing from model-service node time while preserving separate wall/pacing telemetry.
- **Workspace health diagnostics** - capacity utilization, corpus scale, index readiness, estimated vector memory, TTL/idle age and evaluation-history counts are available in the Architecture + API tab and REST diagnostics endpoint.
- **Release readiness** - Standard/Deep apply a transparent checklist across answer quality, grounding, routing, SQL, hard mode, adaptive-budget recall, compression retention and largest-scale stress recall.
No new runtime dependency is introduced in v1.9.
## What is new in v1.8
v1.8 is a context-budget and generation-efficiency release driven by the v1.7 benchmark: Recall@5/Hit@1/MRR were saturated, citation quality reached 100%, but focused QA still carried distractor sources and generation dominated end-to-end latency. The release therefore reduces unnecessary model input without narrowing broad synthesis tasks.
- **Adaptive focused context pruning** - focused local `fact_lookup`/`followup` queries using semantic or hierarchical retrieval keep a conservative three-chunk safety floor and drop only the low-value context tail. Overview, comparison, cross-document and insight/analytical tasks retain broad evidence.
- **Context-budget trace telemetry** - Pipeline Inspector reports chunks, distinct sources, estimated context tokens and reduction percentage before/after pruning, plus the policy reason.
- **Focused prompt budgeting** - ordinary corpus fact lookups no longer include the full session corpus manifest in the generation prompt after routing has already established scope. Broad/mixed tasks retain the manifest. Generation traces expose evidence-context size and estimated prompt tokens.
- **Zero-call context-budget ablation** - Standard/Deep evaluation compare full top-k versus focused pruning on source Precision@5, Recall@5, Hit@1, MRR, context chunks/sources/chars/tokens and median reduction without spending any extra Gemini requests.
- **Grounded-absence hardening** - Hard Mode missing-answer cases pass when either the natural-language absence matcher succeeds or the pipeline explicitly records `grounded_absence=true`, and the table exposes both signals.
- **Reranker diagnostic aligned with runtime policy** - evaluation now states that small-corpus reranking is already skipped and treats the expensive cross-encoder run as an explicit ablation, not as the default recommendation.
- **Focused efficiency telemetry in QA rows** - benchmark rows record pruning use, pre/post context size, manifest inclusion and estimated generation-prompt tokens so latency changes can be tied to actual input-budget changes.
## What is new in v1.7
v1.7 is an evidence-driven correctness, provenance and efficiency release based on the first full v1.6 Hard Mode/profile benchmark. It fixes evaluator/UI artifacts before changing retrieval behavior and turns measured profile/reranker results into explicit runtime policy.
- **Markdown-safe source cards** - retrieved Markdown is rendered as escaped plain text inside uniform source cards, so a snippet beginning with `#` can no longer become a giant UI heading.
- **Graphical latency waterfall** - Pipeline Inspector replaces ASCII `#####` bars with proportional HTML latency bars and exposes the reranker decision reason plus grounded-absence state.
- **Grounded absence handling** - answers such as “the retrieved policy does not mention a dispute fee” are treated as calibrated no-answer responses instead of low-confidence hallucinations. This avoids unnecessary revise calls and gives Hard Mode a robust missing-answer matcher.
- **Markdown-aware citation coverage** - numbered/bulleted claims count even when short, while generic list introductions/headings do not. This fixes false 0% coverage for answers such as the four NIST AI RMF functions.
- **Table-source validity for overviews** - global corpus overviews surface deterministic DuckDB table evidence as `[T#]` sources, so structured claims can be cited validly instead of referring to a table ID absent from the source list.
- **Fresh-vs-saved evaluation provenance** - each saved report gets a run ID and server-boot ID. Fresh execution messages are no longer overwritten by the saved-run selector, and saved-run loading is triggered only by explicit user input.
- **Profile-policy summary** - optional Fast/Balanced/Agentic benchmarking now produces aggregated profile metrics and an evidence-scoped recommendation rather than only six raw rows.
- **Context-efficiency diagnostics** - evaluation now calls out the case where Recall@5 is excellent but Precision@5 is low, separating “found the right source” from “sent too many distractors to generation”.
- **Small-corpus reranker policy tightened** - because v1.6 showed identical source and chunk metrics with/without reranking while adding multi-second latency, the cross-encoder is skipped even in Agentic on small corpora. It remains available for larger corpora.
- **Diagnostic readability** - each `Next:` recommendation starts on its own line.
## What is new in v1.6
v1.6 moves RAGForge from a mostly saturated demo benchmark into harder evaluation and analytical synthesis. The goal is to make the next improvements measurable rather than simply adding more RAG components.
- **Insight Synthesis** - the semantic planner can classify `insight_synthesis` and choose an `analytical` retrieval strategy for questions such as “What does this collection reveal?” or “What trends stand out?”.
- **Document + table analysis in one grounded answer** - analytical retrieval combines source-balanced original document evidence with deterministic DuckDB table summaries/rows. Table-derived claims are cited as `[T#]` alongside `[D#]` documents and `[W#]` web sources.
- **Hard Mode benchmark** - Standard/Deep evaluation now includes paraphrase, distractor, missing-answer, multi-hop, local-freshness, structured filtering, analytical-synthesis and prompt-injection cases. Hard-mode performance contributes to the deterministic score and quality gates.
- **Chunk-level reranker evaluation** - the explicit RRF vs cross-encoder ablation now reports chunk Hit@1/MRR on cases with auditable chunk-content labels, in addition to source-level metrics. This tests whether reranking helps *inside the correct source*, not only whether it finds the right file.
- **Optional profile benchmark** - users can explicitly compare Fast, Balanced and Agentic on a small labeled subset, including accuracy, citation coverage, service latency, estimated LLM calls and reranker/correction usage. It is off by default to protect free-tier quota.
- **Semantic citation attribution** - zero-call citation repair keeps its conservative lexical checks but can use the already-loaded local embedding model as a high-threshold fallback for clearly supported uncited claims.
- **Node latency observability** - query inspection includes a dependency-free node-latency waterfall, and evaluation reports aggregate mean/p50/p95 timings per LangGraph node.
- **Evaluation run history** - each completed run is timestamp-archived inside the workspace. The UI/API expose historical grade/score/citation/hard-mode/latency/request data plus deltas against the previous run of the same depth.
- **v1.5.2 stability retained** - JSON-safe saved reports, table export, incremental Deep, quota pacing, typed Text2SQL evaluation, adaptive reranking and citation cleanup remain intact.
## What is new in v1.5.2
v1.5.2 is a narrow runtime hotfix over v1.5.1. It does not change retrieval, evaluation semantics, dependency pins or the demo benchmark.
- **Quick/Standard/Deep evaluation crash fixed** - `pipeline.py` now explicitly imports the citation normalization and repair helpers introduced in v1.5.1. This fixes `NameError: name 'repair_missing_citations' is not defined` during generated-answer evaluation.
- **Packaging regression coverage** - a dependency-free test now verifies the citation helpers are imported into the pipeline namespace, and the LangGraph-backed helper test invokes the repair path directly when optional runtime dependencies are available.
- **v1.5.1 behavior retained** - saved-report JSON, table export, grouped-citation parsing and conservative citation repair are unchanged.
- **Benchmark compatibility retained** - the benchmark remains `1.5.1` because this hotfix changes runtime wiring only, not benchmark questions or scoring semantics.
## What is new in v1.5
- **Saved evaluation history** - Quick, Standard and Deep reports are stored per workspace with model, benchmark version, corpus version and timestamp. The Evaluation tab can switch between saved runs instantly and compare them side by side without spending Gemini quota again.
- **Incremental Deep evaluation** - when a matching Standard run is saved, Deep reuses that deterministic baseline and runs only the representative judge layer. A Standard → Deep workflow therefore drops from roughly 31 Deep-run requests to about 5 judge requests.
- **Typed Text2SQL evaluation** - benchmark checks compare computed DuckDB scalar values directly, so booleans/numerics are judged as typed values rather than fragile Markdown strings. Correct SQL such as `weekend_support = true` is no longer penalized because of rendering differences.
- **Adaptive reranking** - the cross-encoder remains available, but Fast mode and small/easy corpus queries skip it when the benchmark shows no measurable source-ranking gain. Comparison/cross-document work, larger corpora and Agentic mode can still use it.
- **Zero-call citation repair** - an evidence-aware deterministic pass repairs only clearly supported uncited factual units, improving citation completeness without adding another Gemini call.
- **Planner taxonomy refinement** - structured-data examples explicitly distinguish direct table lookups from cross-row min/max aggregation.
- **Evaluation API history** - FastAPI can list and retrieve saved Quick/Standard/Deep reports, and `POST /api/v1/evaluate/demo` can reuse compatible cached evaluations.
- **v1.4.1 quota protections retained** - rolling RPM pacing, provider-aware 429 backoff, pacing-aware latency, sampled Deep judging and request telemetry remain enabled.
## The retrieval philosophy
RAGForge separates three questions that simple RAG demos often collapse into one:
1. **What does the user mean?** A schema-constrained semantic planner classifies knowledge scope, task type, freshness, retrieval strategy and independent document/web queries.
2. **Where should evidence come from?** Uploaded/private corpus, structured tables, the external web, or a genuinely mixed task.
3. **How should that evidence be retrieved?** Focused semantic chunk search, global source-balanced retrieval, hierarchical source-to-chunk retrieval, analytical document+table synthesis, Text2SQL, or web search.
This matters for ambiguous language. For example, **“What is the current corpus about?”** is session-local: “current” modifies the indexed corpus and does *not* imply current-world freshness. Conversely, **“Compare our NIST document with the latest NIST guidance online”** is a mixed task and legitimately needs both document and web evidence.
The behavior is learned/planned semantically; the application does not hard-code phrases such as “what is the corpus about”. Those phrases exist only as regression tests.
## Why this is more than “chat with a PDF”
### Semantic planning and task-aware routing
- **Balanced/Agentic semantic query analyzer** with structured JSON output
- distinguishes `corpus`, `external`, `mixed`, and `structured_data` knowledge scopes
- task classification: fact lookup, overview, cross-document synthesis, comparison, aggregation, insight synthesis, follow-up
- retrieval strategy selection: `semantic`, `global`, `hierarchical`, `analytical`, `table`, `none`
- separate `document_queries` and `web_queries` instead of blindly sending the same string to every retriever
- explicit `web_relevance = required | useful | irrelevant`
- corpus manifest and recent conversation are supplied to the planner so pronouns/session-local references can be resolved
- manual Auto/Documents/Web/Hybrid/Data(SQL) controls remain available and override the semantic route when selected
### Hierarchical and global retrieval
RAGForge maintains **two retrieval indexes per session**:
1. **Chunk index** - normal evidence chunks for answer generation.
2. **Source-profile index** - one compact, deterministic profile per ingested source, built from source metadata and representative excerpts.
The source-profile index is retrieval-only; generated answers are still cited against original document chunks.
This supports:
- **semantic retrieval** - dense + BM25 over all chunks for focused questions
- **hierarchical retrieval** - source-profile retrieval first, then chunk retrieval restricted to selected sources
- **global/source-balanced retrieval** - one query-relevant representative evidence chunk per source for corpus overviews and broad synthesis
- source diversity for cross-document/comparison tasks, reducing long-document domination
A 48-page PDF therefore cannot monopolize an overview simply because it produced more chunks than four short files.
### Hybrid search
- local CPU embeddings with **FastEmbed / BAAI bge-small-en-v1.5**
- embedded **Qdrant** vector store per user session
- in-memory normalized embedding matrix for efficient source-scoped hierarchical search
- **BM25** lexical retrieval
- **reciprocal-rank fusion (RRF)**
- local **cross-encoder reranking** (`Xenova/ms-marco-MiniLM-L-6-v2`) with an adaptive runtime policy: Fast/small-corpus cases can skip it when the measured latency cost is not justified, while harder/larger cases can retain it
- sentence-aware chunking plus optional **semantic breakpoint chunking**
- source/page metadata
- suspicious retrieved prompt-injection text is down-weighted
### Corrective RAG (CRAG)
The corrective loop no longer means “low score → Google”. It is:
`retrieve → task-aware evidence grade → correct/rewrite retrieval plan → retrieve again → re-grade → conditional web only if web is semantically relevant`
Evidence grading considers:
- dense/BM25 relevance signals rather than raw RRF or reranker logits
- top-hit and top-3 evidence quality
- lexical/semantic retriever agreement
- distinct-source coverage when the task requires breadth
- an optional semantic LLM grader for borderline/Agentic cases
If a corpus-only query remains unsupported after correction, RAGForge can **abstain** instead of contaminating the answer with unrelated internet search.
### Better relevance observability
Cross-encoder logits are useful for ranking but are not calibrated human probabilities. The UI therefore shows:
- retrieval rank
- bounded hybrid relevance signal from dense/BM25 evidence
Raw RRF/dense/sparse/reranker values remain available in the returned source metadata/pipeline trace for debugging.
### Agentic RAG
- **LangGraph** state machine with conditional edges
- history-aware semantic rewriting
- optional **multi-query expansion**
- optional **HyDE** hypothetical-document retrieval in Agentic mode
- query correction/retrieval retry loop
- **Self-RAG-style** answer audit and one bounded revision loop
- response confidence score and full pipeline trace, including whether reranking was used and how many citations were deterministically repaired
- process-level TTL response caching, isolated by session + corpus version
- bounded exponential-backoff retries for transient Gemini API failures
### Ask-the-Web
- free/keyless **DuckDuckGo** fallback
- optional Tavily provider
- native Gemini Google Search provider using a separately configurable grounding submodel (`gemini-2.5-flash-lite` by default)
- independently planned web queries
- query fan-out, parallel page fetching, main-text extraction with Trafilatura, local reranking, Gemini synthesis and URLs in the source panel
- SSRF-oriented URL checks; local/private network targets are rejected
- web permission and web relevance are separate: checking **Allow web fallback** does not force web search
### Documents and data
- PDF, TXT, Markdown, DOCX, PPTX, CSV, XLS/XLSX, JSON, HTML, source-code/text formats, images and **ZIP archives**
- safe ZIP extraction (path traversal, file-count, uncompressed-size and type limits)
- page-aware PDF extraction
- optional **Gemini OCR/document transcription** for scanned PDFs/images
- CSV/XLSX indexed as text **and** loaded into isolated **DuckDB**
- natural-language **Text2SQL** with single-statement read-only SQL validation and row limits
- one-click bundled demo corpus
### Production/demo engineering
- **FastAPI** REST backend + **Gradio** UI in one Docker Space
- optional Bearer auth for API write endpoints
- per-session corpora and in-memory databases; TTL cleanup
- UI + REST per-IP rate limiting
- Prometheus `/metrics`
- health/info/session/status/ingest/query/evaluation endpoints plus saved-evaluation listing/retrieval
- no API keys committed to the repo
- pytest tests + GitHub Actions CI
- pipeline inspector exposes semantic plan, retrieval strategy, source selection, evidence grade, corrective plan, web decision, Self-RAG result and cache hits
## Architecture
```mermaid
flowchart TD
U[User / API] --> G[Input + upload guardrails]
G --> P{Semantic query planner}
P -->|structured data| SQL[Read-only DuckDB Text2SQL]
P -->|external| W[Ask-the-Web]
P -->|corpus / mixed| RS{Retrieval strategy}
RS -->|semantic| C[Chunk index]
RS -->|global| GB[Source-balanced global retrieval]
RS -->|hierarchical| SP[Source-profile index]
SP --> SS[Select relevant sources]
SS --> C
C --> D[Dense FastEmbed/Qdrant]
C --> B[BM25]
D --> F[RRF]
B --> F
F --> RP{Adaptive reranker policy}
RP -->|skip easy/small| E{Task-aware evidence grader}
RP -->|use harder/larger| X[Cross-encoder reranker]
GB --> E
X --> E
E -->|sufficient| A[Gemini generation]
E -->|weak, first attempt| CR[Correct query / strategy]
CR --> RS
E -->|weak + web relevant| W
E -->|weak + web irrelevant| AB[Abstain]
W --> A
A --> V{Self-RAG audit}
V -->|revise once| RV[Faithfulness revision]
RV --> V
V -->|pass| O[Cited answer + sources + trace]
SQL --> O
AB --> O
```
## Deploy on Hugging Face Spaces
1. Create a **Docker** Space.
2. Extract/copy this repository into the Space repo root.
3. In **Settings → Secrets**, add `GEMINI_API_KEY`; optionally `TAVILY_API_KEY` and `APP_API_TOKEN`.
4. Push. The Dockerfile serves `uvicorn` on port `7860`.
5. Leave **Use bundled demo files** checked and either click **Index corpus** or ask a non-Web question; v1.4 can lazily initialize the demo corpus. Then test the Balanced/Auto profile.
The Dockerfile explicitly owns `/tmp/ragforge` and model-cache directories as the non-root runtime user so per-session workspaces can be created safely on Hugging Face Spaces.
> **Public demo key warning:** public visitors consume your server-side Gemini quota. RAGForge adds per-IP limits, but a heavily shared Space should lower limits, add authentication, or require bring-your-own-key.
>
> **Privacy caveat:** use only documents appropriate for the terms of the model/API tier you select.
## Run locally
```bash
cp .env.example .env
python -m venv .venv
source .venv/bin/activate # Windows: .venv\\Scripts\\activate
pip install -r requirements-dev.txt
uvicorn app:app --reload --port 7860
```
Or:
```bash
docker build -t ragforge .
docker run --rm -p 7860:7860 -e GEMINI_API_KEY=YOUR_KEY ragforge
```
## API example
```bash
curl -X POST http://localhost:7860/api/v1/session
curl -X POST http://localhost:7860/api/v1/ingest \
-F session_id=SESSION_ID \
-F files=@demo_documents/acme_cloud_runbook.md
curl -X POST http://localhost:7860/api/v1/query \
-H 'Content-Type: application/json' \
-d '{
"session_id": "SESSION_ID",
"query": "What is this corpus about?",
"config": {"mode": "Auto", "profile": "Balanced", "model": "gemini-3.5-flash-lite"}
}'
curl -X POST http://localhost:7860/api/v1/evaluate/demo \
-H 'Content-Type: application/json' \
-d '{
"session_id": "SESSION_ID",
"level": "Standard",
"model": "gemini-3.5-flash-lite",
"target_rpm": 12,
"reuse_saved": true,
"include_profile_benchmark": false
}'
```
## Pipeline profiles
| Profile | Intended use | Behavior |
|---|---|---|
| **Fast** | cheapest interactive path | narrow deterministic route, single document query, hybrid retrieval, answer generation |
| **Balanced** | default portfolio UX | semantic structured planner, task-aware retrieval, optional borderline evidence grader, one corrective retrieval retry, conditional web, answer |
| **Agentic** | showcases advanced RAG | semantic planner, multi-query, optional HyDE, semantic evidence grading, correction loop, conditional web, answer, Self-RAG verification and bounded revision |
The feature switches remain independent so interviewers can see what each technique changes.
## Evaluation
The built-in **Evaluation** tab is now a layered benchmark instead of a single smoke-test JSON dump. The labels live in `evals/demo_benchmark.json`, so the expected behavior is reviewable rather than hidden in evaluator prompts.
### Standard deterministic metrics
The original demo benchmark became nearly saturated by v1.5. v1.6 deliberately adds harder robustness cases so future changes must improve behavior rather than merely preserve an easy score.
- focused answer-key accuracy
- source Precision@5, Recall@5, MRR and AP@5
- citation validity and sentence-level citation coverage across `[D#]`, `[T#]` and `[W#]` evidence
- planner route, task and retrieval-strategy accuracy
- web-use precision/recall and unnecessary-web rate
- corpus-overview source coverage and no-unnecessary-web regressions
- explicit empty-workspace abstention correctness
- Text2SQL read-only SQL generation/execution checks; SQL routing is measured separately in the planner suite
- service latency p50/p95, pacing/wall time, correction rate, runtime web-use rate, request count and estimated LLM calls
- retrieval ablation comparing Hybrid RRF with Hybrid + local cross-encoder reranking, including chunk Hit@1/MRR on explicitly labeled cases
- hard-mode robustness across paraphrase, distractor, missing-answer, multi-hop, insight-synthesis, structured filtering, local-freshness and prompt-injection cases
- optional Fast/Balanced/Agentic profile benchmark for quality/latency/LLM-call tradeoffs
- aggregated LangGraph node latency mean/p50/p95
- timestamped evaluation history and score/latency deltas inside the current workspace
### Deep LLM-as-judge metrics
Deep mode adds Gemini scores for **faithfulness, answer relevance, completeness and citation support** on a representative labeled sample, reducing free-tier request pressure while retaining diverse judge coverage. These judge scores are kept separate from deterministic metrics because an LLM judge is probabilistic and should not be treated as ground truth. The metric families mirror common RAG evaluation practice: retrieval quality is evaluated separately from generation faithfulness/relevance.
The UI exposes **Quick**, **Standard** and **Deep** modes and renders a score card plus per-layer tables, with the full report still available as JSON. v1.5 saves the latest run of each depth per workspace, provides a side-by-side comparison table, and lets users switch among saved reports without rerunning. A compatible saved Standard report can act as the deterministic baseline for incremental Deep judging. The benchmark is intentionally small and corpus-specific; it is a regression/architecture-validation suite, not a claim of general RAG benchmark performance.
Evaluation defaults to **quota-safe pacing at 12 RPM**. The active Gemini limit is project/model specific, so use the value shown for your project in Google AI Studio and set the evaluation target below it. A Standard run uses typed one-call Text2SQL component checks; a Deep run after a compatible saved Standard normally needs only the representative judge calls rather than repeating the full deterministic benchmark.
## Model and dependency note
The Space pins a compatible deployment set around Gradio 5 (`gradio==5.49.1`, `pydantic==2.11.10`, `google-genai==2.8.0`, `fastapi==0.116.1`) to avoid dependency-resolver conflicts while retaining the Gemini Interactions API used by RAGForge. Model names are still runtime-selectable in the UI.
## ZIP support
ZIP upload is useful for testing a miniature knowledge base, but archives are treated as hostile input. Defaults cap compressed upload size, extracted file count, total uncompressed content and supported types; traversal paths are rejected.
## Privacy and persistence
The UI persists only an opaque session ID in browser local storage. Corpus contents, embeddings, DuckDB tables, chat history and saved evaluation reports remain server-side. A normal browser refresh can reconnect while the Space process is alive; a Hugging Face container restart still removes the in-memory/ephemeral workspace. Demo mode can rebuild automatically, while custom uploads and saved evaluations must be recreated after a restart.
This public-demo build intentionally uses per-session ephemeral storage, embedded Qdrant and in-memory DuckDB. A real multi-tenant deployment should replace these with authenticated object storage, tenant-filtered managed retrieval, durable sessions and governed structured-data access.
## Enterprise migration path
- Qdrant local → Qdrant Cloud / managed vector DB
- in-process TTL cache → Redis
- in-memory DuckDB → governed warehouse / Postgres read replica
- process-local workspace registry → durable session/tenant service
- local metrics → OpenTelemetry + centralized observability
- simple API token → OAuth/OIDC + tenant-aware authorization
- synchronous ingestion → object-storage events + workers
## Project documentation
- `docs/FEATURE_MATRIX.md` - implementation checklist and interview rationale
- `docs/EVALUATION.md` - benchmark methodology, saved-run reuse, typed Text2SQL checks and quota behavior
- `docs/QUERY_PLANNING.md` - semantic planner, hierarchical retrieval and CRAG policy
- `docs/UX_LIFECYCLE.md` - browser/session lifecycle, lazy demo initialization and indexing UX
- `docs/MIGRATION_1.3.md` - v1.2 to v1.3 upgrade notes
- `docs/MIGRATION_1.4.md` - v1.3 to v1.4 upgrade notes
- `docs/MIGRATION_1.4.1.md` - quota-safe evaluation and runtime-view stabilization patch
- `docs/MIGRATION_1.5.md` - saved/incremental evaluation, typed Text2SQL checks and adaptive-reranking upgrade notes
- `docs/ARCHITECTURE_API.md` - live architecture/API surface and endpoint examples
- `docs/MIGRATION_1.2.md` - v1.1 to v1.2 upgrade notes
- `docs/architecture.mmd` - Mermaid architecture source
- `docs/RESUME_BULLETS.md` - resume-ready wording
- `docs/DEMO_DATASETS.md` - corpus-extension ideas
- `docs/SOURCES.md` - architectural references and official implementation docs
- `SECURITY.md` - threat model and residual risk
## License
MIT for this project’s source and synthetic demo files. Bundled/third-party documents retain their original terms.
|