Spaces:
Sleeping
Sleeping
| # DocAgent β Project Overview, Research, Design & How It Works | |
| **Project:** Document Processing AI Agent (chat-based, 40/60 UI) | |
| **Org:** Monkhub Innovations | |
| **One-line:** Upload any document β it parses, classifies, extracts structured | |
| data, summarizes, answers grounded questions, flags anomalies, and exports clean data. | |
| This single document brings together **what we researched, what we chose and why, | |
| the comparison, the feature set, and exactly how the agent works end-to-end.** | |
| For the full deep-dive research see [`RESEARCH.md`](RESEARCH.md); for the layered | |
| architecture record see [`ARCHITECTURE.md`](ARCHITECTURE.md); to deploy see | |
| [`DEPLOY.md`](DEPLOY.md). | |
| --- | |
| ## 1. Research β what we evaluated | |
| We needed a foundation for a **custom, chat-based, production document-processing | |
| agent** with a 40/60 split UI and a **free-LLM-first** model strategy. We evaluated | |
| the five leading systems, one per relevant archetype: | |
| | Archetype | System | Why it represents the category | | |
| |---|---|---| | |
| | Full open-source app | **RAGFlow** | A complete chat-over-documents product with deep understanding, citations, and an agent layer β the closest existing thing to our target. | | |
| | LLM app platform | **Dify** | Reference for clean platform architecture, visual orchestration, observability, polished UI/UX. | | |
| | Extraction toolkit | **Docling** (IBM / LF AI & Data) | Best-in-class open document parsing β layout, tables, OCR, VLM. The extraction *backbone*. | | |
| | Workflow automation | **n8n** | The "glue" archetype: trigger β OCR β LLM extract β validate β route, assembled visually. | | |
| | Agentic framework | **LlamaIndex ADW** | Reference for stateful, multi-step *document agents* and schema-based structured extraction. | | |
| **Method:** sequential, single-researcher web research; each system and source read | |
| in full. Full notes and sources in [`RESEARCH.md`](RESEARCH.md). | |
| --- | |
| ## 2. Comparison table & scoring | |
| **Rubric (1β5; 5 = best)** β scored for *our* specific goal, not in the abstract: | |
| | Code | Parameter | | |
| |---|---| | |
| | A | Feature completeness (OCR, layout, tables, KV, entities, classify, summarize, Q&A, export) | | |
| | B | Architecture quality & extensibility | | |
| | C | Ease of customisation to our UI/flows | | |
| | D | Production-readiness (auth, scaling, observability, deploy) | | |
| | E | UI/UX quality for document work | | |
| | F | Cost / open-source friendliness (license, no per-page cost, self-host) | | |
| | G | Community & documentation | | |
| | System | A | B | C | D | E | F | G | **Total /35** | | |
| |---|---|---|---|---|---|---|---|---| | |
| | **Dify** | 3 | 5 | 5 | 5 | 5 | 4 | 5 | **32** | | |
| | **RAGFlow** | 5 | 4 | 3 | 4 | 4 | 5 | 5 | **30** | | |
| | **Docling** | 5 | 5 | 4 | 4 | 1 | 5 | 5 | **29** | | |
| | **LlamaIndex ADW** | 4 | 5 | 4 | 4 | 2 | 3 | 5 | **27** | | |
| | **n8n workflows** | 3 | 3 | 5 | 4 | 2 | 3 | 5 | **25** | | |
| **How to read it.** Dify wins on raw total, but its lead comes from *platform polish*, | |
| not *document intelligence* (its weakest axis, A=3). RAGFlow and Docling lead on the | |
| axis this project lives or dies on β document depth (A=5). The honest conclusion is | |
| that **no single system is simultaneously the deepest extractor, the cleanest platform, | |
| and a ready-made 40/60 chat UI.** So we composed. | |
| --- | |
| ## 3. What we chose & why | |
| > **Decision: a composed custom stack β not a fork of any one system.** | |
| | Borrowed from | What we took | | |
| |---|---| | |
| | **RAGFlow** | The end-to-end **pipeline blueprint** (deep-understanding-at-ingestion β grounded, citable answers) and citation UX. | | |
| | **Docling** | The **actual extraction engine** β MIT, local, layout + TableFormer tables + OCR, zero per-page cost. | | |
| | **LlamaIndex ADW** | The **agent loop pattern**: *parse β maintain state β retrieve β reason β surface for validation.* | | |
| | **Dify** | **Separation-of-concerns discipline** β thin routes, services, provider abstraction. | | |
| | **Our own** | **FastAPI** backend + **Next.js/React** 40/60 UI tailored to the brief. | | |
| **Why composed instead of forking RAGFlow** (the top *product*): its infra (MinIO + | |
| Elasticsearch + MySQL + Redis) and opinionated UI are too heavy to host our bespoke | |
| 40/60 experience, and the footprint (β₯4 cores / β₯16 GB / β₯50 GB) is unjustified at our | |
| stage. Composing gives us RAGFlow's proven flow, Docling's best-in-class *and free/local* | |
| extraction, LlamaIndex's stateful agent patterns, and Dify's architectural cleanliness β | |
| with none of the licensing, cost, or lock-in downsides. | |
| ### Chosen stack at a glance | |
| | Layer | Choice | Why | | |
| |---|---|---| | |
| | **Extraction** | **Docling** (layout, TableFormer, EasyOCR) | Best open accuracy; MIT; local; $0 per page; offline-capable. | | |
| | **LLM (default)** | **Gemini Flash** | Generous free tier; ~1M-token context for long docs; native vision for scans. | | |
| | **LLM (fallback)** | **GPT-4o-mini** | Cheap, reliable tool-calling; switchable per-request behind a provider interface. | | |
| | **Backend** | **FastAPI** (async, Pydantic v2) | Clean routeβservice layering; SSE streaming; production-ready. | | |
| | **Frontend** | **Next.js (App Router) + React + TS + Tailwind** | 40/60 layout, document viewer with bounding-box overlays, streaming chat. | | |
| | **Vector store** | **Pure-Python numpy cosine** (`.npz`/`.json` on disk) | Dependency-light, no native build/service; chunks carry page+bbox for citations. Swappable for pgvector/Chroma. | | |
| | **Metadata store** | **SQLite** (JSON blobs) | Zero external services; access funnelled through one module β easy Postgres swap. | | |
| **Local-first & ~zero marginal cost:** the whole app runs with only a free Gemini key. | |
| Nothing imports a vendor SDK except `llm/*`, so providers swap by env, per request, or | |
| from the UI. | |
| --- | |
| ## 4. Features β the 7 core functions | |
| | # | Function | What it does | Agent tool | | |
| |---|---|---|---| | |
| | 1 | **Ingest** | Parse PDF, DOCX, PPTX, XLSX, HTML, images with Docling (layout, tables, OCR) β unified Markdown + page images + provenance-carrying chunks. | *(pipeline)* | | |
| | 2 | **Extract** | **Adaptive, template-free** structured data: fields / records / tables / entities, each with a confidence score and best-effort citation. | `get_extracted_data` | | |
| | 3 | **Classify** | Zero-shot document-type detection with confidence + rationale. | `classify_document` | | |
| | 4 | **Summarize** | Map-reduce summary (TL;DR + key points) over chunks. | `summarize_document` | | |
| | 5 | **Q&A** | Conversational, **agentic-RAG** answers **grounded with citations** you can click to highlight on the page. | `query_document` | | |
| | 6 | **Flag anomalies** | Missing fields, low-confidence values, generic arithmetic reconciliation, multi-record sanity, LLM consistency review. | `flag_anomalies` | | |
| | 7 | **Export** | Serialize extracted data to **JSON / CSV / Excel** (records become a table). | *(API)* | | |
| ### The headline feature: adaptive, domain-agnostic extraction | |
| Extraction does **not** rely on per-domain templates. For *any* document β any layout, | |
| any label convention, known type or not β it produces a stable, consistently-keyed shape by: | |
| - **Layout detection** β the model decides whether the doc is one entity | |
| (`single_record`, e.g. an invoice), many similar entities (`multi_record`, e.g. an | |
| employee roster or transaction list), or a header + repeated sub-list | |
| (`mixed`, e.g. invoice + line items). | |
| - **Schema induction** β fields are derived from the actual content (a known type only | |
| *suggests* fields; the model may add/drop to fit reality). | |
| - **Canonical keys** β inconsistent source labels ("Staff Number" / "Personnel ID" / "ID") | |
| are normalised to one `snake_case` key, identical across every record. | |
| - **Grounding** β values come only from the document body (never file metadata); dates and | |
| amounts are normalised (ISO dates, currency); citations are attached by matching values | |
| back to provenance-carrying chunks. | |
| --- | |
| ## 5. How it works β architecture | |
| ``` | |
| ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ | |
| β FRONTEND (Next.js + React + TypeScript + Tailwind) β | |
| β 40% Chat (streaming agent) β 60% Workspace β | |
| β β Tabs: Viewer Β· Fields Β· β | |
| β β Classify Β· Summary Β· Export β | |
| β β Viewer w/ bounding-box cites β | |
| βββββββ¬βββββββββββββββββββββββββββββββββββββββββββββββββββββββββ | |
| β REST + SSE (token streaming) | |
| βββββββΌβββββββββββββββββββββββββββββββββββββββββββββββββββββββββ | |
| β BACKEND (FastAPI, async) β | |
| β api/ documents Β· chat(SSE) Β· actions Β· meta(health) β | |
| β agent/ orchestrator (tool-calling loop) + prompts β | |
| β llm/ provider abstraction: Gemini(default) β OpenAI β | |
| β services/ ingestion(Docling) Β· extraction Β· classification β | |
| β Β· summary Β· qa(RAG) Β· anomaly Β· export Β· β | |
| β vectorstore Β· storage Β· pipeline β | |
| β schemas/ Pydantic contracts (documents, optional hints) β | |
| βββββββ¬ββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββ | |
| β β | |
| βββββββΌββββββ βββββββββββΌβββββββββββ | |
| β Docling β β Storage (/data) β | |
| β (local β β β’ files on disk β | |
| β extract) β β β’ SQLite metadata β | |
| β models β β β’ numpy vectors β | |
| β baked in β β (.npz/.json) β | |
| βββββββββββββ β β’ rendered pages β | |
| ββββββββββββββββββββββ | |
| ``` | |
| **Discipline (borrowed from Dify):** routes are thin β call services; the agent | |
| orchestrates services as tools; LLM access only through `llm/`; storage only through | |
| `services/storage` + `services/vectorstore`; schemas are shared. | |
| --- | |
| ## 6. Working β the two flows | |
| ### Flow A β Ingestion pipeline (runs once per upload, in the background) | |
| Implemented in `services/pipeline.py`. State is written back at every milestone so the | |
| UI can poll status and render results progressively. | |
| ``` | |
| upload β [Stage 1: INGEST]ββfail?βββΊ status=failed (hard fail; only this stage can) | |
| β Docling parse β Markdown + page images + chunks(page,bbox) | |
| βΌ | |
| [index] embed chunks β numpy vector store β | |
| [classify] LLM zero-shot β type + confidence β Stages 2β6 are | |
| [extract] adaptive induction β fields/records/... β best-effort: a | |
| [summarize] map-reduce β TL;DR + key points β failure (e.g. no | |
| [anomalies] rule + LLM checks (needs extraction) β LLM key) degrades | |
| βΌ β gracefully | |
| status=ready (viewer always works, even if AI steps were skipped) | |
| ``` | |
| Key robustness property: **only ingestion can hard-fail a document.** Everything else is | |
| wrapped so a missing API key or a flaky step still yields a `ready` document with a | |
| working viewer and a note about what was skipped. | |
| ### Flow B β Chat agent (per user message) | |
| Implemented in `agent/orchestrator.py` β a bounded **tool-calling loop**, then a | |
| **streamed** final answer. | |
| ``` | |
| user message + history + doc context | |
| β | |
| βΌ | |
| Phase 1: TOOL ROUNDS (up to MAX_TOOL_ROUNDS = 4, non-streamed reasoning) | |
| llm.complete_tools(messages, TOOL_SPECS) | |
| β | |
| βββ model returns tool_calls? ββyesβββΊ run each tool ββΊ append results ββΊ loop | |
| β tools: query_document Β· classify_document Β· | |
| β get_extracted_data Β· summarize_document Β· flag_anomalies | |
| β | |
| βββ no tool_calls (ready to answer) ββΊ exit loop | |
| βΌ | |
| Phase 2: FINAL ANSWER (streamed token-by-token over SSE, temperature 0.4) | |
| β | |
| βΌ | |
| client renders answer + citation chips β click a chip β highlight the | |
| page region (numpy store returned page+bbox with each retrieved chunk) | |
| ``` | |
| `query_document` runs **agentic RAG**: retrieve the most similar chunks, answer from | |
| them, and return grounding passages whose page+bbox the UI turns into clickable | |
| highlights. The orchestrator collects citations in its `ToolContext` so the API can | |
| emit them alongside the streamed answer. | |
| --- | |
| ## 7. End-to-end walkthrough (a user's view) | |
| 1. **Drag a document** into the left rail (or click *New document*). | |
| 2. The pipeline runs: pages render in the viewer first, then **classification, fields, | |
| summary, and anomaly flags** stream into the workspace tabs as each stage finishes. | |
| 3. **Chat** on the left β e.g. *"Summarize this"*, *"What's the total?"*, *"Any missing | |
| fields?"* The agent calls the right tools, then streams a grounded answer. Click any | |
| **citation chip** to highlight its exact source on the page. | |
| 4. **Review & edit** low-confidence fields inline (human-in-the-loop) before exporting. | |
| 5. **Export** from the Export tab as **JSON / CSV / Excel**. | |
| --- | |
| ## 8. Why this design holds up | |
| - **Document depth where it matters** β Docling gives best-in-class layout/table/OCR | |
| locally and for free, the axis on which generic platforms (Dify, n8n) are weakest. | |
| - **Truly document-agnostic** β extraction induces its own schema, so a new document | |
| type needs no code or template changes. | |
| - **Grounded & trustworthy** β every chunk carries page + bbox, so answers and fields | |
| point at their exact source; the single most trust-building UX feature. | |
| - **Cheap and portable** β runs on a free Gemini key with no external services; SQLite | |
| and the numpy store are single-file and swappable for Postgres/pgvector at scale via | |
| one module each. | |
| - **Graceful degradation** β the app stays useful even when AI steps fail. | |
| - **Clean seams for growth** β provider abstraction, storage/vectorstore interfaces, | |
| and a single-orchestrator agent that can grow more tools without a rewrite. | |