File size: 14,612 Bytes
f65e025
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
# DocAgent β€” Project Overview, Research, Design & How It Works

**Project:** Document Processing AI Agent (chat-based, 40/60 UI)
**Org:** Monkhub Innovations
**One-line:** Upload any document β†’ it parses, classifies, extracts structured
data, summarizes, answers grounded questions, flags anomalies, and exports clean data.

This single document brings together **what we researched, what we chose and why,
the comparison, the feature set, and exactly how the agent works end-to-end.**
For the full deep-dive research see [`RESEARCH.md`](RESEARCH.md); for the layered
architecture record see [`ARCHITECTURE.md`](ARCHITECTURE.md); to deploy see
[`DEPLOY.md`](DEPLOY.md).

---

## 1. Research β€” what we evaluated

We needed a foundation for a **custom, chat-based, production document-processing
agent** with a 40/60 split UI and a **free-LLM-first** model strategy. We evaluated
the five leading systems, one per relevant archetype:

| Archetype | System | Why it represents the category |
|---|---|---|
| Full open-source app | **RAGFlow** | A complete chat-over-documents product with deep understanding, citations, and an agent layer β€” the closest existing thing to our target. |
| LLM app platform | **Dify** | Reference for clean platform architecture, visual orchestration, observability, polished UI/UX. |
| Extraction toolkit | **Docling** (IBM / LF AI & Data) | Best-in-class open document parsing β€” layout, tables, OCR, VLM. The extraction *backbone*. |
| Workflow automation | **n8n** | The "glue" archetype: trigger β†’ OCR β†’ LLM extract β†’ validate β†’ route, assembled visually. |
| Agentic framework | **LlamaIndex ADW** | Reference for stateful, multi-step *document agents* and schema-based structured extraction. |

**Method:** sequential, single-researcher web research; each system and source read
in full. Full notes and sources in [`RESEARCH.md`](RESEARCH.md).

---

## 2. Comparison table & scoring

**Rubric (1–5; 5 = best)** β€” scored for *our* specific goal, not in the abstract:

| Code | Parameter |
|---|---|
| A | Feature completeness (OCR, layout, tables, KV, entities, classify, summarize, Q&A, export) |
| B | Architecture quality & extensibility |
| C | Ease of customisation to our UI/flows |
| D | Production-readiness (auth, scaling, observability, deploy) |
| E | UI/UX quality for document work |
| F | Cost / open-source friendliness (license, no per-page cost, self-host) |
| G | Community & documentation |

| System | A | B | C | D | E | F | G | **Total /35** |
|---|---|---|---|---|---|---|---|---|
| **Dify** | 3 | 5 | 5 | 5 | 5 | 4 | 5 | **32** |
| **RAGFlow** | 5 | 4 | 3 | 4 | 4 | 5 | 5 | **30** |
| **Docling** | 5 | 5 | 4 | 4 | 1 | 5 | 5 | **29** |
| **LlamaIndex ADW** | 4 | 5 | 4 | 4 | 2 | 3 | 5 | **27** |
| **n8n workflows** | 3 | 3 | 5 | 4 | 2 | 3 | 5 | **25** |

**How to read it.** Dify wins on raw total, but its lead comes from *platform polish*,
not *document intelligence* (its weakest axis, A=3). RAGFlow and Docling lead on the
axis this project lives or dies on β€” document depth (A=5). The honest conclusion is
that **no single system is simultaneously the deepest extractor, the cleanest platform,
and a ready-made 40/60 chat UI.** So we composed.

---

## 3. What we chose & why

> **Decision: a composed custom stack β€” not a fork of any one system.**

| Borrowed from | What we took |
|---|---|
| **RAGFlow** | The end-to-end **pipeline blueprint** (deep-understanding-at-ingestion β†’ grounded, citable answers) and citation UX. |
| **Docling** | The **actual extraction engine** β€” MIT, local, layout + TableFormer tables + OCR, zero per-page cost. |
| **LlamaIndex ADW** | The **agent loop pattern**: *parse β†’ maintain state β†’ retrieve β†’ reason β†’ surface for validation.* |
| **Dify** | **Separation-of-concerns discipline** β€” thin routes, services, provider abstraction. |
| **Our own** | **FastAPI** backend + **Next.js/React** 40/60 UI tailored to the brief. |

**Why composed instead of forking RAGFlow** (the top *product*): its infra (MinIO +
Elasticsearch + MySQL + Redis) and opinionated UI are too heavy to host our bespoke
40/60 experience, and the footprint (β‰₯4 cores / β‰₯16 GB / β‰₯50 GB) is unjustified at our
stage. Composing gives us RAGFlow's proven flow, Docling's best-in-class *and free/local*
extraction, LlamaIndex's stateful agent patterns, and Dify's architectural cleanliness β€”
with none of the licensing, cost, or lock-in downsides.

### Chosen stack at a glance

| Layer | Choice | Why |
|---|---|---|
| **Extraction** | **Docling** (layout, TableFormer, EasyOCR) | Best open accuracy; MIT; local; $0 per page; offline-capable. |
| **LLM (default)** | **Gemini Flash** | Generous free tier; ~1M-token context for long docs; native vision for scans. |
| **LLM (fallback)** | **GPT-4o-mini** | Cheap, reliable tool-calling; switchable per-request behind a provider interface. |
| **Backend** | **FastAPI** (async, Pydantic v2) | Clean route→service layering; SSE streaming; production-ready. |
| **Frontend** | **Next.js (App Router) + React + TS + Tailwind** | 40/60 layout, document viewer with bounding-box overlays, streaming chat. |
| **Vector store** | **Pure-Python numpy cosine** (`.npz`/`.json` on disk) | Dependency-light, no native build/service; chunks carry page+bbox for citations. Swappable for pgvector/Chroma. |
| **Metadata store** | **SQLite** (JSON blobs) | Zero external services; access funnelled through one module β†’ easy Postgres swap. |

**Local-first & ~zero marginal cost:** the whole app runs with only a free Gemini key.
Nothing imports a vendor SDK except `llm/*`, so providers swap by env, per request, or
from the UI.

---

## 4. Features β€” the 7 core functions

| # | Function | What it does | Agent tool |
|---|---|---|---|
| 1 | **Ingest** | Parse PDF, DOCX, PPTX, XLSX, HTML, images with Docling (layout, tables, OCR) β†’ unified Markdown + page images + provenance-carrying chunks. | *(pipeline)* |
| 2 | **Extract** | **Adaptive, template-free** structured data: fields / records / tables / entities, each with a confidence score and best-effort citation. | `get_extracted_data` |
| 3 | **Classify** | Zero-shot document-type detection with confidence + rationale. | `classify_document` |
| 4 | **Summarize** | Map-reduce summary (TL;DR + key points) over chunks. | `summarize_document` |
| 5 | **Q&A** | Conversational, **agentic-RAG** answers **grounded with citations** you can click to highlight on the page. | `query_document` |
| 6 | **Flag anomalies** | Missing fields, low-confidence values, generic arithmetic reconciliation, multi-record sanity, LLM consistency review. | `flag_anomalies` |
| 7 | **Export** | Serialize extracted data to **JSON / CSV / Excel** (records become a table). | *(API)* |

### The headline feature: adaptive, domain-agnostic extraction

Extraction does **not** rely on per-domain templates. For *any* document β€” any layout,
any label convention, known type or not β€” it produces a stable, consistently-keyed shape by:

- **Layout detection** β€” the model decides whether the doc is one entity
  (`single_record`, e.g. an invoice), many similar entities (`multi_record`, e.g. an
  employee roster or transaction list), or a header + repeated sub-list
  (`mixed`, e.g. invoice + line items).
- **Schema induction** β€” fields are derived from the actual content (a known type only
  *suggests* fields; the model may add/drop to fit reality).
- **Canonical keys** β€” inconsistent source labels ("Staff Number" / "Personnel ID" / "ID")
  are normalised to one `snake_case` key, identical across every record.
- **Grounding** β€” values come only from the document body (never file metadata); dates and
  amounts are normalised (ISO dates, currency); citations are attached by matching values
  back to provenance-carrying chunks.

---

## 5. How it works β€” architecture

```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  FRONTEND  (Next.js + React + TypeScript + Tailwind)          β”‚
β”‚  40% Chat (streaming agent)  β”‚  60% Workspace                 β”‚
β”‚                              β”‚  Tabs: Viewer Β· Fields Β·       β”‚
β”‚                              β”‚  Classify Β· Summary Β· Export   β”‚
β”‚                              β”‚  Viewer w/ bounding-box cites   β”‚
β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
      β”‚  REST + SSE (token streaming)
β”Œβ”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  BACKEND  (FastAPI, async)                                    β”‚
β”‚  api/        documents Β· chat(SSE) Β· actions Β· meta(health)   β”‚
β”‚  agent/      orchestrator (tool-calling loop) + prompts       β”‚
β”‚  llm/        provider abstraction: Gemini(default) ⇄ OpenAI   β”‚
β”‚  services/   ingestion(Docling) Β· extraction Β· classification β”‚
β”‚              Β· summary Β· qa(RAG) Β· anomaly Β· export Β·          β”‚
β”‚              vectorstore Β· storage Β· pipeline                  β”‚
β”‚  schemas/    Pydantic contracts (documents, optional hints)   β”‚
β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
      β”‚                           β”‚
β”Œβ”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Docling   β”‚            β”‚ Storage (/data)    β”‚
β”‚ (local    β”‚            β”‚ β€’ files on disk    β”‚
β”‚  extract) β”‚            β”‚ β€’ SQLite metadata  β”‚
β”‚ models    β”‚            β”‚ β€’ numpy vectors    β”‚
β”‚ baked in  β”‚            β”‚   (.npz/.json)     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜            β”‚ β€’ rendered pages   β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```

**Discipline (borrowed from Dify):** routes are thin β†’ call services; the agent
orchestrates services as tools; LLM access only through `llm/`; storage only through
`services/storage` + `services/vectorstore`; schemas are shared.

---

## 6. Working β€” the two flows

### Flow A β€” Ingestion pipeline (runs once per upload, in the background)

Implemented in `services/pipeline.py`. State is written back at every milestone so the
UI can poll status and render results progressively.

```
upload β†’ [Stage 1: INGEST]──fail?──► status=failed (hard fail; only this stage can)
            β”‚ Docling parse β†’ Markdown + page images + chunks(page,bbox)
            β–Ό
        [index]      embed chunks β†’ numpy vector store        ┐
        [classify]   LLM zero-shot β†’ type + confidence        β”‚ Stages 2–6 are
        [extract]    adaptive induction β†’ fields/records/...   β”‚ best-effort: a
        [summarize]  map-reduce β†’ TL;DR + key points           β”‚ failure (e.g. no
        [anomalies]  rule + LLM checks (needs extraction)      β”‚ LLM key) degrades
            β–Ό                                                  β”˜ gracefully
        status=ready  (viewer always works, even if AI steps were skipped)
```

Key robustness property: **only ingestion can hard-fail a document.** Everything else is
wrapped so a missing API key or a flaky step still yields a `ready` document with a
working viewer and a note about what was skipped.

### Flow B β€” Chat agent (per user message)

Implemented in `agent/orchestrator.py` β€” a bounded **tool-calling loop**, then a
**streamed** final answer.

```
user message + history + doc context
        β”‚
        β–Ό
  Phase 1: TOOL ROUNDS  (up to MAX_TOOL_ROUNDS = 4, non-streamed reasoning)
     llm.complete_tools(messages, TOOL_SPECS)
        β”‚
        β”œβ”€β”€ model returns tool_calls?  ──yes──► run each tool ─► append results ─► loop
        β”‚        tools: query_document Β· classify_document Β·
        β”‚               get_extracted_data Β· summarize_document Β· flag_anomalies
        β”‚
        └── no tool_calls (ready to answer) ─► exit loop
        β–Ό
  Phase 2: FINAL ANSWER  (streamed token-by-token over SSE, temperature 0.4)
        β”‚
        β–Ό
  client renders answer + citation chips β†’ click a chip β†’ highlight the
  page region (numpy store returned page+bbox with each retrieved chunk)
```

`query_document` runs **agentic RAG**: retrieve the most similar chunks, answer from
them, and return grounding passages whose page+bbox the UI turns into clickable
highlights. The orchestrator collects citations in its `ToolContext` so the API can
emit them alongside the streamed answer.

---

## 7. End-to-end walkthrough (a user's view)

1. **Drag a document** into the left rail (or click *New document*).
2. The pipeline runs: pages render in the viewer first, then **classification, fields,
   summary, and anomaly flags** stream into the workspace tabs as each stage finishes.
3. **Chat** on the left β€” e.g. *"Summarize this"*, *"What's the total?"*, *"Any missing
   fields?"* The agent calls the right tools, then streams a grounded answer. Click any
   **citation chip** to highlight its exact source on the page.
4. **Review & edit** low-confidence fields inline (human-in-the-loop) before exporting.
5. **Export** from the Export tab as **JSON / CSV / Excel**.

---

## 8. Why this design holds up

- **Document depth where it matters** β€” Docling gives best-in-class layout/table/OCR
  locally and for free, the axis on which generic platforms (Dify, n8n) are weakest.
- **Truly document-agnostic** β€” extraction induces its own schema, so a new document
  type needs no code or template changes.
- **Grounded & trustworthy** β€” every chunk carries page + bbox, so answers and fields
  point at their exact source; the single most trust-building UX feature.
- **Cheap and portable** β€” runs on a free Gemini key with no external services; SQLite
  and the numpy store are single-file and swappable for Postgres/pgvector at scale via
  one module each.
- **Graceful degradation** β€” the app stays useful even when AI steps fail.
- **Clean seams for growth** β€” provider abstraction, storage/vectorstore interfaces,
  and a single-orchestrator agent that can grow more tools without a rewrite.