File size: 19,827 Bytes
f65e025
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
# Document Processing AI Agent β€” Deep Research Report

**Project:** Production-ready, chat-based Document Processing AI Agent
**Org:** Monkhub Innovations
**Date:** 2026-06-24
**Method:** Sequential, single-researcher web research (no sub-agents). Each system and source read in full, one at a time.

---

## 0. Executive Summary

We evaluated five leading systems for document-processing automation, spanning the four relevant archetypes:

| Archetype | System | Why it represents the category |
|---|---|---|
| Full open-source app | **RAGFlow** (infiniflow) | A complete chat-over-documents product with deep document understanding, citations, and an agent layer β€” the closest existing thing to what we are building. |
| LLM app platform | **Dify** (langgenius) | The reference for clean platform architecture, visual orchestration, observability, and polished UI/UX. |
| Extraction toolkit | **Docling** (IBM / LF AI & Data) | Best-in-class open document parsing (layout, tables, OCR, VLM) β€” the extraction *backbone*. |
| Workflow automation | **n8n** document workflows | The "glue" archetype: trigger β†’ OCR β†’ LLM extract β†’ validate β†’ route, assembled visually. |
| Agentic framework | **LlamaIndex** (Agentic Document Workflows + LlamaParse/Extract) | The reference for stateful, multi-step *document agents* and schema-based structured extraction. |

**Winner / primary reference foundation: RAGFlow**, with the explicit recommendation that we **do not fork it wholesale**. For our build we adopt **RAGFlow's pipeline design as the reference architecture**, use **Docling as the actual extraction engine** (MIT, local, zero per-page cost), borrow **Dify's clean separation-of-concerns and observability**, and structure the agent loop using **LlamaIndex-style agentic-workflow patterns**. Details in Β§3–§6.

---

## 1. The Five Systems β€” Deep Analysis

### 1.1 RAGFlow (infiniflow) β€” *Apache 2.0, ~83.5kβ˜…*

**What it does / core purpose.** Open-source RAG engine fused with agent capabilities that turns unstructured documents into a production "context layer" for LLMs. Philosophy: *"quality in, quality out"* β€” invest heavily in deep document understanding at ingestion so retrieval and answers are grounded and citable.

**End-to-end pipeline.**
1. **Ingest** β€” Word, PPT, Excel, PDF, images, scans, structured data, web pages.
2. **Deep document understanding** β€” the **DeepDoc** module performs layout analysis and structure extraction; can also delegate to MinerU or Docling parsers.
3. **Template-based chunking** β€” intelligent, configurable, *visualized* chunking with human review/intervention.
4. **Embed + index** β€” into Elasticsearch (default) or Infinity (their own engine).
5. **Multi-recall retrieval + fused re-ranking.**
6. **Grounded answer with citations** β€” traceable references to reduce hallucination.

**Architecture.** Frontend (web) β†’ REST API β†’ backend services β†’ DeepDoc / embedding / LLM interface, over a storage tier of **MinIO** (objects), **Elasticsearch or Infinity** (vector + full-text), **MySQL** (metadata), **Redis** (cache). Agent framework adds workflow automation, MCP support, code executor (Python/JS, gVisor sandbox), and agent memory.

**Tech stack / models.** Python + Go backend, TypeScript frontend. Pluggable LLM and embedding providers (OpenAI, DeepSeek, Gemini, local). Docker / Docker Compose / Kubernetes (Helm).

**Key features.** Explainable template chunking; grounded citations with visualization; broad format support; agent templates + MCP; memory; many chat-channel and data-source connectors (Confluence, S3, Notion, Google Drive, Discord, etc.); multi-language UI.

**Strengths.** Deepest end-to-end document pipeline of the five; citation grounding; fully permissive license; mature, very active, well-documented; the most *directly comparable* product to our target.

**Weaknesses.** Heavy footprint (β‰₯4 cores, β‰₯16 GB RAM, β‰₯50 GB disk; x86-first); operational learning curve; opinionated stack that is hard to bend into a *custom* 40/60 UI without significant surgery.

**Relevance.** β˜…β˜…β˜…β˜…β˜… as a **reference** for the pipeline, citation UX, and chunking visualization. Lower as a thing-to-fork because its UI and infra are opinionated.

---

### 1.2 Dify (langgenius) β€” *Apache-2.0-based (with extra conditions), ~146kβ˜…*

**What it does / core purpose.** Open-source **LLM app development platform** β€” bridges prototype to production with a low-code visual canvas plus developer APIs (Backend-as-a-Service).

**End-to-end flow.** Design on a visual canvas β†’ pick model (50+ providers) β†’ ingest documents into a managed RAG pipeline β†’ define agents (function-calling or ReAct) with 50+ built-in tools β†’ test in a prompt IDE β†’ monitor (logs/metrics) β†’ deploy via API or hosted UI.

**Architecture.** Python backend (BaaS) + Next.js/React/TypeScript frontend; **PostgreSQL**; a canvas **workflow engine**; a unified **model-abstraction layer**; out-of-the-box **RAG pipeline**; first-class **observability** (Langfuse, Opik, Arize Phoenix).

**Tech stack / models.** TS (β‰ˆ53%) + Python (β‰ˆ43%). 50+ LLM providers incl. OpenAI, Anthropic, Google, Mistral, local. Docker Compose / K8s / Terraform / CDK. Min 2 cores / 4 GB.

**Key features.** Visual workflow builder; agent framework with tools; managed RAG; prompt IDE with side-by-side model comparison; production observability; complete API surface.

**Strengths.** Cleanest, most extensible *platform* architecture; very polished UI/UX; superb model-provider abstraction (directly relevant to our "free Gemini + GPT-4o-mini fallback" requirement); strong production/observability story; huge community.

**Weaknesses.** Document processing is **generic** β€” basic ingest/chunk, *no* specialized OCR / table-structure / layout intelligence. License has extra conditions beyond pure Apache 2.0. As a platform it can feel heavy if you only need an app.

**Relevance.** β˜…β˜…β˜…β˜…β˜† β€” the best **architectural and UI/UX reference**, and a candidate **model-routing layer**, but its document intelligence is too shallow to be our extraction engine.

---

### 1.3 Docling (IBM Research β†’ LF AI & Data) β€” *MIT, ~62kβ˜…*

**What it does / core purpose.** Open document-processing library that parses diverse formats (with advanced PDF understanding) into a unified, AI-ready representation. It is the **extraction backbone**, not an agent.

**Pipeline.** Ingest (path/URL) β†’ format detect β†’ **layout analysis** (structure + reading order) β†’ content extraction (text/tables/images) β†’ **model inference** for hard elements β†’ unified **DoclingDocument** β†’ export (Markdown / JSON / HTML / DocTags).

**Architecture / models.** Layout analysis; **TableFormer** for table structure; OCR engines for scans; pluggable **VLMs** (built-in **GraniteDocling 258M**); ASR for audio; chart-understanding module. Python, Pydantic v2.

**Formats.** **In:** PDF, DOCX, PPTX, XLSX, HTML, EPUB, images (PNG/TIFF/JPEG), audio, email, LaTeX, Markdown, ODF, XBRL, more. **Out:** Markdown, HTML, lossless JSON, DocTags, domain XML.

**Key features.** Best-in-class layout + table extraction; semantic **chunking**; native adapters for **LangChain, LlamaIndex, CrewAI, Haystack, MCP**; **local / air-gapped** execution (no per-page API cost); CLI + Python.

**Strengths.** Best open extraction accuracy; widest format coverage; permissive MIT; local-first (privacy + zero marginal cost); OpenSSF badge; peer-reviewed; trivially embeddable into any backend.

**Weaknesses.** Not an application β€” no UI, chat, classification, or Q&A on its own (by design). VLM inference adds latency/compute; scaling docs sparse; some domains need fine-tuning.

**Relevance.** β˜…β˜…β˜…β˜…β˜… as the **component we should actually use** for ingest/extraction. Not a competitor to the whole app β€” a building block.

---

### 1.4 n8n Document Workflows β€” *Sustainable Use License (fair-code), source-available*

**What it does / core purpose.** Visual workflow-automation platform; "document processing" emerges from chaining nodes. Hundreds of templates exist (e.g. *Mistral OCR + GPT-4o-mini invoice processing*, *Gemini OCR + Google Sheets*).

**Typical pipeline (3 layers).** **Ingestion** (webhook / email / Drive / upload trigger β†’ download file) β†’ **Intelligence** (OCR node β€” Mistral / Gemini / Google Vision β€” then an **AI Agent / LLM node** with a **structured-output parser** to emit JSON; classification + field extraction) β†’ **Orchestration** (validation against business rules / POs, enrichment, routing to Sheets / SQL / CRM / ERP / Telegram / Gmail). Core design principle: **decouple extraction from delivery** via modular, reusable sub-workflows.

**Strengths.** Fastest way to prototype an extract→validate→route pipeline; 400+ integrations; visual + maintainable; built-in retries/logging; self-hostable.

**Weaknesses.** Not an end-user product β€” **no document-viewer or chat UI** for our use case (you'd build that separately). Logic-in-canvas sprawls at scale; n8n can become an orchestration bottleneck; **fair-code** license (not OSI-open) carries commercial-use restrictions.

**Relevance.** β˜…β˜…β˜…β˜†β˜† β€” excellent **reference for the orchestration pattern** (and a possible internal automation tool), but not the foundation for a polished chat app.

---

### 1.5 LlamaIndex β€” Agentic Document Workflows (ADW) + LlamaParse/Extract β€” *Core MIT; Parse/Extract/Cloud are SaaS with free tier*

**What it does / core purpose.** A leading framework for **document agents**. ADW goes a step beyond both IDP and RAG: a **document agent** that parses, **maintains state across multi-step processes**, retrieves reference material, applies business logic, and produces recommendations (human-in-the-loop by design).

**Architecture / pattern.** Document agent orchestrates: **LlamaParse** (VLM-powered parsing) β†’ **LlamaExtract** (Pydantic-schema structured extraction *with confidence scores*) β†’ indexes/RAG for reference retrieval β†’ **Workflow engine** (event-driven; fan-out/fan-in, self-reflection, human-in-the-loop). Inverts classic RAG: *parse β†’ maintain state β†’ retrieve contextually β†’ reason β†’ surface for validation*.

**Strengths.** Best **agent-orchestration** abstractions; schema-first extraction with confidence; excellent for conversational, stateful Q&A over documents; huge community and docs; MIT core.

**Weaknesses.** The strongest parsing/extraction (LlamaParse/Extract/Cloud) is **paid SaaS** (free tier limited); framework is code-first (no end-user UI provided); you assemble the product yourself.

**Relevance.** β˜…β˜…β˜…β˜…β˜† β€” the best **mental model and orchestration reference** for our agent loop; we can mirror ADW patterns using our own free-LLM + Docling stack instead of the paid services.

---

## 2. Comparison Table & Scoring

**Rubric (1–5; 5 = best).** Scored for the specific goal: *foundation/reference for a custom, chat-based, production document-processing agent with a 40/60 UI and free-LLM-first model strategy.*

| # | Parameter | What a 5 looks like |
|---|---|---|
| A | **Feature completeness (doc processing)** | Native OCR, layout, tables, KV, entities, classification, summarization, Q&A, export |
| B | **Architecture quality & extensibility** | Clean layers, pluggable models/stores, easy to add tools/formats |
| C | **Ease of customisation** | Fast to bend to our UI/flows without fighting the framework |
| D | **Production-readiness** | Auth, scaling, observability, error handling, deploy story |
| E | **UI/UX quality** | Polished, relevant end-user experience for document work |
| F | **Cost / open-source friendliness** | Permissive license, no forced per-page/API cost, self-hostable |
| G | **Community & documentation** | Stars, activity, docs depth |

| System | A | B | C | D | E | F | G | **Total /35** |
|---|---|---|---|---|---|---|---|---|
| **RAGFlow** | 5 | 4 | 3 | 4 | 4 | 5 | 5 | **30** |
| **Dify** | 3 | 5 | 5 | 5 | 5 | 4 | 5 | **32** |
| **Docling** | 5 | 5 | 4 | 4 | 1 | 5 | 5 | **29** |
| **n8n workflows** | 3 | 3 | 5 | 4 | 2 | 3 | 5 | **25** |
| **LlamaIndex ADW** | 4 | 5 | 4 | 4 | 2 | 3 | 5 | **27** |

**Reading the table.** By raw total, **Dify (32)** edges **RAGFlow (30)** β€” but Dify's lead comes from platform polish, *not* document intelligence (its weakest axis, A=3). When we weight by what this project actually needs β€” **document-processing depth (A), architecture (B), and customisation (C)** β€” the picture is:

- **Document depth (A):** RAGFlow & Docling lead (5).
- **Architecture (B):** Dify, Docling, LlamaIndex lead (5).
- **The honest conclusion is a *combination*, not a single fork.** No one system is simultaneously the deepest extractor *and* the cleanest platform *and* a ready-made 40/60 chat UI.

---

## 3. Winner & Why

**Primary reference foundation: RAGFlow.** It is the only system that already *is* the product we're describing β€” chat over documents, deep understanding, grounded citations, agentic layer β€” and it is fully Apache-2.0. It proves the end-to-end pipeline and gives us a battle-tested blueprint for chunking visualization and citation UX.

**But the senior-level recommendation is a composed stack, not a fork**, because RAGFlow's infra and UI are too opinionated to host our bespoke 40/60 experience, and its heaviness is unjustified at our stage:

> **Recommended build = RAGFlow's pipeline blueprint  +  Docling as the extraction engine  +  LlamaIndex-style agentic-workflow orchestration  +  Dify-grade separation-of-concerns & observability  +  our own FastAPI backend and React 40/60 UI.**

This gives us RAGFlow's proven flow, Docling's best-in-class *and free/local* extraction, LlamaIndex's stateful agent patterns, Dify's clean architecture discipline β€” with none of the licensing, cost, or lock-in downsides.

---

## 4. State-of-the-Art Capabilities (what our agent should do)

Modern IDP = **OCR + ICR + NLP + VLMs + agentic AI**. Expected capabilities:

- **Ingest:** PDF (digital + scanned), DOCX, PPTX, XLSX, images (PNG/JPG/TIFF), email, HTML.
- **Parse/extract:** layout & reading order, **tables** (structure-aware), **key–value pairs**, **named entities**, line items, multi-page spanning content.
- **Classify:** document type/category (invoice, contract, form, receipt, report…), zero-shot via LLM.
- **Summarize:** whole-doc and section-level.
- **Conversational Q&A:** grounded, **with citations** back to the source span.
- **Validate / flag:** missing required fields, inconsistencies, anomalies, confidence scoring, **human-in-the-loop** review.
- **Export:** **JSON** (schema-conformant), CSV/Excel (tables), and structured records to downstream systems.

**Field direction:** extraction is shifting from a *spatial* problem to a *semantic reasoning* problem β€” VLMs/LLMs read layout + meaning together and emit JSON directly, which is more robust to layout drift than legacy OCR. Best practice = **hybrid**: deterministic parser (Docling) for structure + LLM/VLM for semantic extraction, classification, and Q&A.

---

## 5. Ideal UI/UX (informs the 40/60 build)

Validated patterns from current document-AI and AI-chat UX research:

- **Split-screen** (chat left / work right) is the established pattern when the AI produces artifacts β€” see Claude Artifacts, Perplexity. Our **40% chat / 60% workspace** split is exactly this, correctly proportioned.
- **Document viewer with bounding-box highlights:** extracted fields and citations should **highlight the exact source region** in the rendered document (each cell/field carries a bounding polygon + confidence). This is the single most trust-building feature.
- **Citations:** numbered inline references linking to expandable source cards / highlighted spans; visually distinguish quoted evidence from synthesis.
- **Workspace panels (the 60%)** β€” multiple pages/tabs encouraged: **Document Viewer**, **Extraction/Fields** (editable table with confidence + flags), **Classification** panel, **Summary**, **Export** controls.
- **Side-panel, no context-switch:** keep the document in view while the AI works.
- **Streaming + accessibility:** stream responses; `aria-live="polite"` on response containers.
- **Human-in-the-loop:** low-confidence fields visibly flagged and inline-editable before export.

---

## 6. Recommended Architecture, LLM & Toolchain Strategy

**Agent design patterns (2025–26).** The field has moved past static RAG toward **reasoning/agentic RAG** and multi-agent patterns. For us: an **orchestrator agent** + **tool-calling** + **agentic RAG** (the agent decides *when/what/how* to retrieve), with structured-output tools and human-in-the-loop. Keep it a single well-orchestrated agent with tools (upload, parse, classify, extract, summarize, query, validate, export) before reaching for multi-agent complexity.

**LLM strategy (matches the brief).**
- **Default (free):** **Google Gemini Flash** family β€” generous free tier, ~1M-token context (great for long documents), native multimodal/vision (can read document images directly).
- **Fallback (switchable):** **GPT-4o-mini** β€” cheap, 128k context, reliable tool-calling.
- **Provider-abstraction layer** so models are swappable per-task (cheap model for classification, stronger for reasoning) β€” mirroring Dify's model layer.

**Toolchain / OCR / orchestration.**
- **Extraction:** **Docling** (layout, TableFormer tables, OCR for scans, DoclingDocument β†’ JSON/Markdown) β€” local, MIT, zero per-page cost. Optional VLM path (Gemini vision) for hard scans.
- **Structured extraction:** **Pydantic schemas + confidence scores** (LlamaExtract pattern), enforced via LLM structured output / tool-calling.
- **Orchestration:** lightweight agent loop in our own backend (LlamaIndex-style workflow patterns; optionally the LlamaIndex Workflows library) β€” avoids the heavy multi-service footprint of RAGFlow.
- **RAG:** chunk (Docling) β†’ embed β†’ vector store (pgvector / Qdrant) β†’ agentic retrieval with citations.

**Proposed stack for the build.**
- **Backend:** **FastAPI** (Python) β€” async, clean separation (routes / services / agent / tools / schemas), production API.
- **Frontend:** **React + TypeScript** (Next.js or Vite) β€” 40/60 layout, document viewer with bounding-box overlays, streaming chat.
- **Storage:** Postgres (+ pgvector) for metadata + vectors; object storage (local/MinIO/S3) for files.
- **Models:** Gemini Flash (default) ⇄ GPT-4o-mini (fallback) behind a provider interface.
- **Extraction:** Docling.

---

## 7. Sources

- https://www.firecrawl.dev/blog/best-open-source-agent-frameworks
- https://github.com/infiniflow/ragflow
- https://github.com/langgenius/dify
- https://github.com/docling-project/docling
- https://github.com/opendatalab/mineru
- https://procycons.com/en/blogs/pdf-data-extraction-benchmark/
- https://n8n.io/workflows/4933-document-parsing-and-data-extraction-with-mistral-ocr-ai-processing-and-multi-channel-delivery/
- https://data-sleek.com/blog/document-workflow-automation-seamless-data-stack-n8n-ai/
- https://www.llamaindex.ai/blog/introducing-agentic-document-workflows
- https://www.llamaindex.ai/workflows
- https://www.everestgrp.com/report/egr-2025-38-r-7283/
- https://scalehub.com/2025-idp-guide/
- https://uxdesign.cc/where-should-ai-sit-in-your-ui-1710a258390e
- https://thefrontkit.com/blogs/ai-chat-ui-best-practices
- https://learn.microsoft.com/en-us/azure/ai-services/document-intelligence/prebuilt/layout
- https://www.azilen.com/blog/agentic-ai-design-patterns/
- https://learn.microsoft.com/en-us/azure/architecture/ai-ml/guide/ai-agent-design-patterns
- https://modal.com/blog/8-top-open-source-ocr-models-compared
- https://www.llamaindex.ai/insights/best-vision-language-models