Spaces:
Running
Running
File size: 18,188 Bytes
f87a697 2417033 f87a697 4da7f83 f87a697 16f0c86 f87a697 2b31b30 f87a697 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 | <!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<meta name="description" content="SLM Document Parser: Local CPU-optimized PDF, DOCX, and text layout structure parser matching target schemas offline using Phi-3.5.">
<link rel="canonical" href="https://www.slmagents.ai/document_parser.html">
<title>SLM Document Parser | Documentation</title>
<link rel="stylesheet" href="style.css">
</head>
<body>
<header>
<div class="container nav-container">
<div style="display: flex; align-items: center; gap: 10px;">
<button class="sidebar-toggle" onclick="toggleSidebar()">
<svg width="24" height="24" fill="none" stroke="currentColor" stroke-width="2" viewBox="0 0 24 24">
<path stroke-linecap="round" stroke-linejoin="round" d="M4 6h16M4 12h16M4 18h16"/>
</svg>
</button>
<a href="index.html" class="logo">
<div class="logo-icon">
<svg width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round">
<path d="M12 2L2 7l10 5 10-5-10-5zM2 17l10 5 10-5M2 12l10 5 10-5"/>
</svg>
</div>
<span>SLM Agents</span>
</a>
</div>
<nav>
<ul>
<li><a href="index.html">Home</a></li>
<li class="dropdown">
<a class="dropdown-trigger" style="color:var(--primary)">
Frameworks
<svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="3" stroke-linecap="round" stroke-linejoin="round"><path d="M6 9l6 6 6-6"/></svg>
</a>
<div class="dropdown-content">
<a href="orchestrator.html">Orchestrator Docs</a>
<a href="rag.html">RAG Docs</a>
<a href="summarizer.html">Summarizer Docs</a>
<a href="sql.html">Text-to-SQL Docs</a>
<a href="cli.html">CLI Docs</a>
<a href="code_interpreter.html">Code Interpreter Docs</a>
<a href="git_repo_manager.html">Git Repo Manager Docs</a>
<a href="json_cleaner.html">JSON Cleaner Docs</a>
<a href="document_parser.html">Document Parser Docs</a>
<a href="vision_parser.html">Vision Parser Docs</a>
<a href="web_agent.html">Web Agent Docs</a>
<a href="web_scraper.html">Web Scraper Docs</a>
<a href="search_orchestrator.html">Search Orchestrator Docs</a>
</div>
</li>
<li><a href="playground.html">Playground</a></li>
<li><a href="index.html#upcoming">Roadmap</a></li>
<li><a href="https://huggingface.co/spaces/spcv/slm-agents" class="btn-hf-nav" target="_blank" style="display: flex; align-items: center; gap: 6px; padding: 7px 14px; border-radius: 8px; background: #fef3c7; border: 1px solid #f59e0b; color: #b45309; font-weight: 700; text-decoration: none; font-size: 0.85rem; transition: all 0.2s;">π€ Hugging Face Space</a></li>
</ul>
</nav>
</div>
</header>
<div class="layout-container">
<aside class="sidebar" id="sidebar">
<div class="sidebar-group">
<div class="sidebar-group-title">Active Libraries</div>
<ul class="sidebar-list" id="sidebar-active-list"></ul>
</div>
</aside>
<main class="main-content">
<div class="doc-content-wrapper">
<section class="doc-section" style="margin-top: 0;">
<div class="breadcrumb" style="font-size: 0.85rem; color: #475569; margin-bottom: 1rem; font-weight: 600;"><a href="index.html" style="color: #4f46e5; text-decoration: none;">Home</a> <span style="margin: 0 6px;">βΊ</span> <span style="color: #0f172a;">SLM Document Parser</span></div>
<div class="badge-pill">π Local Document Extraction</div>
<h1 style="font-size: 2.25rem; color: #0f172a; font-weight: 800; margin-bottom: 0.75rem;">SLM Document Parser</h1>
<p style="font-size: 1.05rem; color: #334155; line-height: 1.65; font-weight: 500; margin: 0;">Extract tabular metrics, metadata structures, and key paragraphs from PDFs, Word docs, and plain text files locally under strict MIT-licensing parameters.</p>
<nav class="doc-nav">
<a href="#overview">Overview</a>
<a href="#install">
<a href="#git">Git Checkout</a>Installation</a>
<a href="#config">Configuration API</a>
<a href="#logs">Verified Logs</a>
</nav>
<h2>π Overview & Capabilities</h2>
<p>The SLM Document Parser leverages Microsoft's MIT-licensed <strong>Phi-3.5-mini-instruct</strong> and <strong>Florence-2-large</strong> models optimized via ONNX Runtime GenAI to parse multi-formatted documents offline. It supports high-fidelity layout preservation for scanned PDFs and legacied formats (<code>.pdf</code>, <code>.docx</code>, <code>.doc</code>, <code>.pptx</code>, <code>.ppt</code>, <code>.txt</code>) to output structured text and semantic RAG database chunks.</p>
</section>
<section class="doc-section" id="workflow">
<h2>π€ Truly Agentic Hybrid Visual Pipeline</h2>
<p>The Document Parser handles layout mapping by combining visual coordinate classifiers with standard extraction loops:</p>
<ul>
<li><strong>LibreOffice Conversion Layer:</strong> Converts complex formats (DOC/PPT/PPTX) into intermediate PDFs in the background using headless <code>soffice</code> conversions.</li>
<li><strong>Page Image Rendering:</strong> Renders PDF page indices to images sequentially using <code>pypdfium2</code>.</li>
<li><strong>Florence-2 OCR & OD:</strong> Scans rendered page layouts, running Object Detection (<code><OD></code>) to locate table and figure boxes. Crops table boxes to OCR them individually, and runs detailed caption generation (<code><DETAILED_CAPTION></code>) for diagrams.</li>
<li><strong>Layout Reconstruction:</strong> Passes text, parsed tables, and diagram captions to Phi-3.5 to synthesize a unified, formatted Markdown page representation.</li>
<li><strong>Semantic Chunker & Linker:</strong> Segments markdown into structured paragraph chunks, extracts heading hierarchies, resolves product entities, and links related sections together.</li>
</ul>
</section>
<section class="doc-section" id="visual-demo">
<h2>π Complex Table Visual OCR Extraction</h2>
<p>When the visual parser encounters complex tables inside a document scan, it crops the table coordinates, routes it to the vision parser, and synthesizes a natural language text description. Below is an example of an input expense table and its corresponding parsed description:</p>
<div style="margin: 20px 0; max-width: 600px; border-radius: 12px; overflow: hidden; border: 1px solid rgba(255, 255, 255, 0.05); box-shadow: 0 10px 30px rgba(0,0,0,0.5);">
<img src="https://raw.githubusercontent.com/t00114218-stack/SLMAgents/main/website/complex_table.png" alt="Input Corporate Expense Table" style="width: 100%; height: auto; display: block;">
</div>
<div class="tip-box">
<strong>Extracted Natural Language Table Description Output:</strong><br>
<pre><code>### Q2 FY2024 Company Expense Summary
The table outlines corporate expenditures across multiple categories:
- **Infrastructure**: Spent $1,195,450 against a $1,250,000 budget (a decrease of 4.36%).
- **AI Compute Nodes**: Spent $3,842,910 against a $3,500,000 budget (an increase of 9.80%).
- **Model Conversion**: Spent $412,300 against a $450,000 budget (a decrease of 8.38%).
- **Data Storage**: Spent $845,600 against a $800,000 budget (an increase of 5.70%).
- **GPU Clusters**: Spent $2,215,800 against a $2,100,000 budget (an increase of 5.51%).</code></pre>
</div>
</section>
<section class="doc-section" id="tuning">
<h2>β‘ CPU Performance Tuning Guidelines</h2>
<p>Follow these configuration rules to optimize latency on standard CPUs:</p>
<ul>
<li><strong>Core Thread Cap:</strong> Set <code>n_threads</code> strictly to the number of physical cores to avoid context switching thread collisions.</li>
<li><strong>Sequential Processing:</strong> Process multi-page images sequentially instead of in parallel batches to keep the memory footprint under 2.0 GB.</li>
<li><strong>OpenMP Settings:</strong> Keep thread counts aligned in your shell environment:
<pre><code>export OMP_NUM_THREADS=4
export MKL_NUM_THREADS=4</code></pre>
</li>
</ul>
</section>
<section class="doc-section" id="api">
<h2>API Reference</h2>
<h3>`SLMDocumentParser` Initialization</h3>
<pre><code class="language-python">from slm_document_parser.document_parser import SLMDocumentParser
parser = SLMDocumentParser(n_ctx=4096, n_threads=4)</code></pre>
<table class="param-table">
<thead>
<tr><th>Parameter</th><th>Type / Default</th><th>Description</th></tr>
</thead>
<tbody>
<tr><td>model_path</td><td>str | None</td><td>Local path to Phi-3.5 weight checkpoints. Defaults to "../../models/phi-3.5-mini-instruct-onnx".</td></tr>
<tr><td>n_ctx</td><td>int | 4096</td><td>Inference token window size. Phi-3.5 supports up to 128K context tokens.</td></tr>
<tr><td>n_threads</td><td>int | 4</td><td>Allocated CPU cores for ORT execution.</td></tr>
<tr><td>system_prompt</td><td>str | None</td><td>Optional custom system prompt instructions overriding the default template.</td></tr>
<tr><td>user_input</td><td>str | None</td><td>Optional additional user-supplied target parameters or variables.</td></tr>
</tbody>
</table>
<h3>`extract_text` Method</h3>
<p>Converts document layouts and images into a single reconstructed Markdown text string:</p>
<pre><code class="language-python">markdown_content = parser.extract_text("financial_report.pdf")</code></pre>
<h3>`chunk_document` Method</h3>
<p>Extracts markdown text and slices it into linked semantic chunks with metadata properties:</p>
<pre><code class="language-python">chunks = parser.chunk_document("financial_report.pdf")
print(chunks[0])</code></pre>
<div class="tip-box">
<strong>Extracted Semantic Chunk Output (JSON):</strong><br>
<pre><code class="language-json">{
"text": "SpaceX successfully launched the Falcon 9 rocket from Cape Canaveral Space Force Station, landing the booster return flight for the 15th time. The mission delivered communication payloads into low Earth orbit.",
"metadata": {
"source": "financial_report.pdf",
"heading": "1. Launch Milestones",
"subheading": "Falcon 9 Performance",
"product": "SpaceX",
"key_terms": ["Falcon 9", "Cape Canaveral", "booster"],
"format": "pdf",
"chunk_index": 0,
"related_chunks": [1, 2]
}
}</code></pre>
</div>
<h3>`parse_and_chunk_stream` Method</h3>
<p>Streaming generator yielding semantic chunks page-by-page as they are processed to reduce first-chunk response times (TTFT):</p>
<pre><code class="language-python">for chunk in parser.parse_and_chunk_stream("multi_page_manual.pdf"):
print(f"Processed Chunk: {chunk['text'][:100]}...")</code></pre>
<h3>`export_chunks_to_excel` Method</h3>
<p>Saves the parsed chunks to an Excel spreadsheet:</p>
<pre><code class="language-python"># Exports to workbook. Set append=True to write to an existing sheet
parser.export_chunks_to_excel(chunks, "rag_dataset.xlsx", append=False)</code></pre>
<div class="tip-box">
<strong>Output Excel Table Layout:</strong><br>
<table class="param-table">
<thead>
<tr><th>Chunk Index</th><th>Source File</th><th>Heading</th><th>Subheading</th><th>Product</th><th>Related Chunks</th><th>Text</th></tr>
</thead>
<tbody>
<tr><td>0</td><td>financial_report.pdf</td><td>1. Launch Milestones</td><td>Falcon 9 Performance</td><td>SpaceX</td><td>1,2</td><td>SpaceX successfully launched the Falcon 9...</td></tr>
<tr><td>system_prompt</td><td>str | None</td><td>Optional custom system prompt instructions overriding the default template.</td></tr>
<tr><td>user_input</td><td>str | None</td><td>Optional additional user-supplied target parameters or variables.</td></tr>
</tbody>
</table>
</div>
</section>
<footer style="margin-top: 3rem; text-align: center; border-top: 1px solid #cbd5e1; padding-top: 2rem; color: #475569; font-size: 0.9rem; font-weight: 500;">
<p>Β© 2026 SLM Agents. Built with Apache 2.0 Permissive Open Source License.</p>
</footer>
</div>
<!-- GIT CHECKOUT -->
<section class="doc-section" id="git">
<h2>π Checkout from GitHub</h2>
<p>Clone only this agent's folder from the monorepo using Git sparse-checkout β no need to download the full repository:</p>
<h3 style="font-size: 1.05rem; color: #0f172a; font-weight: 700; margin-top: 1.5rem; margin-bottom: 0.75rem;">Option 1 β Sparse Checkout (Recommended)</h3>
<div class="code-panel" style="max-width:100%; background: #0f172a; border: 1px solid #1e293b; border-radius: 14px; overflow: hidden; margin: 1rem 0; box-shadow: 0 16px 40px rgba(15, 23, 42, 0.12);">
<div class="code-header" style="background: #1e293b; padding: 10px 16px; display: flex; align-items: center; justify-content: space-between; border-bottom: 1px solid #334155;">
<div class="code-dots"><div class="code-dot"></div><div class="code-dot"></div><div class="code-dot"></div></div>
<div class="code-title" style="color: #94a3b8; font-weight: 700; font-size: 0.8rem; font-family: 'JetBrains Mono', monospace;">Terminal β Git Sparse Checkout</div>
</div>
<div class="code-content" style="display:block; padding: 1.25rem 1.5rem; background: #0f172a;">
<pre style="margin:0; background:#0f172a; color:#f8fafc; font-family:'JetBrains Mono',monospace; font-size:0.88rem; border:none; box-shadow:none; padding:0; line-height: 1.75;"><span style="color:#64748b;"># 1. Create and enter a new directory</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">mkdir</span> <span style="color:#38bdf8;">slm_document_parser</span> <span style="color:#94a3b8;">&&</span> <span style="color:#c084fc; font-weight:700;">cd</span> <span style="color:#38bdf8;">slm_document_parser</span>
<span style="color:#64748b;"># 2. Initialise empty git repo and add remote</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git init</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git remote add origin</span> <span style="color:#38bdf8;">https://github.com/t00114218-stack/SLMAgents.git</span>
<span style="color:#64748b;"># 3. Enable sparse-checkout and set target folder</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git sparse-checkout init</span> <span style="color:#94a3b8;">--cone</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git sparse-checkout set</span> <span style="color:#38bdf8;">slm_document_parser</span>
<span style="color:#64748b;"># 4. Pull only that agent's source</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git pull origin</span> <span style="color:#38bdf8;">main</span></pre>
</div>
</div>
<h3 style="font-size: 1.05rem; color: #0f172a; font-weight: 700; margin-top: 2rem; margin-bottom: 0.75rem;">Option 2 β Full Repository Clone</h3>
<div class="code-panel" style="max-width:100%; background: #0f172a; border: 1px solid #1e293b; border-radius: 14px; overflow: hidden; margin: 1rem 0;">
<div class="code-header" style="background: #1e293b; padding: 10px 16px; display: flex; align-items: center; justify-content: space-between; border-bottom: 1px solid #334155;">
<div class="code-dots"><div class="code-dot"></div><div class="code-dot"></div><div class="code-dot"></div></div>
<div class="code-title" style="color: #94a3b8; font-weight: 700; font-size: 0.8rem; font-family: 'JetBrains Mono', monospace;">Terminal β Full Clone</div>
</div>
<div class="code-content" style="display:block; padding: 1.25rem 1.5rem; background: #0f172a;">
<pre style="margin:0; background:#0f172a; color:#f8fafc; font-family:'JetBrains Mono',monospace; font-size:0.88rem; border:none; box-shadow:none; padding:0; line-height: 1.75;"><span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git clone</span> <span style="color:#38bdf8;">https://github.com/t00114218-stack/SLMAgents.git</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">cd</span> <span style="color:#38bdf8;">SLMAgents/slm_document_parser</span></pre>
</div>
</div>
<p style="margin-top: 1.25rem; font-size: 0.9rem; color: #475569; background: #f8fafc; border: 1px solid #cbd5e1; border-radius: 10px; padding: 1rem 1.25rem;">
π‘ <strong>Tip:</strong> After checkout, install the package locally with <code style="background: #eef2ff; color: #4f46e5; border: 1px solid #c7d2fe; padding: 2px 8px; border-radius: 5px; font-weight: 700;">pip install -e ./slm_document_parser</code> to run in editable mode without publishing to PyPI.
</p>
</section>
</main>
</div>
<script src="app.js"></script>
</body>
</html>
|