File size: 18,188 Bytes
f87a697
 
 
 
 
 
2417033
 
f87a697
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4da7f83
f87a697
 
 
 
 
 
 
 
16f0c86
f87a697
2b31b30
f87a697
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
<!DOCTYPE html>
<html lang="en">
<head>
  <meta charset="UTF-8">
  <meta name="viewport" content="width=device-width, initial-scale=1.0">
  <meta name="description" content="SLM Document Parser: Local CPU-optimized PDF, DOCX, and text layout structure parser matching target schemas offline using Phi-3.5.">
  
  <link rel="canonical" href="https://www.slmagents.ai/document_parser.html">
  <title>SLM Document Parser | Documentation</title>
  <link rel="stylesheet" href="style.css">
  
</head>
<body>
  <header>
    <div class="container nav-container">
      <div style="display: flex; align-items: center; gap: 10px;">
        <button class="sidebar-toggle" onclick="toggleSidebar()">
          <svg width="24" height="24" fill="none" stroke="currentColor" stroke-width="2" viewBox="0 0 24 24">
            <path stroke-linecap="round" stroke-linejoin="round" d="M4 6h16M4 12h16M4 18h16"/>
          </svg>
        </button>
        <a href="index.html" class="logo">
          <div class="logo-icon">
            <svg width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round">
              <path d="M12 2L2 7l10 5 10-5-10-5zM2 17l10 5 10-5M2 12l10 5 10-5"/>
            </svg>
          </div>
          <span>SLM Agents</span>
        </a>
      </div>
      <nav>
        <ul>
          <li><a href="index.html">Home</a></li>
          <li class="dropdown">
            <a class="dropdown-trigger" style="color:var(--primary)">
              Frameworks
              <svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="3" stroke-linecap="round" stroke-linejoin="round"><path d="M6 9l6 6 6-6"/></svg>
            </a>
            <div class="dropdown-content">
              <a href="orchestrator.html">Orchestrator Docs</a>
              <a href="rag.html">RAG Docs</a>
              <a href="summarizer.html">Summarizer Docs</a>
              <a href="sql.html">Text-to-SQL Docs</a>
              <a href="cli.html">CLI Docs</a>
              <a href="code_interpreter.html">Code Interpreter Docs</a>
              <a href="git_repo_manager.html">Git Repo Manager Docs</a>
              <a href="json_cleaner.html">JSON Cleaner Docs</a>
              <a href="document_parser.html">Document Parser Docs</a>
              <a href="vision_parser.html">Vision Parser Docs</a>
              <a href="web_agent.html">Web Agent Docs</a>
              <a href="web_scraper.html">Web Scraper Docs</a>
              <a href="search_orchestrator.html">Search Orchestrator Docs</a>
            </div>
          </li>
          <li><a href="playground.html">Playground</a></li>
          <li><a href="index.html#upcoming">Roadmap</a></li>
          <li><a href="https://huggingface.co/spaces/spcv/slm-agents" class="btn-hf-nav" target="_blank" style="display: flex; align-items: center; gap: 6px; padding: 7px 14px; border-radius: 8px; background: #fef3c7; border: 1px solid #f59e0b; color: #b45309; font-weight: 700; text-decoration: none; font-size: 0.85rem; transition: all 0.2s;">πŸ€— Hugging Face Space</a></li>
        </ul>
      </nav>
    </div>
  </header>

  <div class="layout-container">
    <aside class="sidebar" id="sidebar">
      <div class="sidebar-group">
        <div class="sidebar-group-title">Active Libraries</div>
        <ul class="sidebar-list" id="sidebar-active-list"></ul>
      </div>
    </aside>

    <main class="main-content">
      <div class="doc-content-wrapper">
        <section class="doc-section" style="margin-top: 0;">
          <div class="breadcrumb" style="font-size: 0.85rem; color: #475569; margin-bottom: 1rem; font-weight: 600;"><a href="index.html" style="color: #4f46e5; text-decoration: none;">Home</a> <span style="margin: 0 6px;">β€Ί</span> <span style="color: #0f172a;">SLM Document Parser</span></div>
          <div class="badge-pill">πŸ“‚ Local Document Extraction</div>
          <h1 style="font-size: 2.25rem; color: #0f172a; font-weight: 800; margin-bottom: 0.75rem;">SLM Document Parser</h1>
          <p style="font-size: 1.05rem; color: #334155; line-height: 1.65; font-weight: 500; margin: 0;">Extract tabular metrics, metadata structures, and key paragraphs from PDFs, Word docs, and plain text files locally under strict MIT-licensing parameters.</p>
        
        <nav class="doc-nav">
          <a href="#overview">Overview</a>
          <a href="#install">
          <a href="#git">Git Checkout</a>Installation</a>
          <a href="#config">Configuration API</a>
          <a href="#logs">Verified Logs</a>
        </nav>

        
        <h2>πŸš€ Overview &amp; Capabilities</h2>
        <p>The SLM Document Parser leverages Microsoft's MIT-licensed <strong>Phi-3.5-mini-instruct</strong> and <strong>Florence-2-large</strong> models optimized via ONNX Runtime GenAI to parse multi-formatted documents offline. It supports high-fidelity layout preservation for scanned PDFs and legacied formats (<code>.pdf</code>, <code>.docx</code>, <code>.doc</code>, <code>.pptx</code>, <code>.ppt</code>, <code>.txt</code>) to output structured text and semantic RAG database chunks.</p>
      </section>

        <section class="doc-section" id="workflow">
        <h2>πŸ€– Truly Agentic Hybrid Visual Pipeline</h2>
        <p>The Document Parser handles layout mapping by combining visual coordinate classifiers with standard extraction loops:</p>
        <ul>
          <li><strong>LibreOffice Conversion Layer:</strong> Converts complex formats (DOC/PPT/PPTX) into intermediate PDFs in the background using headless <code>soffice</code> conversions.</li>
          <li><strong>Page Image Rendering:</strong> Renders PDF page indices to images sequentially using <code>pypdfium2</code>.</li>
          <li><strong>Florence-2 OCR &amp; OD:</strong> Scans rendered page layouts, running Object Detection (<code>&lt;OD&gt;</code>) to locate table and figure boxes. Crops table boxes to OCR them individually, and runs detailed caption generation (<code>&lt;DETAILED_CAPTION&gt;</code>) for diagrams.</li>
          <li><strong>Layout Reconstruction:</strong> Passes text, parsed tables, and diagram captions to Phi-3.5 to synthesize a unified, formatted Markdown page representation.</li>
          <li><strong>Semantic Chunker &amp; Linker:</strong> Segments markdown into structured paragraph chunks, extracts heading hierarchies, resolves product entities, and links related sections together.</li>
        </ul>
      </section>

        <section class="doc-section" id="visual-demo">
        <h2>πŸ“Š Complex Table Visual OCR Extraction</h2>
        <p>When the visual parser encounters complex tables inside a document scan, it crops the table coordinates, routes it to the vision parser, and synthesizes a natural language text description. Below is an example of an input expense table and its corresponding parsed description:</p>
        
        <div style="margin: 20px 0; max-width: 600px; border-radius: 12px; overflow: hidden; border: 1px solid rgba(255, 255, 255, 0.05); box-shadow: 0 10px 30px rgba(0,0,0,0.5);">
          <img src="https://raw.githubusercontent.com/t00114218-stack/SLMAgents/main/website/complex_table.png" alt="Input Corporate Expense Table" style="width: 100%; height: auto; display: block;">
        </div>

        <div class="tip-box">
          <strong>Extracted Natural Language Table Description Output:</strong><br>
          <pre><code>### Q2 FY2024 Company Expense Summary
The table outlines corporate expenditures across multiple categories:
- **Infrastructure**: Spent $1,195,450 against a $1,250,000 budget (a decrease of 4.36%).
- **AI Compute Nodes**: Spent $3,842,910 against a $3,500,000 budget (an increase of 9.80%).
- **Model Conversion**: Spent $412,300 against a $450,000 budget (a decrease of 8.38%).
- **Data Storage**: Spent $845,600 against a $800,000 budget (an increase of 5.70%).
- **GPU Clusters**: Spent $2,215,800 against a $2,100,000 budget (an increase of 5.51%).</code></pre>
        </div>
      </section>

        <section class="doc-section" id="tuning">
        <h2>⚑ CPU Performance Tuning Guidelines</h2>
        <p>Follow these configuration rules to optimize latency on standard CPUs:</p>
        <ul>
          <li><strong>Core Thread Cap:</strong> Set <code>n_threads</code> strictly to the number of physical cores to avoid context switching thread collisions.</li>
          <li><strong>Sequential Processing:</strong> Process multi-page images sequentially instead of in parallel batches to keep the memory footprint under 2.0 GB.</li>
          <li><strong>OpenMP Settings:</strong> Keep thread counts aligned in your shell environment:
            <pre><code>export OMP_NUM_THREADS=4
export MKL_NUM_THREADS=4</code></pre>
          </li>
        </ul>
      </section>

        <section class="doc-section" id="api">
        <h2>API Reference</h2>
        
        <h3>`SLMDocumentParser` Initialization</h3>
        <pre><code class="language-python">from slm_document_parser.document_parser import SLMDocumentParser

parser = SLMDocumentParser(n_ctx=4096, n_threads=4)</code></pre>
        
        <table class="param-table">
          <thead>
            <tr><th>Parameter</th><th>Type / Default</th><th>Description</th></tr>
          </thead>
          <tbody>
            <tr><td>model_path</td><td>str | None</td><td>Local path to Phi-3.5 weight checkpoints. Defaults to "../../models/phi-3.5-mini-instruct-onnx".</td></tr>
            <tr><td>n_ctx</td><td>int | 4096</td><td>Inference token window size. Phi-3.5 supports up to 128K context tokens.</td></tr>
            <tr><td>n_threads</td><td>int | 4</td><td>Allocated CPU cores for ORT execution.</td></tr>
                        <tr><td>system_prompt</td><td>str | None</td><td>Optional custom system prompt instructions overriding the default template.</td></tr>
              <tr><td>user_input</td><td>str | None</td><td>Optional additional user-supplied target parameters or variables.</td></tr>
</tbody>
        </table>

        <h3>`extract_text` Method</h3>
        <p>Converts document layouts and images into a single reconstructed Markdown text string:</p>
        <pre><code class="language-python">markdown_content = parser.extract_text("financial_report.pdf")</code></pre>

        <h3>`chunk_document` Method</h3>
        <p>Extracts markdown text and slices it into linked semantic chunks with metadata properties:</p>
        <pre><code class="language-python">chunks = parser.chunk_document("financial_report.pdf")
print(chunks[0])</code></pre>

        <div class="tip-box">
          <strong>Extracted Semantic Chunk Output (JSON):</strong><br>
          <pre><code class="language-json">{
  "text": "SpaceX successfully launched the Falcon 9 rocket from Cape Canaveral Space Force Station, landing the booster return flight for the 15th time. The mission delivered communication payloads into low Earth orbit.",
  "metadata": {
    "source": "financial_report.pdf",
    "heading": "1. Launch Milestones",
    "subheading": "Falcon 9 Performance",
    "product": "SpaceX",
    "key_terms": ["Falcon 9", "Cape Canaveral", "booster"],
    "format": "pdf",
    "chunk_index": 0,
    "related_chunks": [1, 2]
  }
}</code></pre>
        </div>

        <h3>`parse_and_chunk_stream` Method</h3>
        <p>Streaming generator yielding semantic chunks page-by-page as they are processed to reduce first-chunk response times (TTFT):</p>
        <pre><code class="language-python">for chunk in parser.parse_and_chunk_stream("multi_page_manual.pdf"):
    print(f"Processed Chunk: {chunk['text'][:100]}...")</code></pre>

        <h3>`export_chunks_to_excel` Method</h3>
        <p>Saves the parsed chunks to an Excel spreadsheet:</p>
        <pre><code class="language-python"># Exports to workbook. Set append=True to write to an existing sheet
parser.export_chunks_to_excel(chunks, "rag_dataset.xlsx", append=False)</code></pre>

        <div class="tip-box">
          <strong>Output Excel Table Layout:</strong><br>
          <table class="param-table">
            <thead>
              <tr><th>Chunk Index</th><th>Source File</th><th>Heading</th><th>Subheading</th><th>Product</th><th>Related Chunks</th><th>Text</th></tr>
            </thead>
            <tbody>
              <tr><td>0</td><td>financial_report.pdf</td><td>1. Launch Milestones</td><td>Falcon 9 Performance</td><td>SpaceX</td><td>1,2</td><td>SpaceX successfully launched the Falcon 9...</td></tr>
                          <tr><td>system_prompt</td><td>str | None</td><td>Optional custom system prompt instructions overriding the default template.</td></tr>
              <tr><td>user_input</td><td>str | None</td><td>Optional additional user-supplied target parameters or variables.</td></tr>
</tbody>
          </table>
        </div>
      </section>

        <footer style="margin-top: 3rem; text-align: center; border-top: 1px solid #cbd5e1; padding-top: 2rem; color: #475569; font-size: 0.9rem; font-weight: 500;">
          <p>Β© 2026 SLM Agents. Built with Apache 2.0 Permissive Open Source License.</p>
        </footer>
      </div>
    
        <!-- GIT CHECKOUT -->
        <section class="doc-section" id="git">
          <h2>πŸ™ Checkout from GitHub</h2>
          <p>Clone only this agent's folder from the monorepo using Git sparse-checkout β€” no need to download the full repository:</p>

          <h3 style="font-size: 1.05rem; color: #0f172a; font-weight: 700; margin-top: 1.5rem; margin-bottom: 0.75rem;">Option 1 β€” Sparse Checkout (Recommended)</h3>
          <div class="code-panel" style="max-width:100%; background: #0f172a; border: 1px solid #1e293b; border-radius: 14px; overflow: hidden; margin: 1rem 0; box-shadow: 0 16px 40px rgba(15, 23, 42, 0.12);">
            <div class="code-header" style="background: #1e293b; padding: 10px 16px; display: flex; align-items: center; justify-content: space-between; border-bottom: 1px solid #334155;">
              <div class="code-dots"><div class="code-dot"></div><div class="code-dot"></div><div class="code-dot"></div></div>
              <div class="code-title" style="color: #94a3b8; font-weight: 700; font-size: 0.8rem; font-family: 'JetBrains Mono', monospace;">Terminal β€” Git Sparse Checkout</div>
            </div>
            <div class="code-content" style="display:block; padding: 1.25rem 1.5rem; background: #0f172a;">
              <pre style="margin:0; background:#0f172a; color:#f8fafc; font-family:'JetBrains Mono',monospace; font-size:0.88rem; border:none; box-shadow:none; padding:0; line-height: 1.75;"><span style="color:#64748b;"># 1. Create and enter a new directory</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">mkdir</span> <span style="color:#38bdf8;">slm_document_parser</span> <span style="color:#94a3b8;">&amp;&amp;</span> <span style="color:#c084fc; font-weight:700;">cd</span> <span style="color:#38bdf8;">slm_document_parser</span>

<span style="color:#64748b;"># 2. Initialise empty git repo and add remote</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git init</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git remote add origin</span> <span style="color:#38bdf8;">https://github.com/t00114218-stack/SLMAgents.git</span>

<span style="color:#64748b;"># 3. Enable sparse-checkout and set target folder</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git sparse-checkout init</span> <span style="color:#94a3b8;">--cone</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git sparse-checkout set</span> <span style="color:#38bdf8;">slm_document_parser</span>

<span style="color:#64748b;"># 4. Pull only that agent's source</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git pull origin</span> <span style="color:#38bdf8;">main</span></pre>
            </div>
          </div>

          <h3 style="font-size: 1.05rem; color: #0f172a; font-weight: 700; margin-top: 2rem; margin-bottom: 0.75rem;">Option 2 β€” Full Repository Clone</h3>
          <div class="code-panel" style="max-width:100%; background: #0f172a; border: 1px solid #1e293b; border-radius: 14px; overflow: hidden; margin: 1rem 0;">
            <div class="code-header" style="background: #1e293b; padding: 10px 16px; display: flex; align-items: center; justify-content: space-between; border-bottom: 1px solid #334155;">
              <div class="code-dots"><div class="code-dot"></div><div class="code-dot"></div><div class="code-dot"></div></div>
              <div class="code-title" style="color: #94a3b8; font-weight: 700; font-size: 0.8rem; font-family: 'JetBrains Mono', monospace;">Terminal β€” Full Clone</div>
            </div>
            <div class="code-content" style="display:block; padding: 1.25rem 1.5rem; background: #0f172a;">
              <pre style="margin:0; background:#0f172a; color:#f8fafc; font-family:'JetBrains Mono',monospace; font-size:0.88rem; border:none; box-shadow:none; padding:0; line-height: 1.75;"><span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git clone</span> <span style="color:#38bdf8;">https://github.com/t00114218-stack/SLMAgents.git</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">cd</span> <span style="color:#38bdf8;">SLMAgents/slm_document_parser</span></pre>
            </div>
          </div>

          <p style="margin-top: 1.25rem; font-size: 0.9rem; color: #475569; background: #f8fafc; border: 1px solid #cbd5e1; border-radius: 10px; padding: 1rem 1.25rem;">
            πŸ’‘ <strong>Tip:</strong> After checkout, install the package locally with <code style="background: #eef2ff; color: #4f46e5; border: 1px solid #c7d2fe; padding: 2px 8px; border-radius: 5px; font-weight: 700;">pip install -e ./slm_document_parser</code> to run in editable mode without publishing to PyPI.
          </p>
        </section>

</main>
  </div>

  <script src="app.js"></script>
</body>
</html>