File size: 15,042 Bytes
f87a697
 
 
 
 
 
2417033
 
f87a697
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4da7f83
f87a697
 
 
 
 
 
 
 
16f0c86
f87a697
2b31b30
f87a697
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
<!DOCTYPE html>
<html lang="en">
<head>
  <meta charset="UTF-8">
  <meta name="viewport" content="width=device-width, initial-scale=1.0">
  <meta name="description" content="SLM Vision Parser: High-speed local CPU-based diagram, flowchart, and visual table OCR parser using Florence-2.">
  
  <link rel="canonical" href="https://www.slmagents.ai/vision_parser.html">
  <title>SLM Vision Parser | Documentation</title>
  <link rel="stylesheet" href="style.css">
  
</head>
<body>
  <header>
    <div class="container nav-container">
      <div style="display: flex; align-items: center; gap: 10px;">
        <button class="sidebar-toggle" onclick="toggleSidebar()">
          <svg width="24" height="24" fill="none" stroke="currentColor" stroke-width="2" viewBox="0 0 24 24">
            <path stroke-linecap="round" stroke-linejoin="round" d="M4 6h16M4 12h16M4 18h16"/>
          </svg>
        </button>
        <a href="index.html" class="logo">
          <div class="logo-icon">
            <svg width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round">
              <path d="M12 2L2 7l10 5 10-5-10-5zM2 17l10 5 10-5M2 12l10 5 10-5"/>
            </svg>
          </div>
          <span>SLM Agents</span>
        </a>
      </div>
      <nav>
        <ul>
          <li><a href="index.html">Home</a></li>
          <li class="dropdown">
            <a class="dropdown-trigger" style="color:var(--primary)">
              Frameworks
              <svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="3" stroke-linecap="round" stroke-linejoin="round"><path d="M6 9l6 6 6-6"/></svg>
            </a>
            <div class="dropdown-content">
              <a href="orchestrator.html">Orchestrator Docs</a>
              <a href="rag.html">RAG Docs</a>
              <a href="summarizer.html">Summarizer Docs</a>
              <a href="sql.html">Text-to-SQL Docs</a>
              <a href="cli.html">CLI Docs</a>
              <a href="code_interpreter.html">Code Interpreter Docs</a>
              <a href="git_repo_manager.html">Git Repo Manager Docs</a>
              <a href="json_cleaner.html">JSON Cleaner Docs</a>
              <a href="document_parser.html">Document Parser Docs</a>
              <a href="vision_parser.html">Vision Parser Docs</a>
              <a href="web_agent.html">Web Agent Docs</a>
              <a href="web_scraper.html">Web Scraper Docs</a>
              <a href="search_orchestrator.html">Search Orchestrator Docs</a>
            </div>
          </li>
          <li><a href="playground.html">Playground</a></li>
          <li><a href="index.html#upcoming">Roadmap</a></li>
          <li><a href="https://huggingface.co/spaces/spcv/slm-agents" class="btn-hf-nav" target="_blank" style="display: flex; align-items: center; gap: 6px; padding: 7px 14px; border-radius: 8px; background: #fef3c7; border: 1px solid #f59e0b; color: #b45309; font-weight: 700; text-decoration: none; font-size: 0.85rem; transition: all 0.2s;">πŸ€— Hugging Face Space</a></li>
        </ul>
      </nav>
    </div>
  </header>

  <div class="layout-container">
    <aside class="sidebar" id="sidebar">
      <div class="sidebar-group">
        <div class="sidebar-group-title">Active Libraries</div>
        <ul class="sidebar-list" id="sidebar-active-list"></ul>
      </div>
    </aside>

    <main class="main-content">
      <div class="doc-content-wrapper">
        <section class="doc-section" style="margin-top: 0;">
          <div class="breadcrumb" style="font-size: 0.85rem; color: #475569; margin-bottom: 1rem; font-weight: 600;"><a href="index.html" style="color: #4f46e5; text-decoration: none;">Home</a> <span style="margin: 0 6px;">β€Ί</span> <span style="color: #0f172a;">SLM Vision Parser</span></div>
          <div class="badge-pill">πŸ‘οΈ Visual Layout Parsing</div>
          <h1 style="font-size: 2.25rem; color: #0f172a; font-weight: 800; margin-bottom: 0.75rem;">SLM Vision Parser</h1>
          <p style="font-size: 1.05rem; color: #334155; line-height: 1.65; font-weight: 500; margin: 0;">Extract text nodes, bounding boxes, chart labels, and table grids locally on CPU using Microsoft's MIT-licensed Florence-2-large model.</p>
        </section>

        <section class="doc-section">
        <h2>πŸš€ Overview &amp; Capabilities</h2>
        <p>The SLM Vision Parser uses the sequence-to-sequence <strong>Florence-2-large</strong> model to translate image structures directly. It is designed to run offline on CPU under 1.5 GB memory footprint, making it ideal for visual PDF OCR, diagram relationship mappings, and whiteboard extractions.</p>
      </section>

        <section class="doc-section" id="install">
          <h2>πŸ’» Installation</h2>
          <p>Install the local CPU-optimized package using pip:</p>
          <div class="code-panel" style="max-width:100%">
            <div class="code-header"><div class="code-dots"><div class="code-dot"></div><div class="code-dot"></div><div class="code-dot"></div></div><div class="code-title">Terminal</div></div>
            <div class="code-content" style="display:block;padding:1.5rem">
              <pre><code><span class="comment"># Install from PyPI</span>
pip install slm-vision-parser</code></pre>
            </div>
          </div>
        </section>

        <section class="doc-section">
        <h2>πŸ€– Truly Agentic Workflow</h2>
        <p>The Vision Parser maps visual coordinate dimensions into clean layout logs:</p>
        <ul>
          <li><strong>OCR Node Extraction:</strong> Translates raster text strings from diagrams.</li>
          <li><strong>Region Mapping:</strong> Generates bounding box parameters mapping visual coordinates for web rendering.</li>
          <li><strong>Semantic Image Captioning:</strong> Extracts visual flow structures (e.g., arrows, shapes, labels) into a raw description.</li>
          <li><strong>Agentic LLM Refinement:</strong> Pipe Florence-2's raw description or OCR nodes directly into the local Phi-3.5 ONNX model. By guiding the text model with a template prompt, you can synthesize exact custom descriptions (e.g., describing step hierarchies and the exact count of connecting arrows).</li>
        </ul>
      </section>

        <section class="doc-section">
        <h2>⚑ CPU Performance Tuning Guidelines</h2>
        <p>Florence-2 utilizes complex visual transformers. To run model prediction tasks efficiently on CPU under 2.0 GB memory:</p>
        <ul>
          <li><strong>Sequential execution:</strong> Process page scans sequentially rather than in batches to avoid CPU memory thrashing.</li>
          <li><strong>Garbage Collection:</strong> Clear cache blocks after processing high-resolution images:
            <pre><code>import gc
gc.collect()</code></pre>
          </li>
        </ul>
      </section>

        <section class="doc-section">
        <h2>🎯 Accuracy Improvement Tips</h2>
        <div class="tip-box">
          <strong>Tip for Scanned Diagrams:</strong> Contrast is key for visual OCR models. Preprocess low-contrast diagram images by resizing them to a standard width of <code>1024px</code> and scaling contrast values using Pillow before parsing.
        </div>
        <div class="warn-box">
          <strong>Task Tag Wrapper:</strong> Always encapsulate Florence-2 model instructions in explicit brackets (e.g., <code>"&lt;OCR&gt;"</code> or <code>"&lt;DETAILED_CAPTION&gt;"</code>) to trigger correct task headers.
        </div>
      </section>

        <section class="doc-section">
        <h2>API Reference</h2>
        <h3>`SLMVisionParser` Initialization</h3>
        <pre><code class="language-python">from slm_vision_parser.vision_parser import SLMVisionParser

parser = SLMVisionParser()</code></pre>
        
        <table class="param-table">
          <thead>
            <tr><th>Parameter</th><th>Type</th><th>Description</th></tr>
          </thead>
          <tbody>
            <tr><td>model_path</td><td>str</td><td>Local directory containing pre-downloaded Florence-2 checkpoints. Defaults to "../../models/florence-2-large".</td></tr>
                        <tr><td>system_prompt</td><td>str | None</td><td>Optional custom system prompt instructions overriding the default template.</td></tr>
              <tr><td>user_input</td><td>str | None</td><td>Optional additional user-supplied target parameters or variables.</td></tr>
</tbody>
        </table>

        <h3>`parse_image` Method</h3>
        <p>Run coordinate mappings or caption generation by supplying task tags. Below is the input flowchart diagram processed by the vision model:</p>
        
        <div style="margin: 20px 0; max-width: 480px; border-radius: 12px; overflow: hidden; border: 1px solid rgba(255, 255, 255, 0.05); box-shadow: 0 10px 30px rgba(0,0,0,0.5);">
          <img src="https://raw.githubusercontent.com/t00114218-stack/SLMAgents/main/website/flowchart.png" alt="Input Flowchart Diagram" style="width: 100%; height: auto; display: block;">
        </div>

        <pre><code class="language-python">from slm_vision_parser.vision_parser import SLMVisionParser

parser = SLMVisionParser()

# Describe the flowchart diagram in text using the local vision LLM
result = parser.parse_image("flowchart.png", task="&lt;DETAILED_CAPTION&gt;")
print(result)</code></pre>

        <div class="tip-box">
          <strong>Image Description Response Output:</strong><br>
          <pre><code class="language-python">"A flowchart showing a start step ('Start Process') and a next step ('Next Step Link') connected with two arrows from the start step to the next step."</code></pre>
        </div>
      </section>

        <footer style="margin-top: 3rem; text-align: center; border-top: 1px solid #cbd5e1; padding-top: 2rem; color: #475569; font-size: 0.9rem; font-weight: 500;">
          <p>Β© 2026 SLM Agents. Built with Apache 2.0 Permissive Open Source License.</p>
        </footer>
      </div>
    
        <!-- GIT CHECKOUT -->
        <section class="doc-section" id="git">
          <h2>πŸ™ Checkout from GitHub</h2>
          <p>Clone only this agent's folder from the monorepo using Git sparse-checkout β€” no need to download the full repository:</p>

          <h3 style="font-size: 1.05rem; color: #0f172a; font-weight: 700; margin-top: 1.5rem; margin-bottom: 0.75rem;">Option 1 β€” Sparse Checkout (Recommended)</h3>
          <div class="code-panel" style="max-width:100%; background: #0f172a; border: 1px solid #1e293b; border-radius: 14px; overflow: hidden; margin: 1rem 0; box-shadow: 0 16px 40px rgba(15, 23, 42, 0.12);">
            <div class="code-header" style="background: #1e293b; padding: 10px 16px; display: flex; align-items: center; justify-content: space-between; border-bottom: 1px solid #334155;">
              <div class="code-dots"><div class="code-dot"></div><div class="code-dot"></div><div class="code-dot"></div></div>
              <div class="code-title" style="color: #94a3b8; font-weight: 700; font-size: 0.8rem; font-family: 'JetBrains Mono', monospace;">Terminal β€” Git Sparse Checkout</div>
            </div>
            <div class="code-content" style="display:block; padding: 1.25rem 1.5rem; background: #0f172a;">
              <pre style="margin:0; background:#0f172a; color:#f8fafc; font-family:'JetBrains Mono',monospace; font-size:0.88rem; border:none; box-shadow:none; padding:0; line-height: 1.75;"><span style="color:#64748b;"># 1. Create and enter a new directory</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">mkdir</span> <span style="color:#38bdf8;">slm_vision_parser</span> <span style="color:#94a3b8;">&amp;&amp;</span> <span style="color:#c084fc; font-weight:700;">cd</span> <span style="color:#38bdf8;">slm_vision_parser</span>

<span style="color:#64748b;"># 2. Initialise empty git repo and add remote</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git init</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git remote add origin</span> <span style="color:#38bdf8;">https://github.com/t00114218-stack/SLMAgents.git</span>

<span style="color:#64748b;"># 3. Enable sparse-checkout and set target folder</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git sparse-checkout init</span> <span style="color:#94a3b8;">--cone</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git sparse-checkout set</span> <span style="color:#38bdf8;">slm_vision_parser</span>

<span style="color:#64748b;"># 4. Pull only that agent's source</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git pull origin</span> <span style="color:#38bdf8;">main</span></pre>
            </div>
          </div>

          <h3 style="font-size: 1.05rem; color: #0f172a; font-weight: 700; margin-top: 2rem; margin-bottom: 0.75rem;">Option 2 β€” Full Repository Clone</h3>
          <div class="code-panel" style="max-width:100%; background: #0f172a; border: 1px solid #1e293b; border-radius: 14px; overflow: hidden; margin: 1rem 0;">
            <div class="code-header" style="background: #1e293b; padding: 10px 16px; display: flex; align-items: center; justify-content: space-between; border-bottom: 1px solid #334155;">
              <div class="code-dots"><div class="code-dot"></div><div class="code-dot"></div><div class="code-dot"></div></div>
              <div class="code-title" style="color: #94a3b8; font-weight: 700; font-size: 0.8rem; font-family: 'JetBrains Mono', monospace;">Terminal β€” Full Clone</div>
            </div>
            <div class="code-content" style="display:block; padding: 1.25rem 1.5rem; background: #0f172a;">
              <pre style="margin:0; background:#0f172a; color:#f8fafc; font-family:'JetBrains Mono',monospace; font-size:0.88rem; border:none; box-shadow:none; padding:0; line-height: 1.75;"><span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">git clone</span> <span style="color:#38bdf8;">https://github.com/t00114218-stack/SLMAgents.git</span>
<span style="color:#34d399;">$</span> <span style="color:#c084fc; font-weight:700;">cd</span> <span style="color:#38bdf8;">SLMAgents/slm_vision_parser</span></pre>
            </div>
          </div>

          <p style="margin-top: 1.25rem; font-size: 0.9rem; color: #475569; background: #f8fafc; border: 1px solid #cbd5e1; border-radius: 10px; padding: 1rem 1.25rem;">
            πŸ’‘ <strong>Tip:</strong> After checkout, install the package locally with <code style="background: #eef2ff; color: #4f46e5; border: 1px solid #c7d2fe; padding: 2px 8px; border-radius: 5px; font-weight: 700;">pip install -e ./slm_vision_parser</code> to run in editable mode without publishing to PyPI.
          </p>
        </section>

</main>
  </div>

  <script src="app.js"></script>
</body>
</html>