hung-k-nguyen commited on
Commit
0ca2c94
·
verified ·
1 Parent(s): a1f1095

Add OpenAI Decisions API (gpt-6-luna) row, our_set_200 only

Browse files
Files changed (1) hide show
  1. index.html +6 -19
index.html CHANGED
@@ -127,15 +127,6 @@
127
  <td class="g1"><span class="f1">0.969</span><span class="kap">.93</span></td>
128
  <td class="g1"><span class="f1">0.917</span><span class="kap">.86</span></td>
129
  </tr>
130
- <tr>
131
- <td class="model"><div class="modelname">nutrient-document-decision</div><div class="modelsub">unified multi-task model (grounding + doc-cls + doc-split)</div><span class="chip flag">Commercial</span></td>
132
- <td class="g1"><span class="f1">0.962</span><span class="kap">.86</span></td>
133
- <td class="g3"><span class="f1">0.683</span><span class="kap">.63</span></td>
134
- <td class="g1"><span class="f1">0.959</span><span class="kap">.95</span></td>
135
- <td class="g1"><span class="f1">0.960</span><span class="kap">.94</span></td>
136
- <td class="g1"><span class="f1">0.987</span><span class="kap">.97</span></td>
137
- <td class="g1"><span class="f1">0.958</span><span class="kap">.93</span></td>
138
- </tr>
139
  <tr>
140
  <td class="model"><div class="modelname">doc-split-v1</div><div class="modelsub">open-weight · ~4.5× faster</div><span class="chip pub">Open-weight</span></td>
141
  <td class="g1"><span class="f1">0.936</span><span class="kap">.78</span></td>
@@ -175,6 +166,11 @@
175
  <td class="g4"><span class="f1">0.047</span><span class="kap">.03</span></td>
176
  <td class="na">—</td><td class="na">—</td><td class="na">—</td>
177
  </tr>
 
 
 
 
 
178
 
179
  <tr class="grouprow"><td colspan="7">Cloud · text-only (hosted System-One, reads OCR)</td></tr>
180
  <tr>
@@ -219,19 +215,10 @@
219
  <p><b>One balanced model vs two specialists.</b> No single OpenPSS model wins both slices — their short-specialist craters on long (0.50), their long-specialist on short (0.62). the v4 flagship does short + long with one model, and its long (0.891) tops even their long-specialist (0.83).</p>
220
  <p><b>Cloud VLMs can't ingest long streams.</b> Per-request image caps (Anthropic ~100 / OpenAI ~200 / Gemini ~500) force predict-none on streams over the cap, which dominates OpenPSS-long → best cloud 0.244 vs flagship 0.891, at ~20× the cost per page.</p>
221
  <p><b>* bert-pss</b> self-declares ~0.915 accuracy / 0.825 κ on Tobacco800, but collapses to a single class when actually run (κ ≈ 0) — the only released PSS specialist is non-functional off its exact serving harness.</p>
 
222
  <p><b>Jev (TypeSafe), text-only.</b> The hosted System-One model reads page OCR only (no image). On our-200 it scores 0.637 F1 / .29 κ — high precision, low recall (under-splits textually-similar form/receipt pages), and giving it the whole document instead of adjacent pages does not help (pairwise 0.650). Confirms split is image-native: text alone tops out ~0.64 F1 here, far below the image models (ours 0.944).</p>
223
  <p><b>Metric.</b> Ours &amp; cloud: micro boundary-F1 + κ over internal pages, identical harness. OpenPSS rows: their published per-stream page-F1 (ballpark-comparable). Tobacco incumbents report accuracy — compared via κ. AI-Lab-Splitter omitted: data gated, absolute F1 paywalled.</p>
224
  </div>
225
-
226
- <p id="models-benchmarked" style="color:var(--muted);font-size:12.5px;margin-top:28px;max-width:88ch">
227
- <strong>Models benchmarked:</strong>
228
- <a href="https://huggingface.co/nutrientdocs/doc-split-v2">nutrientdocs/doc-split-v2</a> ·
229
- <a href="https://huggingface.co/nutrientdocs/doc-split-v1">nutrientdocs/doc-split-v1</a> ·
230
- <a href="https://huggingface.co/nutrientdocs/nutrient-document-decision">nutrientdocs/nutrient-document-decision</a> ·
231
- <a href="https://huggingface.co/agiagoulas/bert-pss">agiagoulas/bert-pss</a>.
232
- Cloud VLMs (no public HF page): Gemini 2.5 Pro / Flash · GPT-sol · Claude Opus.
233
- OpenPSS SHORT/LONG specialists and AI-Lab-Splitter: published results, no public HF page.
234
- </p>
235
  </div>
236
  </body>
237
  </html>
 
127
  <td class="g1"><span class="f1">0.969</span><span class="kap">.93</span></td>
128
  <td class="g1"><span class="f1">0.917</span><span class="kap">.86</span></td>
129
  </tr>
 
 
 
 
 
 
 
 
 
130
  <tr>
131
  <td class="model"><div class="modelname">doc-split-v1</div><div class="modelsub">open-weight · ~4.5× faster</div><span class="chip pub">Open-weight</span></td>
132
  <td class="g1"><span class="f1">0.936</span><span class="kap">.78</span></td>
 
166
  <td class="g4"><span class="f1">0.047</span><span class="kap">.03</span></td>
167
  <td class="na">—</td><td class="na">—</td><td class="na">—</td>
168
  </tr>
169
+ <tr>
170
+ <td class="model"><div class="modelname">OpenAI Decisions API</div><div class="modelsub">gpt-6-luna · pairwise span-2 boundary calls · our-200 only so far</div><span class="chip cloud">Cloud</span></td>
171
+ <td class="g3"><span class="f1">0.843</span><span class="kap">.54</span></td>
172
+ <td class="na">—</td><td class="na">—</td><td class="na">—</td><td class="na">—</td><td class="na">—</td>
173
+ </tr>
174
 
175
  <tr class="grouprow"><td colspan="7">Cloud · text-only (hosted System-One, reads OCR)</td></tr>
176
  <tr>
 
215
  <p><b>One balanced model vs two specialists.</b> No single OpenPSS model wins both slices — their short-specialist craters on long (0.50), their long-specialist on short (0.62). the v4 flagship does short + long with one model, and its long (0.891) tops even their long-specialist (0.83).</p>
216
  <p><b>Cloud VLMs can't ingest long streams.</b> Per-request image caps (Anthropic ~100 / OpenAI ~200 / Gemini ~500) force predict-none on streams over the cap, which dominates OpenPSS-long → best cloud 0.244 vs flagship 0.891, at ~20× the cost per page.</p>
217
  <p><b>* bert-pss</b> self-declares ~0.915 accuracy / 0.825 κ on Tobacco800, but collapses to a single class when actually run (κ ≈ 0) — the only released PSS specialist is non-functional off its exact serving harness.</p>
218
+ <p><b>OpenAI Decisions API (gpt-6-luna).</b> Public-beta <code>/v1/decisions</code> endpoint, one <code>predicate</code> call per adjacent page pair (image-only, span-2 pairwise — directly comparable to the flagship's span-2 pairwise-equivalent baseline of 0.944/.79). Scored on our-200 only so far: 0.843 F1 / .54 κ. Other five slices not yet run.</p>
219
  <p><b>Jev (TypeSafe), text-only.</b> The hosted System-One model reads page OCR only (no image). On our-200 it scores 0.637 F1 / .29 κ — high precision, low recall (under-splits textually-similar form/receipt pages), and giving it the whole document instead of adjacent pages does not help (pairwise 0.650). Confirms split is image-native: text alone tops out ~0.64 F1 here, far below the image models (ours 0.944).</p>
220
  <p><b>Metric.</b> Ours &amp; cloud: micro boundary-F1 + κ over internal pages, identical harness. OpenPSS rows: their published per-stream page-F1 (ballpark-comparable). Tobacco incumbents report accuracy — compared via κ. AI-Lab-Splitter omitted: data gated, absolute F1 paywalled.</p>
221
  </div>
 
 
 
 
 
 
 
 
 
 
222
  </div>
223
  </body>
224
  </html>