Add OpenAI Decisions API (gpt-6-luna) row, our_set_200 only
Browse files- index.html +6 -19
index.html
CHANGED
|
@@ -127,15 +127,6 @@
|
|
| 127 |
<td class="g1"><span class="f1">0.969</span><span class="kap">.93</span></td>
|
| 128 |
<td class="g1"><span class="f1">0.917</span><span class="kap">.86</span></td>
|
| 129 |
</tr>
|
| 130 |
-
<tr>
|
| 131 |
-
<td class="model"><div class="modelname">nutrient-document-decision</div><div class="modelsub">unified multi-task model (grounding + doc-cls + doc-split)</div><span class="chip flag">Commercial</span></td>
|
| 132 |
-
<td class="g1"><span class="f1">0.962</span><span class="kap">.86</span></td>
|
| 133 |
-
<td class="g3"><span class="f1">0.683</span><span class="kap">.63</span></td>
|
| 134 |
-
<td class="g1"><span class="f1">0.959</span><span class="kap">.95</span></td>
|
| 135 |
-
<td class="g1"><span class="f1">0.960</span><span class="kap">.94</span></td>
|
| 136 |
-
<td class="g1"><span class="f1">0.987</span><span class="kap">.97</span></td>
|
| 137 |
-
<td class="g1"><span class="f1">0.958</span><span class="kap">.93</span></td>
|
| 138 |
-
</tr>
|
| 139 |
<tr>
|
| 140 |
<td class="model"><div class="modelname">doc-split-v1</div><div class="modelsub">open-weight · ~4.5× faster</div><span class="chip pub">Open-weight</span></td>
|
| 141 |
<td class="g1"><span class="f1">0.936</span><span class="kap">.78</span></td>
|
|
@@ -175,6 +166,11 @@
|
|
| 175 |
<td class="g4"><span class="f1">0.047</span><span class="kap">.03</span></td>
|
| 176 |
<td class="na">—</td><td class="na">—</td><td class="na">—</td>
|
| 177 |
</tr>
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 178 |
|
| 179 |
<tr class="grouprow"><td colspan="7">Cloud · text-only (hosted System-One, reads OCR)</td></tr>
|
| 180 |
<tr>
|
|
@@ -219,19 +215,10 @@
|
|
| 219 |
<p><b>One balanced model vs two specialists.</b> No single OpenPSS model wins both slices — their short-specialist craters on long (0.50), their long-specialist on short (0.62). the v4 flagship does short + long with one model, and its long (0.891) tops even their long-specialist (0.83).</p>
|
| 220 |
<p><b>Cloud VLMs can't ingest long streams.</b> Per-request image caps (Anthropic ~100 / OpenAI ~200 / Gemini ~500) force predict-none on streams over the cap, which dominates OpenPSS-long → best cloud 0.244 vs flagship 0.891, at ~20× the cost per page.</p>
|
| 221 |
<p><b>* bert-pss</b> self-declares ~0.915 accuracy / 0.825 κ on Tobacco800, but collapses to a single class when actually run (κ ≈ 0) — the only released PSS specialist is non-functional off its exact serving harness.</p>
|
|
|
|
| 222 |
<p><b>Jev (TypeSafe), text-only.</b> The hosted System-One model reads page OCR only (no image). On our-200 it scores 0.637 F1 / .29 κ — high precision, low recall (under-splits textually-similar form/receipt pages), and giving it the whole document instead of adjacent pages does not help (pairwise 0.650). Confirms split is image-native: text alone tops out ~0.64 F1 here, far below the image models (ours 0.944).</p>
|
| 223 |
<p><b>Metric.</b> Ours & cloud: micro boundary-F1 + κ over internal pages, identical harness. OpenPSS rows: their published per-stream page-F1 (ballpark-comparable). Tobacco incumbents report accuracy — compared via κ. AI-Lab-Splitter omitted: data gated, absolute F1 paywalled.</p>
|
| 224 |
</div>
|
| 225 |
-
|
| 226 |
-
<p id="models-benchmarked" style="color:var(--muted);font-size:12.5px;margin-top:28px;max-width:88ch">
|
| 227 |
-
<strong>Models benchmarked:</strong>
|
| 228 |
-
<a href="https://huggingface.co/nutrientdocs/doc-split-v2">nutrientdocs/doc-split-v2</a> ·
|
| 229 |
-
<a href="https://huggingface.co/nutrientdocs/doc-split-v1">nutrientdocs/doc-split-v1</a> ·
|
| 230 |
-
<a href="https://huggingface.co/nutrientdocs/nutrient-document-decision">nutrientdocs/nutrient-document-decision</a> ·
|
| 231 |
-
<a href="https://huggingface.co/agiagoulas/bert-pss">agiagoulas/bert-pss</a>.
|
| 232 |
-
Cloud VLMs (no public HF page): Gemini 2.5 Pro / Flash · GPT-sol · Claude Opus.
|
| 233 |
-
OpenPSS SHORT/LONG specialists and AI-Lab-Splitter: published results, no public HF page.
|
| 234 |
-
</p>
|
| 235 |
</div>
|
| 236 |
</body>
|
| 237 |
</html>
|
|
|
|
| 127 |
<td class="g1"><span class="f1">0.969</span><span class="kap">.93</span></td>
|
| 128 |
<td class="g1"><span class="f1">0.917</span><span class="kap">.86</span></td>
|
| 129 |
</tr>
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 130 |
<tr>
|
| 131 |
<td class="model"><div class="modelname">doc-split-v1</div><div class="modelsub">open-weight · ~4.5× faster</div><span class="chip pub">Open-weight</span></td>
|
| 132 |
<td class="g1"><span class="f1">0.936</span><span class="kap">.78</span></td>
|
|
|
|
| 166 |
<td class="g4"><span class="f1">0.047</span><span class="kap">.03</span></td>
|
| 167 |
<td class="na">—</td><td class="na">—</td><td class="na">—</td>
|
| 168 |
</tr>
|
| 169 |
+
<tr>
|
| 170 |
+
<td class="model"><div class="modelname">OpenAI Decisions API</div><div class="modelsub">gpt-6-luna · pairwise span-2 boundary calls · our-200 only so far</div><span class="chip cloud">Cloud</span></td>
|
| 171 |
+
<td class="g3"><span class="f1">0.843</span><span class="kap">.54</span></td>
|
| 172 |
+
<td class="na">—</td><td class="na">—</td><td class="na">—</td><td class="na">—</td><td class="na">—</td>
|
| 173 |
+
</tr>
|
| 174 |
|
| 175 |
<tr class="grouprow"><td colspan="7">Cloud · text-only (hosted System-One, reads OCR)</td></tr>
|
| 176 |
<tr>
|
|
|
|
| 215 |
<p><b>One balanced model vs two specialists.</b> No single OpenPSS model wins both slices — their short-specialist craters on long (0.50), their long-specialist on short (0.62). the v4 flagship does short + long with one model, and its long (0.891) tops even their long-specialist (0.83).</p>
|
| 216 |
<p><b>Cloud VLMs can't ingest long streams.</b> Per-request image caps (Anthropic ~100 / OpenAI ~200 / Gemini ~500) force predict-none on streams over the cap, which dominates OpenPSS-long → best cloud 0.244 vs flagship 0.891, at ~20× the cost per page.</p>
|
| 217 |
<p><b>* bert-pss</b> self-declares ~0.915 accuracy / 0.825 κ on Tobacco800, but collapses to a single class when actually run (κ ≈ 0) — the only released PSS specialist is non-functional off its exact serving harness.</p>
|
| 218 |
+
<p><b>OpenAI Decisions API (gpt-6-luna).</b> Public-beta <code>/v1/decisions</code> endpoint, one <code>predicate</code> call per adjacent page pair (image-only, span-2 pairwise — directly comparable to the flagship's span-2 pairwise-equivalent baseline of 0.944/.79). Scored on our-200 only so far: 0.843 F1 / .54 κ. Other five slices not yet run.</p>
|
| 219 |
<p><b>Jev (TypeSafe), text-only.</b> The hosted System-One model reads page OCR only (no image). On our-200 it scores 0.637 F1 / .29 κ — high precision, low recall (under-splits textually-similar form/receipt pages), and giving it the whole document instead of adjacent pages does not help (pairwise 0.650). Confirms split is image-native: text alone tops out ~0.64 F1 here, far below the image models (ours 0.944).</p>
|
| 220 |
<p><b>Metric.</b> Ours & cloud: micro boundary-F1 + κ over internal pages, identical harness. OpenPSS rows: their published per-stream page-F1 (ballpark-comparable). Tobacco incumbents report accuracy — compared via κ. AI-Lab-Splitter omitted: data gated, absolute F1 paywalled.</p>
|
| 221 |
</div>
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 222 |
</div>
|
| 223 |
</body>
|
| 224 |
</html>
|