card: block 5 — per-scope movement on both slices
Browse files
README.md
CHANGED
|
@@ -72,6 +72,9 @@ something the others hid:
|
|
| 72 |
* **4 — composed.** Horos never runs alone; a head-only number is not a product number.
|
| 73 |
It is also where an unqualified 0.500-vs-0.495 claim went out on this card and turned
|
| 74 |
out to be inside the noise (see block 4).
|
|
|
|
|
|
|
|
|
|
| 75 |
|
| 76 |
`swallowed` and `escalated` stay separate throughout: they have opposite fixes, and
|
| 77 |
collapsing them into "not answered" hides which one you have.
|
|
@@ -97,6 +100,8 @@ column is the same benchmark's LLM baseline.
|
|
| 97 |
| **disjoint rate** | 0.206 | 0.217 | — | ≤0.03 ❌ |
|
| 98 |
| per-scope recall ≥ 0.60 | 6 / 14 | 2 / 14 | 9 / 14 | 14/14 ❌ |
|
| 99 |
|
|
|
|
|
|
|
| 100 |
### 2. Real language — a gap finder, not a score
|
| 101 |
|
| 102 |
53 hand-annotated natural phrasings neither version trained on. The labels are one
|
|
@@ -157,6 +162,68 @@ Note the direction of the v1 → v2 trade in this table: composed accuracy up ~7
|
|
| 157 |
composed negatives-abstained down ~6. The escalation path was covering for the head's
|
| 158 |
false-positives, and v2 hands it less to cover.
|
| 159 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 160 |
## Limitations
|
| 161 |
|
| 162 |
* **Confident-none swallowing, now concentrated rather than general.** Overall dead rate
|
|
@@ -175,8 +242,12 @@ false-positives, and v2 hands it less to cover.
|
|
| 175 |
knowing *which* of your data are separate abilities; this round only advanced the
|
| 176 |
first. The escalation path absorbs less of that than it used to — composed
|
| 177 |
negatives-abstained fell 0.972 → 0.909.
|
| 178 |
-
* **8 of 14 scopes are under the 0.60 recall floor** (v1: 12)
|
| 179 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 180 |
* All numbers are synthetic-benchmark. Real-traffic behaviour is being measured in
|
| 181 |
shadow mode.
|
| 182 |
* English only. **No user data, ever** — the loader refuses artifacts whose manifest
|
|
|
|
| 72 |
* **4 — composed.** Horos never runs alone; a head-only number is not a product number.
|
| 73 |
It is also where an unqualified 0.500-vs-0.495 claim went out on this card and turned
|
| 74 |
out to be inside the noise (see block 4).
|
| 75 |
+
* **5 — per-scope.** Every other block is an average or a count, and both let a gain on
|
| 76 |
+
one scope pay for a regression on another without saying so. Only this view names the
|
| 77 |
+
scope that got worse.
|
| 78 |
|
| 79 |
`swallowed` and `escalated` stay separate throughout: they have opposite fixes, and
|
| 80 |
collapsing them into "not answered" hides which one you have.
|
|
|
|
| 100 |
| **disjoint rate** | 0.206 | 0.217 | — | ≤0.03 ❌ |
|
| 101 |
| per-scope recall ≥ 0.60 | 6 / 14 | 2 / 14 | 9 / 14 | 14/14 ❌ |
|
| 102 |
|
| 103 |
+
Which six, and which scope went backwards: block 5.
|
| 104 |
+
|
| 105 |
### 2. Real language — a gap finder, not a score
|
| 106 |
|
| 107 |
53 hand-annotated natural phrasings neither version trained on. The labels are one
|
|
|
|
| 162 |
composed negatives-abstained down ~6. The escalation path was covering for the head's
|
| 163 |
false-positives, and v2 hands it less to cover.
|
| 164 |
|
| 165 |
+
### 5. Per-scope — where it moved, and where it didn't
|
| 166 |
+
|
| 167 |
+
Blocks 1 and 3 report counts ("6 / 14 above the floor", "routed 0.565"). A count cannot
|
| 168 |
+
be audited: it says how many scopes cleared the bar and never which, so a large gain on
|
| 169 |
+
one scope silently pays for a regression on another. Both slices, sorted by movement.
|
| 170 |
+
|
| 171 |
+
**Unseen phrasings** — did the gains reach language nobody wrote down?
|
| 172 |
+
|
| 173 |
+
| scope | n | v1 | v2 | Δ |
|
| 174 |
+
|---|---|---|---|---|
|
| 175 |
+
| `places` | 217 | 30% | 68% | **+37** |
|
| 176 |
+
| `schedule` | 65 | 31% | 65% | **+34** |
|
| 177 |
+
| `work_context` | 420 | 22% | 51% | **+29** |
|
| 178 |
+
| `relationship_context` | 174 | 16% | 42% | **+26** |
|
| 179 |
+
| `public_bio` | 229 | 67% | 85% | +17 |
|
| 180 |
+
| `messages` | 253 | 49% | 66% | +17 |
|
| 181 |
+
| `attention` | 145 | 35% | 46% | +10 |
|
| 182 |
+
| `complexity` | 178 | 39% | 47% | +8 |
|
| 183 |
+
| `activity` | 308 | 61% | 68% | +7 |
|
| 184 |
+
| `resources` | 151 | 28% | 34% | +6 |
|
| 185 |
+
| `availability` | 116 | 21% | 25% | +4 |
|
| 186 |
+
| `contacts` | 119 | 76% | 73% | −3 |
|
| 187 |
+
| `ai_conversations` | 111 | 42% | 39% | −4 |
|
| 188 |
+
| `health` | 259 | 61% | 56% | **−5** |
|
| 189 |
+
|
| 190 |
+
**Gate benchmark** — this is what "6 / 14 above the floor" expands to.
|
| 191 |
+
|
| 192 |
+
| scope | n | v1 | v2 | Δ | ≥0.60 |
|
| 193 |
+
|---|---|---|---|---|---|
|
| 194 |
+
| `contacts` | 58 | 69% | 83% | +14 | ✅ |
|
| 195 |
+
| `public_bio` | 54 | 67% | 81% | +15 | ✅ |
|
| 196 |
+
| `attention` | 73 | 25% | 71% | **+47** | ✅ |
|
| 197 |
+
| `schedule` | 59 | 44% | 69% | +25 | ✅ |
|
| 198 |
+
| `activity` | 78 | 58% | 64% | +6 | ✅ |
|
| 199 |
+
| `messages` | 61 | 51% | 64% | +13 | ✅ |
|
| 200 |
+
| `health` | 116 | 47% | 54% | +8 | ❌ |
|
| 201 |
+
| `places` | 61 | 26% | 49% | +23 | ❌ |
|
| 202 |
+
| `availability` | 62 | 35% | 45% | +10 | ❌ |
|
| 203 |
+
| `resources` | 70 | 27% | 37% | +10 | ❌ |
|
| 204 |
+
| `complexity` | 82 | 35% | 37% | +1 | ❌ |
|
| 205 |
+
| `ai_conversations` | 54 | 43% | 35% | **−7** | ❌ |
|
| 206 |
+
| `relationship_context` | 77 | 17% | 34% | +17 | ❌ |
|
| 207 |
+
| `work_context` | 73 | 25% | 30% | +5 | ❌ |
|
| 208 |
+
|
| 209 |
+
Three things only this view shows:
|
| 210 |
+
|
| 211 |
+
* **`ai_conversations` is down on both slices** (−7, −4) — the one unambiguous regression,
|
| 212 |
+
not a slice artifact. v2's training targeted artifact-concrete phrasings, and questions
|
| 213 |
+
about your own past AI conversations are the scope least like an artifact.
|
| 214 |
+
* **`health` and `contacts` flip sign between slices.** `health` gains 8 on the benchmark
|
| 215 |
+
and loses 5 on unseen phrasings; `contacts` gains 14 and loses 3. The two instruments
|
| 216 |
+
measure genuinely different things, and a card publishing only one of them would report
|
| 217 |
+
either as a clean win.
|
| 218 |
+
* **The four biggest unseen gains are exactly the four scopes v2's authoring targeted**
|
| 219 |
+
(`places`, `schedule`, `work_context`, `relationship_context`, +26 to +37). That is the
|
| 220 |
+
evidence the authoring generalised rather than being memorised — the gains land on
|
| 221 |
+
phrasings of those scopes that nobody wrote down.
|
| 222 |
+
|
| 223 |
+
Floors are still unmet: on the gate slice, 8 of 14 scopes sit under the 0.60 recall bar.
|
| 224 |
+
`work_context` (30%), `complexity` (37%) and `resources` (37%) are the weakest, and
|
| 225 |
+
`relationship_context` at 34% remains the hardest scope in the taxonomy.
|
| 226 |
+
|
| 227 |
## Limitations
|
| 228 |
|
| 229 |
* **Confident-none swallowing, now concentrated rather than general.** Overall dead rate
|
|
|
|
| 242 |
knowing *which* of your data are separate abilities; this round only advanced the
|
| 243 |
first. The escalation path absorbs less of that than it used to — composed
|
| 244 |
negatives-abstained fell 0.972 → 0.909.
|
| 245 |
+
* **8 of 14 scopes are under the 0.60 recall floor** (v1: 12), weakest `work_context`
|
| 246 |
+
30%, `relationship_context` 34%, `complexity` and `resources` 37%. This artifact has
|
| 247 |
+
not cleared its promotion gate; it fronts an LLM in shadow/advisory postures only.
|
| 248 |
+
* **`ai_conversations` regressed on both slices** (−7 gate, −4 unseen) — the one scope
|
| 249 |
+
v2 made unambiguously worse. Questions about your own past AI conversations are the
|
| 250 |
+
scope least like the artifact-concrete register v2 was trained to fix.
|
| 251 |
* All numbers are synthetic-benchmark. Real-traffic behaviour is being measured in
|
| 252 |
shadow mode.
|
| 253 |
* English only. **No user data, ever** — the loader refuses artifacts whose manifest
|