jonny commited on
Commit
1b2a2fb
·
verified ·
1 Parent(s): 1dd7c3e

card: block 5 — per-scope movement on both slices

Browse files
Files changed (1) hide show
  1. README.md +73 -2
README.md CHANGED
@@ -72,6 +72,9 @@ something the others hid:
72
  * **4 — composed.** Horos never runs alone; a head-only number is not a product number.
73
  It is also where an unqualified 0.500-vs-0.495 claim went out on this card and turned
74
  out to be inside the noise (see block 4).
 
 
 
75
 
76
  `swallowed` and `escalated` stay separate throughout: they have opposite fixes, and
77
  collapsing them into "not answered" hides which one you have.
@@ -97,6 +100,8 @@ column is the same benchmark's LLM baseline.
97
  | **disjoint rate** | 0.206 | 0.217 | — | ≤0.03 ❌ |
98
  | per-scope recall ≥ 0.60 | 6 / 14 | 2 / 14 | 9 / 14 | 14/14 ❌ |
99
 
 
 
100
  ### 2. Real language — a gap finder, not a score
101
 
102
  53 hand-annotated natural phrasings neither version trained on. The labels are one
@@ -157,6 +162,68 @@ Note the direction of the v1 → v2 trade in this table: composed accuracy up ~7
157
  composed negatives-abstained down ~6. The escalation path was covering for the head's
158
  false-positives, and v2 hands it less to cover.
159
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
160
  ## Limitations
161
 
162
  * **Confident-none swallowing, now concentrated rather than general.** Overall dead rate
@@ -175,8 +242,12 @@ false-positives, and v2 hands it less to cover.
175
  knowing *which* of your data are separate abilities; this round only advanced the
176
  first. The escalation path absorbs less of that than it used to — composed
177
  negatives-abstained fell 0.972 → 0.909.
178
- * **8 of 14 scopes are under the 0.60 recall floor** (v1: 12). This artifact has not
179
- cleared its promotion gate; it fronts an LLM in shadow/advisory postures only.
 
 
 
 
180
  * All numbers are synthetic-benchmark. Real-traffic behaviour is being measured in
181
  shadow mode.
182
  * English only. **No user data, ever** — the loader refuses artifacts whose manifest
 
72
  * **4 — composed.** Horos never runs alone; a head-only number is not a product number.
73
  It is also where an unqualified 0.500-vs-0.495 claim went out on this card and turned
74
  out to be inside the noise (see block 4).
75
+ * **5 — per-scope.** Every other block is an average or a count, and both let a gain on
76
+ one scope pay for a regression on another without saying so. Only this view names the
77
+ scope that got worse.
78
 
79
  `swallowed` and `escalated` stay separate throughout: they have opposite fixes, and
80
  collapsing them into "not answered" hides which one you have.
 
100
  | **disjoint rate** | 0.206 | 0.217 | — | ≤0.03 ❌ |
101
  | per-scope recall ≥ 0.60 | 6 / 14 | 2 / 14 | 9 / 14 | 14/14 ❌ |
102
 
103
+ Which six, and which scope went backwards: block 5.
104
+
105
  ### 2. Real language — a gap finder, not a score
106
 
107
  53 hand-annotated natural phrasings neither version trained on. The labels are one
 
162
  composed negatives-abstained down ~6. The escalation path was covering for the head's
163
  false-positives, and v2 hands it less to cover.
164
 
165
+ ### 5. Per-scope — where it moved, and where it didn't
166
+
167
+ Blocks 1 and 3 report counts ("6 / 14 above the floor", "routed 0.565"). A count cannot
168
+ be audited: it says how many scopes cleared the bar and never which, so a large gain on
169
+ one scope silently pays for a regression on another. Both slices, sorted by movement.
170
+
171
+ **Unseen phrasings** — did the gains reach language nobody wrote down?
172
+
173
+ | scope | n | v1 | v2 | Δ |
174
+ |---|---|---|---|---|
175
+ | `places` | 217 | 30% | 68% | **+37** |
176
+ | `schedule` | 65 | 31% | 65% | **+34** |
177
+ | `work_context` | 420 | 22% | 51% | **+29** |
178
+ | `relationship_context` | 174 | 16% | 42% | **+26** |
179
+ | `public_bio` | 229 | 67% | 85% | +17 |
180
+ | `messages` | 253 | 49% | 66% | +17 |
181
+ | `attention` | 145 | 35% | 46% | +10 |
182
+ | `complexity` | 178 | 39% | 47% | +8 |
183
+ | `activity` | 308 | 61% | 68% | +7 |
184
+ | `resources` | 151 | 28% | 34% | +6 |
185
+ | `availability` | 116 | 21% | 25% | +4 |
186
+ | `contacts` | 119 | 76% | 73% | −3 |
187
+ | `ai_conversations` | 111 | 42% | 39% | −4 |
188
+ | `health` | 259 | 61% | 56% | **−5** |
189
+
190
+ **Gate benchmark** — this is what "6 / 14 above the floor" expands to.
191
+
192
+ | scope | n | v1 | v2 | Δ | ≥0.60 |
193
+ |---|---|---|---|---|---|
194
+ | `contacts` | 58 | 69% | 83% | +14 | ✅ |
195
+ | `public_bio` | 54 | 67% | 81% | +15 | ✅ |
196
+ | `attention` | 73 | 25% | 71% | **+47** | ✅ |
197
+ | `schedule` | 59 | 44% | 69% | +25 | ✅ |
198
+ | `activity` | 78 | 58% | 64% | +6 | ✅ |
199
+ | `messages` | 61 | 51% | 64% | +13 | ✅ |
200
+ | `health` | 116 | 47% | 54% | +8 | ❌ |
201
+ | `places` | 61 | 26% | 49% | +23 | ❌ |
202
+ | `availability` | 62 | 35% | 45% | +10 | ❌ |
203
+ | `resources` | 70 | 27% | 37% | +10 | ❌ |
204
+ | `complexity` | 82 | 35% | 37% | +1 | ❌ |
205
+ | `ai_conversations` | 54 | 43% | 35% | **−7** | ❌ |
206
+ | `relationship_context` | 77 | 17% | 34% | +17 | ❌ |
207
+ | `work_context` | 73 | 25% | 30% | +5 | ❌ |
208
+
209
+ Three things only this view shows:
210
+
211
+ * **`ai_conversations` is down on both slices** (−7, −4) — the one unambiguous regression,
212
+ not a slice artifact. v2's training targeted artifact-concrete phrasings, and questions
213
+ about your own past AI conversations are the scope least like an artifact.
214
+ * **`health` and `contacts` flip sign between slices.** `health` gains 8 on the benchmark
215
+ and loses 5 on unseen phrasings; `contacts` gains 14 and loses 3. The two instruments
216
+ measure genuinely different things, and a card publishing only one of them would report
217
+ either as a clean win.
218
+ * **The four biggest unseen gains are exactly the four scopes v2's authoring targeted**
219
+ (`places`, `schedule`, `work_context`, `relationship_context`, +26 to +37). That is the
220
+ evidence the authoring generalised rather than being memorised — the gains land on
221
+ phrasings of those scopes that nobody wrote down.
222
+
223
+ Floors are still unmet: on the gate slice, 8 of 14 scopes sit under the 0.60 recall bar.
224
+ `work_context` (30%), `complexity` (37%) and `resources` (37%) are the weakest, and
225
+ `relationship_context` at 34% remains the hardest scope in the taxonomy.
226
+
227
  ## Limitations
228
 
229
  * **Confident-none swallowing, now concentrated rather than general.** Overall dead rate
 
242
  knowing *which* of your data are separate abilities; this round only advanced the
243
  first. The escalation path absorbs less of that than it used to — composed
244
  negatives-abstained fell 0.972 → 0.909.
245
+ * **8 of 14 scopes are under the 0.60 recall floor** (v1: 12), weakest `work_context`
246
+ 30%, `relationship_context` 34%, `complexity` and `resources` 37%. This artifact has
247
+ not cleared its promotion gate; it fronts an LLM in shadow/advisory postures only.
248
+ * **`ai_conversations` regressed on both slices** (−7 gate, −4 unseen) — the one scope
249
+ v2 made unambiguously worse. Questions about your own past AI conversations are the
250
+ scope least like the artifact-concrete register v2 was trained to fix.
251
  * All numbers are synthetic-benchmark. Real-traffic behaviour is being measured in
252
  shadow mode.
253
  * English only. **No user data, ever** — the loader refuses artifacts whose manifest