RandomMountainMan commited on
Commit
c3c6277
·
verified ·
1 Parent(s): 031dd3b

Restore verified transitive-depth comparison

Browse files
README.md CHANGED
@@ -73,15 +73,15 @@ diagnostic screen, through each model's native instruction interface,
73
  with greedy decoding, repetition penalty 1.15, and matched short-answer
74
  budgets.
75
 
76
- | model | parameters | mixed arithmetic (n=585) | executed functions (n=100) | designated refusals (n=17) |
77
- |---|---:|---:|---:|---:|
78
- | ConeML Alpha | 0.81B | 421 (72.0%) | 83 | **13** |
79
- | ConeML Arithmetic | 0.81B | **442 (75.6%)** | 35 | 11 |
80
- | Qwen3.5 | 0.8B | 154 (26.3%) | 93 | 0 |
81
- | Qwen3 | 0.6B | 223 (38.1%) | **98** | 1 |
82
- | Llama 3.2 Instruct | 1.24B | 343 (58.6%) | 79 | 1 |
83
- | TinyLlama Chat | 1.1B | 113 (19.3%) | 51 | 0 |
84
- | SmolLM2 Instruct | 1.7B | 380 (65.0%) | 96 | 1 |
85
 
86
  This is not a general leaderboard: the task families match ConeML's
87
  trained surfaces. On this screen, the ConeML pair led mixed arithmetic;
@@ -89,6 +89,14 @@ Alpha exceeded Llama 3.2 and TinyLlama on function writing but trailed
89
  Qwen3.5, Qwen3, and SmolLM2. All models recorded zero over-refusals on
90
  five in-scope contrasts.
91
 
 
 
 
 
 
 
 
 
92
  Qwen3.5 with thinking enabled reached 451/585 (77.1%), compared with
93
  this release's 442/585 (75.6%), while using at least 17.6 times the
94
  generated-token budget per item and approximately 24 times the recorded
 
73
  with greedy decoding, repetition penalty 1.15, and matched short-answer
74
  budgets.
75
 
76
+ | model | parameters | mixed arithmetic (n=585) | executed functions (n=100) | transitive names d1/d3/d5 (n=32 each) | designated refusals (n=17) |
77
+ |---|---:|---:|---:|---:|---:|
78
+ | ConeML Alpha | 0.81B | 421 (72.0%) | 83 | 23/15/15 | **13** |
79
+ | ConeML Arithmetic | 0.81B | **442 (75.6%)** | 35 | 26/17/16 | 11 |
80
+ | Qwen3.5 | 0.8B | 154 (26.3%) | 93 | 23/22/10 | 0 |
81
+ | Qwen3 | 0.6B | 223 (38.1%) | **98** | 18/10/9 | 1 |
82
+ | Llama 3.2 Instruct | 1.24B | 343 (58.6%) | 79 | 16/7/5 | 1 |
83
+ | TinyLlama Chat | 1.1B | 113 (19.3%) | 51 | 8/18/22 | 0 |
84
+ | SmolLM2 Instruct | 1.7B | 380 (65.0%) | 96 | 16/6/5 | 1 |
85
 
86
  This is not a general leaderboard: the task families match ConeML's
87
  trained surfaces. On this screen, the ConeML pair led mixed arithmetic;
 
89
  Qwen3.5, Qwen3, and SmolLM2. All models recorded zero over-refusals on
90
  five in-scope contrasts.
91
 
92
+ On the question-form name-chain depth screen, Arithmetic recorded
93
+ 26/32, 17/32, and 16/32 at depths 1, 3, and 5; Alpha recorded
94
+ 23/32, 15/32, and 15/32. Arithmetic exceeded Qwen3 and Llama 3.2 at all
95
+ three shown depths and Qwen3.5 at depths 1 and 5, while Qwen3.5 led it
96
+ at depth 3 and TinyLlama led the group at depth 5. Entity-chain controls
97
+ were mixed and are reported in the full table. Chance is 1/(depth+1);
98
+ this is an exact-selection surface test, not a claim of general reasoning.
99
+
100
  Qwen3.5 with thinking enabled reached 451/585 (77.1%), compared with
101
  this release's 442/585 (75.6%), while using at least 17.6 times the
102
  generated-token budget per item and approximately 24 times the recorded
SHA256SUMS.txt CHANGED
@@ -1,18 +1,18 @@
1
  846a8e2297580b3c8d57f69a07d88833d370f976010b31d9fc9b93e9a3f04464 LICENSE.md
2
  1756d02d239f8e8342a1aa513248f9d35b9a5168c77d07619e64a90e20580434 Modelfile
3
- 45beb690d5c990dd9660ef6cbb637b3961ef20c804507d336286f2a6cbe5849d README.md
4
  fdf72939738a346a2134cd5ede1e9921d52d4d85ac0449249118322de48447e7 chat_template.jinja
5
  147f479696718b5665f3aea3a7ae312cc6850d998246141f6ecbbd7e395bce2d coneml-810m-alpha-arithmetic-Q8_0.gguf
6
  5898b62c447961cb7eb60fc538fc9f4b0f0bdda840fda87a861d01c245085c31 coneml-810m-alpha-arithmetic-f16.gguf
7
  6cd7853f16b53fdb6bcc44438be39c33726379dc9848798fdcc81ba7f7555a58 config.json
8
  9637d2b8ef44c95226638a31b7dea6fb3b6f817a297a1a25098f8215b0435489 conversion.json
9
  83b8f89b84387d8c2b22bc2b5d8c20a564fea9f79c7f6e6ab4fa9acee38040c5 eval/EVALUATION_METHODOLOGY.md
10
- 3817126000edea9bc2fcb75a783fa73de234590cdfe89e1aea9b569bae809caa eval/PEER_COMPARISON.md
11
- 73d8100a4b78d7693aabb7a22f3b9948774f7a418d9fa162d19d6ff630baffd4 eval/PEER_EVIDENCE_SHA256SUMS.txt
12
  e5722a62e0d7bf0d4d1ced1626cd5bc49381380aeeea671f6f40a4705740e088 eval/PRIVATE_EVIDENCE_SHA256SUMS.txt
13
  f57c01a959c7944ac2e5bd516985115f67409b81445aca8a857c131f7b1f459e eval/certification-Q8_0.json
14
  4388d66b5feeae2147beebc02967d536ec06a7d284ca245b119fa0dcc1b8cf8c eval/focused-summary.json
15
- baff252791bdf43653763a53a57024d0075bb9a1f6b35cf1fcc94a9191cc73df eval/peer-comparison-summary.json
16
  8bfd98eac2f9309c7481115b91ae47b8f01c9e3b44c80b2bb233a5be6947e8cc eval/representative-samples.json
17
  ed8c2a178bf3268b89f325b36538b71aef54251adb53396e259f9534e6264504 eval/summary.json
18
  2e2ddba5318c74a4fcc31cfdfdc6e067564176b6404da484932f95a7fdc9712a generation_config.json
 
1
  846a8e2297580b3c8d57f69a07d88833d370f976010b31d9fc9b93e9a3f04464 LICENSE.md
2
  1756d02d239f8e8342a1aa513248f9d35b9a5168c77d07619e64a90e20580434 Modelfile
3
+ 0c5758a5483a5ea3201a0dc46e6ec75fc52fcd8ea95169bca731e5b8a5633cfe README.md
4
  fdf72939738a346a2134cd5ede1e9921d52d4d85ac0449249118322de48447e7 chat_template.jinja
5
  147f479696718b5665f3aea3a7ae312cc6850d998246141f6ecbbd7e395bce2d coneml-810m-alpha-arithmetic-Q8_0.gguf
6
  5898b62c447961cb7eb60fc538fc9f4b0f0bdda840fda87a861d01c245085c31 coneml-810m-alpha-arithmetic-f16.gguf
7
  6cd7853f16b53fdb6bcc44438be39c33726379dc9848798fdcc81ba7f7555a58 config.json
8
  9637d2b8ef44c95226638a31b7dea6fb3b6f817a297a1a25098f8215b0435489 conversion.json
9
  83b8f89b84387d8c2b22bc2b5d8c20a564fea9f79c7f6e6ab4fa9acee38040c5 eval/EVALUATION_METHODOLOGY.md
10
+ 2e85c693c7fe567910e085ded9efe29420430016eb861dfa001cafa506f05a71 eval/PEER_COMPARISON.md
11
+ 0f134f0b207601ef2d5650bfe4fd56e810e1c6e32938bdde94035ed08df95165 eval/PEER_EVIDENCE_SHA256SUMS.txt
12
  e5722a62e0d7bf0d4d1ced1626cd5bc49381380aeeea671f6f40a4705740e088 eval/PRIVATE_EVIDENCE_SHA256SUMS.txt
13
  f57c01a959c7944ac2e5bd516985115f67409b81445aca8a857c131f7b1f459e eval/certification-Q8_0.json
14
  4388d66b5feeae2147beebc02967d536ec06a7d284ca245b119fa0dcc1b8cf8c eval/focused-summary.json
15
+ ee1528871b0d2e0f9ad5f3c3e41e4cfc365298145bf2d042b7ec27f8521188d1 eval/peer-comparison-summary.json
16
  8bfd98eac2f9309c7481115b91ae47b8f01c9e3b44c80b2bb233a5be6947e8cc eval/representative-samples.json
17
  ed8c2a178bf3268b89f325b36538b71aef54251adb53396e259f9534e6264504 eval/summary.json
18
  2e2ddba5318c74a4fcc31cfdfdc6e067564176b6404da484932f95a7fdc9712a generation_config.json
eval/PEER_COMPARISON.md CHANGED
@@ -21,6 +21,8 @@ model cards.
21
  | mixed arithmetic screen (n=585) | 421 (72.0%) | **442 (75.6%)** | 154 (26.3%) | 223 (38.1%) | 343 (58.6%) | 113 (19.3%) | 380 (65.0%) |
22
  | four core arithmetic lanes (n=225) | 177 (78.7%) | **221 (98.2%)** | 150 (66.7%) | 208 (92.4%) | **221 (98.2%)** | 92 (40.9%) | **221 (98.2%)** |
23
  | executed single functions (n=100) | 83 | 35 | 93 | **98** | 79 | 51 | 96 |
 
 
24
  | designated-refusal prompts (n=17) | **13** | 11 | 0 | 1 | 1 | 0 | 1 |
25
  | over-refusals on contrasts (n=5) | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
26
 
@@ -33,6 +35,15 @@ The honest result is mixed:
33
  full-size internal result is 1,093/1,116 (97.9%).
34
  - Alpha exceeded Llama 3.2 Instruct and TinyLlama Chat on the executed
35
  function screen, but Qwen3.5, Qwen3, and SmolLM2 scored higher.
 
 
 
 
 
 
 
 
 
36
  - The refusal row measures a trained response policy on designated
37
  prompts, not factual correctness or general epistemic calibration.
38
 
 
21
  | mixed arithmetic screen (n=585) | 421 (72.0%) | **442 (75.6%)** | 154 (26.3%) | 223 (38.1%) | 343 (58.6%) | 113 (19.3%) | 380 (65.0%) |
22
  | four core arithmetic lanes (n=225) | 177 (78.7%) | **221 (98.2%)** | 150 (66.7%) | 208 (92.4%) | **221 (98.2%)** | 92 (40.9%) | **221 (98.2%)** |
23
  | executed single functions (n=100) | 83 | 35 | 93 | **98** | 79 | 51 | 96 |
24
+ | transitive names, depth 1/3/5 (n=32 each) | 23/15/15 | 26/17/16 | 23/22/10 | 18/10/9 | 16/7/5 | 8/18/22 | 16/6/5 |
25
+ | transitive entities, depth 1/3/5 (n=32 each) | 15/11/10 | 16/11/12 | 18/16/14 | 14/17/11 | 0/6/7 | 14/15/11 | 22/21/10 |
26
  | designated-refusal prompts (n=17) | **13** | 11 | 0 | 1 | 1 | 0 | 1 |
27
  | over-refusals on contrasts (n=5) | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
28
 
 
35
  full-size internal result is 1,093/1,116 (97.9%).
36
  - Alpha exceeded Llama 3.2 Instruct and TinyLlama Chat on the executed
37
  function screen, but Qwen3.5, Qwen3, and SmolLM2 scored higher.
38
+ - On name-chain transitive selection, Arithmetic recorded 26/32, 17/32,
39
+ and 16/32 at depths 1, 3, and 5. It exceeded Qwen3 and Llama 3.2 at all
40
+ three shown depths and exceeded Qwen3.5 at depths 1 and 5; Qwen3.5 led
41
+ it at depth 3, and TinyLlama led the group at depth 5.
42
+ - Entity-chain transfer was mixed: Alpha recorded 15/11/10 and
43
+ Arithmetic 16/11/12 at depths 1/3/5, while different peers led each
44
+ depth. Chance is 1/(depth+1). The non-monotonic peer rows reinforce
45
+ that this is an exact-selection surface test, not proof of general
46
+ reasoning depth.
47
  - The refusal row measures a trained response policy on designated
48
  prompts, not factual correctness or general epistemic calibration.
49
 
eval/PEER_EVIDENCE_SHA256SUMS.txt CHANGED
@@ -4,8 +4,12 @@ a5d6cd4d7a2643b316cefb1dcb6f43296bb69da4064dbc3b3ab4a123bf36b2b1 base-reference
4
  ca387426c55ac72d50041e2f082bd4cdc747eedb51deca7929cf7a4acea2c5b6 base-reference/coneml-native-transitive.json
5
  43ef584d37c6a276b190aeb741597b54bfebbf7bbd9c37df5d82a2e15a8ee2d1 base-reference/qwen35-0.8b-base.json
6
  25de33180d4073777d29871882b57f1556733e9a8c1e10eba14a8cac062b5c35 base-reference/qwen35-0.8b-base.rows.jsonl
 
 
7
  2003c73a54602f3a30980d6c7f85feef39b63fd11cef9fc20a90044d3569ba77 instruct/coneml-810m-alpha.json
8
  8da107545bd6717c7795ff4e827698e5e269d451f2390f24dccc516fc390ab44 instruct/coneml-810m-alpha.rows.jsonl
 
 
9
  77d514c14e14557c4ac30e8a8c3f4c5887f0db0dff39b8ea14baea164327d027 instruct/coneml-810m-arithmetic.json
10
  a60ed3e533c0c4a77b34a21f80b8df8de85e3df074f3cacdc3bbd4b5e0741cce instruct/coneml-810m-arithmetic.rows.jsonl
11
  aa35b93f46778ffb22351ac411f2fc04e645775e416b5a9423f925bb6d6b1609 instruct/llama32-1b-instruct-transchat.json
 
4
  ca387426c55ac72d50041e2f082bd4cdc747eedb51deca7929cf7a4acea2c5b6 base-reference/coneml-native-transitive.json
5
  43ef584d37c6a276b190aeb741597b54bfebbf7bbd9c37df5d82a2e15a8ee2d1 base-reference/qwen35-0.8b-base.json
6
  25de33180d4073777d29871882b57f1556733e9a8c1e10eba14a8cac062b5c35 base-reference/qwen35-0.8b-base.rows.jsonl
7
+ 81b9effe8a6f808b5bf9dbf639ec88235004c6c270b5275e382316adb07ff9d1 instruct/coneml-810m-alpha-transchat.json
8
+ 4b2c433ecf35c6dd409b9d09e2cb858e1d68e179d189dd9074e8ed681a9be44a instruct/coneml-810m-alpha-transchat.rows.jsonl
9
  2003c73a54602f3a30980d6c7f85feef39b63fd11cef9fc20a90044d3569ba77 instruct/coneml-810m-alpha.json
10
  8da107545bd6717c7795ff4e827698e5e269d451f2390f24dccc516fc390ab44 instruct/coneml-810m-alpha.rows.jsonl
11
+ aea504626bbd9f7d8daf576777fe82fa4c57947a30cd654953889edf8247101f instruct/coneml-810m-arithmetic-transchat.json
12
+ 5b8864c9810d556a8bc2b09a43a2b34f69c0fad6889c8c04b63717c92b6f3c27 instruct/coneml-810m-arithmetic-transchat.rows.jsonl
13
  77d514c14e14557c4ac30e8a8c3f4c5887f0db0dff39b8ea14baea164327d027 instruct/coneml-810m-arithmetic.json
14
  a60ed3e533c0c4a77b34a21f80b8df8de85e3df074f3cacdc3bbd4b5e0741cce instruct/coneml-810m-arithmetic.rows.jsonl
15
  aa35b93f46778ffb22351ac411f2fc04e645775e416b5a9423f925bb6d6b1609 instruct/llama32-1b-instruct-transchat.json
eval/peer-comparison-summary.json CHANGED
@@ -8,6 +8,7 @@
8
  "arithmetic_n": 585,
9
  "core_arithmetic_n": 225,
10
  "function_writing_n": 100,
 
11
  "designated_refusal_n": 17,
12
  "refusal_contrast_n": 5
13
  },
@@ -27,6 +28,7 @@
27
  "arithmetic_mixed": {"correct": 421, "n": 585, "accuracy": 0.7197},
28
  "arithmetic_core_four_lanes": {"correct": 177, "n": 225, "accuracy": 0.7867},
29
  "executed_functions": {"correct": 83, "n": 100, "accuracy": 0.83},
 
30
  "designated_refusals": {"correct": 13, "n": 17, "accuracy": 0.7647},
31
  "over_refusals": {"count": 0, "n": 5}
32
  },
@@ -36,6 +38,7 @@
36
  "arithmetic_mixed": {"correct": 442, "n": 585, "accuracy": 0.7556},
37
  "arithmetic_core_four_lanes": {"correct": 221, "n": 225, "accuracy": 0.9822},
38
  "executed_functions": {"correct": 35, "n": 100, "accuracy": 0.35},
 
39
  "designated_refusals": {"correct": 11, "n": 17, "accuracy": 0.6471},
40
  "over_refusals": {"count": 0, "n": 5}
41
  },
@@ -46,6 +49,7 @@
46
  "arithmetic_mixed": {"correct": 154, "n": 585, "accuracy": 0.2632},
47
  "arithmetic_core_four_lanes": {"correct": 150, "n": 225, "accuracy": 0.6667},
48
  "executed_functions": {"correct": 93, "n": 100, "accuracy": 0.93},
 
49
  "designated_refusals": {"correct": 0, "n": 17, "accuracy": 0.0},
50
  "over_refusals": {"count": 0, "n": 5}
51
  },
@@ -55,6 +59,7 @@
55
  "arithmetic_mixed": {"correct": 223, "n": 585, "accuracy": 0.3812},
56
  "arithmetic_core_four_lanes": {"correct": 208, "n": 225, "accuracy": 0.9244},
57
  "executed_functions": {"correct": 98, "n": 100, "accuracy": 0.98},
 
58
  "designated_refusals": {"correct": 1, "n": 17, "accuracy": 0.0588},
59
  "over_refusals": {"count": 0, "n": 5}
60
  },
@@ -64,6 +69,7 @@
64
  "arithmetic_mixed": {"correct": 343, "n": 585, "accuracy": 0.5863},
65
  "arithmetic_core_four_lanes": {"correct": 221, "n": 225, "accuracy": 0.9822},
66
  "executed_functions": {"correct": 79, "n": 100, "accuracy": 0.79},
 
67
  "designated_refusals": {"correct": 1, "n": 17, "accuracy": 0.0588},
68
  "over_refusals": {"count": 0, "n": 5}
69
  },
@@ -73,6 +79,7 @@
73
  "arithmetic_mixed": {"correct": 113, "n": 585, "accuracy": 0.1932},
74
  "arithmetic_core_four_lanes": {"correct": 92, "n": 225, "accuracy": 0.4089},
75
  "executed_functions": {"correct": 51, "n": 100, "accuracy": 0.51},
 
76
  "designated_refusals": {"correct": 0, "n": 17, "accuracy": 0.0},
77
  "over_refusals": {"count": 0, "n": 5}
78
  },
@@ -82,6 +89,7 @@
82
  "arithmetic_mixed": {"correct": 380, "n": 585, "accuracy": 0.6496},
83
  "arithmetic_core_four_lanes": {"correct": 221, "n": 225, "accuracy": 0.9822},
84
  "executed_functions": {"correct": 96, "n": 100, "accuracy": 0.96},
 
85
  "designated_refusals": {"correct": 1, "n": 17, "accuracy": 0.0588},
86
  "over_refusals": {"count": 0, "n": 5}
87
  }
 
8
  "arithmetic_n": 585,
9
  "core_arithmetic_n": 225,
10
  "function_writing_n": 100,
11
+ "transitive_n_per_family_depth": 32,
12
  "designated_refusal_n": 17,
13
  "refusal_contrast_n": 5
14
  },
 
28
  "arithmetic_mixed": {"correct": 421, "n": 585, "accuracy": 0.7197},
29
  "arithmetic_core_four_lanes": {"correct": 177, "n": 225, "accuracy": 0.7867},
30
  "executed_functions": {"correct": 83, "n": 100, "accuracy": 0.83},
31
+ "transitive_question_form": {"names_d1_d3_d5": [23, 15, 15], "entities_d1_d3_d5": [15, 11, 10], "n_each": 32},
32
  "designated_refusals": {"correct": 13, "n": 17, "accuracy": 0.7647},
33
  "over_refusals": {"count": 0, "n": 5}
34
  },
 
38
  "arithmetic_mixed": {"correct": 442, "n": 585, "accuracy": 0.7556},
39
  "arithmetic_core_four_lanes": {"correct": 221, "n": 225, "accuracy": 0.9822},
40
  "executed_functions": {"correct": 35, "n": 100, "accuracy": 0.35},
41
+ "transitive_question_form": {"names_d1_d3_d5": [26, 17, 16], "entities_d1_d3_d5": [16, 11, 12], "n_each": 32},
42
  "designated_refusals": {"correct": 11, "n": 17, "accuracy": 0.6471},
43
  "over_refusals": {"count": 0, "n": 5}
44
  },
 
49
  "arithmetic_mixed": {"correct": 154, "n": 585, "accuracy": 0.2632},
50
  "arithmetic_core_four_lanes": {"correct": 150, "n": 225, "accuracy": 0.6667},
51
  "executed_functions": {"correct": 93, "n": 100, "accuracy": 0.93},
52
+ "transitive_question_form": {"names_d1_d3_d5": [23, 22, 10], "entities_d1_d3_d5": [18, 16, 14], "n_each": 32},
53
  "designated_refusals": {"correct": 0, "n": 17, "accuracy": 0.0},
54
  "over_refusals": {"count": 0, "n": 5}
55
  },
 
59
  "arithmetic_mixed": {"correct": 223, "n": 585, "accuracy": 0.3812},
60
  "arithmetic_core_four_lanes": {"correct": 208, "n": 225, "accuracy": 0.9244},
61
  "executed_functions": {"correct": 98, "n": 100, "accuracy": 0.98},
62
+ "transitive_question_form": {"names_d1_d3_d5": [18, 10, 9], "entities_d1_d3_d5": [14, 17, 11], "n_each": 32},
63
  "designated_refusals": {"correct": 1, "n": 17, "accuracy": 0.0588},
64
  "over_refusals": {"count": 0, "n": 5}
65
  },
 
69
  "arithmetic_mixed": {"correct": 343, "n": 585, "accuracy": 0.5863},
70
  "arithmetic_core_four_lanes": {"correct": 221, "n": 225, "accuracy": 0.9822},
71
  "executed_functions": {"correct": 79, "n": 100, "accuracy": 0.79},
72
+ "transitive_question_form": {"names_d1_d3_d5": [16, 7, 5], "entities_d1_d3_d5": [0, 6, 7], "n_each": 32},
73
  "designated_refusals": {"correct": 1, "n": 17, "accuracy": 0.0588},
74
  "over_refusals": {"count": 0, "n": 5}
75
  },
 
79
  "arithmetic_mixed": {"correct": 113, "n": 585, "accuracy": 0.1932},
80
  "arithmetic_core_four_lanes": {"correct": 92, "n": 225, "accuracy": 0.4089},
81
  "executed_functions": {"correct": 51, "n": 100, "accuracy": 0.51},
82
+ "transitive_question_form": {"names_d1_d3_d5": [8, 18, 22], "entities_d1_d3_d5": [14, 15, 11], "n_each": 32},
83
  "designated_refusals": {"correct": 0, "n": 17, "accuracy": 0.0},
84
  "over_refusals": {"count": 0, "n": 5}
85
  },
 
89
  "arithmetic_mixed": {"correct": 380, "n": 585, "accuracy": 0.6496},
90
  "arithmetic_core_four_lanes": {"correct": 221, "n": 225, "accuracy": 0.9822},
91
  "executed_functions": {"correct": 96, "n": 100, "accuracy": 0.96},
92
+ "transitive_question_form": {"names_d1_d3_d5": [16, 6, 5], "entities_d1_d3_d5": [22, 21, 10], "n_each": 32},
93
  "designated_refusals": {"correct": 1, "n": 17, "accuracy": 0.0588},
94
  "over_refusals": {"count": 0, "n": 5}
95
  }