Add CHI-Bench eval results — agent harness: Hermes

#46
Files changed (1) hide show
  1. .eval_results/chi-bench.yaml +40 -0
.eval_results/chi-bench.yaml ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Place at .eval_results/chi-bench.yaml in the moonshotai/Kimi-K2.6 model repo.
2
+ # Submit via the model's Community tab as a PR; shows "community-provided" until merged.
3
+ # Values are pass@1 (%) for the best-performing harness for this model: Hermes.
4
+ # (OpenAI Agents SDK is close at 15.1 overall.)
5
+ - dataset:
6
+ id: actava/chi-bench
7
+ task_id: chi_bench
8
+ value: 15.6
9
+ date: "2026-05-08"
10
+ source:
11
+ url: https://arxiv.org/abs/2605.16679
12
+ name: CHI-Bench
13
+ notes: "Harness: Hermes; Protocol: 75 tasks x 3 trials; Metric: pass@1 (%)"
14
+ - dataset:
15
+ id: actava/chi-bench
16
+ task_id: prior_authorization
17
+ value: 18.7
18
+ date: "2026-05-08"
19
+ source:
20
+ url: https://arxiv.org/abs/2605.16679
21
+ name: CHI-Bench
22
+ notes: "Harness: Hermes; Protocol: 75 tasks x 3 trials; Metric: pass@1 (%)"
23
+ - dataset:
24
+ id: actava/chi-bench
25
+ task_id: utilization_management
26
+ value: 21.3
27
+ date: "2026-05-08"
28
+ source:
29
+ url: https://arxiv.org/abs/2605.16679
30
+ name: CHI-Bench
31
+ notes: "Harness: Hermes; Protocol: 75 tasks x 3 trials; Metric: pass@1 (%)"
32
+ - dataset:
33
+ id: actava/chi-bench
34
+ task_id: care_management
35
+ value: 6.7
36
+ date: "2026-05-08"
37
+ source:
38
+ url: https://arxiv.org/abs/2605.16679
39
+ name: CHI-Bench
40
+ notes: "Harness: Hermes; Protocol: 75 tasks x 3 trials; Metric: pass@1 (%)"