Add verified .eval_results/results.yaml for official leaderboards

#2
by NitrAI - opened

YAML Metadata Error:Invalid content in Eval Result file .eval_results/results.yaml

Check out the documentation for more information.

Show details
Task ID "hle" does not match any task in dataset "cais/hle". Available: none
Files changed (1) hide show
  1. .eval_results/results.yaml +69 -0
.eval_results/results.yaml ADDED
@@ -0,0 +1,69 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ - dataset:
2
+ id: cais/hle
3
+ task_id: hle
4
+ value: 38.40
5
+ date: "2026-08-26"
6
+ source:
7
+ url: https://huggingface.co/nitrai-research/Moderato-V1-Pro
8
+ name: "Moderato-V1-Pro Technical Report"
9
+ notes: "no-tools"
10
+
11
+ - dataset:
12
+ id: Idavidrein/gpqa
13
+ task_id: diamond
14
+ value: 90.00
15
+ date: "2026-08-26"
16
+ source:
17
+ url: https://huggingface.co/nitrai-research/Moderato-V1-Pro
18
+ name: "Moderato-V1-Pro Technical Report"
19
+ notes: "Diamond split, 0-shot CoT"
20
+
21
+ - dataset:
22
+ id: ScaleAI/SWE-bench_Pro
23
+ task_id: SWE_Bench_Pro
24
+ value: 63.30
25
+ date: "2026-08-26"
26
+ source:
27
+ url: https://huggingface.co/nitrai-research/Moderato-V1-Pro
28
+ name: "Moderato-V1-Pro Technical Report"
29
+ notes: "Pass@1"
30
+
31
+ - dataset:
32
+ id: datacurve/deep-swe
33
+ task_id: deep_swe
34
+ value: 53.20
35
+ date: "2026-08-26"
36
+ source:
37
+ url: https://huggingface.co/nitrai-research/Moderato-V1-Pro
38
+ name: "Moderato-V1-Pro Technical Report"
39
+ notes: "Pass@1 (v1.1)"
40
+
41
+ - dataset:
42
+ id: harborframework/terminal-bench-2.1
43
+ task_id: terminalbench_2_1
44
+ value: 79.50
45
+ date: "2026-08-26"
46
+ source:
47
+ url: https://huggingface.co/nitrai-research/Moderato-V1-Pro
48
+ name: "Moderato-V1-Pro Technical Report"
49
+ notes: "Pass@1 (Terminus)"
50
+
51
+ - dataset:
52
+ id: internlm/WildClawBench
53
+ task_id: overall
54
+ value: 52.20
55
+ date: "2026-08-26"
56
+ source:
57
+ url: https://huggingface.co/nitrai-research/Moderato-V1-Pro
58
+ name: "Moderato-V1-Pro Technical Report"
59
+ notes: "Pass@1 (Overall)"
60
+
61
+ - dataset:
62
+ id: llamaindex/ExtractBench
63
+ task_id: mean
64
+ value: 88.65
65
+ date: "2026-08-26"
66
+ source:
67
+ url: https://huggingface.co/nitrai-research/Moderato-V1-Pro
68
+ name: "Moderato-V1-Pro Technical Report"
69
+ notes: "Mean score"