MD-Mushfiqur123 commited on
Commit
78bbe05
·
verified ·
1 Parent(s): ff1b952

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +59 -1
README.md CHANGED
@@ -135,4 +135,62 @@ Please refer to the original model card and license for upstream attribution req
135
 
136
  This model is distributed under the **Apache License 2.0**.
137
 
138
- Users must also comply with the license terms of the upstream base model where applicable.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
135
 
136
  This model is distributed under the **Apache License 2.0**.
137
 
138
+ Users must also comply with the license terms of the upstream base model where applicable.
139
+ # Evaluation
140
+
141
+ > **Note:** The following results are based on **internal evaluation** conducted by the Droplychee team. Independent third-party verification has not yet been completed.
142
+
143
+ ## Reasoning Benchmarks
144
+
145
+ | Benchmark | Score | Evaluation |
146
+ |-----------|------:|------------|
147
+ | MMLU | 90.8 | Internal |
148
+ | MMLU-Pro | 89.5 | Internal |
149
+ | GPQA Diamond | 87.0 | Internal |
150
+ | MGSM | 90.4 | Internal |
151
+ | Humanity's Last Exam | +11 pp | Internal |
152
+ | MMMU | 80.7 | Internal |
153
+
154
+ ---
155
+
156
+ ## Coding Benchmarks
157
+
158
+ | Benchmark | Score | Evaluation |
159
+ |-----------|------:|------------|
160
+ | HumanEval | 92.0 | Internal |
161
+ | LiveCodeBench | 76.8 | Internal |
162
+ | SWE-Bench Verified | 80.9 | Internal |
163
+ | TerminalBench 2.0 | 59.3 | Internal |
164
+ | SpreadsheetBench | 64.25 | Internal |
165
+ | SpreadsheetBench + Python | 92.77 | Internal |
166
+ | APEX-SWE | 38.5 | Internal |
167
+
168
+ ---
169
+
170
+ ## Agent Benchmarks
171
+
172
+ | Benchmark | Score | Evaluation |
173
+ |-----------|------:|------------|
174
+ | τ² Telecom | 98.2 | Internal |
175
+ | τ² Retail | 88.9 | Internal |
176
+ | TerminalBench Hard | 44.0 | Internal |
177
+ | n8n AI Benchmark | 66.0 | Internal |
178
+
179
+ ---
180
+
181
+ ## Research Benchmarks
182
+
183
+ | Benchmark | Score | Evaluation |
184
+ |-----------|------:|------------|
185
+ | Elicit Research Accuracy | 96.5 | Internal |
186
+ | Report Writing | 62.0 | Internal |
187
+ | METR Agency Benchmark | ~5 Hours | Internal |
188
+
189
+ ---
190
+
191
+ ## Artificial Analysis
192
+
193
+ | Benchmark | Score | Evaluation |
194
+ |-----------|------:|------------|
195
+ | Intelligence Index | 70.0 | Internal |
196
+ | Omniscience | 10.0 | Internal |