Make benchmark comparison grug
Browse filesRewrite the shared 27B/35B benchmark section in the model cards' established Grug voice without changing any benchmark values.
README.md
CHANGED
|
@@ -107,12 +107,12 @@ saving and stability. MATH-500 show the trade win on hard unseen ground.
|
|
| 107 |
- known wart fixed in v2.1: think sometimes too thin on hard problems;
|
| 108 |
over-verification in long agent sessions.
|
| 109 |
|
| 110 |
-
## 27b
|
| 111 |
|
| 112 |
-
HumanEval and sanitized MBPP
|
| 113 |
-
|
| 114 |
|
| 115 |
-
|
|
| 116 |
|---|---:|---:|
|
| 117 |
| HumanEval (164) | **87.2** | 80.5 |
|
| 118 |
| MBPP sanitized (100) | 85.0 | **88.0** |
|
|
|
|
| 107 |
- known wart fixed in v2.1: think sometimes too thin on hard problems;
|
| 108 |
over-verification in long agent sessions.
|
| 109 |
|
| 110 |
+
## 27b and 35b hunt same prey
|
| 111 |
|
| 112 |
+
both grug hunt HumanEval and sanitized MBPP. number show pass@1 percent.
|
| 113 |
+
bold grug win that hunt.
|
| 114 |
|
| 115 |
+
| hunt | [grug-27b](https://huggingface.co/ProCreations/grug-27b) v2.1 | [grug-35b](https://huggingface.co/ProCreations/grug-35b) rebuilt |
|
| 116 |
|---|---:|---:|
|
| 117 |
| HumanEval (164) | **87.2** | 80.5 |
|
| 118 |
| MBPP sanitized (100) | 85.0 | **88.0** |
|