Yixin Song commited on
Update README.md
Browse files
README.md
CHANGED
|
@@ -10,7 +10,14 @@ SmallThinker brings powerful, private, and low-latency AI directly to your perso
|
|
| 10 |
without relying on the cloud.
|
| 11 |
|
| 12 |
## Performance
|
| 13 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 14 |
|
| 15 |
For the MMLU evaluation, we use a 0-shot CoT setting.
|
| 16 |
|
|
|
|
| 10 |
without relying on the cloud.
|
| 11 |
|
| 12 |
## Performance
|
| 13 |
+
| Model | MMLU | GPQA-diamond | MATH-500 | IFEVAL | LIVEBENCH | HUMANEVAL | Average |
|
| 14 |
+
|------------------------------|-------|--------------|----------|--------|-----------|-----------|---------|
|
| 15 |
+
| SmallThinker-21BA3B-Instruct | 84.43 | 55.05 | 82.4 | 85.77 | 60.3 | 89.63 | 76.26 |
|
| 16 |
+
| Gemma3-12b-it | 78.52 | 34.85 | 82.4 | 74.68 | 44.5 | 82.93 | 66.31 |
|
| 17 |
+
| Qwen3-14B | 84.82 | 50 | 84.6 | 85.21 | 59.5 | 88.41 | 75.42 |
|
| 18 |
+
| Qwen3-30BA3B | 85.1 | 44.4 | 84.4 | 84.29 | 58.8 | 90.24 | 74.54 |
|
| 19 |
+
| Qwen3-8B | 81.79 | 38.89 | 81.6 | 83.92 | 49.5 | 85.9 | 70.26 |
|
| 20 |
+
| Phi-4-14B | 84.58 | 55.45 | 80.2 | 63.22 | 42.4 | 87.2 | 68.84 |
|
| 21 |
|
| 22 |
For the MMLU evaluation, we use a 0-shot CoT setting.
|
| 23 |
|