YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Merged from all available midtraining checkpoints using mergekit-evolve, optimizing for arc_easy and arc_challenge.
| Tasks | Version | Filter | n-shot | Metric | Value | Stderr | ||
|---|---|---|---|---|---|---|---|---|
| arc_challenge | 1 | none | 0 | acc | ↑ | 0.2056 | ± | 0.0118 |
| none | 0 | acc_norm | ↑ | 0.2415 | ± | 0.0125 | ||
| arc_easy | 1 | none | 0 | acc | ↑ | 0.5063 | ± | 0.0103 |
| none | 0 | acc_norm | ↑ | 0.4848 | ± | 0.0103 | ||
| hellaswag | 1 | none | 0 | acc | ↑ | 0.2901 | ± | 0.0045 |
| none | 0 | acc_norm | ↑ | 0.3173 | ± | 0.0046 | ||
| piqa | 1 | none | 0 | acc | ↑ | 0.6224 | ± | 0.0113 |
| none | 0 | acc_norm | ↑ | 0.6148 | ± | 0.0114 |
BananaMind Base Bench 1.1
Overall Elo: 994
Accuracy: 179/350 (51.14%)
Weighted accuracy: 47.20%
language_completion: Elo 1278 | 46/50 (92.00%) | weighted 90.45%
commonsense: Elo 1022 | 32/50 (64.00%) | weighted 59.31%
world_knowledge: Elo 1040 | 32/50 (64.00%) | weighted 61.76%
context_tracking: Elo 824 | 14/50 (28.00%) | weighted 27.39%
quantitative: Elo 828 | 11/50 (22.00%) | weighted 22.40%
logical_reasoning: Elo 985 | 20/50 (40.00%) | weighted 34.80%
code_completion: Elo 1076 | 24/50 (48.00%) | weighted 47.45%
====================================================================
/root/best_midtrain_merge (149,935,616 params) RESULTS
====================================================================
Raw continuation accuracy 40.00%
Length-normalized accuracy 40.00%
Primary (acc_norm) 40.00%
====================================================================
- Downloads last month
- 26
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support