Add zero-shot benchmark results (ARC, HellaSwag, SciQ, PIQA) and full-split val ppl

#1
No description provided.

Merging my own benchmark-results PR: adds the zero-shot ARC/HellaSwag/SciQ/PIQA table (all at/below chance, expected for a 24M children's-story model) and clarifies the val-ppl methodology (fixed-window 8.76 vs full-split random-window 12.39).

Compactbot changed pull request status to merged

Sign up or log in to comment