pretrain / README.md
wenhuach's picture
Create README.md
044bc6a verified
|
Raw
History Blame Contribute Delete
2.48 kB
Evaluating: hellaswag_zeroshot (0-shot, type: multiple_choice)... accuracy: 0.5511 | centered: 0.4014 | time: 58.63s
Evaluating: jeopardy (10-shot, type: language_modeling)... accuracy: 0.2518 | centered: 0.2518 | time: 12.50s
Evaluating: bigbench_qa_wikidata (10-shot, type: language_modeling)... accuracy: 0.5716 | centered: 0.5716 | time: 119.92s
Evaluating: arc_easy (10-shot, type: multiple_choice)... accuracy: 0.6860 | centered: 0.5814 | time: 14.57s
Evaluating: arc_challenge (10-shot, type: multiple_choice)... accuracy: 0.3985 | centered: 0.1980 | time: 7.18s
Evaluating: copa (0-shot, type: multiple_choice)... accuracy: 0.7000 | centered: 0.4000 | time: 0.58s
Evaluating: commonsense_qa (10-shot, type: multiple_choice)... accuracy: 0.2473 | centered: 0.0592 | time: 7.57s
Evaluating: piqa (10-shot, type: multiple_choice)... accuracy: 0.7323 | centered: 0.4646 | time: 11.02s
Evaluating: openbook_qa (0-shot, type: multiple_choice)... accuracy: 0.3880 | centered: 0.1840 | time: 2.90s
Evaluating: lambada_openai (0-shot, type: language_modeling)... accuracy: 0.4679 | centered: 0.4679 | time: 29.68s
Evaluating: hellaswag (10-shot, type: multiple_choice)... accuracy: 0.5595 | centered: 0.4126 | time: 65.66s
Evaluating: winograd (0-shot, type: schema)... accuracy: 0.7289 | centered: 0.4579 | time: 1.53s
Evaluating: winogrande (0-shot, type: schema)... accuracy: 0.5699 | centered: 0.1397 | time: 7.12s
Evaluating: bigbench_dyck_languages (10-shot, type: language_modeling)... accuracy: 0.1340 | centered: 0.1340 | time: 5.99s
Evaluating: agi_eval_lsat_ar (3-shot, type: multiple_choice)... accuracy: 0.2130 | centered: 0.0163 | time: 1.46s
Evaluating: bigbench_cs_algorithms (10-shot, type: language_modeling)... accuracy: 0.4114 | centered: 0.4114 | time: 7.78s
Evaluating: bigbench_operators (10-shot, type: language_modeling)... accuracy: 0.2000 | centered: 0.2000 | time: 1.23s
Evaluating: bigbench_repeat_copy_logic (10-shot, type: language_modeling)... accuracy: 0.0000 | centered: 0.0000 | time: 0.19s
Evaluating: squad (10-shot, type: language_modeling)... accuracy: 0.2967 | centered: 0.2967 | time: 67.21s
Evaluating: coqa (0-shot, type: language_modeling)... accuracy: 0.3024 | centered: 0.3024 | time: 47.20s
Evaluating: boolq (10-shot, type: multiple_choice)... accuracy: 0.5361 | centered: -0.2208 | time: 21.41s
Evaluating: bigbench_language_identification (10-shot, type: multiple_choice)... accuracy: 0.2500 | centered: 0.1749 | time: 77.11s