Hoglet-33 commited on
Commit
37597da
·
verified ·
1 Parent(s): 0b9c345

Fixed the benchmark scores.

Browse files

I fixed the benchmark scores to reflect what the model actually scored I accidentally used the same scores for base Pebble-10M.

Files changed (1) hide show
  1. README.md +6 -6
README.md CHANGED
@@ -55,12 +55,12 @@ Pebble-10M-Chat was evaluated on several commonsense and arithmetic benchmarks.
55
 
56
  | Benchmark | Accuracy | Random Baseline |
57
  |-----------------|----------|-----------------|
58
- | PIQA | 58.43% | 50.00% |
59
- | ARC-Easy | 37.29% | 25.00% |
60
- | ARC-Challenge | 18.60% | 25.00% |
61
- | HellaSwag | 26.81% | 25.00% |
62
- | ArithMark-2.0 | 27.64% | 25.00% |
63
- | ArithMark-3.0 | 32.80% | 25.00% |
64
 
65
  ### Evaluation Notes
66
 
 
55
 
56
  | Benchmark | Accuracy | Random Baseline |
57
  |-----------------|----------|-----------------|
58
+ | PIQA | 51.41% | 50.00% |
59
+ | ARC-Easy | 26.09% | 25.00% |
60
+ | ARC-Challenge | 20.22% | 25.00% |
61
+ | HellaSwag | 25.30% | 25.00% |
62
+ | ArithMark-2.0 | 27.28% | 25.00% |
63
+ | ArithMark-3.0 | 27.50% | 25.00% |
64
 
65
  ### Evaluation Notes
66