Medyassino commited on
Commit
f4162f7
·
verified ·
1 Parent(s): 7061f9e

Add files using upload-large-folder tool

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. hf_cache/hub/CACHEDIR.TAG +4 -0
  2. hf_cache/hub/datasets--EdinburghNLP--xsum/.no_exist/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/.huggingface.yaml +0 -0
  3. hf_cache/hub/datasets--EdinburghNLP--xsum/.no_exist/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/dataset_infos.json +0 -0
  4. hf_cache/hub/datasets--EdinburghNLP--xsum/.no_exist/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/xsum.py +0 -0
  5. hf_cache/hub/datasets--EdinburghNLP--xsum/blobs/fee00d9f711981e883d7a06af05d4ee18b7fe5d9 +222 -0
  6. hf_cache/hub/datasets--EdinburghNLP--xsum/refs/main +1 -0
  7. hf_cache/hub/datasets--EdinburghNLP--xsum/snapshots/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/README.md +222 -0
  8. hf_cache/hub/datasets--HuggingFaceFW--fineweb-2/blobs/67fa474443e4c2a3a8f1b9f19502898a1d86ef29 +0 -0
  9. hf_cache/hub/datasets--HuggingFaceFW--fineweb-2/refs/main +1 -0
  10. hf_cache/hub/datasets--HuggingFaceFW--fineweb-2/snapshots/af9c13333eb981300149d5ca60a8e9d659b276b9/README.md +0 -0
  11. hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/.no_exist/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/.huggingface.yaml +0 -0
  12. hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/.no_exist/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/dataset_infos.json +0 -0
  13. hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/.no_exist/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/fineweb-edu.py +0 -0
  14. hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/blobs/16e64d907529c7f742578a2fee4b72039d64211d +651 -0
  15. hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/refs/main +1 -0
  16. hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/snapshots/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/README.md +651 -0
  17. hf_cache/hub/datasets--Rowan--hellaswag/.no_exist/218ec52e09a7e7462a5400043bb9a69a41d06b76/.huggingface.yaml +0 -0
  18. hf_cache/hub/datasets--Rowan--hellaswag/.no_exist/218ec52e09a7e7462a5400043bb9a69a41d06b76/dataset_infos.json +0 -0
  19. hf_cache/hub/datasets--Rowan--hellaswag/.no_exist/218ec52e09a7e7462a5400043bb9a69a41d06b76/hellaswag.py +0 -0
  20. hf_cache/hub/datasets--Rowan--hellaswag/blobs/29f11d90eb3a5b319cfe8ce2a4e78d9f2a1aea3f +218 -0
  21. hf_cache/hub/datasets--Rowan--hellaswag/refs/main +1 -0
  22. hf_cache/hub/datasets--Rowan--hellaswag/snapshots/218ec52e09a7e7462a5400043bb9a69a41d06b76/README.md +218 -0
  23. hf_cache/hub/datasets--Salesforce--wikitext/.no_exist/b08601e04326c79dfdd32d625aee71d232d685c3/.huggingface.yaml +0 -0
  24. hf_cache/hub/datasets--Salesforce--wikitext/.no_exist/b08601e04326c79dfdd32d625aee71d232d685c3/dataset_infos.json +0 -0
  25. hf_cache/hub/datasets--Salesforce--wikitext/.no_exist/b08601e04326c79dfdd32d625aee71d232d685c3/wikitext.py +0 -0
  26. hf_cache/hub/datasets--Salesforce--wikitext/blobs/2a4fec2bc8df76c9d4da1c8e8865b625eb221c76 +344 -0
  27. hf_cache/hub/datasets--Salesforce--wikitext/refs/main +1 -0
  28. hf_cache/hub/datasets--Salesforce--wikitext/snapshots/b08601e04326c79dfdd32d625aee71d232d685c3/README.md +344 -0
  29. hf_cache/hub/datasets--abisee--cnn_dailymail/.no_exist/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/.huggingface.yaml +0 -0
  30. hf_cache/hub/datasets--abisee--cnn_dailymail/.no_exist/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/cnn_dailymail.py +0 -0
  31. hf_cache/hub/datasets--abisee--cnn_dailymail/.no_exist/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/dataset_infos.json +0 -0
  32. hf_cache/hub/datasets--abisee--cnn_dailymail/blobs/feadf7d245b4f6818e20b4cf65841d03b5703d47 +305 -0
  33. hf_cache/hub/datasets--abisee--cnn_dailymail/refs/main +1 -0
  34. hf_cache/hub/datasets--abisee--cnn_dailymail/snapshots/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/README.md +305 -0
  35. hf_cache/hub/datasets--allenai--openbookqa/.no_exist/388097ea7776314e93a529163e0fea805b8a6454/.huggingface.yaml +0 -0
  36. hf_cache/hub/datasets--allenai--openbookqa/.no_exist/388097ea7776314e93a529163e0fea805b8a6454/dataset_infos.json +0 -0
  37. hf_cache/hub/datasets--allenai--openbookqa/.no_exist/388097ea7776314e93a529163e0fea805b8a6454/openbookqa.py +0 -0
  38. hf_cache/hub/datasets--allenai--openbookqa/blobs/08128898cc0433b97f7a9c9ff09c5054c8587e3e +301 -0
  39. hf_cache/hub/datasets--allenai--openbookqa/refs/main +1 -0
  40. hf_cache/hub/datasets--allenai--openbookqa/snapshots/388097ea7776314e93a529163e0fea805b8a6454/README.md +301 -0
  41. hf_cache/hub/datasets--allenai--sciq/.no_exist/2c94ad3e1aafab77146f384e23536f97a4849815/.huggingface.yaml +0 -0
  42. hf_cache/hub/datasets--allenai--sciq/.no_exist/2c94ad3e1aafab77146f384e23536f97a4849815/dataset_infos.json +0 -0
  43. hf_cache/hub/datasets--allenai--sciq/.no_exist/2c94ad3e1aafab77146f384e23536f97a4849815/sciq.py +0 -0
  44. hf_cache/hub/datasets--allenai--sciq/blobs/c644057869cabcde87a2b5ab9665ec0d0bd1405b +216 -0
  45. hf_cache/hub/datasets--allenai--sciq/refs/main +1 -0
  46. hf_cache/hub/datasets--allenai--sciq/snapshots/2c94ad3e1aafab77146f384e23536f97a4849815/README.md +216 -0
  47. hf_cache/hub/datasets--cais--mmlu/.no_exist/c30699e8356da336a370243923dbaf21066bb9fe/.huggingface.yaml +0 -0
  48. hf_cache/hub/datasets--cais--mmlu/.no_exist/c30699e8356da336a370243923dbaf21066bb9fe/mmlu.py +0 -0
  49. hf_cache/hub/datasets--cais--mmlu/blobs/08de94c560ad7420252bff6e4729f1d1683def4f +2299 -0
  50. hf_cache/hub/datasets--cais--mmlu/blobs/e133c92a3269e646dfd034d5c94ad41a31c7b194 +0 -0
hf_cache/hub/CACHEDIR.TAG ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ Signature: 8a477f597d28d172789f06886806bc55
2
+ # This file is a cache directory tag created by huggingface_hub.
3
+ # For information about cache directory tags, see:
4
+ # https://bford.info/cachedir/
hf_cache/hub/datasets--EdinburghNLP--xsum/.no_exist/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/.huggingface.yaml ADDED
File without changes
hf_cache/hub/datasets--EdinburghNLP--xsum/.no_exist/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/dataset_infos.json ADDED
File without changes
hf_cache/hub/datasets--EdinburghNLP--xsum/.no_exist/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/xsum.py ADDED
File without changes
hf_cache/hub/datasets--EdinburghNLP--xsum/blobs/fee00d9f711981e883d7a06af05d4ee18b7fe5d9 ADDED
@@ -0,0 +1,222 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ annotations_creators:
3
+ - found
4
+ language_creators:
5
+ - found
6
+ language:
7
+ - en
8
+ license:
9
+ - unknown
10
+ multilinguality:
11
+ - monolingual
12
+ pretty_name: Extreme Summarization (XSum)
13
+ paperswithcode_id: xsum
14
+ size_categories:
15
+ - 100K<n<1M
16
+ source_datasets:
17
+ - original
18
+ task_categories:
19
+ - summarization
20
+ task_ids:
21
+ - news-articles-summarization
22
+ dataset_info:
23
+ features:
24
+ - name: document
25
+ dtype: string
26
+ - name: summary
27
+ dtype: string
28
+ - name: id
29
+ dtype: string
30
+ splits:
31
+ - name: train
32
+ num_bytes: 479206363
33
+ num_examples: 204045
34
+ - name: validation
35
+ num_bytes: 26292877
36
+ num_examples: 11332
37
+ - name: test
38
+ num_bytes: 26756141
39
+ num_examples: 11334
40
+ download_size: 332791351
41
+ dataset_size: 532255381
42
+ configs:
43
+ - config_name: default
44
+ data_files:
45
+ - split: train
46
+ path: data/train-*
47
+ - split: validation
48
+ path: data/validation-*
49
+ - split: test
50
+ path: data/test-*
51
+ train-eval-index:
52
+ - config: default
53
+ task: summarization
54
+ task_id: summarization
55
+ splits:
56
+ train_split: train
57
+ eval_split: test
58
+ col_mapping:
59
+ document: text
60
+ summary: target
61
+ metrics:
62
+ - type: rouge
63
+ name: Rouge
64
+ ---
65
+
66
+ # Dataset Card for "xsum"
67
+
68
+ ## Table of Contents
69
+ - [Dataset Description](#dataset-description)
70
+ - [Dataset Summary](#dataset-summary)
71
+ - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
72
+ - [Languages](#languages)
73
+ - [Dataset Structure](#dataset-structure)
74
+ - [Data Instances](#data-instances)
75
+ - [Data Fields](#data-fields)
76
+ - [Data Splits](#data-splits)
77
+ - [Dataset Creation](#dataset-creation)
78
+ - [Curation Rationale](#curation-rationale)
79
+ - [Source Data](#source-data)
80
+ - [Annotations](#annotations)
81
+ - [Personal and Sensitive Information](#personal-and-sensitive-information)
82
+ - [Considerations for Using the Data](#considerations-for-using-the-data)
83
+ - [Social Impact of Dataset](#social-impact-of-dataset)
84
+ - [Discussion of Biases](#discussion-of-biases)
85
+ - [Other Known Limitations](#other-known-limitations)
86
+ - [Additional Information](#additional-information)
87
+ - [Dataset Curators](#dataset-curators)
88
+ - [Licensing Information](#licensing-information)
89
+ - [Citation Information](#citation-information)
90
+ - [Contributions](#contributions)
91
+
92
+ ## Dataset Description
93
+
94
+ - **Homepage:**
95
+ - **Repository:** https://github.com/EdinburghNLP/XSum
96
+ - **Paper:** [Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization](https://arxiv.org/abs/1808.08745)
97
+ - **Point of Contact:** [Shashi Narayan](mailto:shashi.narayan@ed.ac.uk)
98
+ - **Size of downloaded dataset files:** 257.30 MB
99
+ - **Size of the generated dataset:** 532.26 MB
100
+ - **Total amount of disk used:** 789.56 MB
101
+
102
+ ### Dataset Summary
103
+
104
+ Extreme Summarization (XSum) Dataset.
105
+
106
+ There are three features:
107
+ - document: Input news article.
108
+ - summary: One sentence summary of the article.
109
+ - id: BBC ID of the article.
110
+
111
+ ### Supported Tasks and Leaderboards
112
+
113
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
114
+
115
+ ### Languages
116
+
117
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
118
+
119
+ ## Dataset Structure
120
+
121
+ ### Data Instances
122
+
123
+ #### default
124
+
125
+ - **Size of downloaded dataset files:** 257.30 MB
126
+ - **Size of the generated dataset:** 532.26 MB
127
+ - **Total amount of disk used:** 789.56 MB
128
+
129
+ An example of 'validation' looks as follows.
130
+ ```
131
+ {
132
+ "document": "some-body",
133
+ "id": "29750031",
134
+ "summary": "some-sentence"
135
+ }
136
+ ```
137
+
138
+ ### Data Fields
139
+
140
+ The data fields are the same among all splits.
141
+
142
+ #### default
143
+ - `document`: a `string` feature.
144
+ - `summary`: a `string` feature.
145
+ - `id`: a `string` feature.
146
+
147
+ ### Data Splits
148
+
149
+ | name |train |validation|test |
150
+ |-------|-----:|---------:|----:|
151
+ |default|204045| 11332|11334|
152
+
153
+ ## Dataset Creation
154
+
155
+ ### Curation Rationale
156
+
157
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
158
+
159
+ ### Source Data
160
+
161
+ #### Initial Data Collection and Normalization
162
+
163
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
164
+
165
+ #### Who are the source language producers?
166
+
167
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
168
+
169
+ ### Annotations
170
+
171
+ #### Annotation process
172
+
173
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
174
+
175
+ #### Who are the annotators?
176
+
177
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
178
+
179
+ ### Personal and Sensitive Information
180
+
181
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
182
+
183
+ ## Considerations for Using the Data
184
+
185
+ ### Social Impact of Dataset
186
+
187
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
188
+
189
+ ### Discussion of Biases
190
+
191
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
192
+
193
+ ### Other Known Limitations
194
+
195
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
196
+
197
+ ## Additional Information
198
+
199
+ ### Dataset Curators
200
+
201
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
202
+
203
+ ### Licensing Information
204
+
205
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
206
+
207
+ ### Citation Information
208
+
209
+ ```
210
+ @article{Narayan2018DontGM,
211
+ title={Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization},
212
+ author={Shashi Narayan and Shay B. Cohen and Mirella Lapata},
213
+ journal={ArXiv},
214
+ year={2018},
215
+ volume={abs/1808.08745}
216
+ }
217
+ ```
218
+
219
+
220
+ ### Contributions
221
+
222
+ Thanks to [@thomwolf](https://github.com/thomwolf), [@lewtun](https://github.com/lewtun), [@mariamabarham](https://github.com/mariamabarham), [@jbragg](https://github.com/jbragg), [@lhoestq](https://github.com/lhoestq), [@patrickvonplaten](https://github.com/patrickvonplaten) for adding this dataset.
hf_cache/hub/datasets--EdinburghNLP--xsum/refs/main ADDED
@@ -0,0 +1 @@
 
 
1
+ 7d4d486c2f8ef850b1a11aead99b894ff3dd7da9
hf_cache/hub/datasets--EdinburghNLP--xsum/snapshots/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/README.md ADDED
@@ -0,0 +1,222 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ annotations_creators:
3
+ - found
4
+ language_creators:
5
+ - found
6
+ language:
7
+ - en
8
+ license:
9
+ - unknown
10
+ multilinguality:
11
+ - monolingual
12
+ pretty_name: Extreme Summarization (XSum)
13
+ paperswithcode_id: xsum
14
+ size_categories:
15
+ - 100K<n<1M
16
+ source_datasets:
17
+ - original
18
+ task_categories:
19
+ - summarization
20
+ task_ids:
21
+ - news-articles-summarization
22
+ dataset_info:
23
+ features:
24
+ - name: document
25
+ dtype: string
26
+ - name: summary
27
+ dtype: string
28
+ - name: id
29
+ dtype: string
30
+ splits:
31
+ - name: train
32
+ num_bytes: 479206363
33
+ num_examples: 204045
34
+ - name: validation
35
+ num_bytes: 26292877
36
+ num_examples: 11332
37
+ - name: test
38
+ num_bytes: 26756141
39
+ num_examples: 11334
40
+ download_size: 332791351
41
+ dataset_size: 532255381
42
+ configs:
43
+ - config_name: default
44
+ data_files:
45
+ - split: train
46
+ path: data/train-*
47
+ - split: validation
48
+ path: data/validation-*
49
+ - split: test
50
+ path: data/test-*
51
+ train-eval-index:
52
+ - config: default
53
+ task: summarization
54
+ task_id: summarization
55
+ splits:
56
+ train_split: train
57
+ eval_split: test
58
+ col_mapping:
59
+ document: text
60
+ summary: target
61
+ metrics:
62
+ - type: rouge
63
+ name: Rouge
64
+ ---
65
+
66
+ # Dataset Card for "xsum"
67
+
68
+ ## Table of Contents
69
+ - [Dataset Description](#dataset-description)
70
+ - [Dataset Summary](#dataset-summary)
71
+ - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
72
+ - [Languages](#languages)
73
+ - [Dataset Structure](#dataset-structure)
74
+ - [Data Instances](#data-instances)
75
+ - [Data Fields](#data-fields)
76
+ - [Data Splits](#data-splits)
77
+ - [Dataset Creation](#dataset-creation)
78
+ - [Curation Rationale](#curation-rationale)
79
+ - [Source Data](#source-data)
80
+ - [Annotations](#annotations)
81
+ - [Personal and Sensitive Information](#personal-and-sensitive-information)
82
+ - [Considerations for Using the Data](#considerations-for-using-the-data)
83
+ - [Social Impact of Dataset](#social-impact-of-dataset)
84
+ - [Discussion of Biases](#discussion-of-biases)
85
+ - [Other Known Limitations](#other-known-limitations)
86
+ - [Additional Information](#additional-information)
87
+ - [Dataset Curators](#dataset-curators)
88
+ - [Licensing Information](#licensing-information)
89
+ - [Citation Information](#citation-information)
90
+ - [Contributions](#contributions)
91
+
92
+ ## Dataset Description
93
+
94
+ - **Homepage:**
95
+ - **Repository:** https://github.com/EdinburghNLP/XSum
96
+ - **Paper:** [Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization](https://arxiv.org/abs/1808.08745)
97
+ - **Point of Contact:** [Shashi Narayan](mailto:shashi.narayan@ed.ac.uk)
98
+ - **Size of downloaded dataset files:** 257.30 MB
99
+ - **Size of the generated dataset:** 532.26 MB
100
+ - **Total amount of disk used:** 789.56 MB
101
+
102
+ ### Dataset Summary
103
+
104
+ Extreme Summarization (XSum) Dataset.
105
+
106
+ There are three features:
107
+ - document: Input news article.
108
+ - summary: One sentence summary of the article.
109
+ - id: BBC ID of the article.
110
+
111
+ ### Supported Tasks and Leaderboards
112
+
113
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
114
+
115
+ ### Languages
116
+
117
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
118
+
119
+ ## Dataset Structure
120
+
121
+ ### Data Instances
122
+
123
+ #### default
124
+
125
+ - **Size of downloaded dataset files:** 257.30 MB
126
+ - **Size of the generated dataset:** 532.26 MB
127
+ - **Total amount of disk used:** 789.56 MB
128
+
129
+ An example of 'validation' looks as follows.
130
+ ```
131
+ {
132
+ "document": "some-body",
133
+ "id": "29750031",
134
+ "summary": "some-sentence"
135
+ }
136
+ ```
137
+
138
+ ### Data Fields
139
+
140
+ The data fields are the same among all splits.
141
+
142
+ #### default
143
+ - `document`: a `string` feature.
144
+ - `summary`: a `string` feature.
145
+ - `id`: a `string` feature.
146
+
147
+ ### Data Splits
148
+
149
+ | name |train |validation|test |
150
+ |-------|-----:|---------:|----:|
151
+ |default|204045| 11332|11334|
152
+
153
+ ## Dataset Creation
154
+
155
+ ### Curation Rationale
156
+
157
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
158
+
159
+ ### Source Data
160
+
161
+ #### Initial Data Collection and Normalization
162
+
163
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
164
+
165
+ #### Who are the source language producers?
166
+
167
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
168
+
169
+ ### Annotations
170
+
171
+ #### Annotation process
172
+
173
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
174
+
175
+ #### Who are the annotators?
176
+
177
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
178
+
179
+ ### Personal and Sensitive Information
180
+
181
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
182
+
183
+ ## Considerations for Using the Data
184
+
185
+ ### Social Impact of Dataset
186
+
187
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
188
+
189
+ ### Discussion of Biases
190
+
191
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
192
+
193
+ ### Other Known Limitations
194
+
195
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
196
+
197
+ ## Additional Information
198
+
199
+ ### Dataset Curators
200
+
201
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
202
+
203
+ ### Licensing Information
204
+
205
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
206
+
207
+ ### Citation Information
208
+
209
+ ```
210
+ @article{Narayan2018DontGM,
211
+ title={Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization},
212
+ author={Shashi Narayan and Shay B. Cohen and Mirella Lapata},
213
+ journal={ArXiv},
214
+ year={2018},
215
+ volume={abs/1808.08745}
216
+ }
217
+ ```
218
+
219
+
220
+ ### Contributions
221
+
222
+ Thanks to [@thomwolf](https://github.com/thomwolf), [@lewtun](https://github.com/lewtun), [@mariamabarham](https://github.com/mariamabarham), [@jbragg](https://github.com/jbragg), [@lhoestq](https://github.com/lhoestq), [@patrickvonplaten](https://github.com/patrickvonplaten) for adding this dataset.
hf_cache/hub/datasets--HuggingFaceFW--fineweb-2/blobs/67fa474443e4c2a3a8f1b9f19502898a1d86ef29 ADDED
The diff for this file is too large to render. See raw diff
 
hf_cache/hub/datasets--HuggingFaceFW--fineweb-2/refs/main ADDED
@@ -0,0 +1 @@
 
 
1
+ af9c13333eb981300149d5ca60a8e9d659b276b9
hf_cache/hub/datasets--HuggingFaceFW--fineweb-2/snapshots/af9c13333eb981300149d5ca60a8e9d659b276b9/README.md ADDED
The diff for this file is too large to render. See raw diff
 
hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/.no_exist/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/.huggingface.yaml ADDED
File without changes
hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/.no_exist/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/dataset_infos.json ADDED
File without changes
hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/.no_exist/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/fineweb-edu.py ADDED
File without changes
hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/blobs/16e64d907529c7f742578a2fee4b72039d64211d ADDED
@@ -0,0 +1,651 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: odc-by
3
+ task_categories:
4
+ - text-generation
5
+ language:
6
+ - en
7
+ pretty_name: FineWeb-Edu
8
+ size_categories:
9
+ - n>1T
10
+ configs:
11
+ - config_name: default
12
+ data_files:
13
+ - split: train
14
+ path: data/*/*
15
+ features:
16
+ - name: text
17
+ dtype: string
18
+ - name: id
19
+ dtype: string
20
+ - name: dump
21
+ dtype: string
22
+ - name: url
23
+ dtype: string
24
+ - name: date
25
+ dtype: string
26
+ - name: file_path
27
+ dtype: string
28
+ - name: language
29
+ dtype: string
30
+ - name: language_score
31
+ dtype: float64
32
+ - name: token_count
33
+ dtype: int64
34
+ - name: score
35
+ dtype: float64
36
+ - name: int_score
37
+ dtype: int64
38
+ - config_name: sample-10BT
39
+ data_files:
40
+ - split: train
41
+ path: sample/10BT/*
42
+ - config_name: sample-100BT
43
+ data_files:
44
+ - split: train
45
+ path: sample/100BT/*
46
+ - config_name: sample-350BT
47
+ data_files:
48
+ - split: train
49
+ path: sample/350BT/*
50
+ - config_name: CC-MAIN-2025-05
51
+ data_files:
52
+ - split: train
53
+ path: data/CC-MAIN-2025-05/*
54
+ - config_name: CC-MAIN-2025-08
55
+ data_files:
56
+ - split: train
57
+ path: data/CC-MAIN-2025-08/*
58
+ - config_name: CC-MAIN-2025-13
59
+ data_files:
60
+ - split: train
61
+ path: data/CC-MAIN-2025-13/*
62
+ - config_name: CC-MAIN-2025-18
63
+ data_files:
64
+ - split: train
65
+ path: data/CC-MAIN-2025-18/*
66
+ - config_name: CC-MAIN-2025-21
67
+ data_files:
68
+ - split: train
69
+ path: data/CC-MAIN-2025-21/*
70
+ - config_name: CC-MAIN-2025-26
71
+ data_files:
72
+ - split: train
73
+ path: data/CC-MAIN-2025-26/*
74
+ - config_name: CC-MAIN-2024-51
75
+ data_files:
76
+ - split: train
77
+ path: data/CC-MAIN-2024-51/*
78
+ - config_name: CC-MAIN-2024-46
79
+ data_files:
80
+ - split: train
81
+ path: data/CC-MAIN-2024-46/*
82
+ - config_name: CC-MAIN-2024-42
83
+ data_files:
84
+ - split: train
85
+ path: data/CC-MAIN-2024-42/*
86
+ - config_name: CC-MAIN-2024-38
87
+ data_files:
88
+ - split: train
89
+ path: data/CC-MAIN-2024-38/*
90
+ - config_name: CC-MAIN-2024-33
91
+ data_files:
92
+ - split: train
93
+ path: data/CC-MAIN-2024-33/*
94
+ - config_name: CC-MAIN-2024-30
95
+ data_files:
96
+ - split: train
97
+ path: data/CC-MAIN-2024-30/*
98
+ - config_name: CC-MAIN-2024-26
99
+ data_files:
100
+ - split: train
101
+ path: data/CC-MAIN-2024-26/*
102
+ - config_name: CC-MAIN-2024-22
103
+ data_files:
104
+ - split: train
105
+ path: data/CC-MAIN-2024-22/*
106
+ - config_name: CC-MAIN-2024-18
107
+ data_files:
108
+ - split: train
109
+ path: data/CC-MAIN-2024-18/*
110
+ - config_name: CC-MAIN-2024-10
111
+ data_files:
112
+ - split: train
113
+ path: data/CC-MAIN-2024-10/*
114
+ - config_name: CC-MAIN-2023-50
115
+ data_files:
116
+ - split: train
117
+ path: data/CC-MAIN-2023-50/*
118
+ - config_name: CC-MAIN-2023-40
119
+ data_files:
120
+ - split: train
121
+ path: data/CC-MAIN-2023-40/*
122
+ - config_name: CC-MAIN-2023-23
123
+ data_files:
124
+ - split: train
125
+ path: data/CC-MAIN-2023-23/*
126
+ - config_name: CC-MAIN-2023-14
127
+ data_files:
128
+ - split: train
129
+ path: data/CC-MAIN-2023-14/*
130
+ - config_name: CC-MAIN-2023-06
131
+ data_files:
132
+ - split: train
133
+ path: data/CC-MAIN-2023-06/*
134
+ - config_name: CC-MAIN-2022-49
135
+ data_files:
136
+ - split: train
137
+ path: data/CC-MAIN-2022-49/*
138
+ - config_name: CC-MAIN-2022-40
139
+ data_files:
140
+ - split: train
141
+ path: data/CC-MAIN-2022-40/*
142
+ - config_name: CC-MAIN-2022-33
143
+ data_files:
144
+ - split: train
145
+ path: data/CC-MAIN-2022-33/*
146
+ - config_name: CC-MAIN-2022-27
147
+ data_files:
148
+ - split: train
149
+ path: data/CC-MAIN-2022-27/*
150
+ - config_name: CC-MAIN-2022-21
151
+ data_files:
152
+ - split: train
153
+ path: data/CC-MAIN-2022-21/*
154
+ - config_name: CC-MAIN-2022-05
155
+ data_files:
156
+ - split: train
157
+ path: data/CC-MAIN-2022-05/*
158
+ - config_name: CC-MAIN-2021-49
159
+ data_files:
160
+ - split: train
161
+ path: data/CC-MAIN-2021-49/*
162
+ - config_name: CC-MAIN-2021-43
163
+ data_files:
164
+ - split: train
165
+ path: data/CC-MAIN-2021-43/*
166
+ - config_name: CC-MAIN-2021-39
167
+ data_files:
168
+ - split: train
169
+ path: data/CC-MAIN-2021-39/*
170
+ - config_name: CC-MAIN-2021-31
171
+ data_files:
172
+ - split: train
173
+ path: data/CC-MAIN-2021-31/*
174
+ - config_name: CC-MAIN-2021-25
175
+ data_files:
176
+ - split: train
177
+ path: data/CC-MAIN-2021-25/*
178
+ - config_name: CC-MAIN-2021-21
179
+ data_files:
180
+ - split: train
181
+ path: data/CC-MAIN-2021-21/*
182
+ - config_name: CC-MAIN-2021-17
183
+ data_files:
184
+ - split: train
185
+ path: data/CC-MAIN-2021-17/*
186
+ - config_name: CC-MAIN-2021-10
187
+ data_files:
188
+ - split: train
189
+ path: data/CC-MAIN-2021-10/*
190
+ - config_name: CC-MAIN-2021-04
191
+ data_files:
192
+ - split: train
193
+ path: data/CC-MAIN-2021-04/*
194
+ - config_name: CC-MAIN-2020-50
195
+ data_files:
196
+ - split: train
197
+ path: data/CC-MAIN-2020-50/*
198
+ - config_name: CC-MAIN-2020-45
199
+ data_files:
200
+ - split: train
201
+ path: data/CC-MAIN-2020-45/*
202
+ - config_name: CC-MAIN-2020-40
203
+ data_files:
204
+ - split: train
205
+ path: data/CC-MAIN-2020-40/*
206
+ - config_name: CC-MAIN-2020-34
207
+ data_files:
208
+ - split: train
209
+ path: data/CC-MAIN-2020-34/*
210
+ - config_name: CC-MAIN-2020-29
211
+ data_files:
212
+ - split: train
213
+ path: data/CC-MAIN-2020-29/*
214
+ - config_name: CC-MAIN-2020-24
215
+ data_files:
216
+ - split: train
217
+ path: data/CC-MAIN-2020-24/*
218
+ - config_name: CC-MAIN-2020-16
219
+ data_files:
220
+ - split: train
221
+ path: data/CC-MAIN-2020-16/*
222
+ - config_name: CC-MAIN-2020-10
223
+ data_files:
224
+ - split: train
225
+ path: data/CC-MAIN-2020-10/*
226
+ - config_name: CC-MAIN-2020-05
227
+ data_files:
228
+ - split: train
229
+ path: data/CC-MAIN-2020-05/*
230
+ - config_name: CC-MAIN-2019-51
231
+ data_files:
232
+ - split: train
233
+ path: data/CC-MAIN-2019-51/*
234
+ - config_name: CC-MAIN-2019-47
235
+ data_files:
236
+ - split: train
237
+ path: data/CC-MAIN-2019-47/*
238
+ - config_name: CC-MAIN-2019-43
239
+ data_files:
240
+ - split: train
241
+ path: data/CC-MAIN-2019-43/*
242
+ - config_name: CC-MAIN-2019-39
243
+ data_files:
244
+ - split: train
245
+ path: data/CC-MAIN-2019-39/*
246
+ - config_name: CC-MAIN-2019-35
247
+ data_files:
248
+ - split: train
249
+ path: data/CC-MAIN-2019-35/*
250
+ - config_name: CC-MAIN-2019-30
251
+ data_files:
252
+ - split: train
253
+ path: data/CC-MAIN-2019-30/*
254
+ - config_name: CC-MAIN-2019-26
255
+ data_files:
256
+ - split: train
257
+ path: data/CC-MAIN-2019-26/*
258
+ - config_name: CC-MAIN-2019-22
259
+ data_files:
260
+ - split: train
261
+ path: data/CC-MAIN-2019-22/*
262
+ - config_name: CC-MAIN-2019-18
263
+ data_files:
264
+ - split: train
265
+ path: data/CC-MAIN-2019-18/*
266
+ - config_name: CC-MAIN-2019-13
267
+ data_files:
268
+ - split: train
269
+ path: data/CC-MAIN-2019-13/*
270
+ - config_name: CC-MAIN-2019-09
271
+ data_files:
272
+ - split: train
273
+ path: data/CC-MAIN-2019-09/*
274
+ - config_name: CC-MAIN-2019-04
275
+ data_files:
276
+ - split: train
277
+ path: data/CC-MAIN-2019-04/*
278
+ - config_name: CC-MAIN-2018-51
279
+ data_files:
280
+ - split: train
281
+ path: data/CC-MAIN-2018-51/*
282
+ - config_name: CC-MAIN-2018-47
283
+ data_files:
284
+ - split: train
285
+ path: data/CC-MAIN-2018-47/*
286
+ - config_name: CC-MAIN-2018-43
287
+ data_files:
288
+ - split: train
289
+ path: data/CC-MAIN-2018-43/*
290
+ - config_name: CC-MAIN-2018-39
291
+ data_files:
292
+ - split: train
293
+ path: data/CC-MAIN-2018-39/*
294
+ - config_name: CC-MAIN-2018-34
295
+ data_files:
296
+ - split: train
297
+ path: data/CC-MAIN-2018-34/*
298
+ - config_name: CC-MAIN-2018-30
299
+ data_files:
300
+ - split: train
301
+ path: data/CC-MAIN-2018-30/*
302
+ - config_name: CC-MAIN-2018-26
303
+ data_files:
304
+ - split: train
305
+ path: data/CC-MAIN-2018-26/*
306
+ - config_name: CC-MAIN-2018-22
307
+ data_files:
308
+ - split: train
309
+ path: data/CC-MAIN-2018-22/*
310
+ - config_name: CC-MAIN-2018-17
311
+ data_files:
312
+ - split: train
313
+ path: data/CC-MAIN-2018-17/*
314
+ - config_name: CC-MAIN-2018-13
315
+ data_files:
316
+ - split: train
317
+ path: data/CC-MAIN-2018-13/*
318
+ - config_name: CC-MAIN-2018-09
319
+ data_files:
320
+ - split: train
321
+ path: data/CC-MAIN-2018-09/*
322
+ - config_name: CC-MAIN-2018-05
323
+ data_files:
324
+ - split: train
325
+ path: data/CC-MAIN-2018-05/*
326
+ - config_name: CC-MAIN-2017-51
327
+ data_files:
328
+ - split: train
329
+ path: data/CC-MAIN-2017-51/*
330
+ - config_name: CC-MAIN-2017-47
331
+ data_files:
332
+ - split: train
333
+ path: data/CC-MAIN-2017-47/*
334
+ - config_name: CC-MAIN-2017-43
335
+ data_files:
336
+ - split: train
337
+ path: data/CC-MAIN-2017-43/*
338
+ - config_name: CC-MAIN-2017-39
339
+ data_files:
340
+ - split: train
341
+ path: data/CC-MAIN-2017-39/*
342
+ - config_name: CC-MAIN-2017-34
343
+ data_files:
344
+ - split: train
345
+ path: data/CC-MAIN-2017-34/*
346
+ - config_name: CC-MAIN-2017-30
347
+ data_files:
348
+ - split: train
349
+ path: data/CC-MAIN-2017-30/*
350
+ - config_name: CC-MAIN-2017-26
351
+ data_files:
352
+ - split: train
353
+ path: data/CC-MAIN-2017-26/*
354
+ - config_name: CC-MAIN-2017-22
355
+ data_files:
356
+ - split: train
357
+ path: data/CC-MAIN-2017-22/*
358
+ - config_name: CC-MAIN-2017-17
359
+ data_files:
360
+ - split: train
361
+ path: data/CC-MAIN-2017-17/*
362
+ - config_name: CC-MAIN-2017-13
363
+ data_files:
364
+ - split: train
365
+ path: data/CC-MAIN-2017-13/*
366
+ - config_name: CC-MAIN-2017-09
367
+ data_files:
368
+ - split: train
369
+ path: data/CC-MAIN-2017-09/*
370
+ - config_name: CC-MAIN-2017-04
371
+ data_files:
372
+ - split: train
373
+ path: data/CC-MAIN-2017-04/*
374
+ - config_name: CC-MAIN-2016-50
375
+ data_files:
376
+ - split: train
377
+ path: data/CC-MAIN-2016-50/*
378
+ - config_name: CC-MAIN-2016-44
379
+ data_files:
380
+ - split: train
381
+ path: data/CC-MAIN-2016-44/*
382
+ - config_name: CC-MAIN-2016-40
383
+ data_files:
384
+ - split: train
385
+ path: data/CC-MAIN-2016-40/*
386
+ - config_name: CC-MAIN-2016-36
387
+ data_files:
388
+ - split: train
389
+ path: data/CC-MAIN-2016-36/*
390
+ - config_name: CC-MAIN-2016-30
391
+ data_files:
392
+ - split: train
393
+ path: data/CC-MAIN-2016-30/*
394
+ - config_name: CC-MAIN-2016-26
395
+ data_files:
396
+ - split: train
397
+ path: data/CC-MAIN-2016-26/*
398
+ - config_name: CC-MAIN-2016-22
399
+ data_files:
400
+ - split: train
401
+ path: data/CC-MAIN-2016-22/*
402
+ - config_name: CC-MAIN-2016-18
403
+ data_files:
404
+ - split: train
405
+ path: data/CC-MAIN-2016-18/*
406
+ - config_name: CC-MAIN-2016-07
407
+ data_files:
408
+ - split: train
409
+ path: data/CC-MAIN-2016-07/*
410
+ - config_name: CC-MAIN-2015-48
411
+ data_files:
412
+ - split: train
413
+ path: data/CC-MAIN-2015-48/*
414
+ - config_name: CC-MAIN-2015-40
415
+ data_files:
416
+ - split: train
417
+ path: data/CC-MAIN-2015-40/*
418
+ - config_name: CC-MAIN-2015-35
419
+ data_files:
420
+ - split: train
421
+ path: data/CC-MAIN-2015-35/*
422
+ - config_name: CC-MAIN-2015-32
423
+ data_files:
424
+ - split: train
425
+ path: data/CC-MAIN-2015-32/*
426
+ - config_name: CC-MAIN-2015-27
427
+ data_files:
428
+ - split: train
429
+ path: data/CC-MAIN-2015-27/*
430
+ - config_name: CC-MAIN-2015-22
431
+ data_files:
432
+ - split: train
433
+ path: data/CC-MAIN-2015-22/*
434
+ - config_name: CC-MAIN-2015-18
435
+ data_files:
436
+ - split: train
437
+ path: data/CC-MAIN-2015-18/*
438
+ - config_name: CC-MAIN-2015-14
439
+ data_files:
440
+ - split: train
441
+ path: data/CC-MAIN-2015-14/*
442
+ - config_name: CC-MAIN-2015-11
443
+ data_files:
444
+ - split: train
445
+ path: data/CC-MAIN-2015-11/*
446
+ - config_name: CC-MAIN-2015-06
447
+ data_files:
448
+ - split: train
449
+ path: data/CC-MAIN-2015-06/*
450
+ - config_name: CC-MAIN-2014-52
451
+ data_files:
452
+ - split: train
453
+ path: data/CC-MAIN-2014-52/*
454
+ - config_name: CC-MAIN-2014-49
455
+ data_files:
456
+ - split: train
457
+ path: data/CC-MAIN-2014-49/*
458
+ - config_name: CC-MAIN-2014-42
459
+ data_files:
460
+ - split: train
461
+ path: data/CC-MAIN-2014-42/*
462
+ - config_name: CC-MAIN-2014-41
463
+ data_files:
464
+ - split: train
465
+ path: data/CC-MAIN-2014-41/*
466
+ - config_name: CC-MAIN-2014-35
467
+ data_files:
468
+ - split: train
469
+ path: data/CC-MAIN-2014-35/*
470
+ - config_name: CC-MAIN-2014-23
471
+ data_files:
472
+ - split: train
473
+ path: data/CC-MAIN-2014-23/*
474
+ - config_name: CC-MAIN-2014-15
475
+ data_files:
476
+ - split: train
477
+ path: data/CC-MAIN-2014-15/*
478
+ - config_name: CC-MAIN-2014-10
479
+ data_files:
480
+ - split: train
481
+ path: data/CC-MAIN-2014-10/*
482
+ - config_name: CC-MAIN-2013-48
483
+ data_files:
484
+ - split: train
485
+ path: data/CC-MAIN-2013-48/*
486
+ - config_name: CC-MAIN-2013-20
487
+ data_files:
488
+ - split: train
489
+ path: data/CC-MAIN-2013-20/*
490
+ ---
491
+
492
+ # 📚 FineWeb-Edu
493
+ <center>
494
+ <img src="https://cdn-uploads.huggingface.co/production/uploads/61c141342aac764ce1654e43/wwRnEQydH9qdRtFofIE-A.png" alt="FineWeb-Edu: The finest collection of educational content the web has to offer">
495
+ </center>
496
+
497
+ > 1.3 trillion tokens of the finest educational data the 🌐 web has to offer
498
+
499
+ **Paper:** https://arxiv.org/abs/2406.17557
500
+
501
+ ## What is it?
502
+
503
+ 📚 FineWeb-Edu dataset consists of **1.3T tokens** and **5.4T tokens** ([FineWeb-Edu-score-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu-score-2)) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version.
504
+
505
+ To enhance FineWeb's quality, we developed an [educational quality classifier](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier) using annotations generated by LLama3-70B-Instruct. We then used this classifier to retain only the most educational web pages. FineWeb-Edu outperforms FineWeb on popular benchmarks and shows the power of classifiers trained on synthetic data.
506
+
507
+ The [Dataset Curation](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu#dataset-curation) section details the process for creating the dataset.
508
+
509
+ ![image/png](https://cdn-uploads.huggingface.co/production/uploads/61c141342aac764ce1654e43/QqXOM8h_ZjjhuCv71xmV7.png)
510
+
511
+ You can find a deduplicated version of FineWeb-edu in [SmolLM-Corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus). We find that the deduplication of this dataset doesn't have any impact on model performance in our ablation setup (1.8B trained on 350B tokens).
512
+
513
+ ## What is being released?
514
+
515
+ Along with the dataset, which includes all filtered CommonCrawl dumps since 2013, we also release the educational classifier used for the filtering as well as the code for training it and running inference at: https://github.com/huggingface/cosmopedia/tree/main/classification
516
+
517
+ ## Changelog
518
+ _Previous versions remain available in the branch `version name`._
519
+
520
+ - **v1.4.0 (11-07-2025):** Added 6 new snapshots: `CC-MAIN-2025-05`, `CC-MAIN-2025-08`, `CC-MAIN-2025-13`, `CC-MAIN-2025-18`, `CC-MAIN-2025-21`, and `CC-MAIN-2025-26` (January to June 2025)
521
+ - **v1.3.0 (31-01-2025):** Fixed an issue with some dumps where some documents hadn't been processed: `CC-MAIN-2024-10`, `CC-MAIN-2024-18`, `CC-MAIN-2024-22`, `CC-MAIN-2024-26`, `CC-MAIN-2024-30`, `CC-MAIN-2024-33`, `CC-MAIN-2024-38`, `CC-MAIN-2024-42`, `CC-MAIN-2024-46` -- they now contain more data (~35B additional tokens).
522
+ - **v1.2.0 (03-01-2025):** Added 9 new snapshots: `CC-MAIN-2024-18`, `CC-MAIN-2024-22`, `CC-MAIN-2024-26`, `CC-MAIN-2024-30`, `CC-MAIN-2024-33`, `CC-MAIN-2024-38`, `CC-MAIN-2024-42`, `CC-MAIN-2024-46`, `CC-MAIN-2024-51`, covering April to December 2024.
523
+ - **v1.0.0 (02-06-2024):** Initial version
524
+
525
+
526
+ ## How to load the dataset
527
+ Similarily to FineWeb, You can load the full dataset or a specific crawl/dump. Dumps have the format `CC-MAIN-(year)-(week number)`.
528
+
529
+ ### (Smaller) sample versions
530
+ Along with config `default` (all the data), and the configs for each individual dump, you can also download the following configs:
531
+ - `sample-350BT`: a subset randomly sampled from the whole dataset of around 350B gpt2 tokens
532
+ - `sample-100BT`: a subset randomly sampled from the whole dataset of around 100B gpt2 tokens
533
+ - `sample-10BT`: a subset randomly sampled from the whole dataset of around 10B gpt2 tokens
534
+
535
+ `sample-10BT` was sampled from `sample-100BT` which in turn was sampled from `sample-350BT`.
536
+
537
+ ### Using 🏭 [`datatrove`](https://github.com/huggingface/datatrove/)
538
+
539
+ ```python
540
+ from datatrove.pipeline.readers import ParquetReader
541
+
542
+ # limit determines how many documents will be streamed (remove for all)
543
+ data_reader = ParquetReader("hf://datasets/HuggingFaceFW/fineweb-edu", glob_pattern="data/*/*.parquet", limit=1000)
544
+ # or to fetch a specific dump CC-MAIN-2024-10, eplace "CC-MAIN-2024-10" with "sample/100BT" to use the 100BT sample
545
+ data_reader = ParquetReader("hf://datasets/HuggingFaceFW/fineweb-edu/CC-MAIN-2024-10", limit=1000)
546
+ for document in data_reader():
547
+ # do something with document
548
+ print(document)
549
+
550
+ ###############################
551
+ # OR for a processing pipeline:
552
+ ###############################
553
+
554
+ from datatrove.executor import LocalPipelineExecutor
555
+ from datatrove.pipeline.readers import ParquetReader
556
+ from datatrove.pipeline.filters import LambdaFilter
557
+ from datatrove.pipeline.writers import JsonlWriter
558
+
559
+ pipeline_exec = LocalPipelineExecutor(
560
+ pipeline=[
561
+ # replace "CC-MAIN-2024-10" with "sample/100BT" to use the 100BT sample
562
+ ParquetReader("hf://datasets/HuggingFaceFW/fineweb-edu/CC-MAIN-2024-10", limit=1000),
563
+ LambdaFilter(lambda doc: "hugging" in doc.text),
564
+ JsonlWriter("some-output-path")
565
+ ],
566
+ tasks=10
567
+ )
568
+ pipeline_exec.run()
569
+ ```
570
+
571
+ ### Using `datasets`
572
+
573
+ ```python
574
+ from datasets import load_dataset
575
+ # use name="sample-10BT" to use the 10BT sample
576
+ fw = load_dataset("HuggingFaceFW/fineweb-edu", name="CC-MAIN-2024-10", split="train", streaming=True)
577
+ ```
578
+
579
+ ## Dataset curation
580
+ A new approach has recently emerged for filtering LLM training datasets: using synthetic data to develop classifiers for identifying educational content. This technique was used in the trainings of [LLama3](https://ai.meta.com/blog/meta-llama-3-meta-ai-responsibility/) and [Phi3](https://arxiv.org/abs/2404.14219), but its large-scale impact on web data filtering hasn't been fully explored or published.
581
+
582
+ The highly popular Phi3 models were trained on 3.3 and 4.8 trillion tokens, with the paper stating: “Our training data consists of heavily filtered publicly available web data (according to the 'educational level') from various open internet sources, as well as synthetic LLM-generated data". Similarly, the LLama3 blog post notes: “We found that previous generations of Llama are good at identifying high-quality data, so we used Llama 2 to help build the text-quality classifiers that are powering Llama 3.” However these classifiers and filtered datasets are not publicly available. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by [LLama3-70B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct) to create FineWeb-Edu.
583
+
584
+ ### Annotation
585
+ We used [Llama3-70B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct) to score 500k FineWeb samples for their educational quality on a scale from 0 to 5.
586
+
587
+ We explored various prompts and found that the additive scale by [Yuan et al.](https://arxiv.org/pdf/2401.10020) worked best. To avoid the LLM favoring highly technical pages like arXiv abstracts and submissions, we focused on grade-school and middle-school level knowledge. By setting a threshold of 3 (on a scale of 0 to 5) during the filtering process, we were able to also retain some high-level educational pages. The final prompt can be found [here](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier/blob/main/utils/prompt.txt).
588
+
589
+ We also experimented with different LLMs: Llama3-70B-Instruct, Mixtral-8x-7B-Instruct, and Mixtral-8x22B-Instruct. Llama 3 and Mixtral-8x22B produced similar scores, while Mixtral-8x7B tended to be more generous, not fully adhering to the score scale. Verga et al. suggest using multiple LLMs as juries. We tried averaging the scores from the three models, but this shifted the distribution to the right due to the higher scores from Mixtral-8x7B. Training on a dataset filtered with a classifier using jury annotations performed worse than using a classifier based on Llama3 annotations. We hypothesize that the jury-based approach retains more low-quality samples.
590
+
591
+ ### Classifier training
592
+ We fine-tuned a Bert-like regression model using these annotations, based on [Snowflake-arctic-embed](https://huggingface.co/Snowflake/snowflake-arctic-embed-m). When converted to a binary classification using a score of 3 as a threshold for keeping and removing files, the model achieved an F1 score of 82%. The classification of FineWeb 15T tokens took 6k H100 GPU hours.
593
+
594
+ The classifier is available at: [HuggingFaceFW/fineweb-edu-classifier/](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier/)
595
+
596
+ ### Filtering and results
597
+ **Note**: You can find more details about the ablations and results in the FineWeb [blog post](https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1).
598
+
599
+ We investigated the impact of using different thresholds for the filtering and found that threshold 3 gave the best overall results. Although using a threshold higher than 3 improves performance on knowledge and reasoning intensive benchmarks, it significantly degrades performance on HellaSwag and PIQA.
600
+
601
+ We then built 📚 FineWeb-Edu by filtering out samples with scores lower than 3. This removed 92% of the dataset, leaving us with 1.3T educational tokens. Our ablation demonstrated that this refined dataset surpasses 🍷 FineWeb and all other open web datasets, with remarkable improvements on educational benchmarks such as MMLU, ARC, and OpenBookQA. The plot below compares FineWeb-Edu to other web datasets:
602
+
603
+ ![image/png](https://cdn-uploads.huggingface.co/production/uploads/61c141342aac764ce1654e43/hJlyTgDzZpYuxO9LUm0PF.png)
604
+
605
+ To retain more tokens, we also experimented with a less strict threshold of 2 instead of 3. While being less performant than using threshold 3, it still outperformed FineWeb and it preserved 5.4T tokens. We release these two dataset as [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) and [FineWeb-Edu-score-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu-score-2) along with the [classifier](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier).
606
+
607
+ You will find all the ablation models in [this collection](https://huggingface.co/collections/HuggingFaceFW/ablation-models-662457b0d213e8c14fe47f32). The FineWeb-Edu ablation model (trained on 350B tokens) is available at [https://huggingface.co/HuggingFaceFW/ablation-model-fineweb-edu](https://huggingface.co/HuggingFaceFW/ablation-model-fineweb-edu).
608
+
609
+ ## Considerations for Using the Data
610
+ This section is copied from the parent dataset: [FineWeb](https://huggingface.co/datasets/HuggingFaceFW/fineweb).
611
+
612
+ ### Social Impact of Dataset
613
+
614
+ With the release of this dataset we aim to make model training more accessible to the machine learning community at large.
615
+
616
+ While multiple open-weights models with strong performance have been publicly released in the past, more often than not these releases are not accompanied by the corresponding training dataset. This is unfortunate as the dataset specificities and characteristics have been demonstrated to have a very large impact and role in the performances of the models. As the creation of a high quality training dataset is a fundamental requirement to training an LLM capable of excelling at downstream tasks, with 🍷 FineWeb we (a) not only make the dataset creation process more transparent, by sharing our entire processing setup including the codebase used, we also (b) help alleviate the costs of dataset curation, both in time and in compute, for model creators by publicly releasing our dataset with the community.
617
+
618
+ ### Discussion of Biases
619
+
620
+ Efforts were made to minimize the amount of NSFW and toxic content present in the dataset by employing filtering on the URL level. However, there are still a significant number of documents present in the final dataset that could be considered toxic or contain harmful content. As 🍷 FineWeb was sourced from the web as a whole, any harmful biases typically present in it may be reproduced on our dataset.
621
+
622
+ We deliberately avoided using machine learning filtering methods that define text quality based on the similarity to a “gold” source such as wikipedia or toxicity classifiers as these methods have been known to [disproportionately remove content in specific dialects](https://aclanthology.org/D16-1120/) and [overclassify as toxic text related to specific social identities](https://arxiv.org/pdf/2109.07445.pdf), respectively.
623
+
624
+ ### Other Known Limitations
625
+
626
+ As a consequence of some of the filtering steps applied, it is likely that code content is not prevalent in our dataset. If you are training a model that should also perform code tasks, we recommend you use 🍷 FineWeb with a code dataset, such as [The Stack v2](https://huggingface.co/datasets/bigcode/the-stack-v2). You should also probably consider complementing 🍷 FineWeb with specialized curated sources (such as Wikipedia, for example) as they will likely have better formatting than the wikipedia content included in 🍷 FineWeb (we did not tailor the processing to individual websites).
627
+
628
+ ## Additional Information
629
+
630
+ ### Licensing Information
631
+
632
+ The dataset is released under the **Open Data Commons Attribution License (ODC-By) v1.0** [license](https://opendatacommons.org/licenses/by/1-0/). The use of this dataset is also subject to [CommonCrawl's Terms of Use](https://commoncrawl.org/terms-of-use).
633
+
634
+ ### Future work
635
+
636
+ We plan to work on better educational classifier to improve the quality of FineWeb-Edu.
637
+
638
+ ### Citation Information
639
+
640
+ You can cite our paper https://arxiv.org/abs/2406.17557 or this dataset:
641
+
642
+ ```
643
+ @misc{lozhkov2024fineweb-edu,
644
+ author = { Lozhkov, Anton and Ben Allal, Loubna and von Werra, Leandro and Wolf, Thomas },
645
+ title = { FineWeb-Edu: the Finest Collection of Educational Content },
646
+ year = 2024,
647
+ url = { https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu },
648
+ doi = { 10.57967/hf/2497 },
649
+ publisher = { Hugging Face }
650
+ }
651
+ ```
hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/refs/main ADDED
@@ -0,0 +1 @@
 
 
1
+ 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9
hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/snapshots/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/README.md ADDED
@@ -0,0 +1,651 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: odc-by
3
+ task_categories:
4
+ - text-generation
5
+ language:
6
+ - en
7
+ pretty_name: FineWeb-Edu
8
+ size_categories:
9
+ - n>1T
10
+ configs:
11
+ - config_name: default
12
+ data_files:
13
+ - split: train
14
+ path: data/*/*
15
+ features:
16
+ - name: text
17
+ dtype: string
18
+ - name: id
19
+ dtype: string
20
+ - name: dump
21
+ dtype: string
22
+ - name: url
23
+ dtype: string
24
+ - name: date
25
+ dtype: string
26
+ - name: file_path
27
+ dtype: string
28
+ - name: language
29
+ dtype: string
30
+ - name: language_score
31
+ dtype: float64
32
+ - name: token_count
33
+ dtype: int64
34
+ - name: score
35
+ dtype: float64
36
+ - name: int_score
37
+ dtype: int64
38
+ - config_name: sample-10BT
39
+ data_files:
40
+ - split: train
41
+ path: sample/10BT/*
42
+ - config_name: sample-100BT
43
+ data_files:
44
+ - split: train
45
+ path: sample/100BT/*
46
+ - config_name: sample-350BT
47
+ data_files:
48
+ - split: train
49
+ path: sample/350BT/*
50
+ - config_name: CC-MAIN-2025-05
51
+ data_files:
52
+ - split: train
53
+ path: data/CC-MAIN-2025-05/*
54
+ - config_name: CC-MAIN-2025-08
55
+ data_files:
56
+ - split: train
57
+ path: data/CC-MAIN-2025-08/*
58
+ - config_name: CC-MAIN-2025-13
59
+ data_files:
60
+ - split: train
61
+ path: data/CC-MAIN-2025-13/*
62
+ - config_name: CC-MAIN-2025-18
63
+ data_files:
64
+ - split: train
65
+ path: data/CC-MAIN-2025-18/*
66
+ - config_name: CC-MAIN-2025-21
67
+ data_files:
68
+ - split: train
69
+ path: data/CC-MAIN-2025-21/*
70
+ - config_name: CC-MAIN-2025-26
71
+ data_files:
72
+ - split: train
73
+ path: data/CC-MAIN-2025-26/*
74
+ - config_name: CC-MAIN-2024-51
75
+ data_files:
76
+ - split: train
77
+ path: data/CC-MAIN-2024-51/*
78
+ - config_name: CC-MAIN-2024-46
79
+ data_files:
80
+ - split: train
81
+ path: data/CC-MAIN-2024-46/*
82
+ - config_name: CC-MAIN-2024-42
83
+ data_files:
84
+ - split: train
85
+ path: data/CC-MAIN-2024-42/*
86
+ - config_name: CC-MAIN-2024-38
87
+ data_files:
88
+ - split: train
89
+ path: data/CC-MAIN-2024-38/*
90
+ - config_name: CC-MAIN-2024-33
91
+ data_files:
92
+ - split: train
93
+ path: data/CC-MAIN-2024-33/*
94
+ - config_name: CC-MAIN-2024-30
95
+ data_files:
96
+ - split: train
97
+ path: data/CC-MAIN-2024-30/*
98
+ - config_name: CC-MAIN-2024-26
99
+ data_files:
100
+ - split: train
101
+ path: data/CC-MAIN-2024-26/*
102
+ - config_name: CC-MAIN-2024-22
103
+ data_files:
104
+ - split: train
105
+ path: data/CC-MAIN-2024-22/*
106
+ - config_name: CC-MAIN-2024-18
107
+ data_files:
108
+ - split: train
109
+ path: data/CC-MAIN-2024-18/*
110
+ - config_name: CC-MAIN-2024-10
111
+ data_files:
112
+ - split: train
113
+ path: data/CC-MAIN-2024-10/*
114
+ - config_name: CC-MAIN-2023-50
115
+ data_files:
116
+ - split: train
117
+ path: data/CC-MAIN-2023-50/*
118
+ - config_name: CC-MAIN-2023-40
119
+ data_files:
120
+ - split: train
121
+ path: data/CC-MAIN-2023-40/*
122
+ - config_name: CC-MAIN-2023-23
123
+ data_files:
124
+ - split: train
125
+ path: data/CC-MAIN-2023-23/*
126
+ - config_name: CC-MAIN-2023-14
127
+ data_files:
128
+ - split: train
129
+ path: data/CC-MAIN-2023-14/*
130
+ - config_name: CC-MAIN-2023-06
131
+ data_files:
132
+ - split: train
133
+ path: data/CC-MAIN-2023-06/*
134
+ - config_name: CC-MAIN-2022-49
135
+ data_files:
136
+ - split: train
137
+ path: data/CC-MAIN-2022-49/*
138
+ - config_name: CC-MAIN-2022-40
139
+ data_files:
140
+ - split: train
141
+ path: data/CC-MAIN-2022-40/*
142
+ - config_name: CC-MAIN-2022-33
143
+ data_files:
144
+ - split: train
145
+ path: data/CC-MAIN-2022-33/*
146
+ - config_name: CC-MAIN-2022-27
147
+ data_files:
148
+ - split: train
149
+ path: data/CC-MAIN-2022-27/*
150
+ - config_name: CC-MAIN-2022-21
151
+ data_files:
152
+ - split: train
153
+ path: data/CC-MAIN-2022-21/*
154
+ - config_name: CC-MAIN-2022-05
155
+ data_files:
156
+ - split: train
157
+ path: data/CC-MAIN-2022-05/*
158
+ - config_name: CC-MAIN-2021-49
159
+ data_files:
160
+ - split: train
161
+ path: data/CC-MAIN-2021-49/*
162
+ - config_name: CC-MAIN-2021-43
163
+ data_files:
164
+ - split: train
165
+ path: data/CC-MAIN-2021-43/*
166
+ - config_name: CC-MAIN-2021-39
167
+ data_files:
168
+ - split: train
169
+ path: data/CC-MAIN-2021-39/*
170
+ - config_name: CC-MAIN-2021-31
171
+ data_files:
172
+ - split: train
173
+ path: data/CC-MAIN-2021-31/*
174
+ - config_name: CC-MAIN-2021-25
175
+ data_files:
176
+ - split: train
177
+ path: data/CC-MAIN-2021-25/*
178
+ - config_name: CC-MAIN-2021-21
179
+ data_files:
180
+ - split: train
181
+ path: data/CC-MAIN-2021-21/*
182
+ - config_name: CC-MAIN-2021-17
183
+ data_files:
184
+ - split: train
185
+ path: data/CC-MAIN-2021-17/*
186
+ - config_name: CC-MAIN-2021-10
187
+ data_files:
188
+ - split: train
189
+ path: data/CC-MAIN-2021-10/*
190
+ - config_name: CC-MAIN-2021-04
191
+ data_files:
192
+ - split: train
193
+ path: data/CC-MAIN-2021-04/*
194
+ - config_name: CC-MAIN-2020-50
195
+ data_files:
196
+ - split: train
197
+ path: data/CC-MAIN-2020-50/*
198
+ - config_name: CC-MAIN-2020-45
199
+ data_files:
200
+ - split: train
201
+ path: data/CC-MAIN-2020-45/*
202
+ - config_name: CC-MAIN-2020-40
203
+ data_files:
204
+ - split: train
205
+ path: data/CC-MAIN-2020-40/*
206
+ - config_name: CC-MAIN-2020-34
207
+ data_files:
208
+ - split: train
209
+ path: data/CC-MAIN-2020-34/*
210
+ - config_name: CC-MAIN-2020-29
211
+ data_files:
212
+ - split: train
213
+ path: data/CC-MAIN-2020-29/*
214
+ - config_name: CC-MAIN-2020-24
215
+ data_files:
216
+ - split: train
217
+ path: data/CC-MAIN-2020-24/*
218
+ - config_name: CC-MAIN-2020-16
219
+ data_files:
220
+ - split: train
221
+ path: data/CC-MAIN-2020-16/*
222
+ - config_name: CC-MAIN-2020-10
223
+ data_files:
224
+ - split: train
225
+ path: data/CC-MAIN-2020-10/*
226
+ - config_name: CC-MAIN-2020-05
227
+ data_files:
228
+ - split: train
229
+ path: data/CC-MAIN-2020-05/*
230
+ - config_name: CC-MAIN-2019-51
231
+ data_files:
232
+ - split: train
233
+ path: data/CC-MAIN-2019-51/*
234
+ - config_name: CC-MAIN-2019-47
235
+ data_files:
236
+ - split: train
237
+ path: data/CC-MAIN-2019-47/*
238
+ - config_name: CC-MAIN-2019-43
239
+ data_files:
240
+ - split: train
241
+ path: data/CC-MAIN-2019-43/*
242
+ - config_name: CC-MAIN-2019-39
243
+ data_files:
244
+ - split: train
245
+ path: data/CC-MAIN-2019-39/*
246
+ - config_name: CC-MAIN-2019-35
247
+ data_files:
248
+ - split: train
249
+ path: data/CC-MAIN-2019-35/*
250
+ - config_name: CC-MAIN-2019-30
251
+ data_files:
252
+ - split: train
253
+ path: data/CC-MAIN-2019-30/*
254
+ - config_name: CC-MAIN-2019-26
255
+ data_files:
256
+ - split: train
257
+ path: data/CC-MAIN-2019-26/*
258
+ - config_name: CC-MAIN-2019-22
259
+ data_files:
260
+ - split: train
261
+ path: data/CC-MAIN-2019-22/*
262
+ - config_name: CC-MAIN-2019-18
263
+ data_files:
264
+ - split: train
265
+ path: data/CC-MAIN-2019-18/*
266
+ - config_name: CC-MAIN-2019-13
267
+ data_files:
268
+ - split: train
269
+ path: data/CC-MAIN-2019-13/*
270
+ - config_name: CC-MAIN-2019-09
271
+ data_files:
272
+ - split: train
273
+ path: data/CC-MAIN-2019-09/*
274
+ - config_name: CC-MAIN-2019-04
275
+ data_files:
276
+ - split: train
277
+ path: data/CC-MAIN-2019-04/*
278
+ - config_name: CC-MAIN-2018-51
279
+ data_files:
280
+ - split: train
281
+ path: data/CC-MAIN-2018-51/*
282
+ - config_name: CC-MAIN-2018-47
283
+ data_files:
284
+ - split: train
285
+ path: data/CC-MAIN-2018-47/*
286
+ - config_name: CC-MAIN-2018-43
287
+ data_files:
288
+ - split: train
289
+ path: data/CC-MAIN-2018-43/*
290
+ - config_name: CC-MAIN-2018-39
291
+ data_files:
292
+ - split: train
293
+ path: data/CC-MAIN-2018-39/*
294
+ - config_name: CC-MAIN-2018-34
295
+ data_files:
296
+ - split: train
297
+ path: data/CC-MAIN-2018-34/*
298
+ - config_name: CC-MAIN-2018-30
299
+ data_files:
300
+ - split: train
301
+ path: data/CC-MAIN-2018-30/*
302
+ - config_name: CC-MAIN-2018-26
303
+ data_files:
304
+ - split: train
305
+ path: data/CC-MAIN-2018-26/*
306
+ - config_name: CC-MAIN-2018-22
307
+ data_files:
308
+ - split: train
309
+ path: data/CC-MAIN-2018-22/*
310
+ - config_name: CC-MAIN-2018-17
311
+ data_files:
312
+ - split: train
313
+ path: data/CC-MAIN-2018-17/*
314
+ - config_name: CC-MAIN-2018-13
315
+ data_files:
316
+ - split: train
317
+ path: data/CC-MAIN-2018-13/*
318
+ - config_name: CC-MAIN-2018-09
319
+ data_files:
320
+ - split: train
321
+ path: data/CC-MAIN-2018-09/*
322
+ - config_name: CC-MAIN-2018-05
323
+ data_files:
324
+ - split: train
325
+ path: data/CC-MAIN-2018-05/*
326
+ - config_name: CC-MAIN-2017-51
327
+ data_files:
328
+ - split: train
329
+ path: data/CC-MAIN-2017-51/*
330
+ - config_name: CC-MAIN-2017-47
331
+ data_files:
332
+ - split: train
333
+ path: data/CC-MAIN-2017-47/*
334
+ - config_name: CC-MAIN-2017-43
335
+ data_files:
336
+ - split: train
337
+ path: data/CC-MAIN-2017-43/*
338
+ - config_name: CC-MAIN-2017-39
339
+ data_files:
340
+ - split: train
341
+ path: data/CC-MAIN-2017-39/*
342
+ - config_name: CC-MAIN-2017-34
343
+ data_files:
344
+ - split: train
345
+ path: data/CC-MAIN-2017-34/*
346
+ - config_name: CC-MAIN-2017-30
347
+ data_files:
348
+ - split: train
349
+ path: data/CC-MAIN-2017-30/*
350
+ - config_name: CC-MAIN-2017-26
351
+ data_files:
352
+ - split: train
353
+ path: data/CC-MAIN-2017-26/*
354
+ - config_name: CC-MAIN-2017-22
355
+ data_files:
356
+ - split: train
357
+ path: data/CC-MAIN-2017-22/*
358
+ - config_name: CC-MAIN-2017-17
359
+ data_files:
360
+ - split: train
361
+ path: data/CC-MAIN-2017-17/*
362
+ - config_name: CC-MAIN-2017-13
363
+ data_files:
364
+ - split: train
365
+ path: data/CC-MAIN-2017-13/*
366
+ - config_name: CC-MAIN-2017-09
367
+ data_files:
368
+ - split: train
369
+ path: data/CC-MAIN-2017-09/*
370
+ - config_name: CC-MAIN-2017-04
371
+ data_files:
372
+ - split: train
373
+ path: data/CC-MAIN-2017-04/*
374
+ - config_name: CC-MAIN-2016-50
375
+ data_files:
376
+ - split: train
377
+ path: data/CC-MAIN-2016-50/*
378
+ - config_name: CC-MAIN-2016-44
379
+ data_files:
380
+ - split: train
381
+ path: data/CC-MAIN-2016-44/*
382
+ - config_name: CC-MAIN-2016-40
383
+ data_files:
384
+ - split: train
385
+ path: data/CC-MAIN-2016-40/*
386
+ - config_name: CC-MAIN-2016-36
387
+ data_files:
388
+ - split: train
389
+ path: data/CC-MAIN-2016-36/*
390
+ - config_name: CC-MAIN-2016-30
391
+ data_files:
392
+ - split: train
393
+ path: data/CC-MAIN-2016-30/*
394
+ - config_name: CC-MAIN-2016-26
395
+ data_files:
396
+ - split: train
397
+ path: data/CC-MAIN-2016-26/*
398
+ - config_name: CC-MAIN-2016-22
399
+ data_files:
400
+ - split: train
401
+ path: data/CC-MAIN-2016-22/*
402
+ - config_name: CC-MAIN-2016-18
403
+ data_files:
404
+ - split: train
405
+ path: data/CC-MAIN-2016-18/*
406
+ - config_name: CC-MAIN-2016-07
407
+ data_files:
408
+ - split: train
409
+ path: data/CC-MAIN-2016-07/*
410
+ - config_name: CC-MAIN-2015-48
411
+ data_files:
412
+ - split: train
413
+ path: data/CC-MAIN-2015-48/*
414
+ - config_name: CC-MAIN-2015-40
415
+ data_files:
416
+ - split: train
417
+ path: data/CC-MAIN-2015-40/*
418
+ - config_name: CC-MAIN-2015-35
419
+ data_files:
420
+ - split: train
421
+ path: data/CC-MAIN-2015-35/*
422
+ - config_name: CC-MAIN-2015-32
423
+ data_files:
424
+ - split: train
425
+ path: data/CC-MAIN-2015-32/*
426
+ - config_name: CC-MAIN-2015-27
427
+ data_files:
428
+ - split: train
429
+ path: data/CC-MAIN-2015-27/*
430
+ - config_name: CC-MAIN-2015-22
431
+ data_files:
432
+ - split: train
433
+ path: data/CC-MAIN-2015-22/*
434
+ - config_name: CC-MAIN-2015-18
435
+ data_files:
436
+ - split: train
437
+ path: data/CC-MAIN-2015-18/*
438
+ - config_name: CC-MAIN-2015-14
439
+ data_files:
440
+ - split: train
441
+ path: data/CC-MAIN-2015-14/*
442
+ - config_name: CC-MAIN-2015-11
443
+ data_files:
444
+ - split: train
445
+ path: data/CC-MAIN-2015-11/*
446
+ - config_name: CC-MAIN-2015-06
447
+ data_files:
448
+ - split: train
449
+ path: data/CC-MAIN-2015-06/*
450
+ - config_name: CC-MAIN-2014-52
451
+ data_files:
452
+ - split: train
453
+ path: data/CC-MAIN-2014-52/*
454
+ - config_name: CC-MAIN-2014-49
455
+ data_files:
456
+ - split: train
457
+ path: data/CC-MAIN-2014-49/*
458
+ - config_name: CC-MAIN-2014-42
459
+ data_files:
460
+ - split: train
461
+ path: data/CC-MAIN-2014-42/*
462
+ - config_name: CC-MAIN-2014-41
463
+ data_files:
464
+ - split: train
465
+ path: data/CC-MAIN-2014-41/*
466
+ - config_name: CC-MAIN-2014-35
467
+ data_files:
468
+ - split: train
469
+ path: data/CC-MAIN-2014-35/*
470
+ - config_name: CC-MAIN-2014-23
471
+ data_files:
472
+ - split: train
473
+ path: data/CC-MAIN-2014-23/*
474
+ - config_name: CC-MAIN-2014-15
475
+ data_files:
476
+ - split: train
477
+ path: data/CC-MAIN-2014-15/*
478
+ - config_name: CC-MAIN-2014-10
479
+ data_files:
480
+ - split: train
481
+ path: data/CC-MAIN-2014-10/*
482
+ - config_name: CC-MAIN-2013-48
483
+ data_files:
484
+ - split: train
485
+ path: data/CC-MAIN-2013-48/*
486
+ - config_name: CC-MAIN-2013-20
487
+ data_files:
488
+ - split: train
489
+ path: data/CC-MAIN-2013-20/*
490
+ ---
491
+
492
+ # 📚 FineWeb-Edu
493
+ <center>
494
+ <img src="https://cdn-uploads.huggingface.co/production/uploads/61c141342aac764ce1654e43/wwRnEQydH9qdRtFofIE-A.png" alt="FineWeb-Edu: The finest collection of educational content the web has to offer">
495
+ </center>
496
+
497
+ > 1.3 trillion tokens of the finest educational data the 🌐 web has to offer
498
+
499
+ **Paper:** https://arxiv.org/abs/2406.17557
500
+
501
+ ## What is it?
502
+
503
+ 📚 FineWeb-Edu dataset consists of **1.3T tokens** and **5.4T tokens** ([FineWeb-Edu-score-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu-score-2)) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version.
504
+
505
+ To enhance FineWeb's quality, we developed an [educational quality classifier](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier) using annotations generated by LLama3-70B-Instruct. We then used this classifier to retain only the most educational web pages. FineWeb-Edu outperforms FineWeb on popular benchmarks and shows the power of classifiers trained on synthetic data.
506
+
507
+ The [Dataset Curation](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu#dataset-curation) section details the process for creating the dataset.
508
+
509
+ ![image/png](https://cdn-uploads.huggingface.co/production/uploads/61c141342aac764ce1654e43/QqXOM8h_ZjjhuCv71xmV7.png)
510
+
511
+ You can find a deduplicated version of FineWeb-edu in [SmolLM-Corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus). We find that the deduplication of this dataset doesn't have any impact on model performance in our ablation setup (1.8B trained on 350B tokens).
512
+
513
+ ## What is being released?
514
+
515
+ Along with the dataset, which includes all filtered CommonCrawl dumps since 2013, we also release the educational classifier used for the filtering as well as the code for training it and running inference at: https://github.com/huggingface/cosmopedia/tree/main/classification
516
+
517
+ ## Changelog
518
+ _Previous versions remain available in the branch `version name`._
519
+
520
+ - **v1.4.0 (11-07-2025):** Added 6 new snapshots: `CC-MAIN-2025-05`, `CC-MAIN-2025-08`, `CC-MAIN-2025-13`, `CC-MAIN-2025-18`, `CC-MAIN-2025-21`, and `CC-MAIN-2025-26` (January to June 2025)
521
+ - **v1.3.0 (31-01-2025):** Fixed an issue with some dumps where some documents hadn't been processed: `CC-MAIN-2024-10`, `CC-MAIN-2024-18`, `CC-MAIN-2024-22`, `CC-MAIN-2024-26`, `CC-MAIN-2024-30`, `CC-MAIN-2024-33`, `CC-MAIN-2024-38`, `CC-MAIN-2024-42`, `CC-MAIN-2024-46` -- they now contain more data (~35B additional tokens).
522
+ - **v1.2.0 (03-01-2025):** Added 9 new snapshots: `CC-MAIN-2024-18`, `CC-MAIN-2024-22`, `CC-MAIN-2024-26`, `CC-MAIN-2024-30`, `CC-MAIN-2024-33`, `CC-MAIN-2024-38`, `CC-MAIN-2024-42`, `CC-MAIN-2024-46`, `CC-MAIN-2024-51`, covering April to December 2024.
523
+ - **v1.0.0 (02-06-2024):** Initial version
524
+
525
+
526
+ ## How to load the dataset
527
+ Similarily to FineWeb, You can load the full dataset or a specific crawl/dump. Dumps have the format `CC-MAIN-(year)-(week number)`.
528
+
529
+ ### (Smaller) sample versions
530
+ Along with config `default` (all the data), and the configs for each individual dump, you can also download the following configs:
531
+ - `sample-350BT`: a subset randomly sampled from the whole dataset of around 350B gpt2 tokens
532
+ - `sample-100BT`: a subset randomly sampled from the whole dataset of around 100B gpt2 tokens
533
+ - `sample-10BT`: a subset randomly sampled from the whole dataset of around 10B gpt2 tokens
534
+
535
+ `sample-10BT` was sampled from `sample-100BT` which in turn was sampled from `sample-350BT`.
536
+
537
+ ### Using 🏭 [`datatrove`](https://github.com/huggingface/datatrove/)
538
+
539
+ ```python
540
+ from datatrove.pipeline.readers import ParquetReader
541
+
542
+ # limit determines how many documents will be streamed (remove for all)
543
+ data_reader = ParquetReader("hf://datasets/HuggingFaceFW/fineweb-edu", glob_pattern="data/*/*.parquet", limit=1000)
544
+ # or to fetch a specific dump CC-MAIN-2024-10, eplace "CC-MAIN-2024-10" with "sample/100BT" to use the 100BT sample
545
+ data_reader = ParquetReader("hf://datasets/HuggingFaceFW/fineweb-edu/CC-MAIN-2024-10", limit=1000)
546
+ for document in data_reader():
547
+ # do something with document
548
+ print(document)
549
+
550
+ ###############################
551
+ # OR for a processing pipeline:
552
+ ###############################
553
+
554
+ from datatrove.executor import LocalPipelineExecutor
555
+ from datatrove.pipeline.readers import ParquetReader
556
+ from datatrove.pipeline.filters import LambdaFilter
557
+ from datatrove.pipeline.writers import JsonlWriter
558
+
559
+ pipeline_exec = LocalPipelineExecutor(
560
+ pipeline=[
561
+ # replace "CC-MAIN-2024-10" with "sample/100BT" to use the 100BT sample
562
+ ParquetReader("hf://datasets/HuggingFaceFW/fineweb-edu/CC-MAIN-2024-10", limit=1000),
563
+ LambdaFilter(lambda doc: "hugging" in doc.text),
564
+ JsonlWriter("some-output-path")
565
+ ],
566
+ tasks=10
567
+ )
568
+ pipeline_exec.run()
569
+ ```
570
+
571
+ ### Using `datasets`
572
+
573
+ ```python
574
+ from datasets import load_dataset
575
+ # use name="sample-10BT" to use the 10BT sample
576
+ fw = load_dataset("HuggingFaceFW/fineweb-edu", name="CC-MAIN-2024-10", split="train", streaming=True)
577
+ ```
578
+
579
+ ## Dataset curation
580
+ A new approach has recently emerged for filtering LLM training datasets: using synthetic data to develop classifiers for identifying educational content. This technique was used in the trainings of [LLama3](https://ai.meta.com/blog/meta-llama-3-meta-ai-responsibility/) and [Phi3](https://arxiv.org/abs/2404.14219), but its large-scale impact on web data filtering hasn't been fully explored or published.
581
+
582
+ The highly popular Phi3 models were trained on 3.3 and 4.8 trillion tokens, with the paper stating: “Our training data consists of heavily filtered publicly available web data (according to the 'educational level') from various open internet sources, as well as synthetic LLM-generated data". Similarly, the LLama3 blog post notes: “We found that previous generations of Llama are good at identifying high-quality data, so we used Llama 2 to help build the text-quality classifiers that are powering Llama 3.” However these classifiers and filtered datasets are not publicly available. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by [LLama3-70B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct) to create FineWeb-Edu.
583
+
584
+ ### Annotation
585
+ We used [Llama3-70B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct) to score 500k FineWeb samples for their educational quality on a scale from 0 to 5.
586
+
587
+ We explored various prompts and found that the additive scale by [Yuan et al.](https://arxiv.org/pdf/2401.10020) worked best. To avoid the LLM favoring highly technical pages like arXiv abstracts and submissions, we focused on grade-school and middle-school level knowledge. By setting a threshold of 3 (on a scale of 0 to 5) during the filtering process, we were able to also retain some high-level educational pages. The final prompt can be found [here](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier/blob/main/utils/prompt.txt).
588
+
589
+ We also experimented with different LLMs: Llama3-70B-Instruct, Mixtral-8x-7B-Instruct, and Mixtral-8x22B-Instruct. Llama 3 and Mixtral-8x22B produced similar scores, while Mixtral-8x7B tended to be more generous, not fully adhering to the score scale. Verga et al. suggest using multiple LLMs as juries. We tried averaging the scores from the three models, but this shifted the distribution to the right due to the higher scores from Mixtral-8x7B. Training on a dataset filtered with a classifier using jury annotations performed worse than using a classifier based on Llama3 annotations. We hypothesize that the jury-based approach retains more low-quality samples.
590
+
591
+ ### Classifier training
592
+ We fine-tuned a Bert-like regression model using these annotations, based on [Snowflake-arctic-embed](https://huggingface.co/Snowflake/snowflake-arctic-embed-m). When converted to a binary classification using a score of 3 as a threshold for keeping and removing files, the model achieved an F1 score of 82%. The classification of FineWeb 15T tokens took 6k H100 GPU hours.
593
+
594
+ The classifier is available at: [HuggingFaceFW/fineweb-edu-classifier/](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier/)
595
+
596
+ ### Filtering and results
597
+ **Note**: You can find more details about the ablations and results in the FineWeb [blog post](https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1).
598
+
599
+ We investigated the impact of using different thresholds for the filtering and found that threshold 3 gave the best overall results. Although using a threshold higher than 3 improves performance on knowledge and reasoning intensive benchmarks, it significantly degrades performance on HellaSwag and PIQA.
600
+
601
+ We then built 📚 FineWeb-Edu by filtering out samples with scores lower than 3. This removed 92% of the dataset, leaving us with 1.3T educational tokens. Our ablation demonstrated that this refined dataset surpasses 🍷 FineWeb and all other open web datasets, with remarkable improvements on educational benchmarks such as MMLU, ARC, and OpenBookQA. The plot below compares FineWeb-Edu to other web datasets:
602
+
603
+ ![image/png](https://cdn-uploads.huggingface.co/production/uploads/61c141342aac764ce1654e43/hJlyTgDzZpYuxO9LUm0PF.png)
604
+
605
+ To retain more tokens, we also experimented with a less strict threshold of 2 instead of 3. While being less performant than using threshold 3, it still outperformed FineWeb and it preserved 5.4T tokens. We release these two dataset as [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) and [FineWeb-Edu-score-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu-score-2) along with the [classifier](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier).
606
+
607
+ You will find all the ablation models in [this collection](https://huggingface.co/collections/HuggingFaceFW/ablation-models-662457b0d213e8c14fe47f32). The FineWeb-Edu ablation model (trained on 350B tokens) is available at [https://huggingface.co/HuggingFaceFW/ablation-model-fineweb-edu](https://huggingface.co/HuggingFaceFW/ablation-model-fineweb-edu).
608
+
609
+ ## Considerations for Using the Data
610
+ This section is copied from the parent dataset: [FineWeb](https://huggingface.co/datasets/HuggingFaceFW/fineweb).
611
+
612
+ ### Social Impact of Dataset
613
+
614
+ With the release of this dataset we aim to make model training more accessible to the machine learning community at large.
615
+
616
+ While multiple open-weights models with strong performance have been publicly released in the past, more often than not these releases are not accompanied by the corresponding training dataset. This is unfortunate as the dataset specificities and characteristics have been demonstrated to have a very large impact and role in the performances of the models. As the creation of a high quality training dataset is a fundamental requirement to training an LLM capable of excelling at downstream tasks, with 🍷 FineWeb we (a) not only make the dataset creation process more transparent, by sharing our entire processing setup including the codebase used, we also (b) help alleviate the costs of dataset curation, both in time and in compute, for model creators by publicly releasing our dataset with the community.
617
+
618
+ ### Discussion of Biases
619
+
620
+ Efforts were made to minimize the amount of NSFW and toxic content present in the dataset by employing filtering on the URL level. However, there are still a significant number of documents present in the final dataset that could be considered toxic or contain harmful content. As 🍷 FineWeb was sourced from the web as a whole, any harmful biases typically present in it may be reproduced on our dataset.
621
+
622
+ We deliberately avoided using machine learning filtering methods that define text quality based on the similarity to a “gold” source such as wikipedia or toxicity classifiers as these methods have been known to [disproportionately remove content in specific dialects](https://aclanthology.org/D16-1120/) and [overclassify as toxic text related to specific social identities](https://arxiv.org/pdf/2109.07445.pdf), respectively.
623
+
624
+ ### Other Known Limitations
625
+
626
+ As a consequence of some of the filtering steps applied, it is likely that code content is not prevalent in our dataset. If you are training a model that should also perform code tasks, we recommend you use 🍷 FineWeb with a code dataset, such as [The Stack v2](https://huggingface.co/datasets/bigcode/the-stack-v2). You should also probably consider complementing 🍷 FineWeb with specialized curated sources (such as Wikipedia, for example) as they will likely have better formatting than the wikipedia content included in 🍷 FineWeb (we did not tailor the processing to individual websites).
627
+
628
+ ## Additional Information
629
+
630
+ ### Licensing Information
631
+
632
+ The dataset is released under the **Open Data Commons Attribution License (ODC-By) v1.0** [license](https://opendatacommons.org/licenses/by/1-0/). The use of this dataset is also subject to [CommonCrawl's Terms of Use](https://commoncrawl.org/terms-of-use).
633
+
634
+ ### Future work
635
+
636
+ We plan to work on better educational classifier to improve the quality of FineWeb-Edu.
637
+
638
+ ### Citation Information
639
+
640
+ You can cite our paper https://arxiv.org/abs/2406.17557 or this dataset:
641
+
642
+ ```
643
+ @misc{lozhkov2024fineweb-edu,
644
+ author = { Lozhkov, Anton and Ben Allal, Loubna and von Werra, Leandro and Wolf, Thomas },
645
+ title = { FineWeb-Edu: the Finest Collection of Educational Content },
646
+ year = 2024,
647
+ url = { https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu },
648
+ doi = { 10.57967/hf/2497 },
649
+ publisher = { Hugging Face }
650
+ }
651
+ ```
hf_cache/hub/datasets--Rowan--hellaswag/.no_exist/218ec52e09a7e7462a5400043bb9a69a41d06b76/.huggingface.yaml ADDED
File without changes
hf_cache/hub/datasets--Rowan--hellaswag/.no_exist/218ec52e09a7e7462a5400043bb9a69a41d06b76/dataset_infos.json ADDED
File without changes
hf_cache/hub/datasets--Rowan--hellaswag/.no_exist/218ec52e09a7e7462a5400043bb9a69a41d06b76/hellaswag.py ADDED
File without changes
hf_cache/hub/datasets--Rowan--hellaswag/blobs/29f11d90eb3a5b319cfe8ce2a4e78d9f2a1aea3f ADDED
@@ -0,0 +1,218 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ paperswithcode_id: hellaswag
5
+ pretty_name: HellaSwag
6
+ dataset_info:
7
+ features:
8
+ - name: ind
9
+ dtype: int32
10
+ - name: activity_label
11
+ dtype: string
12
+ - name: ctx_a
13
+ dtype: string
14
+ - name: ctx_b
15
+ dtype: string
16
+ - name: ctx
17
+ dtype: string
18
+ - name: endings
19
+ sequence: string
20
+ - name: source_id
21
+ dtype: string
22
+ - name: split
23
+ dtype: string
24
+ - name: split_type
25
+ dtype: string
26
+ - name: label
27
+ dtype: string
28
+ splits:
29
+ - name: train
30
+ num_bytes: 43232624
31
+ num_examples: 39905
32
+ - name: test
33
+ num_bytes: 10791853
34
+ num_examples: 10003
35
+ - name: validation
36
+ num_bytes: 11175717
37
+ num_examples: 10042
38
+ download_size: 36793872
39
+ dataset_size: 65200194
40
+ configs:
41
+ - config_name: default
42
+ data_files:
43
+ - split: train
44
+ path: data/train-*
45
+ - split: test
46
+ path: data/test-*
47
+ - split: validation
48
+ path: data/validation-*
49
+ ---
50
+
51
+ # Dataset Card for "hellaswag"
52
+
53
+ ## Table of Contents
54
+ - [Dataset Description](#dataset-description)
55
+ - [Dataset Summary](#dataset-summary)
56
+ - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
57
+ - [Languages](#languages)
58
+ - [Dataset Structure](#dataset-structure)
59
+ - [Data Instances](#data-instances)
60
+ - [Data Fields](#data-fields)
61
+ - [Data Splits](#data-splits)
62
+ - [Dataset Creation](#dataset-creation)
63
+ - [Curation Rationale](#curation-rationale)
64
+ - [Source Data](#source-data)
65
+ - [Annotations](#annotations)
66
+ - [Personal and Sensitive Information](#personal-and-sensitive-information)
67
+ - [Considerations for Using the Data](#considerations-for-using-the-data)
68
+ - [Social Impact of Dataset](#social-impact-of-dataset)
69
+ - [Discussion of Biases](#discussion-of-biases)
70
+ - [Other Known Limitations](#other-known-limitations)
71
+ - [Additional Information](#additional-information)
72
+ - [Dataset Curators](#dataset-curators)
73
+ - [Licensing Information](#licensing-information)
74
+ - [Citation Information](#citation-information)
75
+ - [Contributions](#contributions)
76
+
77
+ ## Dataset Description
78
+
79
+ - **Homepage:** [https://rowanzellers.com/hellaswag/](https://rowanzellers.com/hellaswag/)
80
+ - **Repository:** [https://github.com/rowanz/hellaswag/](https://github.com/rowanz/hellaswag/)
81
+ - **Paper:** [HellaSwag: Can a Machine Really Finish Your Sentence?](https://arxiv.org/abs/1905.07830)
82
+ - **Point of Contact:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
83
+ - **Size of downloaded dataset files:** 71.49 MB
84
+ - **Size of the generated dataset:** 65.32 MB
85
+ - **Total amount of disk used:** 136.81 MB
86
+
87
+ ### Dataset Summary
88
+
89
+ HellaSwag: Can a Machine Really Finish Your Sentence? is a new dataset for commonsense NLI. A paper was published at ACL2019.
90
+
91
+ ### Supported Tasks and Leaderboards
92
+
93
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
94
+
95
+ ### Languages
96
+
97
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
98
+
99
+ ## Dataset Structure
100
+
101
+ ### Data Instances
102
+
103
+ #### default
104
+
105
+ - **Size of downloaded dataset files:** 71.49 MB
106
+ - **Size of the generated dataset:** 65.32 MB
107
+ - **Total amount of disk used:** 136.81 MB
108
+
109
+ An example of 'train' looks as follows.
110
+ ```
111
+ This example was too long and was cropped:
112
+
113
+ {
114
+ "activity_label": "Removing ice from car",
115
+ "ctx": "Then, the man writes over the snow covering the window of a car, and a woman wearing winter clothes smiles. then",
116
+ "ctx_a": "Then, the man writes over the snow covering the window of a car, and a woman wearing winter clothes smiles.",
117
+ "ctx_b": "then",
118
+ "endings": "[\", the man adds wax to the windshield and cuts it.\", \", a person board a ski lift, while two men supporting the head of the per...",
119
+ "ind": 4,
120
+ "label": "3",
121
+ "source_id": "activitynet~v_-1IBHYS3L-Y",
122
+ "split": "train",
123
+ "split_type": "indomain"
124
+ }
125
+ ```
126
+
127
+ ### Data Fields
128
+
129
+ The data fields are the same among all splits.
130
+
131
+ #### default
132
+ - `ind`: a `int32` feature.
133
+ - `activity_label`: a `string` feature.
134
+ - `ctx_a`: a `string` feature.
135
+ - `ctx_b`: a `string` feature.
136
+ - `ctx`: a `string` feature.
137
+ - `endings`: a `list` of `string` features.
138
+ - `source_id`: a `string` feature.
139
+ - `split`: a `string` feature.
140
+ - `split_type`: a `string` feature.
141
+ - `label`: a `string` feature.
142
+
143
+ ### Data Splits
144
+
145
+ | name |train|validation|test |
146
+ |-------|----:|---------:|----:|
147
+ |default|39905| 10042|10003|
148
+
149
+ ## Dataset Creation
150
+
151
+ ### Curation Rationale
152
+
153
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
154
+
155
+ ### Source Data
156
+
157
+ #### Initial Data Collection and Normalization
158
+
159
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
160
+
161
+ #### Who are the source language producers?
162
+
163
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
164
+
165
+ ### Annotations
166
+
167
+ #### Annotation process
168
+
169
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
170
+
171
+ #### Who are the annotators?
172
+
173
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
174
+
175
+ ### Personal and Sensitive Information
176
+
177
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
178
+
179
+ ## Considerations for Using the Data
180
+
181
+ ### Social Impact of Dataset
182
+
183
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
184
+
185
+ ### Discussion of Biases
186
+
187
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
188
+
189
+ ### Other Known Limitations
190
+
191
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
192
+
193
+ ## Additional Information
194
+
195
+ ### Dataset Curators
196
+
197
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
198
+
199
+ ### Licensing Information
200
+
201
+ MIT https://github.com/rowanz/hellaswag/blob/master/LICENSE
202
+
203
+ ### Citation Information
204
+
205
+ ```
206
+ @inproceedings{zellers2019hellaswag,
207
+ title={HellaSwag: Can a Machine Really Finish Your Sentence?},
208
+ author={Zellers, Rowan and Holtzman, Ari and Bisk, Yonatan and Farhadi, Ali and Choi, Yejin},
209
+ booktitle ={Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics},
210
+ year={2019}
211
+ }
212
+
213
+ ```
214
+
215
+
216
+ ### Contributions
217
+
218
+ Thanks to [@albertvillanova](https://github.com/albertvillanova), [@mariamabarham](https://github.com/mariamabarham), [@thomwolf](https://github.com/thomwolf), [@patrickvonplaten](https://github.com/patrickvonplaten), [@lewtun](https://github.com/lewtun) for adding this dataset.
hf_cache/hub/datasets--Rowan--hellaswag/refs/main ADDED
@@ -0,0 +1 @@
 
 
1
+ 218ec52e09a7e7462a5400043bb9a69a41d06b76
hf_cache/hub/datasets--Rowan--hellaswag/snapshots/218ec52e09a7e7462a5400043bb9a69a41d06b76/README.md ADDED
@@ -0,0 +1,218 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ paperswithcode_id: hellaswag
5
+ pretty_name: HellaSwag
6
+ dataset_info:
7
+ features:
8
+ - name: ind
9
+ dtype: int32
10
+ - name: activity_label
11
+ dtype: string
12
+ - name: ctx_a
13
+ dtype: string
14
+ - name: ctx_b
15
+ dtype: string
16
+ - name: ctx
17
+ dtype: string
18
+ - name: endings
19
+ sequence: string
20
+ - name: source_id
21
+ dtype: string
22
+ - name: split
23
+ dtype: string
24
+ - name: split_type
25
+ dtype: string
26
+ - name: label
27
+ dtype: string
28
+ splits:
29
+ - name: train
30
+ num_bytes: 43232624
31
+ num_examples: 39905
32
+ - name: test
33
+ num_bytes: 10791853
34
+ num_examples: 10003
35
+ - name: validation
36
+ num_bytes: 11175717
37
+ num_examples: 10042
38
+ download_size: 36793872
39
+ dataset_size: 65200194
40
+ configs:
41
+ - config_name: default
42
+ data_files:
43
+ - split: train
44
+ path: data/train-*
45
+ - split: test
46
+ path: data/test-*
47
+ - split: validation
48
+ path: data/validation-*
49
+ ---
50
+
51
+ # Dataset Card for "hellaswag"
52
+
53
+ ## Table of Contents
54
+ - [Dataset Description](#dataset-description)
55
+ - [Dataset Summary](#dataset-summary)
56
+ - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
57
+ - [Languages](#languages)
58
+ - [Dataset Structure](#dataset-structure)
59
+ - [Data Instances](#data-instances)
60
+ - [Data Fields](#data-fields)
61
+ - [Data Splits](#data-splits)
62
+ - [Dataset Creation](#dataset-creation)
63
+ - [Curation Rationale](#curation-rationale)
64
+ - [Source Data](#source-data)
65
+ - [Annotations](#annotations)
66
+ - [Personal and Sensitive Information](#personal-and-sensitive-information)
67
+ - [Considerations for Using the Data](#considerations-for-using-the-data)
68
+ - [Social Impact of Dataset](#social-impact-of-dataset)
69
+ - [Discussion of Biases](#discussion-of-biases)
70
+ - [Other Known Limitations](#other-known-limitations)
71
+ - [Additional Information](#additional-information)
72
+ - [Dataset Curators](#dataset-curators)
73
+ - [Licensing Information](#licensing-information)
74
+ - [Citation Information](#citation-information)
75
+ - [Contributions](#contributions)
76
+
77
+ ## Dataset Description
78
+
79
+ - **Homepage:** [https://rowanzellers.com/hellaswag/](https://rowanzellers.com/hellaswag/)
80
+ - **Repository:** [https://github.com/rowanz/hellaswag/](https://github.com/rowanz/hellaswag/)
81
+ - **Paper:** [HellaSwag: Can a Machine Really Finish Your Sentence?](https://arxiv.org/abs/1905.07830)
82
+ - **Point of Contact:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
83
+ - **Size of downloaded dataset files:** 71.49 MB
84
+ - **Size of the generated dataset:** 65.32 MB
85
+ - **Total amount of disk used:** 136.81 MB
86
+
87
+ ### Dataset Summary
88
+
89
+ HellaSwag: Can a Machine Really Finish Your Sentence? is a new dataset for commonsense NLI. A paper was published at ACL2019.
90
+
91
+ ### Supported Tasks and Leaderboards
92
+
93
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
94
+
95
+ ### Languages
96
+
97
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
98
+
99
+ ## Dataset Structure
100
+
101
+ ### Data Instances
102
+
103
+ #### default
104
+
105
+ - **Size of downloaded dataset files:** 71.49 MB
106
+ - **Size of the generated dataset:** 65.32 MB
107
+ - **Total amount of disk used:** 136.81 MB
108
+
109
+ An example of 'train' looks as follows.
110
+ ```
111
+ This example was too long and was cropped:
112
+
113
+ {
114
+ "activity_label": "Removing ice from car",
115
+ "ctx": "Then, the man writes over the snow covering the window of a car, and a woman wearing winter clothes smiles. then",
116
+ "ctx_a": "Then, the man writes over the snow covering the window of a car, and a woman wearing winter clothes smiles.",
117
+ "ctx_b": "then",
118
+ "endings": "[\", the man adds wax to the windshield and cuts it.\", \", a person board a ski lift, while two men supporting the head of the per...",
119
+ "ind": 4,
120
+ "label": "3",
121
+ "source_id": "activitynet~v_-1IBHYS3L-Y",
122
+ "split": "train",
123
+ "split_type": "indomain"
124
+ }
125
+ ```
126
+
127
+ ### Data Fields
128
+
129
+ The data fields are the same among all splits.
130
+
131
+ #### default
132
+ - `ind`: a `int32` feature.
133
+ - `activity_label`: a `string` feature.
134
+ - `ctx_a`: a `string` feature.
135
+ - `ctx_b`: a `string` feature.
136
+ - `ctx`: a `string` feature.
137
+ - `endings`: a `list` of `string` features.
138
+ - `source_id`: a `string` feature.
139
+ - `split`: a `string` feature.
140
+ - `split_type`: a `string` feature.
141
+ - `label`: a `string` feature.
142
+
143
+ ### Data Splits
144
+
145
+ | name |train|validation|test |
146
+ |-------|----:|---------:|----:|
147
+ |default|39905| 10042|10003|
148
+
149
+ ## Dataset Creation
150
+
151
+ ### Curation Rationale
152
+
153
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
154
+
155
+ ### Source Data
156
+
157
+ #### Initial Data Collection and Normalization
158
+
159
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
160
+
161
+ #### Who are the source language producers?
162
+
163
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
164
+
165
+ ### Annotations
166
+
167
+ #### Annotation process
168
+
169
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
170
+
171
+ #### Who are the annotators?
172
+
173
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
174
+
175
+ ### Personal and Sensitive Information
176
+
177
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
178
+
179
+ ## Considerations for Using the Data
180
+
181
+ ### Social Impact of Dataset
182
+
183
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
184
+
185
+ ### Discussion of Biases
186
+
187
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
188
+
189
+ ### Other Known Limitations
190
+
191
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
192
+
193
+ ## Additional Information
194
+
195
+ ### Dataset Curators
196
+
197
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
198
+
199
+ ### Licensing Information
200
+
201
+ MIT https://github.com/rowanz/hellaswag/blob/master/LICENSE
202
+
203
+ ### Citation Information
204
+
205
+ ```
206
+ @inproceedings{zellers2019hellaswag,
207
+ title={HellaSwag: Can a Machine Really Finish Your Sentence?},
208
+ author={Zellers, Rowan and Holtzman, Ari and Bisk, Yonatan and Farhadi, Ali and Choi, Yejin},
209
+ booktitle ={Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics},
210
+ year={2019}
211
+ }
212
+
213
+ ```
214
+
215
+
216
+ ### Contributions
217
+
218
+ Thanks to [@albertvillanova](https://github.com/albertvillanova), [@mariamabarham](https://github.com/mariamabarham), [@thomwolf](https://github.com/thomwolf), [@patrickvonplaten](https://github.com/patrickvonplaten), [@lewtun](https://github.com/lewtun) for adding this dataset.
hf_cache/hub/datasets--Salesforce--wikitext/.no_exist/b08601e04326c79dfdd32d625aee71d232d685c3/.huggingface.yaml ADDED
File without changes
hf_cache/hub/datasets--Salesforce--wikitext/.no_exist/b08601e04326c79dfdd32d625aee71d232d685c3/dataset_infos.json ADDED
File without changes
hf_cache/hub/datasets--Salesforce--wikitext/.no_exist/b08601e04326c79dfdd32d625aee71d232d685c3/wikitext.py ADDED
File without changes
hf_cache/hub/datasets--Salesforce--wikitext/blobs/2a4fec2bc8df76c9d4da1c8e8865b625eb221c76 ADDED
@@ -0,0 +1,344 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ annotations_creators:
3
+ - no-annotation
4
+ language_creators:
5
+ - crowdsourced
6
+ language:
7
+ - en
8
+ license:
9
+ - cc-by-sa-3.0
10
+ - gfdl
11
+ multilinguality:
12
+ - monolingual
13
+ size_categories:
14
+ - 1M<n<10M
15
+ source_datasets:
16
+ - original
17
+ task_categories:
18
+ - text-generation
19
+ - fill-mask
20
+ task_ids:
21
+ - language-modeling
22
+ - masked-language-modeling
23
+ paperswithcode_id: wikitext-2
24
+ pretty_name: WikiText
25
+ dataset_info:
26
+ - config_name: wikitext-103-raw-v1
27
+ features:
28
+ - name: text
29
+ dtype: string
30
+ splits:
31
+ - name: test
32
+ num_bytes: 1305088
33
+ num_examples: 4358
34
+ - name: train
35
+ num_bytes: 546500949
36
+ num_examples: 1801350
37
+ - name: validation
38
+ num_bytes: 1159288
39
+ num_examples: 3760
40
+ download_size: 315466397
41
+ dataset_size: 548965325
42
+ - config_name: wikitext-103-v1
43
+ features:
44
+ - name: text
45
+ dtype: string
46
+ splits:
47
+ - name: test
48
+ num_bytes: 1295575
49
+ num_examples: 4358
50
+ - name: train
51
+ num_bytes: 545141915
52
+ num_examples: 1801350
53
+ - name: validation
54
+ num_bytes: 1154751
55
+ num_examples: 3760
56
+ download_size: 313093838
57
+ dataset_size: 547592241
58
+ - config_name: wikitext-2-raw-v1
59
+ features:
60
+ - name: text
61
+ dtype: string
62
+ splits:
63
+ - name: test
64
+ num_bytes: 1305088
65
+ num_examples: 4358
66
+ - name: train
67
+ num_bytes: 11061717
68
+ num_examples: 36718
69
+ - name: validation
70
+ num_bytes: 1159288
71
+ num_examples: 3760
72
+ download_size: 7747362
73
+ dataset_size: 13526093
74
+ - config_name: wikitext-2-v1
75
+ features:
76
+ - name: text
77
+ dtype: string
78
+ splits:
79
+ - name: test
80
+ num_bytes: 1270947
81
+ num_examples: 4358
82
+ - name: train
83
+ num_bytes: 10918118
84
+ num_examples: 36718
85
+ - name: validation
86
+ num_bytes: 1134123
87
+ num_examples: 3760
88
+ download_size: 7371282
89
+ dataset_size: 13323188
90
+ configs:
91
+ - config_name: wikitext-103-raw-v1
92
+ data_files:
93
+ - split: test
94
+ path: wikitext-103-raw-v1/test-*
95
+ - split: train
96
+ path: wikitext-103-raw-v1/train-*
97
+ - split: validation
98
+ path: wikitext-103-raw-v1/validation-*
99
+ - config_name: wikitext-103-v1
100
+ data_files:
101
+ - split: test
102
+ path: wikitext-103-v1/test-*
103
+ - split: train
104
+ path: wikitext-103-v1/train-*
105
+ - split: validation
106
+ path: wikitext-103-v1/validation-*
107
+ - config_name: wikitext-2-raw-v1
108
+ data_files:
109
+ - split: test
110
+ path: wikitext-2-raw-v1/test-*
111
+ - split: train
112
+ path: wikitext-2-raw-v1/train-*
113
+ - split: validation
114
+ path: wikitext-2-raw-v1/validation-*
115
+ - config_name: wikitext-2-v1
116
+ data_files:
117
+ - split: test
118
+ path: wikitext-2-v1/test-*
119
+ - split: train
120
+ path: wikitext-2-v1/train-*
121
+ - split: validation
122
+ path: wikitext-2-v1/validation-*
123
+ ---
124
+
125
+ # Dataset Card for "wikitext"
126
+
127
+ ## Table of Contents
128
+ - [Dataset Description](#dataset-description)
129
+ - [Dataset Summary](#dataset-summary)
130
+ - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
131
+ - [Languages](#languages)
132
+ - [Dataset Structure](#dataset-structure)
133
+ - [Data Instances](#data-instances)
134
+ - [Data Fields](#data-fields)
135
+ - [Data Splits](#data-splits)
136
+ - [Dataset Creation](#dataset-creation)
137
+ - [Curation Rationale](#curation-rationale)
138
+ - [Source Data](#source-data)
139
+ - [Annotations](#annotations)
140
+ - [Personal and Sensitive Information](#personal-and-sensitive-information)
141
+ - [Considerations for Using the Data](#considerations-for-using-the-data)
142
+ - [Social Impact of Dataset](#social-impact-of-dataset)
143
+ - [Discussion of Biases](#discussion-of-biases)
144
+ - [Other Known Limitations](#other-known-limitations)
145
+ - [Additional Information](#additional-information)
146
+ - [Dataset Curators](#dataset-curators)
147
+ - [Licensing Information](#licensing-information)
148
+ - [Citation Information](#citation-information)
149
+ - [Contributions](#contributions)
150
+
151
+ ## Dataset Description
152
+
153
+ - **Homepage:** [https://blog.einstein.ai/the-wikitext-long-term-dependency-language-modeling-dataset/](https://blog.einstein.ai/the-wikitext-long-term-dependency-language-modeling-dataset/)
154
+ - **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
155
+ - **Paper:** [Pointer Sentinel Mixture Models](https://arxiv.org/abs/1609.07843)
156
+ - **Point of Contact:** [Stephen Merity](mailto:smerity@salesforce.com)
157
+ - **Size of downloaded dataset files:** 391.41 MB
158
+ - **Size of the generated dataset:** 1.12 GB
159
+ - **Total amount of disk used:** 1.52 GB
160
+
161
+ ### Dataset Summary
162
+
163
+ The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
164
+ Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.
165
+
166
+ Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over
167
+ 110 times larger. The WikiText dataset also features a far larger vocabulary and retains the original case, punctuation
168
+ and numbers - all of which are removed in PTB. As it is composed of full articles, the dataset is well suited for models
169
+ that can take advantage of long term dependencies.
170
+
171
+ Each subset comes in two different variants:
172
+ - Raw (for character level work) contain the raw tokens, before the addition of the <unk> (unknown) tokens.
173
+ - Non-raw (for word level work) contain only the tokens in their vocabulary (wiki.train.tokens, wiki.valid.tokens, and wiki.test.tokens).
174
+ The out-of-vocabulary tokens have been replaced with the the <unk> token.
175
+
176
+
177
+ ### Supported Tasks and Leaderboards
178
+
179
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
180
+
181
+ ### Languages
182
+
183
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
184
+
185
+ ## Dataset Structure
186
+
187
+ ### Data Instances
188
+
189
+ #### wikitext-103-raw-v1
190
+
191
+ - **Size of downloaded dataset files:** 191.98 MB
192
+ - **Size of the generated dataset:** 549.42 MB
193
+ - **Total amount of disk used:** 741.41 MB
194
+
195
+ An example of 'validation' looks as follows.
196
+ ```
197
+ This example was too long and was cropped:
198
+
199
+ {
200
+ "text": "\" The gold dollar or gold one @-@ dollar piece was a coin struck as a regular issue by the United States Bureau of the Mint from..."
201
+ }
202
+ ```
203
+
204
+ #### wikitext-103-v1
205
+
206
+ - **Size of downloaded dataset files:** 190.23 MB
207
+ - **Size of the generated dataset:** 548.05 MB
208
+ - **Total amount of disk used:** 738.27 MB
209
+
210
+ An example of 'train' looks as follows.
211
+ ```
212
+ This example was too long and was cropped:
213
+
214
+ {
215
+ "text": "\" Senjō no Valkyria 3 : <unk> Chronicles ( Japanese : 戦場のヴァルキュリア3 , lit . Valkyria of the Battlefield 3 ) , commonly referred to..."
216
+ }
217
+ ```
218
+
219
+ #### wikitext-2-raw-v1
220
+
221
+ - **Size of downloaded dataset files:** 4.72 MB
222
+ - **Size of the generated dataset:** 13.54 MB
223
+ - **Total amount of disk used:** 18.26 MB
224
+
225
+ An example of 'train' looks as follows.
226
+ ```
227
+ This example was too long and was cropped:
228
+
229
+ {
230
+ "text": "\" The Sinclair Scientific Programmable was introduced in 1975 , with the same case as the Sinclair Oxford . It was larger than t..."
231
+ }
232
+ ```
233
+
234
+ #### wikitext-2-v1
235
+
236
+ - **Size of downloaded dataset files:** 4.48 MB
237
+ - **Size of the generated dataset:** 13.34 MB
238
+ - **Total amount of disk used:** 17.82 MB
239
+
240
+ An example of 'train' looks as follows.
241
+ ```
242
+ This example was too long and was cropped:
243
+
244
+ {
245
+ "text": "\" Senjō no Valkyria 3 : <unk> Chronicles ( Japanese : 戦場のヴァルキュリア3 , lit . Valkyria of the Battlefield 3 ) , commonly referred to..."
246
+ }
247
+ ```
248
+
249
+ ### Data Fields
250
+
251
+ The data fields are the same among all splits.
252
+
253
+ #### wikitext-103-raw-v1
254
+ - `text`: a `string` feature.
255
+
256
+ #### wikitext-103-v1
257
+ - `text`: a `string` feature.
258
+
259
+ #### wikitext-2-raw-v1
260
+ - `text`: a `string` feature.
261
+
262
+ #### wikitext-2-v1
263
+ - `text`: a `string` feature.
264
+
265
+ ### Data Splits
266
+
267
+ | name | train |validation|test|
268
+ |-------------------|------:|---------:|---:|
269
+ |wikitext-103-raw-v1|1801350| 3760|4358|
270
+ |wikitext-103-v1 |1801350| 3760|4358|
271
+ |wikitext-2-raw-v1 | 36718| 3760|4358|
272
+ |wikitext-2-v1 | 36718| 3760|4358|
273
+
274
+ ## Dataset Creation
275
+
276
+ ### Curation Rationale
277
+
278
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
279
+
280
+ ### Source Data
281
+
282
+ #### Initial Data Collection and Normalization
283
+
284
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
285
+
286
+ #### Who are the source language producers?
287
+
288
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
289
+
290
+ ### Annotations
291
+
292
+ #### Annotation process
293
+
294
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
295
+
296
+ #### Who are the annotators?
297
+
298
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
299
+
300
+ ### Personal and Sensitive Information
301
+
302
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
303
+
304
+ ## Considerations for Using the Data
305
+
306
+ ### Social Impact of Dataset
307
+
308
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
309
+
310
+ ### Discussion of Biases
311
+
312
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
313
+
314
+ ### Other Known Limitations
315
+
316
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
317
+
318
+ ## Additional Information
319
+
320
+ ### Dataset Curators
321
+
322
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
323
+
324
+ ### Licensing Information
325
+
326
+ The dataset is available under the [Creative Commons Attribution-ShareAlike License (CC BY-SA 4.0)](https://creativecommons.org/licenses/by-sa/4.0/).
327
+
328
+ ### Citation Information
329
+
330
+ ```
331
+ @misc{merity2016pointer,
332
+ title={Pointer Sentinel Mixture Models},
333
+ author={Stephen Merity and Caiming Xiong and James Bradbury and Richard Socher},
334
+ year={2016},
335
+ eprint={1609.07843},
336
+ archivePrefix={arXiv},
337
+ primaryClass={cs.CL}
338
+ }
339
+ ```
340
+
341
+
342
+ ### Contributions
343
+
344
+ Thanks to [@thomwolf](https://github.com/thomwolf), [@lewtun](https://github.com/lewtun), [@patrickvonplaten](https://github.com/patrickvonplaten), [@mariamabarham](https://github.com/mariamabarham) for adding this dataset.
hf_cache/hub/datasets--Salesforce--wikitext/refs/main ADDED
@@ -0,0 +1 @@
 
 
1
+ b08601e04326c79dfdd32d625aee71d232d685c3
hf_cache/hub/datasets--Salesforce--wikitext/snapshots/b08601e04326c79dfdd32d625aee71d232d685c3/README.md ADDED
@@ -0,0 +1,344 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ annotations_creators:
3
+ - no-annotation
4
+ language_creators:
5
+ - crowdsourced
6
+ language:
7
+ - en
8
+ license:
9
+ - cc-by-sa-3.0
10
+ - gfdl
11
+ multilinguality:
12
+ - monolingual
13
+ size_categories:
14
+ - 1M<n<10M
15
+ source_datasets:
16
+ - original
17
+ task_categories:
18
+ - text-generation
19
+ - fill-mask
20
+ task_ids:
21
+ - language-modeling
22
+ - masked-language-modeling
23
+ paperswithcode_id: wikitext-2
24
+ pretty_name: WikiText
25
+ dataset_info:
26
+ - config_name: wikitext-103-raw-v1
27
+ features:
28
+ - name: text
29
+ dtype: string
30
+ splits:
31
+ - name: test
32
+ num_bytes: 1305088
33
+ num_examples: 4358
34
+ - name: train
35
+ num_bytes: 546500949
36
+ num_examples: 1801350
37
+ - name: validation
38
+ num_bytes: 1159288
39
+ num_examples: 3760
40
+ download_size: 315466397
41
+ dataset_size: 548965325
42
+ - config_name: wikitext-103-v1
43
+ features:
44
+ - name: text
45
+ dtype: string
46
+ splits:
47
+ - name: test
48
+ num_bytes: 1295575
49
+ num_examples: 4358
50
+ - name: train
51
+ num_bytes: 545141915
52
+ num_examples: 1801350
53
+ - name: validation
54
+ num_bytes: 1154751
55
+ num_examples: 3760
56
+ download_size: 313093838
57
+ dataset_size: 547592241
58
+ - config_name: wikitext-2-raw-v1
59
+ features:
60
+ - name: text
61
+ dtype: string
62
+ splits:
63
+ - name: test
64
+ num_bytes: 1305088
65
+ num_examples: 4358
66
+ - name: train
67
+ num_bytes: 11061717
68
+ num_examples: 36718
69
+ - name: validation
70
+ num_bytes: 1159288
71
+ num_examples: 3760
72
+ download_size: 7747362
73
+ dataset_size: 13526093
74
+ - config_name: wikitext-2-v1
75
+ features:
76
+ - name: text
77
+ dtype: string
78
+ splits:
79
+ - name: test
80
+ num_bytes: 1270947
81
+ num_examples: 4358
82
+ - name: train
83
+ num_bytes: 10918118
84
+ num_examples: 36718
85
+ - name: validation
86
+ num_bytes: 1134123
87
+ num_examples: 3760
88
+ download_size: 7371282
89
+ dataset_size: 13323188
90
+ configs:
91
+ - config_name: wikitext-103-raw-v1
92
+ data_files:
93
+ - split: test
94
+ path: wikitext-103-raw-v1/test-*
95
+ - split: train
96
+ path: wikitext-103-raw-v1/train-*
97
+ - split: validation
98
+ path: wikitext-103-raw-v1/validation-*
99
+ - config_name: wikitext-103-v1
100
+ data_files:
101
+ - split: test
102
+ path: wikitext-103-v1/test-*
103
+ - split: train
104
+ path: wikitext-103-v1/train-*
105
+ - split: validation
106
+ path: wikitext-103-v1/validation-*
107
+ - config_name: wikitext-2-raw-v1
108
+ data_files:
109
+ - split: test
110
+ path: wikitext-2-raw-v1/test-*
111
+ - split: train
112
+ path: wikitext-2-raw-v1/train-*
113
+ - split: validation
114
+ path: wikitext-2-raw-v1/validation-*
115
+ - config_name: wikitext-2-v1
116
+ data_files:
117
+ - split: test
118
+ path: wikitext-2-v1/test-*
119
+ - split: train
120
+ path: wikitext-2-v1/train-*
121
+ - split: validation
122
+ path: wikitext-2-v1/validation-*
123
+ ---
124
+
125
+ # Dataset Card for "wikitext"
126
+
127
+ ## Table of Contents
128
+ - [Dataset Description](#dataset-description)
129
+ - [Dataset Summary](#dataset-summary)
130
+ - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
131
+ - [Languages](#languages)
132
+ - [Dataset Structure](#dataset-structure)
133
+ - [Data Instances](#data-instances)
134
+ - [Data Fields](#data-fields)
135
+ - [Data Splits](#data-splits)
136
+ - [Dataset Creation](#dataset-creation)
137
+ - [Curation Rationale](#curation-rationale)
138
+ - [Source Data](#source-data)
139
+ - [Annotations](#annotations)
140
+ - [Personal and Sensitive Information](#personal-and-sensitive-information)
141
+ - [Considerations for Using the Data](#considerations-for-using-the-data)
142
+ - [Social Impact of Dataset](#social-impact-of-dataset)
143
+ - [Discussion of Biases](#discussion-of-biases)
144
+ - [Other Known Limitations](#other-known-limitations)
145
+ - [Additional Information](#additional-information)
146
+ - [Dataset Curators](#dataset-curators)
147
+ - [Licensing Information](#licensing-information)
148
+ - [Citation Information](#citation-information)
149
+ - [Contributions](#contributions)
150
+
151
+ ## Dataset Description
152
+
153
+ - **Homepage:** [https://blog.einstein.ai/the-wikitext-long-term-dependency-language-modeling-dataset/](https://blog.einstein.ai/the-wikitext-long-term-dependency-language-modeling-dataset/)
154
+ - **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
155
+ - **Paper:** [Pointer Sentinel Mixture Models](https://arxiv.org/abs/1609.07843)
156
+ - **Point of Contact:** [Stephen Merity](mailto:smerity@salesforce.com)
157
+ - **Size of downloaded dataset files:** 391.41 MB
158
+ - **Size of the generated dataset:** 1.12 GB
159
+ - **Total amount of disk used:** 1.52 GB
160
+
161
+ ### Dataset Summary
162
+
163
+ The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
164
+ Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.
165
+
166
+ Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over
167
+ 110 times larger. The WikiText dataset also features a far larger vocabulary and retains the original case, punctuation
168
+ and numbers - all of which are removed in PTB. As it is composed of full articles, the dataset is well suited for models
169
+ that can take advantage of long term dependencies.
170
+
171
+ Each subset comes in two different variants:
172
+ - Raw (for character level work) contain the raw tokens, before the addition of the <unk> (unknown) tokens.
173
+ - Non-raw (for word level work) contain only the tokens in their vocabulary (wiki.train.tokens, wiki.valid.tokens, and wiki.test.tokens).
174
+ The out-of-vocabulary tokens have been replaced with the the <unk> token.
175
+
176
+
177
+ ### Supported Tasks and Leaderboards
178
+
179
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
180
+
181
+ ### Languages
182
+
183
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
184
+
185
+ ## Dataset Structure
186
+
187
+ ### Data Instances
188
+
189
+ #### wikitext-103-raw-v1
190
+
191
+ - **Size of downloaded dataset files:** 191.98 MB
192
+ - **Size of the generated dataset:** 549.42 MB
193
+ - **Total amount of disk used:** 741.41 MB
194
+
195
+ An example of 'validation' looks as follows.
196
+ ```
197
+ This example was too long and was cropped:
198
+
199
+ {
200
+ "text": "\" The gold dollar or gold one @-@ dollar piece was a coin struck as a regular issue by the United States Bureau of the Mint from..."
201
+ }
202
+ ```
203
+
204
+ #### wikitext-103-v1
205
+
206
+ - **Size of downloaded dataset files:** 190.23 MB
207
+ - **Size of the generated dataset:** 548.05 MB
208
+ - **Total amount of disk used:** 738.27 MB
209
+
210
+ An example of 'train' looks as follows.
211
+ ```
212
+ This example was too long and was cropped:
213
+
214
+ {
215
+ "text": "\" Senjō no Valkyria 3 : <unk> Chronicles ( Japanese : 戦場のヴァルキュリア3 , lit . Valkyria of the Battlefield 3 ) , commonly referred to..."
216
+ }
217
+ ```
218
+
219
+ #### wikitext-2-raw-v1
220
+
221
+ - **Size of downloaded dataset files:** 4.72 MB
222
+ - **Size of the generated dataset:** 13.54 MB
223
+ - **Total amount of disk used:** 18.26 MB
224
+
225
+ An example of 'train' looks as follows.
226
+ ```
227
+ This example was too long and was cropped:
228
+
229
+ {
230
+ "text": "\" The Sinclair Scientific Programmable was introduced in 1975 , with the same case as the Sinclair Oxford . It was larger than t..."
231
+ }
232
+ ```
233
+
234
+ #### wikitext-2-v1
235
+
236
+ - **Size of downloaded dataset files:** 4.48 MB
237
+ - **Size of the generated dataset:** 13.34 MB
238
+ - **Total amount of disk used:** 17.82 MB
239
+
240
+ An example of 'train' looks as follows.
241
+ ```
242
+ This example was too long and was cropped:
243
+
244
+ {
245
+ "text": "\" Senjō no Valkyria 3 : <unk> Chronicles ( Japanese : 戦場のヴァルキュリア3 , lit . Valkyria of the Battlefield 3 ) , commonly referred to..."
246
+ }
247
+ ```
248
+
249
+ ### Data Fields
250
+
251
+ The data fields are the same among all splits.
252
+
253
+ #### wikitext-103-raw-v1
254
+ - `text`: a `string` feature.
255
+
256
+ #### wikitext-103-v1
257
+ - `text`: a `string` feature.
258
+
259
+ #### wikitext-2-raw-v1
260
+ - `text`: a `string` feature.
261
+
262
+ #### wikitext-2-v1
263
+ - `text`: a `string` feature.
264
+
265
+ ### Data Splits
266
+
267
+ | name | train |validation|test|
268
+ |-------------------|------:|---------:|---:|
269
+ |wikitext-103-raw-v1|1801350| 3760|4358|
270
+ |wikitext-103-v1 |1801350| 3760|4358|
271
+ |wikitext-2-raw-v1 | 36718| 3760|4358|
272
+ |wikitext-2-v1 | 36718| 3760|4358|
273
+
274
+ ## Dataset Creation
275
+
276
+ ### Curation Rationale
277
+
278
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
279
+
280
+ ### Source Data
281
+
282
+ #### Initial Data Collection and Normalization
283
+
284
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
285
+
286
+ #### Who are the source language producers?
287
+
288
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
289
+
290
+ ### Annotations
291
+
292
+ #### Annotation process
293
+
294
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
295
+
296
+ #### Who are the annotators?
297
+
298
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
299
+
300
+ ### Personal and Sensitive Information
301
+
302
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
303
+
304
+ ## Considerations for Using the Data
305
+
306
+ ### Social Impact of Dataset
307
+
308
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
309
+
310
+ ### Discussion of Biases
311
+
312
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
313
+
314
+ ### Other Known Limitations
315
+
316
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
317
+
318
+ ## Additional Information
319
+
320
+ ### Dataset Curators
321
+
322
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
323
+
324
+ ### Licensing Information
325
+
326
+ The dataset is available under the [Creative Commons Attribution-ShareAlike License (CC BY-SA 4.0)](https://creativecommons.org/licenses/by-sa/4.0/).
327
+
328
+ ### Citation Information
329
+
330
+ ```
331
+ @misc{merity2016pointer,
332
+ title={Pointer Sentinel Mixture Models},
333
+ author={Stephen Merity and Caiming Xiong and James Bradbury and Richard Socher},
334
+ year={2016},
335
+ eprint={1609.07843},
336
+ archivePrefix={arXiv},
337
+ primaryClass={cs.CL}
338
+ }
339
+ ```
340
+
341
+
342
+ ### Contributions
343
+
344
+ Thanks to [@thomwolf](https://github.com/thomwolf), [@lewtun](https://github.com/lewtun), [@patrickvonplaten](https://github.com/patrickvonplaten), [@mariamabarham](https://github.com/mariamabarham) for adding this dataset.
hf_cache/hub/datasets--abisee--cnn_dailymail/.no_exist/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/.huggingface.yaml ADDED
File without changes
hf_cache/hub/datasets--abisee--cnn_dailymail/.no_exist/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/cnn_dailymail.py ADDED
File without changes
hf_cache/hub/datasets--abisee--cnn_dailymail/.no_exist/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/dataset_infos.json ADDED
File without changes
hf_cache/hub/datasets--abisee--cnn_dailymail/blobs/feadf7d245b4f6818e20b4cf65841d03b5703d47 ADDED
@@ -0,0 +1,305 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ annotations_creators:
3
+ - no-annotation
4
+ language_creators:
5
+ - found
6
+ language:
7
+ - en
8
+ license:
9
+ - apache-2.0
10
+ multilinguality:
11
+ - monolingual
12
+ size_categories:
13
+ - 100K<n<1M
14
+ source_datasets:
15
+ - original
16
+ task_categories:
17
+ - summarization
18
+ task_ids:
19
+ - news-articles-summarization
20
+ paperswithcode_id: cnn-daily-mail-1
21
+ pretty_name: CNN / Daily Mail
22
+ dataset_info:
23
+ - config_name: 1.0.0
24
+ features:
25
+ - name: article
26
+ dtype: string
27
+ - name: highlights
28
+ dtype: string
29
+ - name: id
30
+ dtype: string
31
+ splits:
32
+ - name: train
33
+ num_bytes: 1261703785
34
+ num_examples: 287113
35
+ - name: validation
36
+ num_bytes: 57732412
37
+ num_examples: 13368
38
+ - name: test
39
+ num_bytes: 49925732
40
+ num_examples: 11490
41
+ download_size: 836927248
42
+ dataset_size: 1369361929
43
+ - config_name: 2.0.0
44
+ features:
45
+ - name: article
46
+ dtype: string
47
+ - name: highlights
48
+ dtype: string
49
+ - name: id
50
+ dtype: string
51
+ splits:
52
+ - name: train
53
+ num_bytes: 1261703785
54
+ num_examples: 287113
55
+ - name: validation
56
+ num_bytes: 57732412
57
+ num_examples: 13368
58
+ - name: test
59
+ num_bytes: 49925732
60
+ num_examples: 11490
61
+ download_size: 837094602
62
+ dataset_size: 1369361929
63
+ - config_name: 3.0.0
64
+ features:
65
+ - name: article
66
+ dtype: string
67
+ - name: highlights
68
+ dtype: string
69
+ - name: id
70
+ dtype: string
71
+ splits:
72
+ - name: train
73
+ num_bytes: 1261703785
74
+ num_examples: 287113
75
+ - name: validation
76
+ num_bytes: 57732412
77
+ num_examples: 13368
78
+ - name: test
79
+ num_bytes: 49925732
80
+ num_examples: 11490
81
+ download_size: 837094602
82
+ dataset_size: 1369361929
83
+ configs:
84
+ - config_name: 1.0.0
85
+ data_files:
86
+ - split: train
87
+ path: 1.0.0/train-*
88
+ - split: validation
89
+ path: 1.0.0/validation-*
90
+ - split: test
91
+ path: 1.0.0/test-*
92
+ - config_name: 2.0.0
93
+ data_files:
94
+ - split: train
95
+ path: 2.0.0/train-*
96
+ - split: validation
97
+ path: 2.0.0/validation-*
98
+ - split: test
99
+ path: 2.0.0/test-*
100
+ - config_name: 3.0.0
101
+ data_files:
102
+ - split: train
103
+ path: 3.0.0/train-*
104
+ - split: validation
105
+ path: 3.0.0/validation-*
106
+ - split: test
107
+ path: 3.0.0/test-*
108
+ train-eval-index:
109
+ - config: 3.0.0
110
+ task: summarization
111
+ task_id: summarization
112
+ splits:
113
+ eval_split: test
114
+ col_mapping:
115
+ article: text
116
+ highlights: target
117
+ ---
118
+ # Dataset Card for CNN Dailymail Dataset
119
+
120
+ ## Table of Contents
121
+ - [Dataset Description](#dataset-description)
122
+ - [Dataset Summary](#dataset-summary)
123
+ - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
124
+ - [Languages](#languages)
125
+ - [Dataset Structure](#dataset-structure)
126
+ - [Data Instances](#data-instances)
127
+ - [Data Fields](#data-fields)
128
+ - [Data Splits](#data-splits)
129
+ - [Dataset Creation](#dataset-creation)
130
+ - [Curation Rationale](#curation-rationale)
131
+ - [Source Data](#source-data)
132
+ - [Annotations](#annotations)
133
+ - [Personal and Sensitive Information](#personal-and-sensitive-information)
134
+ - [Considerations for Using the Data](#considerations-for-using-the-data)
135
+ - [Social Impact of Dataset](#social-impact-of-dataset)
136
+ - [Discussion of Biases](#discussion-of-biases)
137
+ - [Other Known Limitations](#other-known-limitations)
138
+ - [Additional Information](#additional-information)
139
+ - [Dataset Curators](#dataset-curators)
140
+ - [Licensing Information](#licensing-information)
141
+ - [Citation Information](#citation-information)
142
+ - [Contributions](#contributions)
143
+
144
+ ## Dataset Description
145
+
146
+ - **Homepage:**
147
+ - **Repository:** [CNN / DailyMail Dataset repository](https://github.com/abisee/cnn-dailymail)
148
+ - **Paper:** [Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond](https://papers.nips.cc/paper/5945-teaching-machines-to-read-and-comprehend.pdf), [Get To The Point: Summarization with Pointer-Generator Networks](https://www.aclweb.org/anthology/K16-1028.pdf)
149
+ - **Leaderboard:** [Papers with Code leaderboard for CNN / Dailymail Dataset](https://paperswithcode.com/sota/document-summarization-on-cnn-daily-mail)
150
+ - **Point of Contact:** [Abigail See](mailto:abisee@stanford.edu)
151
+
152
+ ### Dataset Summary
153
+
154
+ The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering.
155
+
156
+ ### Supported Tasks and Leaderboards
157
+
158
+ - 'summarization': [Versions 2.0.0 and 3.0.0 of the CNN / DailyMail Dataset](https://www.aclweb.org/anthology/K16-1028.pdf) can be used to train a model for abstractive and extractive summarization ([Version 1.0.0](https://papers.nips.cc/paper/5945-teaching-machines-to-read-and-comprehend.pdf) was developed for machine reading and comprehension and abstractive question answering). The model performance is measured by how high the output summary's [ROUGE](https://huggingface.co/metrics/rouge) score for a given article is when compared to the highlight as written by the original article author. [Zhong et al (2020)](https://www.aclweb.org/anthology/2020.acl-main.552.pdf) report a ROUGE-1 score of 44.41 when testing a model trained for extractive summarization. See the [Papers With Code leaderboard](https://paperswithcode.com/sota/document-summarization-on-cnn-daily-mail) for more models.
159
+
160
+ ### Languages
161
+
162
+ The BCP-47 code for English as generally spoken in the United States is en-US and the BCP-47 code for English as generally spoken in the United Kingdom is en-GB. It is unknown if other varieties of English are represented in the data.
163
+
164
+ ## Dataset Structure
165
+
166
+ ### Data Instances
167
+
168
+ For each instance, there is a string for the article, a string for the highlights, and a string for the id. See the [CNN / Daily Mail dataset viewer](https://huggingface.co/datasets/viewer/?dataset=cnn_dailymail&config=3.0.0) to explore more examples.
169
+
170
+ ```
171
+ {'id': '0054d6d30dbcad772e20b22771153a2a9cbeaf62',
172
+ 'article': '(CNN) -- An American woman died aboard a cruise ship that docked at Rio de Janeiro on Tuesday, the same ship on which 86 passengers previously fell ill, according to the state-run Brazilian news agency, Agencia Brasil. The American tourist died aboard the MS Veendam, owned by cruise operator Holland America. Federal Police told Agencia Brasil that forensic doctors were investigating her death. The ship's doctors told police that the woman was elderly and suffered from diabetes and hypertension, according the agency. The other passengers came down with diarrhea prior to her death during an earlier part of the trip, the ship's doctors said. The Veendam left New York 36 days ago for a South America tour.'
173
+ 'highlights': 'The elderly woman suffered from diabetes and hypertension, ship's doctors say .\nPreviously, 86 passengers had fallen ill on the ship, Agencia Brasil says .'}
174
+ ```
175
+
176
+ The average token count for the articles and the highlights are provided below:
177
+
178
+ | Feature | Mean Token Count |
179
+ | ---------- | ---------------- |
180
+ | Article | 781 |
181
+ | Highlights | 56 |
182
+
183
+ ### Data Fields
184
+
185
+ - `id`: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from
186
+ - `article`: a string containing the body of the news article
187
+ - `highlights`: a string containing the highlight of the article as written by the article author
188
+
189
+ ### Data Splits
190
+
191
+ The CNN/DailyMail dataset has 3 splits: _train_, _validation_, and _test_. Below are the statistics for Version 3.0.0 of the dataset.
192
+
193
+ | Dataset Split | Number of Instances in Split |
194
+ | ------------- | ------------------------------------------- |
195
+ | Train | 287,113 |
196
+ | Validation | 13,368 |
197
+ | Test | 11,490 |
198
+
199
+ ## Dataset Creation
200
+
201
+ ### Curation Rationale
202
+
203
+ Version 1.0.0 aimed to support supervised neural methodologies for machine reading and question answering with a large amount of real natural language training data and released about 313k unique articles and nearly 1M Cloze style questions to go with the articles. Versions 2.0.0 and 3.0.0 changed the structure of the dataset to support summarization rather than question answering. Version 3.0.0 provided a non-anonymized version of the data, whereas both the previous versions were preprocessed to replace named entities with unique identifier labels.
204
+
205
+ ### Source Data
206
+
207
+ #### Initial Data Collection and Normalization
208
+
209
+ The data consists of news articles and highlight sentences. In the question answering setting of the data, the articles are used as the context and entities are hidden one at a time in the highlight sentences, producing Cloze style questions where the goal of the model is to correctly guess which entity in the context has been hidden in the highlight. In the summarization setting, the highlight sentences are concatenated to form a summary of the article. The CNN articles were written between April 2007 and April 2015. The Daily Mail articles were written between June 2010 and April 2015.
210
+
211
+ The code for the original data collection is available at <https://github.com/deepmind/rc-data>. The articles were downloaded using archives of <www.cnn.com> and <www.dailymail.co.uk> on the Wayback Machine. Articles were not included in the Version 1.0.0 collection if they exceeded 2000 tokens. Due to accessibility issues with the Wayback Machine, Kyunghyun Cho has made the datasets available at <https://cs.nyu.edu/~kcho/DMQA/>. An updated version of the code that does not anonymize the data is available at <https://github.com/abisee/cnn-dailymail>.
212
+
213
+ Hermann et al provided their own tokenization script. The script provided by See uses the PTBTokenizer. It also lowercases the text and adds periods to lines missing them.
214
+
215
+ #### Who are the source language producers?
216
+
217
+ The text was written by journalists at CNN and the Daily Mail.
218
+
219
+ ### Annotations
220
+
221
+ The dataset does not contain any additional annotations.
222
+
223
+ #### Annotation process
224
+
225
+ [N/A]
226
+
227
+ #### Who are the annotators?
228
+
229
+ [N/A]
230
+
231
+ ### Personal and Sensitive Information
232
+
233
+ Version 3.0 is not anonymized, so individuals' names can be found in the dataset. Information about the original author is not included in the dataset.
234
+
235
+ ## Considerations for Using the Data
236
+
237
+ ### Social Impact of Dataset
238
+
239
+ The purpose of this dataset is to help develop models that can summarize long paragraphs of text in one or two sentences.
240
+
241
+ This task is useful for efficiently presenting information given a large quantity of text. It should be made clear that any summarizations produced by models trained on this dataset are reflective of the language used in the articles, but are in fact automatically generated.
242
+
243
+ ### Discussion of Biases
244
+
245
+ [Bordia and Bowman (2019)](https://www.aclweb.org/anthology/N19-3002.pdf) explore measuring gender bias and debiasing techniques in the CNN / Dailymail dataset, the Penn Treebank, and WikiText-2. They find the CNN / Dailymail dataset to have a slightly lower gender bias based on their metric compared to the other datasets, but still show evidence of gender bias when looking at words such as 'fragile'.
246
+
247
+ Because the articles were written by and for people in the US and the UK, they will likely present specifically US and UK perspectives and feature events that are considered relevant to those populations during the time that the articles were published.
248
+
249
+ ### Other Known Limitations
250
+
251
+ News articles have been shown to conform to writing conventions in which important information is primarily presented in the first third of the article [(Kryściński et al, 2019)](https://www.aclweb.org/anthology/D19-1051.pdf). [Chen et al (2016)](https://www.aclweb.org/anthology/P16-1223.pdf) conducted a manual study of 100 random instances of the first version of the dataset and found 25% of the samples to be difficult even for humans to answer correctly due to ambiguity and coreference errors.
252
+
253
+ It should also be noted that machine-generated summarizations, even when extractive, may differ in truth values when compared to the original articles.
254
+
255
+ ## Additional Information
256
+
257
+ ### Dataset Curators
258
+
259
+ The data was originally collected by Karl Moritz Hermann, Tomáš Kočiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom of Google DeepMind. Tomáš Kočiský and Phil Blunsom are also affiliated with the University of Oxford. They released scripts to collect and process the data into the question answering format.
260
+
261
+ Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, and Bing Xiang of IMB Watson and Çağlar Gu̇lçehre of Université de Montréal modified Hermann et al's collection scripts to restore the data to a summary format. They also produced both anonymized and non-anonymized versions.
262
+
263
+ The code for the non-anonymized version is made publicly available by Abigail See of Stanford University, Peter J. Liu of Google Brain and Christopher D. Manning of Stanford University at <https://github.com/abisee/cnn-dailymail>. The work at Stanford University was supported by the DARPA DEFT ProgramAFRL contract no. FA8750-13-2-0040.
264
+
265
+ ### Licensing Information
266
+
267
+ The CNN / Daily Mail dataset version 1.0.0 is released under the [Apache-2.0 License](http://www.apache.org/licenses/LICENSE-2.0).
268
+
269
+ ### Citation Information
270
+
271
+ ```
272
+ @inproceedings{see-etal-2017-get,
273
+ title = "Get To The Point: Summarization with Pointer-Generator Networks",
274
+ author = "See, Abigail and
275
+ Liu, Peter J. and
276
+ Manning, Christopher D.",
277
+ booktitle = "Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
278
+ month = jul,
279
+ year = "2017",
280
+ address = "Vancouver, Canada",
281
+ publisher = "Association for Computational Linguistics",
282
+ url = "https://www.aclweb.org/anthology/P17-1099",
283
+ doi = "10.18653/v1/P17-1099",
284
+ pages = "1073--1083",
285
+ abstract = "Neural sequence-to-sequence models have provided a viable new approach for abstractive text summarization (meaning they are not restricted to simply selecting and rearranging passages from the original text). However, these models have two shortcomings: they are liable to reproduce factual details inaccurately, and they tend to repeat themselves. In this work we propose a novel architecture that augments the standard sequence-to-sequence attentional model in two orthogonal ways. First, we use a hybrid pointer-generator network that can copy words from the source text via pointing, which aids accurate reproduction of information, while retaining the ability to produce novel words through the generator. Second, we use coverage to keep track of what has been summarized, which discourages repetition. We apply our model to the CNN / Daily Mail summarization task, outperforming the current abstractive state-of-the-art by at least 2 ROUGE points.",
286
+ }
287
+ ```
288
+
289
+ ```
290
+ @inproceedings{DBLP:conf/nips/HermannKGEKSB15,
291
+ author={Karl Moritz Hermann and Tomás Kociský and Edward Grefenstette and Lasse Espeholt and Will Kay and Mustafa Suleyman and Phil Blunsom},
292
+ title={Teaching Machines to Read and Comprehend},
293
+ year={2015},
294
+ cdate={1420070400000},
295
+ pages={1693-1701},
296
+ url={http://papers.nips.cc/paper/5945-teaching-machines-to-read-and-comprehend},
297
+ booktitle={NIPS},
298
+ crossref={conf/nips/2015}
299
+ }
300
+
301
+ ```
302
+
303
+ ### Contributions
304
+
305
+ Thanks to [@thomwolf](https://github.com/thomwolf), [@lewtun](https://github.com/lewtun), [@jplu](https://github.com/jplu), [@jbragg](https://github.com/jbragg), [@patrickvonplaten](https://github.com/patrickvonplaten) and [@mcmillanmajora](https://github.com/mcmillanmajora) for adding this dataset.
hf_cache/hub/datasets--abisee--cnn_dailymail/refs/main ADDED
@@ -0,0 +1 @@
 
 
1
+ 96df5e686bee6baa90b8bee7c28b81fa3fa6223d
hf_cache/hub/datasets--abisee--cnn_dailymail/snapshots/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/README.md ADDED
@@ -0,0 +1,305 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ annotations_creators:
3
+ - no-annotation
4
+ language_creators:
5
+ - found
6
+ language:
7
+ - en
8
+ license:
9
+ - apache-2.0
10
+ multilinguality:
11
+ - monolingual
12
+ size_categories:
13
+ - 100K<n<1M
14
+ source_datasets:
15
+ - original
16
+ task_categories:
17
+ - summarization
18
+ task_ids:
19
+ - news-articles-summarization
20
+ paperswithcode_id: cnn-daily-mail-1
21
+ pretty_name: CNN / Daily Mail
22
+ dataset_info:
23
+ - config_name: 1.0.0
24
+ features:
25
+ - name: article
26
+ dtype: string
27
+ - name: highlights
28
+ dtype: string
29
+ - name: id
30
+ dtype: string
31
+ splits:
32
+ - name: train
33
+ num_bytes: 1261703785
34
+ num_examples: 287113
35
+ - name: validation
36
+ num_bytes: 57732412
37
+ num_examples: 13368
38
+ - name: test
39
+ num_bytes: 49925732
40
+ num_examples: 11490
41
+ download_size: 836927248
42
+ dataset_size: 1369361929
43
+ - config_name: 2.0.0
44
+ features:
45
+ - name: article
46
+ dtype: string
47
+ - name: highlights
48
+ dtype: string
49
+ - name: id
50
+ dtype: string
51
+ splits:
52
+ - name: train
53
+ num_bytes: 1261703785
54
+ num_examples: 287113
55
+ - name: validation
56
+ num_bytes: 57732412
57
+ num_examples: 13368
58
+ - name: test
59
+ num_bytes: 49925732
60
+ num_examples: 11490
61
+ download_size: 837094602
62
+ dataset_size: 1369361929
63
+ - config_name: 3.0.0
64
+ features:
65
+ - name: article
66
+ dtype: string
67
+ - name: highlights
68
+ dtype: string
69
+ - name: id
70
+ dtype: string
71
+ splits:
72
+ - name: train
73
+ num_bytes: 1261703785
74
+ num_examples: 287113
75
+ - name: validation
76
+ num_bytes: 57732412
77
+ num_examples: 13368
78
+ - name: test
79
+ num_bytes: 49925732
80
+ num_examples: 11490
81
+ download_size: 837094602
82
+ dataset_size: 1369361929
83
+ configs:
84
+ - config_name: 1.0.0
85
+ data_files:
86
+ - split: train
87
+ path: 1.0.0/train-*
88
+ - split: validation
89
+ path: 1.0.0/validation-*
90
+ - split: test
91
+ path: 1.0.0/test-*
92
+ - config_name: 2.0.0
93
+ data_files:
94
+ - split: train
95
+ path: 2.0.0/train-*
96
+ - split: validation
97
+ path: 2.0.0/validation-*
98
+ - split: test
99
+ path: 2.0.0/test-*
100
+ - config_name: 3.0.0
101
+ data_files:
102
+ - split: train
103
+ path: 3.0.0/train-*
104
+ - split: validation
105
+ path: 3.0.0/validation-*
106
+ - split: test
107
+ path: 3.0.0/test-*
108
+ train-eval-index:
109
+ - config: 3.0.0
110
+ task: summarization
111
+ task_id: summarization
112
+ splits:
113
+ eval_split: test
114
+ col_mapping:
115
+ article: text
116
+ highlights: target
117
+ ---
118
+ # Dataset Card for CNN Dailymail Dataset
119
+
120
+ ## Table of Contents
121
+ - [Dataset Description](#dataset-description)
122
+ - [Dataset Summary](#dataset-summary)
123
+ - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
124
+ - [Languages](#languages)
125
+ - [Dataset Structure](#dataset-structure)
126
+ - [Data Instances](#data-instances)
127
+ - [Data Fields](#data-fields)
128
+ - [Data Splits](#data-splits)
129
+ - [Dataset Creation](#dataset-creation)
130
+ - [Curation Rationale](#curation-rationale)
131
+ - [Source Data](#source-data)
132
+ - [Annotations](#annotations)
133
+ - [Personal and Sensitive Information](#personal-and-sensitive-information)
134
+ - [Considerations for Using the Data](#considerations-for-using-the-data)
135
+ - [Social Impact of Dataset](#social-impact-of-dataset)
136
+ - [Discussion of Biases](#discussion-of-biases)
137
+ - [Other Known Limitations](#other-known-limitations)
138
+ - [Additional Information](#additional-information)
139
+ - [Dataset Curators](#dataset-curators)
140
+ - [Licensing Information](#licensing-information)
141
+ - [Citation Information](#citation-information)
142
+ - [Contributions](#contributions)
143
+
144
+ ## Dataset Description
145
+
146
+ - **Homepage:**
147
+ - **Repository:** [CNN / DailyMail Dataset repository](https://github.com/abisee/cnn-dailymail)
148
+ - **Paper:** [Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond](https://papers.nips.cc/paper/5945-teaching-machines-to-read-and-comprehend.pdf), [Get To The Point: Summarization with Pointer-Generator Networks](https://www.aclweb.org/anthology/K16-1028.pdf)
149
+ - **Leaderboard:** [Papers with Code leaderboard for CNN / Dailymail Dataset](https://paperswithcode.com/sota/document-summarization-on-cnn-daily-mail)
150
+ - **Point of Contact:** [Abigail See](mailto:abisee@stanford.edu)
151
+
152
+ ### Dataset Summary
153
+
154
+ The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering.
155
+
156
+ ### Supported Tasks and Leaderboards
157
+
158
+ - 'summarization': [Versions 2.0.0 and 3.0.0 of the CNN / DailyMail Dataset](https://www.aclweb.org/anthology/K16-1028.pdf) can be used to train a model for abstractive and extractive summarization ([Version 1.0.0](https://papers.nips.cc/paper/5945-teaching-machines-to-read-and-comprehend.pdf) was developed for machine reading and comprehension and abstractive question answering). The model performance is measured by how high the output summary's [ROUGE](https://huggingface.co/metrics/rouge) score for a given article is when compared to the highlight as written by the original article author. [Zhong et al (2020)](https://www.aclweb.org/anthology/2020.acl-main.552.pdf) report a ROUGE-1 score of 44.41 when testing a model trained for extractive summarization. See the [Papers With Code leaderboard](https://paperswithcode.com/sota/document-summarization-on-cnn-daily-mail) for more models.
159
+
160
+ ### Languages
161
+
162
+ The BCP-47 code for English as generally spoken in the United States is en-US and the BCP-47 code for English as generally spoken in the United Kingdom is en-GB. It is unknown if other varieties of English are represented in the data.
163
+
164
+ ## Dataset Structure
165
+
166
+ ### Data Instances
167
+
168
+ For each instance, there is a string for the article, a string for the highlights, and a string for the id. See the [CNN / Daily Mail dataset viewer](https://huggingface.co/datasets/viewer/?dataset=cnn_dailymail&config=3.0.0) to explore more examples.
169
+
170
+ ```
171
+ {'id': '0054d6d30dbcad772e20b22771153a2a9cbeaf62',
172
+ 'article': '(CNN) -- An American woman died aboard a cruise ship that docked at Rio de Janeiro on Tuesday, the same ship on which 86 passengers previously fell ill, according to the state-run Brazilian news agency, Agencia Brasil. The American tourist died aboard the MS Veendam, owned by cruise operator Holland America. Federal Police told Agencia Brasil that forensic doctors were investigating her death. The ship's doctors told police that the woman was elderly and suffered from diabetes and hypertension, according the agency. The other passengers came down with diarrhea prior to her death during an earlier part of the trip, the ship's doctors said. The Veendam left New York 36 days ago for a South America tour.'
173
+ 'highlights': 'The elderly woman suffered from diabetes and hypertension, ship's doctors say .\nPreviously, 86 passengers had fallen ill on the ship, Agencia Brasil says .'}
174
+ ```
175
+
176
+ The average token count for the articles and the highlights are provided below:
177
+
178
+ | Feature | Mean Token Count |
179
+ | ---------- | ---------------- |
180
+ | Article | 781 |
181
+ | Highlights | 56 |
182
+
183
+ ### Data Fields
184
+
185
+ - `id`: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from
186
+ - `article`: a string containing the body of the news article
187
+ - `highlights`: a string containing the highlight of the article as written by the article author
188
+
189
+ ### Data Splits
190
+
191
+ The CNN/DailyMail dataset has 3 splits: _train_, _validation_, and _test_. Below are the statistics for Version 3.0.0 of the dataset.
192
+
193
+ | Dataset Split | Number of Instances in Split |
194
+ | ------------- | ------------------------------------------- |
195
+ | Train | 287,113 |
196
+ | Validation | 13,368 |
197
+ | Test | 11,490 |
198
+
199
+ ## Dataset Creation
200
+
201
+ ### Curation Rationale
202
+
203
+ Version 1.0.0 aimed to support supervised neural methodologies for machine reading and question answering with a large amount of real natural language training data and released about 313k unique articles and nearly 1M Cloze style questions to go with the articles. Versions 2.0.0 and 3.0.0 changed the structure of the dataset to support summarization rather than question answering. Version 3.0.0 provided a non-anonymized version of the data, whereas both the previous versions were preprocessed to replace named entities with unique identifier labels.
204
+
205
+ ### Source Data
206
+
207
+ #### Initial Data Collection and Normalization
208
+
209
+ The data consists of news articles and highlight sentences. In the question answering setting of the data, the articles are used as the context and entities are hidden one at a time in the highlight sentences, producing Cloze style questions where the goal of the model is to correctly guess which entity in the context has been hidden in the highlight. In the summarization setting, the highlight sentences are concatenated to form a summary of the article. The CNN articles were written between April 2007 and April 2015. The Daily Mail articles were written between June 2010 and April 2015.
210
+
211
+ The code for the original data collection is available at <https://github.com/deepmind/rc-data>. The articles were downloaded using archives of <www.cnn.com> and <www.dailymail.co.uk> on the Wayback Machine. Articles were not included in the Version 1.0.0 collection if they exceeded 2000 tokens. Due to accessibility issues with the Wayback Machine, Kyunghyun Cho has made the datasets available at <https://cs.nyu.edu/~kcho/DMQA/>. An updated version of the code that does not anonymize the data is available at <https://github.com/abisee/cnn-dailymail>.
212
+
213
+ Hermann et al provided their own tokenization script. The script provided by See uses the PTBTokenizer. It also lowercases the text and adds periods to lines missing them.
214
+
215
+ #### Who are the source language producers?
216
+
217
+ The text was written by journalists at CNN and the Daily Mail.
218
+
219
+ ### Annotations
220
+
221
+ The dataset does not contain any additional annotations.
222
+
223
+ #### Annotation process
224
+
225
+ [N/A]
226
+
227
+ #### Who are the annotators?
228
+
229
+ [N/A]
230
+
231
+ ### Personal and Sensitive Information
232
+
233
+ Version 3.0 is not anonymized, so individuals' names can be found in the dataset. Information about the original author is not included in the dataset.
234
+
235
+ ## Considerations for Using the Data
236
+
237
+ ### Social Impact of Dataset
238
+
239
+ The purpose of this dataset is to help develop models that can summarize long paragraphs of text in one or two sentences.
240
+
241
+ This task is useful for efficiently presenting information given a large quantity of text. It should be made clear that any summarizations produced by models trained on this dataset are reflective of the language used in the articles, but are in fact automatically generated.
242
+
243
+ ### Discussion of Biases
244
+
245
+ [Bordia and Bowman (2019)](https://www.aclweb.org/anthology/N19-3002.pdf) explore measuring gender bias and debiasing techniques in the CNN / Dailymail dataset, the Penn Treebank, and WikiText-2. They find the CNN / Dailymail dataset to have a slightly lower gender bias based on their metric compared to the other datasets, but still show evidence of gender bias when looking at words such as 'fragile'.
246
+
247
+ Because the articles were written by and for people in the US and the UK, they will likely present specifically US and UK perspectives and feature events that are considered relevant to those populations during the time that the articles were published.
248
+
249
+ ### Other Known Limitations
250
+
251
+ News articles have been shown to conform to writing conventions in which important information is primarily presented in the first third of the article [(Kryściński et al, 2019)](https://www.aclweb.org/anthology/D19-1051.pdf). [Chen et al (2016)](https://www.aclweb.org/anthology/P16-1223.pdf) conducted a manual study of 100 random instances of the first version of the dataset and found 25% of the samples to be difficult even for humans to answer correctly due to ambiguity and coreference errors.
252
+
253
+ It should also be noted that machine-generated summarizations, even when extractive, may differ in truth values when compared to the original articles.
254
+
255
+ ## Additional Information
256
+
257
+ ### Dataset Curators
258
+
259
+ The data was originally collected by Karl Moritz Hermann, Tomáš Kočiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom of Google DeepMind. Tomáš Kočiský and Phil Blunsom are also affiliated with the University of Oxford. They released scripts to collect and process the data into the question answering format.
260
+
261
+ Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, and Bing Xiang of IMB Watson and Çağlar Gu̇lçehre of Université de Montréal modified Hermann et al's collection scripts to restore the data to a summary format. They also produced both anonymized and non-anonymized versions.
262
+
263
+ The code for the non-anonymized version is made publicly available by Abigail See of Stanford University, Peter J. Liu of Google Brain and Christopher D. Manning of Stanford University at <https://github.com/abisee/cnn-dailymail>. The work at Stanford University was supported by the DARPA DEFT ProgramAFRL contract no. FA8750-13-2-0040.
264
+
265
+ ### Licensing Information
266
+
267
+ The CNN / Daily Mail dataset version 1.0.0 is released under the [Apache-2.0 License](http://www.apache.org/licenses/LICENSE-2.0).
268
+
269
+ ### Citation Information
270
+
271
+ ```
272
+ @inproceedings{see-etal-2017-get,
273
+ title = "Get To The Point: Summarization with Pointer-Generator Networks",
274
+ author = "See, Abigail and
275
+ Liu, Peter J. and
276
+ Manning, Christopher D.",
277
+ booktitle = "Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
278
+ month = jul,
279
+ year = "2017",
280
+ address = "Vancouver, Canada",
281
+ publisher = "Association for Computational Linguistics",
282
+ url = "https://www.aclweb.org/anthology/P17-1099",
283
+ doi = "10.18653/v1/P17-1099",
284
+ pages = "1073--1083",
285
+ abstract = "Neural sequence-to-sequence models have provided a viable new approach for abstractive text summarization (meaning they are not restricted to simply selecting and rearranging passages from the original text). However, these models have two shortcomings: they are liable to reproduce factual details inaccurately, and they tend to repeat themselves. In this work we propose a novel architecture that augments the standard sequence-to-sequence attentional model in two orthogonal ways. First, we use a hybrid pointer-generator network that can copy words from the source text via pointing, which aids accurate reproduction of information, while retaining the ability to produce novel words through the generator. Second, we use coverage to keep track of what has been summarized, which discourages repetition. We apply our model to the CNN / Daily Mail summarization task, outperforming the current abstractive state-of-the-art by at least 2 ROUGE points.",
286
+ }
287
+ ```
288
+
289
+ ```
290
+ @inproceedings{DBLP:conf/nips/HermannKGEKSB15,
291
+ author={Karl Moritz Hermann and Tomás Kociský and Edward Grefenstette and Lasse Espeholt and Will Kay and Mustafa Suleyman and Phil Blunsom},
292
+ title={Teaching Machines to Read and Comprehend},
293
+ year={2015},
294
+ cdate={1420070400000},
295
+ pages={1693-1701},
296
+ url={http://papers.nips.cc/paper/5945-teaching-machines-to-read-and-comprehend},
297
+ booktitle={NIPS},
298
+ crossref={conf/nips/2015}
299
+ }
300
+
301
+ ```
302
+
303
+ ### Contributions
304
+
305
+ Thanks to [@thomwolf](https://github.com/thomwolf), [@lewtun](https://github.com/lewtun), [@jplu](https://github.com/jplu), [@jbragg](https://github.com/jbragg), [@patrickvonplaten](https://github.com/patrickvonplaten) and [@mcmillanmajora](https://github.com/mcmillanmajora) for adding this dataset.
hf_cache/hub/datasets--allenai--openbookqa/.no_exist/388097ea7776314e93a529163e0fea805b8a6454/.huggingface.yaml ADDED
File without changes
hf_cache/hub/datasets--allenai--openbookqa/.no_exist/388097ea7776314e93a529163e0fea805b8a6454/dataset_infos.json ADDED
File without changes
hf_cache/hub/datasets--allenai--openbookqa/.no_exist/388097ea7776314e93a529163e0fea805b8a6454/openbookqa.py ADDED
File without changes
hf_cache/hub/datasets--allenai--openbookqa/blobs/08128898cc0433b97f7a9c9ff09c5054c8587e3e ADDED
@@ -0,0 +1,301 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ annotations_creators:
3
+ - crowdsourced
4
+ - expert-generated
5
+ language_creators:
6
+ - expert-generated
7
+ language:
8
+ - en
9
+ license:
10
+ - unknown
11
+ multilinguality:
12
+ - monolingual
13
+ size_categories:
14
+ - 1K<n<10K
15
+ source_datasets:
16
+ - original
17
+ task_categories:
18
+ - question-answering
19
+ task_ids:
20
+ - open-domain-qa
21
+ paperswithcode_id: openbookqa
22
+ pretty_name: OpenBookQA
23
+ dataset_info:
24
+ - config_name: additional
25
+ features:
26
+ - name: id
27
+ dtype: string
28
+ - name: question_stem
29
+ dtype: string
30
+ - name: choices
31
+ sequence:
32
+ - name: text
33
+ dtype: string
34
+ - name: label
35
+ dtype: string
36
+ - name: answerKey
37
+ dtype: string
38
+ - name: fact1
39
+ dtype: string
40
+ - name: humanScore
41
+ dtype: float32
42
+ - name: clarity
43
+ dtype: float32
44
+ - name: turkIdAnonymized
45
+ dtype: string
46
+ splits:
47
+ - name: train
48
+ num_bytes: 1288577
49
+ num_examples: 4957
50
+ - name: validation
51
+ num_bytes: 135916
52
+ num_examples: 500
53
+ - name: test
54
+ num_bytes: 130701
55
+ num_examples: 500
56
+ download_size: 783789
57
+ dataset_size: 1555194
58
+ - config_name: main
59
+ features:
60
+ - name: id
61
+ dtype: string
62
+ - name: question_stem
63
+ dtype: string
64
+ - name: choices
65
+ sequence:
66
+ - name: text
67
+ dtype: string
68
+ - name: label
69
+ dtype: string
70
+ - name: answerKey
71
+ dtype: string
72
+ splits:
73
+ - name: train
74
+ num_bytes: 895386
75
+ num_examples: 4957
76
+ - name: validation
77
+ num_bytes: 95428
78
+ num_examples: 500
79
+ - name: test
80
+ num_bytes: 91759
81
+ num_examples: 500
82
+ download_size: 609613
83
+ dataset_size: 1082573
84
+ configs:
85
+ - config_name: additional
86
+ data_files:
87
+ - split: train
88
+ path: additional/train-*
89
+ - split: validation
90
+ path: additional/validation-*
91
+ - split: test
92
+ path: additional/test-*
93
+ - config_name: main
94
+ data_files:
95
+ - split: train
96
+ path: main/train-*
97
+ - split: validation
98
+ path: main/validation-*
99
+ - split: test
100
+ path: main/test-*
101
+ default: true
102
+ ---
103
+
104
+ # Dataset Card for OpenBookQA
105
+
106
+ ## Table of Contents
107
+ - [Dataset Description](#dataset-description)
108
+ - [Dataset Summary](#dataset-summary)
109
+ - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
110
+ - [Languages](#languages)
111
+ - [Dataset Structure](#dataset-structure)
112
+ - [Data Instances](#data-instances)
113
+ - [Data Fields](#data-fields)
114
+ - [Data Splits](#data-splits)
115
+ - [Dataset Creation](#dataset-creation)
116
+ - [Curation Rationale](#curation-rationale)
117
+ - [Source Data](#source-data)
118
+ - [Annotations](#annotations)
119
+ - [Personal and Sensitive Information](#personal-and-sensitive-information)
120
+ - [Considerations for Using the Data](#considerations-for-using-the-data)
121
+ - [Social Impact of Dataset](#social-impact-of-dataset)
122
+ - [Discussion of Biases](#discussion-of-biases)
123
+ - [Other Known Limitations](#other-known-limitations)
124
+ - [Additional Information](#additional-information)
125
+ - [Dataset Curators](#dataset-curators)
126
+ - [Licensing Information](#licensing-information)
127
+ - [Citation Information](#citation-information)
128
+ - [Contributions](#contributions)
129
+
130
+ ## Dataset Description
131
+
132
+ - **Homepage:** [https://allenai.org/data/open-book-qa](https://allenai.org/data/open-book-qa)
133
+ - **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
134
+ - **Paper:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
135
+ - **Point of Contact:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
136
+ - **Size of downloaded dataset files:** 2.89 MB
137
+ - **Size of the generated dataset:** 2.88 MB
138
+ - **Total amount of disk used:** 5.78 MB
139
+
140
+ ### Dataset Summary
141
+
142
+ OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic
143
+ (with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In
144
+ particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge,
145
+ and rich text comprehension.
146
+ OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of
147
+ a subject.
148
+
149
+ ### Supported Tasks and Leaderboards
150
+
151
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
152
+
153
+ ### Languages
154
+
155
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
156
+
157
+ ## Dataset Structure
158
+
159
+ ### Data Instances
160
+
161
+ #### main
162
+
163
+ - **Size of downloaded dataset files:** 1.45 MB
164
+ - **Size of the generated dataset:** 1.45 MB
165
+ - **Total amount of disk used:** 2.88 MB
166
+
167
+ An example of 'train' looks as follows:
168
+ ```
169
+ {'id': '7-980',
170
+ 'question_stem': 'The sun is responsible for',
171
+ 'choices': {'text': ['puppies learning new tricks',
172
+ 'children growing up and getting old',
173
+ 'flowers wilting in a vase',
174
+ 'plants sprouting, blooming and wilting'],
175
+ 'label': ['A', 'B', 'C', 'D']},
176
+ 'answerKey': 'D'}
177
+ ```
178
+
179
+ #### additional
180
+
181
+ - **Size of downloaded dataset files:** 1.45 MB
182
+ - **Size of the generated dataset:** 1.45 MB
183
+ - **Total amount of disk used:** 2.88 MB
184
+
185
+ An example of 'train' looks as follows:
186
+ ```
187
+ {'id': '7-980',
188
+ 'question_stem': 'The sun is responsible for',
189
+ 'choices': {'text': ['puppies learning new tricks',
190
+ 'children growing up and getting old',
191
+ 'flowers wilting in a vase',
192
+ 'plants sprouting, blooming and wilting'],
193
+ 'label': ['A', 'B', 'C', 'D']},
194
+ 'answerKey': 'D',
195
+ 'fact1': 'the sun is the source of energy for physical cycles on Earth',
196
+ 'humanScore': 1.0,
197
+ 'clarity': 2.0,
198
+ 'turkIdAnonymized': 'b356d338b7'}
199
+ ```
200
+
201
+ ### Data Fields
202
+
203
+ The data fields are the same among all splits.
204
+
205
+ #### main
206
+ - `id`: a `string` feature.
207
+ - `question_stem`: a `string` feature.
208
+ - `choices`: a dictionary feature containing:
209
+ - `text`: a `string` feature.
210
+ - `label`: a `string` feature.
211
+ - `answerKey`: a `string` feature.
212
+
213
+ #### additional
214
+ - `id`: a `string` feature.
215
+ - `question_stem`: a `string` feature.
216
+ - `choices`: a dictionary feature containing:
217
+ - `text`: a `string` feature.
218
+ - `label`: a `string` feature.
219
+ - `answerKey`: a `string` feature.
220
+ - `fact1` (`str`): oOriginating common knowledge core fact associated to the question.
221
+ - `humanScore` (`float`): Human accuracy score.
222
+ - `clarity` (`float`): Clarity score.
223
+ - `turkIdAnonymized` (`str`): Anonymized crowd-worker ID.
224
+
225
+ ### Data Splits
226
+
227
+ | name | train | validation | test |
228
+ |------------|------:|-----------:|-----:|
229
+ | main | 4957 | 500 | 500 |
230
+ | additional | 4957 | 500 | 500 |
231
+
232
+ ## Dataset Creation
233
+
234
+ ### Curation Rationale
235
+
236
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
237
+
238
+ ### Source Data
239
+
240
+ #### Initial Data Collection and Normalization
241
+
242
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
243
+
244
+ #### Who are the source language producers?
245
+
246
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
247
+
248
+ ### Annotations
249
+
250
+ #### Annotation process
251
+
252
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
253
+
254
+ #### Who are the annotators?
255
+
256
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
257
+
258
+ ### Personal and Sensitive Information
259
+
260
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
261
+
262
+ ## Considerations for Using the Data
263
+
264
+ ### Social Impact of Dataset
265
+
266
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
267
+
268
+ ### Discussion of Biases
269
+
270
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
271
+
272
+ ### Other Known Limitations
273
+
274
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
275
+
276
+ ## Additional Information
277
+
278
+ ### Dataset Curators
279
+
280
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
281
+
282
+ ### Licensing Information
283
+
284
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
285
+
286
+ ### Citation Information
287
+
288
+ ```
289
+ @inproceedings{OpenBookQA2018,
290
+ title={Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering},
291
+ author={Todor Mihaylov and Peter Clark and Tushar Khot and Ashish Sabharwal},
292
+ booktitle={EMNLP},
293
+ year={2018}
294
+ }
295
+
296
+ ```
297
+
298
+
299
+ ### Contributions
300
+
301
+ Thanks to [@thomwolf](https://github.com/thomwolf), [@patrickvonplaten](https://github.com/patrickvonplaten), [@lewtun](https://github.com/lewtun) for adding this dataset.
hf_cache/hub/datasets--allenai--openbookqa/refs/main ADDED
@@ -0,0 +1 @@
 
 
1
+ 388097ea7776314e93a529163e0fea805b8a6454
hf_cache/hub/datasets--allenai--openbookqa/snapshots/388097ea7776314e93a529163e0fea805b8a6454/README.md ADDED
@@ -0,0 +1,301 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ annotations_creators:
3
+ - crowdsourced
4
+ - expert-generated
5
+ language_creators:
6
+ - expert-generated
7
+ language:
8
+ - en
9
+ license:
10
+ - unknown
11
+ multilinguality:
12
+ - monolingual
13
+ size_categories:
14
+ - 1K<n<10K
15
+ source_datasets:
16
+ - original
17
+ task_categories:
18
+ - question-answering
19
+ task_ids:
20
+ - open-domain-qa
21
+ paperswithcode_id: openbookqa
22
+ pretty_name: OpenBookQA
23
+ dataset_info:
24
+ - config_name: additional
25
+ features:
26
+ - name: id
27
+ dtype: string
28
+ - name: question_stem
29
+ dtype: string
30
+ - name: choices
31
+ sequence:
32
+ - name: text
33
+ dtype: string
34
+ - name: label
35
+ dtype: string
36
+ - name: answerKey
37
+ dtype: string
38
+ - name: fact1
39
+ dtype: string
40
+ - name: humanScore
41
+ dtype: float32
42
+ - name: clarity
43
+ dtype: float32
44
+ - name: turkIdAnonymized
45
+ dtype: string
46
+ splits:
47
+ - name: train
48
+ num_bytes: 1288577
49
+ num_examples: 4957
50
+ - name: validation
51
+ num_bytes: 135916
52
+ num_examples: 500
53
+ - name: test
54
+ num_bytes: 130701
55
+ num_examples: 500
56
+ download_size: 783789
57
+ dataset_size: 1555194
58
+ - config_name: main
59
+ features:
60
+ - name: id
61
+ dtype: string
62
+ - name: question_stem
63
+ dtype: string
64
+ - name: choices
65
+ sequence:
66
+ - name: text
67
+ dtype: string
68
+ - name: label
69
+ dtype: string
70
+ - name: answerKey
71
+ dtype: string
72
+ splits:
73
+ - name: train
74
+ num_bytes: 895386
75
+ num_examples: 4957
76
+ - name: validation
77
+ num_bytes: 95428
78
+ num_examples: 500
79
+ - name: test
80
+ num_bytes: 91759
81
+ num_examples: 500
82
+ download_size: 609613
83
+ dataset_size: 1082573
84
+ configs:
85
+ - config_name: additional
86
+ data_files:
87
+ - split: train
88
+ path: additional/train-*
89
+ - split: validation
90
+ path: additional/validation-*
91
+ - split: test
92
+ path: additional/test-*
93
+ - config_name: main
94
+ data_files:
95
+ - split: train
96
+ path: main/train-*
97
+ - split: validation
98
+ path: main/validation-*
99
+ - split: test
100
+ path: main/test-*
101
+ default: true
102
+ ---
103
+
104
+ # Dataset Card for OpenBookQA
105
+
106
+ ## Table of Contents
107
+ - [Dataset Description](#dataset-description)
108
+ - [Dataset Summary](#dataset-summary)
109
+ - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
110
+ - [Languages](#languages)
111
+ - [Dataset Structure](#dataset-structure)
112
+ - [Data Instances](#data-instances)
113
+ - [Data Fields](#data-fields)
114
+ - [Data Splits](#data-splits)
115
+ - [Dataset Creation](#dataset-creation)
116
+ - [Curation Rationale](#curation-rationale)
117
+ - [Source Data](#source-data)
118
+ - [Annotations](#annotations)
119
+ - [Personal and Sensitive Information](#personal-and-sensitive-information)
120
+ - [Considerations for Using the Data](#considerations-for-using-the-data)
121
+ - [Social Impact of Dataset](#social-impact-of-dataset)
122
+ - [Discussion of Biases](#discussion-of-biases)
123
+ - [Other Known Limitations](#other-known-limitations)
124
+ - [Additional Information](#additional-information)
125
+ - [Dataset Curators](#dataset-curators)
126
+ - [Licensing Information](#licensing-information)
127
+ - [Citation Information](#citation-information)
128
+ - [Contributions](#contributions)
129
+
130
+ ## Dataset Description
131
+
132
+ - **Homepage:** [https://allenai.org/data/open-book-qa](https://allenai.org/data/open-book-qa)
133
+ - **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
134
+ - **Paper:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
135
+ - **Point of Contact:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
136
+ - **Size of downloaded dataset files:** 2.89 MB
137
+ - **Size of the generated dataset:** 2.88 MB
138
+ - **Total amount of disk used:** 5.78 MB
139
+
140
+ ### Dataset Summary
141
+
142
+ OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic
143
+ (with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In
144
+ particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge,
145
+ and rich text comprehension.
146
+ OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of
147
+ a subject.
148
+
149
+ ### Supported Tasks and Leaderboards
150
+
151
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
152
+
153
+ ### Languages
154
+
155
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
156
+
157
+ ## Dataset Structure
158
+
159
+ ### Data Instances
160
+
161
+ #### main
162
+
163
+ - **Size of downloaded dataset files:** 1.45 MB
164
+ - **Size of the generated dataset:** 1.45 MB
165
+ - **Total amount of disk used:** 2.88 MB
166
+
167
+ An example of 'train' looks as follows:
168
+ ```
169
+ {'id': '7-980',
170
+ 'question_stem': 'The sun is responsible for',
171
+ 'choices': {'text': ['puppies learning new tricks',
172
+ 'children growing up and getting old',
173
+ 'flowers wilting in a vase',
174
+ 'plants sprouting, blooming and wilting'],
175
+ 'label': ['A', 'B', 'C', 'D']},
176
+ 'answerKey': 'D'}
177
+ ```
178
+
179
+ #### additional
180
+
181
+ - **Size of downloaded dataset files:** 1.45 MB
182
+ - **Size of the generated dataset:** 1.45 MB
183
+ - **Total amount of disk used:** 2.88 MB
184
+
185
+ An example of 'train' looks as follows:
186
+ ```
187
+ {'id': '7-980',
188
+ 'question_stem': 'The sun is responsible for',
189
+ 'choices': {'text': ['puppies learning new tricks',
190
+ 'children growing up and getting old',
191
+ 'flowers wilting in a vase',
192
+ 'plants sprouting, blooming and wilting'],
193
+ 'label': ['A', 'B', 'C', 'D']},
194
+ 'answerKey': 'D',
195
+ 'fact1': 'the sun is the source of energy for physical cycles on Earth',
196
+ 'humanScore': 1.0,
197
+ 'clarity': 2.0,
198
+ 'turkIdAnonymized': 'b356d338b7'}
199
+ ```
200
+
201
+ ### Data Fields
202
+
203
+ The data fields are the same among all splits.
204
+
205
+ #### main
206
+ - `id`: a `string` feature.
207
+ - `question_stem`: a `string` feature.
208
+ - `choices`: a dictionary feature containing:
209
+ - `text`: a `string` feature.
210
+ - `label`: a `string` feature.
211
+ - `answerKey`: a `string` feature.
212
+
213
+ #### additional
214
+ - `id`: a `string` feature.
215
+ - `question_stem`: a `string` feature.
216
+ - `choices`: a dictionary feature containing:
217
+ - `text`: a `string` feature.
218
+ - `label`: a `string` feature.
219
+ - `answerKey`: a `string` feature.
220
+ - `fact1` (`str`): oOriginating common knowledge core fact associated to the question.
221
+ - `humanScore` (`float`): Human accuracy score.
222
+ - `clarity` (`float`): Clarity score.
223
+ - `turkIdAnonymized` (`str`): Anonymized crowd-worker ID.
224
+
225
+ ### Data Splits
226
+
227
+ | name | train | validation | test |
228
+ |------------|------:|-----------:|-----:|
229
+ | main | 4957 | 500 | 500 |
230
+ | additional | 4957 | 500 | 500 |
231
+
232
+ ## Dataset Creation
233
+
234
+ ### Curation Rationale
235
+
236
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
237
+
238
+ ### Source Data
239
+
240
+ #### Initial Data Collection and Normalization
241
+
242
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
243
+
244
+ #### Who are the source language producers?
245
+
246
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
247
+
248
+ ### Annotations
249
+
250
+ #### Annotation process
251
+
252
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
253
+
254
+ #### Who are the annotators?
255
+
256
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
257
+
258
+ ### Personal and Sensitive Information
259
+
260
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
261
+
262
+ ## Considerations for Using the Data
263
+
264
+ ### Social Impact of Dataset
265
+
266
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
267
+
268
+ ### Discussion of Biases
269
+
270
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
271
+
272
+ ### Other Known Limitations
273
+
274
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
275
+
276
+ ## Additional Information
277
+
278
+ ### Dataset Curators
279
+
280
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
281
+
282
+ ### Licensing Information
283
+
284
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
285
+
286
+ ### Citation Information
287
+
288
+ ```
289
+ @inproceedings{OpenBookQA2018,
290
+ title={Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering},
291
+ author={Todor Mihaylov and Peter Clark and Tushar Khot and Ashish Sabharwal},
292
+ booktitle={EMNLP},
293
+ year={2018}
294
+ }
295
+
296
+ ```
297
+
298
+
299
+ ### Contributions
300
+
301
+ Thanks to [@thomwolf](https://github.com/thomwolf), [@patrickvonplaten](https://github.com/patrickvonplaten), [@lewtun](https://github.com/lewtun) for adding this dataset.
hf_cache/hub/datasets--allenai--sciq/.no_exist/2c94ad3e1aafab77146f384e23536f97a4849815/.huggingface.yaml ADDED
File without changes
hf_cache/hub/datasets--allenai--sciq/.no_exist/2c94ad3e1aafab77146f384e23536f97a4849815/dataset_infos.json ADDED
File without changes
hf_cache/hub/datasets--allenai--sciq/.no_exist/2c94ad3e1aafab77146f384e23536f97a4849815/sciq.py ADDED
File without changes
hf_cache/hub/datasets--allenai--sciq/blobs/c644057869cabcde87a2b5ab9665ec0d0bd1405b ADDED
@@ -0,0 +1,216 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ annotations_creators:
3
+ - no-annotation
4
+ language_creators:
5
+ - crowdsourced
6
+ language:
7
+ - en
8
+ license:
9
+ - cc-by-nc-3.0
10
+ multilinguality:
11
+ - monolingual
12
+ size_categories:
13
+ - 10K<n<100K
14
+ source_datasets:
15
+ - original
16
+ task_categories:
17
+ - question-answering
18
+ task_ids:
19
+ - closed-domain-qa
20
+ paperswithcode_id: sciq
21
+ pretty_name: SciQ
22
+ dataset_info:
23
+ features:
24
+ - name: question
25
+ dtype: string
26
+ - name: distractor3
27
+ dtype: string
28
+ - name: distractor1
29
+ dtype: string
30
+ - name: distractor2
31
+ dtype: string
32
+ - name: correct_answer
33
+ dtype: string
34
+ - name: support
35
+ dtype: string
36
+ splits:
37
+ - name: train
38
+ num_bytes: 6546183
39
+ num_examples: 11679
40
+ - name: validation
41
+ num_bytes: 554120
42
+ num_examples: 1000
43
+ - name: test
44
+ num_bytes: 563927
45
+ num_examples: 1000
46
+ download_size: 4674410
47
+ dataset_size: 7664230
48
+ configs:
49
+ - config_name: default
50
+ data_files:
51
+ - split: train
52
+ path: data/train-*
53
+ - split: validation
54
+ path: data/validation-*
55
+ - split: test
56
+ path: data/test-*
57
+ ---
58
+
59
+ # Dataset Card for "sciq"
60
+
61
+ ## Table of Contents
62
+ - [Dataset Description](#dataset-description)
63
+ - [Dataset Summary](#dataset-summary)
64
+ - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
65
+ - [Languages](#languages)
66
+ - [Dataset Structure](#dataset-structure)
67
+ - [Data Instances](#data-instances)
68
+ - [Data Fields](#data-fields)
69
+ - [Data Splits](#data-splits)
70
+ - [Dataset Creation](#dataset-creation)
71
+ - [Curation Rationale](#curation-rationale)
72
+ - [Source Data](#source-data)
73
+ - [Annotations](#annotations)
74
+ - [Personal and Sensitive Information](#personal-and-sensitive-information)
75
+ - [Considerations for Using the Data](#considerations-for-using-the-data)
76
+ - [Social Impact of Dataset](#social-impact-of-dataset)
77
+ - [Discussion of Biases](#discussion-of-biases)
78
+ - [Other Known Limitations](#other-known-limitations)
79
+ - [Additional Information](#additional-information)
80
+ - [Dataset Curators](#dataset-curators)
81
+ - [Licensing Information](#licensing-information)
82
+ - [Citation Information](#citation-information)
83
+ - [Contributions](#contributions)
84
+
85
+ ## Dataset Description
86
+
87
+ - **Homepage:** [https://allenai.org/data/sciq](https://allenai.org/data/sciq)
88
+ - **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
89
+ - **Paper:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
90
+ - **Point of Contact:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
91
+ - **Size of downloaded dataset files:** 2.82 MB
92
+ - **Size of the generated dataset:** 7.68 MB
93
+ - **Total amount of disk used:** 10.50 MB
94
+
95
+ ### Dataset Summary
96
+
97
+ The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others. The questions are in multiple-choice format with 4 answer options each. For the majority of the questions, an additional paragraph with supporting evidence for the correct answer is provided.
98
+
99
+ ### Supported Tasks and Leaderboards
100
+
101
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
102
+
103
+ ### Languages
104
+
105
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
106
+
107
+ ## Dataset Structure
108
+
109
+ ### Data Instances
110
+
111
+ #### default
112
+
113
+ - **Size of downloaded dataset files:** 2.82 MB
114
+ - **Size of the generated dataset:** 7.68 MB
115
+ - **Total amount of disk used:** 10.50 MB
116
+
117
+ An example of 'train' looks as follows.
118
+ ```
119
+ This example was too long and was cropped:
120
+
121
+ {
122
+ "correct_answer": "coriolis effect",
123
+ "distractor1": "muon effect",
124
+ "distractor2": "centrifugal effect",
125
+ "distractor3": "tropical effect",
126
+ "question": "What phenomenon makes global winds blow northeast to southwest or the reverse in the northern hemisphere and northwest to southeast or the reverse in the southern hemisphere?",
127
+ "support": "\"Without Coriolis Effect the global winds would blow north to south or south to north. But Coriolis makes them blow northeast to..."
128
+ }
129
+ ```
130
+
131
+ ### Data Fields
132
+
133
+ The data fields are the same among all splits.
134
+
135
+ #### default
136
+ - `question`: a `string` feature.
137
+ - `distractor3`: a `string` feature.
138
+ - `distractor1`: a `string` feature.
139
+ - `distractor2`: a `string` feature.
140
+ - `correct_answer`: a `string` feature.
141
+ - `support`: a `string` feature.
142
+
143
+ ### Data Splits
144
+
145
+ | name |train|validation|test|
146
+ |-------|----:|---------:|---:|
147
+ |default|11679| 1000|1000|
148
+
149
+ ## Dataset Creation
150
+
151
+ ### Curation Rationale
152
+
153
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
154
+
155
+ ### Source Data
156
+
157
+ #### Initial Data Collection and Normalization
158
+
159
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
160
+
161
+ #### Who are the source language producers?
162
+
163
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
164
+
165
+ ### Annotations
166
+
167
+ #### Annotation process
168
+
169
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
170
+
171
+ #### Who are the annotators?
172
+
173
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
174
+
175
+ ### Personal and Sensitive Information
176
+
177
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
178
+
179
+ ## Considerations for Using the Data
180
+
181
+ ### Social Impact of Dataset
182
+
183
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
184
+
185
+ ### Discussion of Biases
186
+
187
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
188
+
189
+ ### Other Known Limitations
190
+
191
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
192
+
193
+ ## Additional Information
194
+
195
+ ### Dataset Curators
196
+
197
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
198
+
199
+ ### Licensing Information
200
+
201
+ The dataset is licensed under the [Creative Commons Attribution-NonCommercial 3.0 Unported License](http://creativecommons.org/licenses/by-nc/3.0/).
202
+
203
+ ### Citation Information
204
+
205
+ ```
206
+ @inproceedings{SciQ,
207
+ title={Crowdsourcing Multiple Choice Science Questions},
208
+ author={Johannes Welbl, Nelson F. Liu, Matt Gardner},
209
+ year={2017},
210
+ journal={arXiv:1707.06209v1}
211
+ }
212
+ ```
213
+
214
+ ### Contributions
215
+
216
+ Thanks to [@patrickvonplaten](https://github.com/patrickvonplaten), [@lewtun](https://github.com/lewtun), [@thomwolf](https://github.com/thomwolf) for adding this dataset.
hf_cache/hub/datasets--allenai--sciq/refs/main ADDED
@@ -0,0 +1 @@
 
 
1
+ 2c94ad3e1aafab77146f384e23536f97a4849815
hf_cache/hub/datasets--allenai--sciq/snapshots/2c94ad3e1aafab77146f384e23536f97a4849815/README.md ADDED
@@ -0,0 +1,216 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ annotations_creators:
3
+ - no-annotation
4
+ language_creators:
5
+ - crowdsourced
6
+ language:
7
+ - en
8
+ license:
9
+ - cc-by-nc-3.0
10
+ multilinguality:
11
+ - monolingual
12
+ size_categories:
13
+ - 10K<n<100K
14
+ source_datasets:
15
+ - original
16
+ task_categories:
17
+ - question-answering
18
+ task_ids:
19
+ - closed-domain-qa
20
+ paperswithcode_id: sciq
21
+ pretty_name: SciQ
22
+ dataset_info:
23
+ features:
24
+ - name: question
25
+ dtype: string
26
+ - name: distractor3
27
+ dtype: string
28
+ - name: distractor1
29
+ dtype: string
30
+ - name: distractor2
31
+ dtype: string
32
+ - name: correct_answer
33
+ dtype: string
34
+ - name: support
35
+ dtype: string
36
+ splits:
37
+ - name: train
38
+ num_bytes: 6546183
39
+ num_examples: 11679
40
+ - name: validation
41
+ num_bytes: 554120
42
+ num_examples: 1000
43
+ - name: test
44
+ num_bytes: 563927
45
+ num_examples: 1000
46
+ download_size: 4674410
47
+ dataset_size: 7664230
48
+ configs:
49
+ - config_name: default
50
+ data_files:
51
+ - split: train
52
+ path: data/train-*
53
+ - split: validation
54
+ path: data/validation-*
55
+ - split: test
56
+ path: data/test-*
57
+ ---
58
+
59
+ # Dataset Card for "sciq"
60
+
61
+ ## Table of Contents
62
+ - [Dataset Description](#dataset-description)
63
+ - [Dataset Summary](#dataset-summary)
64
+ - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
65
+ - [Languages](#languages)
66
+ - [Dataset Structure](#dataset-structure)
67
+ - [Data Instances](#data-instances)
68
+ - [Data Fields](#data-fields)
69
+ - [Data Splits](#data-splits)
70
+ - [Dataset Creation](#dataset-creation)
71
+ - [Curation Rationale](#curation-rationale)
72
+ - [Source Data](#source-data)
73
+ - [Annotations](#annotations)
74
+ - [Personal and Sensitive Information](#personal-and-sensitive-information)
75
+ - [Considerations for Using the Data](#considerations-for-using-the-data)
76
+ - [Social Impact of Dataset](#social-impact-of-dataset)
77
+ - [Discussion of Biases](#discussion-of-biases)
78
+ - [Other Known Limitations](#other-known-limitations)
79
+ - [Additional Information](#additional-information)
80
+ - [Dataset Curators](#dataset-curators)
81
+ - [Licensing Information](#licensing-information)
82
+ - [Citation Information](#citation-information)
83
+ - [Contributions](#contributions)
84
+
85
+ ## Dataset Description
86
+
87
+ - **Homepage:** [https://allenai.org/data/sciq](https://allenai.org/data/sciq)
88
+ - **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
89
+ - **Paper:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
90
+ - **Point of Contact:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
91
+ - **Size of downloaded dataset files:** 2.82 MB
92
+ - **Size of the generated dataset:** 7.68 MB
93
+ - **Total amount of disk used:** 10.50 MB
94
+
95
+ ### Dataset Summary
96
+
97
+ The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others. The questions are in multiple-choice format with 4 answer options each. For the majority of the questions, an additional paragraph with supporting evidence for the correct answer is provided.
98
+
99
+ ### Supported Tasks and Leaderboards
100
+
101
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
102
+
103
+ ### Languages
104
+
105
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
106
+
107
+ ## Dataset Structure
108
+
109
+ ### Data Instances
110
+
111
+ #### default
112
+
113
+ - **Size of downloaded dataset files:** 2.82 MB
114
+ - **Size of the generated dataset:** 7.68 MB
115
+ - **Total amount of disk used:** 10.50 MB
116
+
117
+ An example of 'train' looks as follows.
118
+ ```
119
+ This example was too long and was cropped:
120
+
121
+ {
122
+ "correct_answer": "coriolis effect",
123
+ "distractor1": "muon effect",
124
+ "distractor2": "centrifugal effect",
125
+ "distractor3": "tropical effect",
126
+ "question": "What phenomenon makes global winds blow northeast to southwest or the reverse in the northern hemisphere and northwest to southeast or the reverse in the southern hemisphere?",
127
+ "support": "\"Without Coriolis Effect the global winds would blow north to south or south to north. But Coriolis makes them blow northeast to..."
128
+ }
129
+ ```
130
+
131
+ ### Data Fields
132
+
133
+ The data fields are the same among all splits.
134
+
135
+ #### default
136
+ - `question`: a `string` feature.
137
+ - `distractor3`: a `string` feature.
138
+ - `distractor1`: a `string` feature.
139
+ - `distractor2`: a `string` feature.
140
+ - `correct_answer`: a `string` feature.
141
+ - `support`: a `string` feature.
142
+
143
+ ### Data Splits
144
+
145
+ | name |train|validation|test|
146
+ |-------|----:|---------:|---:|
147
+ |default|11679| 1000|1000|
148
+
149
+ ## Dataset Creation
150
+
151
+ ### Curation Rationale
152
+
153
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
154
+
155
+ ### Source Data
156
+
157
+ #### Initial Data Collection and Normalization
158
+
159
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
160
+
161
+ #### Who are the source language producers?
162
+
163
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
164
+
165
+ ### Annotations
166
+
167
+ #### Annotation process
168
+
169
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
170
+
171
+ #### Who are the annotators?
172
+
173
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
174
+
175
+ ### Personal and Sensitive Information
176
+
177
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
178
+
179
+ ## Considerations for Using the Data
180
+
181
+ ### Social Impact of Dataset
182
+
183
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
184
+
185
+ ### Discussion of Biases
186
+
187
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
188
+
189
+ ### Other Known Limitations
190
+
191
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
192
+
193
+ ## Additional Information
194
+
195
+ ### Dataset Curators
196
+
197
+ [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
198
+
199
+ ### Licensing Information
200
+
201
+ The dataset is licensed under the [Creative Commons Attribution-NonCommercial 3.0 Unported License](http://creativecommons.org/licenses/by-nc/3.0/).
202
+
203
+ ### Citation Information
204
+
205
+ ```
206
+ @inproceedings{SciQ,
207
+ title={Crowdsourcing Multiple Choice Science Questions},
208
+ author={Johannes Welbl, Nelson F. Liu, Matt Gardner},
209
+ year={2017},
210
+ journal={arXiv:1707.06209v1}
211
+ }
212
+ ```
213
+
214
+ ### Contributions
215
+
216
+ Thanks to [@patrickvonplaten](https://github.com/patrickvonplaten), [@lewtun](https://github.com/lewtun), [@thomwolf](https://github.com/thomwolf) for adding this dataset.
hf_cache/hub/datasets--cais--mmlu/.no_exist/c30699e8356da336a370243923dbaf21066bb9fe/.huggingface.yaml ADDED
File without changes
hf_cache/hub/datasets--cais--mmlu/.no_exist/c30699e8356da336a370243923dbaf21066bb9fe/mmlu.py ADDED
File without changes
hf_cache/hub/datasets--cais--mmlu/blobs/08de94c560ad7420252bff6e4729f1d1683def4f ADDED
@@ -0,0 +1,2299 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ annotations_creators:
3
+ - no-annotation
4
+ language_creators:
5
+ - expert-generated
6
+ language:
7
+ - en
8
+ license:
9
+ - mit
10
+ multilinguality:
11
+ - monolingual
12
+ size_categories:
13
+ - 10K<n<100K
14
+ source_datasets:
15
+ - original
16
+ task_categories:
17
+ - question-answering
18
+ task_ids:
19
+ - multiple-choice-qa
20
+ paperswithcode_id: mmlu
21
+ pretty_name: Measuring Massive Multitask Language Understanding
22
+ language_bcp47:
23
+ - en-US
24
+ dataset_info:
25
+ - config_name: abstract_algebra
26
+ features:
27
+ - name: question
28
+ dtype: string
29
+ - name: subject
30
+ dtype: string
31
+ - name: choices
32
+ sequence: string
33
+ - name: answer
34
+ dtype:
35
+ class_label:
36
+ names:
37
+ '0': A
38
+ '1': B
39
+ '2': C
40
+ '3': D
41
+ splits:
42
+ - name: test
43
+ num_bytes: 49618.6654322746
44
+ num_examples: 100
45
+ - name: validation
46
+ num_bytes: 5485.515349444808
47
+ num_examples: 11
48
+ - name: dev
49
+ num_bytes: 2199.1754385964914
50
+ num_examples: 5
51
+ download_size: 17143
52
+ dataset_size: 57303.3562203159
53
+ - config_name: all
54
+ features:
55
+ - name: question
56
+ dtype: string
57
+ - name: subject
58
+ dtype: string
59
+ - name: choices
60
+ sequence: string
61
+ - name: answer
62
+ dtype:
63
+ class_label:
64
+ names:
65
+ '0': A
66
+ '1': B
67
+ '2': C
68
+ '3': D
69
+ splits:
70
+ - name: test
71
+ num_bytes: 6967453
72
+ num_examples: 14042
73
+ - name: validation
74
+ num_bytes: 763484
75
+ num_examples: 1531
76
+ - name: dev
77
+ num_bytes: 125353
78
+ num_examples: 285
79
+ - name: auxiliary_train
80
+ num_bytes: 161000625
81
+ num_examples: 99842
82
+ download_size: 51503402
83
+ dataset_size: 168856915
84
+ - config_name: anatomy
85
+ features:
86
+ - name: question
87
+ dtype: string
88
+ - name: subject
89
+ dtype: string
90
+ - name: choices
91
+ sequence: string
92
+ - name: answer
93
+ dtype:
94
+ class_label:
95
+ names:
96
+ '0': A
97
+ '1': B
98
+ '2': C
99
+ '3': D
100
+ splits:
101
+ - name: test
102
+ num_bytes: 66985.19833357072
103
+ num_examples: 135
104
+ - name: validation
105
+ num_bytes: 6981.5649902024825
106
+ num_examples: 14
107
+ - name: dev
108
+ num_bytes: 2199.1754385964914
109
+ num_examples: 5
110
+ download_size: 28864
111
+ dataset_size: 76165.9387623697
112
+ - config_name: astronomy
113
+ features:
114
+ - name: question
115
+ dtype: string
116
+ - name: subject
117
+ dtype: string
118
+ - name: choices
119
+ sequence: string
120
+ - name: answer
121
+ dtype:
122
+ class_label:
123
+ names:
124
+ '0': A
125
+ '1': B
126
+ '2': C
127
+ '3': D
128
+ splits:
129
+ - name: test
130
+ num_bytes: 75420.3714570574
131
+ num_examples: 152
132
+ - name: validation
133
+ num_bytes: 7978.931417374265
134
+ num_examples: 16
135
+ - name: dev
136
+ num_bytes: 2199.1754385964914
137
+ num_examples: 5
138
+ download_size: 39316
139
+ dataset_size: 85598.47831302814
140
+ - config_name: auxiliary_train
141
+ features:
142
+ - name: train
143
+ struct:
144
+ - name: answer
145
+ dtype: int64
146
+ - name: choices
147
+ sequence: string
148
+ - name: question
149
+ dtype: string
150
+ - name: subject
151
+ dtype: string
152
+ splits:
153
+ - name: train
154
+ num_bytes: 161000625
155
+ num_examples: 99842
156
+ download_size: 47518592
157
+ dataset_size: 161000625
158
+ - config_name: business_ethics
159
+ features:
160
+ - name: question
161
+ dtype: string
162
+ - name: subject
163
+ dtype: string
164
+ - name: choices
165
+ sequence: string
166
+ - name: answer
167
+ dtype:
168
+ class_label:
169
+ names:
170
+ '0': A
171
+ '1': B
172
+ '2': C
173
+ '3': D
174
+ splits:
175
+ - name: test
176
+ num_bytes: 49618.6654322746
177
+ num_examples: 100
178
+ - name: validation
179
+ num_bytes: 5485.515349444808
180
+ num_examples: 11
181
+ - name: dev
182
+ num_bytes: 2199.1754385964914
183
+ num_examples: 5
184
+ download_size: 31619
185
+ dataset_size: 57303.3562203159
186
+ - config_name: clinical_knowledge
187
+ features:
188
+ - name: question
189
+ dtype: string
190
+ - name: subject
191
+ dtype: string
192
+ - name: choices
193
+ sequence: string
194
+ - name: answer
195
+ dtype:
196
+ class_label:
197
+ names:
198
+ '0': A
199
+ '1': B
200
+ '2': C
201
+ '3': D
202
+ splits:
203
+ - name: test
204
+ num_bytes: 131489.4633955277
205
+ num_examples: 265
206
+ - name: validation
207
+ num_bytes: 14461.813193990856
208
+ num_examples: 29
209
+ - name: dev
210
+ num_bytes: 2199.1754385964914
211
+ num_examples: 5
212
+ download_size: 51655
213
+ dataset_size: 148150.45202811505
214
+ - config_name: college_biology
215
+ features:
216
+ - name: question
217
+ dtype: string
218
+ - name: subject
219
+ dtype: string
220
+ - name: choices
221
+ sequence: string
222
+ - name: answer
223
+ dtype:
224
+ class_label:
225
+ names:
226
+ '0': A
227
+ '1': B
228
+ '2': C
229
+ '3': D
230
+ splits:
231
+ - name: test
232
+ num_bytes: 71450.87822247542
233
+ num_examples: 144
234
+ - name: validation
235
+ num_bytes: 7978.931417374265
236
+ num_examples: 16
237
+ - name: dev
238
+ num_bytes: 2199.1754385964914
239
+ num_examples: 5
240
+ download_size: 43017
241
+ dataset_size: 81628.98507844617
242
+ - config_name: college_chemistry
243
+ features:
244
+ - name: question
245
+ dtype: string
246
+ - name: subject
247
+ dtype: string
248
+ - name: choices
249
+ sequence: string
250
+ - name: answer
251
+ dtype:
252
+ class_label:
253
+ names:
254
+ '0': A
255
+ '1': B
256
+ '2': C
257
+ '3': D
258
+ splits:
259
+ - name: test
260
+ num_bytes: 49618.6654322746
261
+ num_examples: 100
262
+ - name: validation
263
+ num_bytes: 3989.4657086871325
264
+ num_examples: 8
265
+ - name: dev
266
+ num_bytes: 2199.1754385964914
267
+ num_examples: 5
268
+ download_size: 26781
269
+ dataset_size: 55807.30657955822
270
+ - config_name: college_computer_science
271
+ features:
272
+ - name: question
273
+ dtype: string
274
+ - name: subject
275
+ dtype: string
276
+ - name: choices
277
+ sequence: string
278
+ - name: answer
279
+ dtype:
280
+ class_label:
281
+ names:
282
+ '0': A
283
+ '1': B
284
+ '2': C
285
+ '3': D
286
+ splits:
287
+ - name: test
288
+ num_bytes: 49618.6654322746
289
+ num_examples: 100
290
+ - name: validation
291
+ num_bytes: 5485.515349444808
292
+ num_examples: 11
293
+ - name: dev
294
+ num_bytes: 2199.1754385964914
295
+ num_examples: 5
296
+ download_size: 41132
297
+ dataset_size: 57303.3562203159
298
+ - config_name: college_mathematics
299
+ features:
300
+ - name: question
301
+ dtype: string
302
+ - name: subject
303
+ dtype: string
304
+ - name: choices
305
+ sequence: string
306
+ - name: answer
307
+ dtype:
308
+ class_label:
309
+ names:
310
+ '0': A
311
+ '1': B
312
+ '2': C
313
+ '3': D
314
+ splits:
315
+ - name: test
316
+ num_bytes: 49618.6654322746
317
+ num_examples: 100
318
+ - name: validation
319
+ num_bytes: 5485.515349444808
320
+ num_examples: 11
321
+ - name: dev
322
+ num_bytes: 2199.1754385964914
323
+ num_examples: 5
324
+ download_size: 26779
325
+ dataset_size: 57303.3562203159
326
+ - config_name: college_medicine
327
+ features:
328
+ - name: question
329
+ dtype: string
330
+ - name: subject
331
+ dtype: string
332
+ - name: choices
333
+ sequence: string
334
+ - name: answer
335
+ dtype:
336
+ class_label:
337
+ names:
338
+ '0': A
339
+ '1': B
340
+ '2': C
341
+ '3': D
342
+ splits:
343
+ - name: test
344
+ num_bytes: 85840.29119783506
345
+ num_examples: 173
346
+ - name: validation
347
+ num_bytes: 10971.030698889615
348
+ num_examples: 22
349
+ - name: dev
350
+ num_bytes: 2199.1754385964914
351
+ num_examples: 5
352
+ download_size: 56303
353
+ dataset_size: 99010.49733532117
354
+ - config_name: college_physics
355
+ features:
356
+ - name: question
357
+ dtype: string
358
+ - name: subject
359
+ dtype: string
360
+ - name: choices
361
+ sequence: string
362
+ - name: answer
363
+ dtype:
364
+ class_label:
365
+ names:
366
+ '0': A
367
+ '1': B
368
+ '2': C
369
+ '3': D
370
+ splits:
371
+ - name: test
372
+ num_bytes: 50611.0387409201
373
+ num_examples: 102
374
+ - name: validation
375
+ num_bytes: 5485.515349444808
376
+ num_examples: 11
377
+ - name: dev
378
+ num_bytes: 2199.1754385964914
379
+ num_examples: 5
380
+ download_size: 29539
381
+ dataset_size: 58295.7295289614
382
+ - config_name: computer_security
383
+ features:
384
+ - name: question
385
+ dtype: string
386
+ - name: subject
387
+ dtype: string
388
+ - name: choices
389
+ sequence: string
390
+ - name: answer
391
+ dtype:
392
+ class_label:
393
+ names:
394
+ '0': A
395
+ '1': B
396
+ '2': C
397
+ '3': D
398
+ splits:
399
+ - name: test
400
+ num_bytes: 49618.6654322746
401
+ num_examples: 100
402
+ - name: validation
403
+ num_bytes: 5485.515349444808
404
+ num_examples: 11
405
+ - name: dev
406
+ num_bytes: 2199.1754385964914
407
+ num_examples: 5
408
+ download_size: 30150
409
+ dataset_size: 57303.3562203159
410
+ - config_name: conceptual_physics
411
+ features:
412
+ - name: question
413
+ dtype: string
414
+ - name: subject
415
+ dtype: string
416
+ - name: choices
417
+ sequence: string
418
+ - name: answer
419
+ dtype:
420
+ class_label:
421
+ names:
422
+ '0': A
423
+ '1': B
424
+ '2': C
425
+ '3': D
426
+ splits:
427
+ - name: test
428
+ num_bytes: 116603.86376584532
429
+ num_examples: 235
430
+ - name: validation
431
+ num_bytes: 12965.76355323318
432
+ num_examples: 26
433
+ - name: dev
434
+ num_bytes: 2199.1754385964914
435
+ num_examples: 5
436
+ download_size: 34968
437
+ dataset_size: 131768.802757675
438
+ - config_name: econometrics
439
+ features:
440
+ - name: question
441
+ dtype: string
442
+ - name: subject
443
+ dtype: string
444
+ - name: choices
445
+ sequence: string
446
+ - name: answer
447
+ dtype:
448
+ class_label:
449
+ names:
450
+ '0': A
451
+ '1': B
452
+ '2': C
453
+ '3': D
454
+ splits:
455
+ - name: test
456
+ num_bytes: 56565.27859279305
457
+ num_examples: 114
458
+ - name: validation
459
+ num_bytes: 5984.198563030699
460
+ num_examples: 12
461
+ - name: dev
462
+ num_bytes: 2199.1754385964914
463
+ num_examples: 5
464
+ download_size: 36040
465
+ dataset_size: 64748.652594420244
466
+ - config_name: electrical_engineering
467
+ features:
468
+ - name: question
469
+ dtype: string
470
+ - name: subject
471
+ dtype: string
472
+ - name: choices
473
+ sequence: string
474
+ - name: answer
475
+ dtype:
476
+ class_label:
477
+ names:
478
+ '0': A
479
+ '1': B
480
+ '2': C
481
+ '3': D
482
+ splits:
483
+ - name: test
484
+ num_bytes: 71947.06487679818
485
+ num_examples: 145
486
+ - name: validation
487
+ num_bytes: 7978.931417374265
488
+ num_examples: 16
489
+ - name: dev
490
+ num_bytes: 2199.1754385964914
491
+ num_examples: 5
492
+ download_size: 26746
493
+ dataset_size: 82125.17173276893
494
+ - config_name: elementary_mathematics
495
+ features:
496
+ - name: question
497
+ dtype: string
498
+ - name: subject
499
+ dtype: string
500
+ - name: choices
501
+ sequence: string
502
+ - name: answer
503
+ dtype:
504
+ class_label:
505
+ names:
506
+ '0': A
507
+ '1': B
508
+ '2': C
509
+ '3': D
510
+ splits:
511
+ - name: test
512
+ num_bytes: 187558.555333998
513
+ num_examples: 378
514
+ - name: validation
515
+ num_bytes: 20446.011757021555
516
+ num_examples: 41
517
+ - name: dev
518
+ num_bytes: 2199.1754385964914
519
+ num_examples: 5
520
+ download_size: 54987
521
+ dataset_size: 210203.74252961605
522
+ - config_name: formal_logic
523
+ features:
524
+ - name: question
525
+ dtype: string
526
+ - name: subject
527
+ dtype: string
528
+ - name: choices
529
+ sequence: string
530
+ - name: answer
531
+ dtype:
532
+ class_label:
533
+ names:
534
+ '0': A
535
+ '1': B
536
+ '2': C
537
+ '3': D
538
+ splits:
539
+ - name: test
540
+ num_bytes: 62519.518444666
541
+ num_examples: 126
542
+ - name: validation
543
+ num_bytes: 6981.5649902024825
544
+ num_examples: 14
545
+ - name: dev
546
+ num_bytes: 2199.1754385964914
547
+ num_examples: 5
548
+ download_size: 32884
549
+ dataset_size: 71700.25887346498
550
+ - config_name: global_facts
551
+ features:
552
+ - name: question
553
+ dtype: string
554
+ - name: subject
555
+ dtype: string
556
+ - name: choices
557
+ sequence: string
558
+ - name: answer
559
+ dtype:
560
+ class_label:
561
+ names:
562
+ '0': A
563
+ '1': B
564
+ '2': C
565
+ '3': D
566
+ splits:
567
+ - name: test
568
+ num_bytes: 49618.6654322746
569
+ num_examples: 100
570
+ - name: validation
571
+ num_bytes: 4986.8321358589155
572
+ num_examples: 10
573
+ - name: dev
574
+ num_bytes: 2199.1754385964914
575
+ num_examples: 5
576
+ download_size: 19258
577
+ dataset_size: 56804.67300673001
578
+ - config_name: high_school_biology
579
+ features:
580
+ - name: question
581
+ dtype: string
582
+ - name: subject
583
+ dtype: string
584
+ - name: choices
585
+ sequence: string
586
+ - name: answer
587
+ dtype:
588
+ class_label:
589
+ names:
590
+ '0': A
591
+ '1': B
592
+ '2': C
593
+ '3': D
594
+ splits:
595
+ - name: test
596
+ num_bytes: 153817.86284005127
597
+ num_examples: 310
598
+ - name: validation
599
+ num_bytes: 15957.86283474853
600
+ num_examples: 32
601
+ - name: dev
602
+ num_bytes: 2199.1754385964914
603
+ num_examples: 5
604
+ download_size: 78216
605
+ dataset_size: 171974.90111339628
606
+ - config_name: high_school_chemistry
607
+ features:
608
+ - name: question
609
+ dtype: string
610
+ - name: subject
611
+ dtype: string
612
+ - name: choices
613
+ sequence: string
614
+ - name: answer
615
+ dtype:
616
+ class_label:
617
+ names:
618
+ '0': A
619
+ '1': B
620
+ '2': C
621
+ '3': D
622
+ splits:
623
+ - name: test
624
+ num_bytes: 100725.89082751745
625
+ num_examples: 203
626
+ - name: validation
627
+ num_bytes: 10971.030698889615
628
+ num_examples: 22
629
+ - name: dev
630
+ num_bytes: 2199.1754385964914
631
+ num_examples: 5
632
+ download_size: 45799
633
+ dataset_size: 113896.09696500355
634
+ - config_name: high_school_computer_science
635
+ features:
636
+ - name: question
637
+ dtype: string
638
+ - name: subject
639
+ dtype: string
640
+ - name: choices
641
+ sequence: string
642
+ - name: answer
643
+ dtype:
644
+ class_label:
645
+ names:
646
+ '0': A
647
+ '1': B
648
+ '2': C
649
+ '3': D
650
+ splits:
651
+ - name: test
652
+ num_bytes: 49618.6654322746
653
+ num_examples: 100
654
+ - name: validation
655
+ num_bytes: 4488.148922273024
656
+ num_examples: 9
657
+ - name: dev
658
+ num_bytes: 2199.1754385964914
659
+ num_examples: 5
660
+ download_size: 39072
661
+ dataset_size: 56305.989793144116
662
+ - config_name: high_school_european_history
663
+ features:
664
+ - name: question
665
+ dtype: string
666
+ - name: subject
667
+ dtype: string
668
+ - name: choices
669
+ sequence: string
670
+ - name: answer
671
+ dtype:
672
+ class_label:
673
+ names:
674
+ '0': A
675
+ '1': B
676
+ '2': C
677
+ '3': D
678
+ splits:
679
+ - name: test
680
+ num_bytes: 81870.79796325309
681
+ num_examples: 165
682
+ - name: validation
683
+ num_bytes: 8976.297844546049
684
+ num_examples: 18
685
+ - name: dev
686
+ num_bytes: 2199.1754385964914
687
+ num_examples: 5
688
+ download_size: 196270
689
+ dataset_size: 93046.27124639563
690
+ - config_name: high_school_geography
691
+ features:
692
+ - name: question
693
+ dtype: string
694
+ - name: subject
695
+ dtype: string
696
+ - name: choices
697
+ sequence: string
698
+ - name: answer
699
+ dtype:
700
+ class_label:
701
+ names:
702
+ '0': A
703
+ '1': B
704
+ '2': C
705
+ '3': D
706
+ splits:
707
+ - name: test
708
+ num_bytes: 98244.95755590372
709
+ num_examples: 198
710
+ - name: validation
711
+ num_bytes: 10971.030698889615
712
+ num_examples: 22
713
+ - name: dev
714
+ num_bytes: 2199.1754385964914
715
+ num_examples: 5
716
+ download_size: 38255
717
+ dataset_size: 111415.16369338983
718
+ - config_name: high_school_government_and_politics
719
+ features:
720
+ - name: question
721
+ dtype: string
722
+ - name: subject
723
+ dtype: string
724
+ - name: choices
725
+ sequence: string
726
+ - name: answer
727
+ dtype:
728
+ class_label:
729
+ names:
730
+ '0': A
731
+ '1': B
732
+ '2': C
733
+ '3': D
734
+ splits:
735
+ - name: test
736
+ num_bytes: 95764.02428428999
737
+ num_examples: 193
738
+ - name: validation
739
+ num_bytes: 10472.347485303722
740
+ num_examples: 21
741
+ - name: dev
742
+ num_bytes: 2199.1754385964914
743
+ num_examples: 5
744
+ download_size: 52963
745
+ dataset_size: 108435.5472081902
746
+ - config_name: high_school_macroeconomics
747
+ features:
748
+ - name: question
749
+ dtype: string
750
+ - name: subject
751
+ dtype: string
752
+ - name: choices
753
+ sequence: string
754
+ - name: answer
755
+ dtype:
756
+ class_label:
757
+ names:
758
+ '0': A
759
+ '1': B
760
+ '2': C
761
+ '3': D
762
+ splits:
763
+ - name: test
764
+ num_bytes: 193512.79518587096
765
+ num_examples: 390
766
+ - name: validation
767
+ num_bytes: 21443.378184193338
768
+ num_examples: 43
769
+ - name: dev
770
+ num_bytes: 2199.1754385964914
771
+ num_examples: 5
772
+ download_size: 68758
773
+ dataset_size: 217155.34880866078
774
+ - config_name: high_school_mathematics
775
+ features:
776
+ - name: question
777
+ dtype: string
778
+ - name: subject
779
+ dtype: string
780
+ - name: choices
781
+ sequence: string
782
+ - name: answer
783
+ dtype:
784
+ class_label:
785
+ names:
786
+ '0': A
787
+ '1': B
788
+ '2': C
789
+ '3': D
790
+ splits:
791
+ - name: test
792
+ num_bytes: 133970.39666714144
793
+ num_examples: 270
794
+ - name: validation
795
+ num_bytes: 14461.813193990856
796
+ num_examples: 29
797
+ - name: dev
798
+ num_bytes: 2199.1754385964914
799
+ num_examples: 5
800
+ download_size: 45210
801
+ dataset_size: 150631.38529972878
802
+ - config_name: high_school_microeconomics
803
+ features:
804
+ - name: question
805
+ dtype: string
806
+ - name: subject
807
+ dtype: string
808
+ - name: choices
809
+ sequence: string
810
+ - name: answer
811
+ dtype:
812
+ class_label:
813
+ names:
814
+ '0': A
815
+ '1': B
816
+ '2': C
817
+ '3': D
818
+ splits:
819
+ - name: test
820
+ num_bytes: 118092.42372881356
821
+ num_examples: 238
822
+ - name: validation
823
+ num_bytes: 12965.76355323318
824
+ num_examples: 26
825
+ - name: dev
826
+ num_bytes: 2199.1754385964914
827
+ num_examples: 5
828
+ download_size: 49885
829
+ dataset_size: 133257.36272064323
830
+ - config_name: high_school_physics
831
+ features:
832
+ - name: question
833
+ dtype: string
834
+ - name: subject
835
+ dtype: string
836
+ - name: choices
837
+ sequence: string
838
+ - name: answer
839
+ dtype:
840
+ class_label:
841
+ names:
842
+ '0': A
843
+ '1': B
844
+ '2': C
845
+ '3': D
846
+ splits:
847
+ - name: test
848
+ num_bytes: 74924.18480273466
849
+ num_examples: 151
850
+ - name: validation
851
+ num_bytes: 8477.614630960157
852
+ num_examples: 17
853
+ - name: dev
854
+ num_bytes: 2199.1754385964914
855
+ num_examples: 5
856
+ download_size: 45483
857
+ dataset_size: 85600.9748722913
858
+ - config_name: high_school_psychology
859
+ features:
860
+ - name: question
861
+ dtype: string
862
+ - name: subject
863
+ dtype: string
864
+ - name: choices
865
+ sequence: string
866
+ - name: answer
867
+ dtype:
868
+ class_label:
869
+ names:
870
+ '0': A
871
+ '1': B
872
+ '2': C
873
+ '3': D
874
+ splits:
875
+ - name: test
876
+ num_bytes: 270421.7266058966
877
+ num_examples: 545
878
+ - name: validation
879
+ num_bytes: 29920.992815153495
880
+ num_examples: 60
881
+ - name: dev
882
+ num_bytes: 2199.1754385964914
883
+ num_examples: 5
884
+ download_size: 113158
885
+ dataset_size: 302541.8948596466
886
+ - config_name: high_school_statistics
887
+ features:
888
+ - name: question
889
+ dtype: string
890
+ - name: subject
891
+ dtype: string
892
+ - name: choices
893
+ sequence: string
894
+ - name: answer
895
+ dtype:
896
+ class_label:
897
+ names:
898
+ '0': A
899
+ '1': B
900
+ '2': C
901
+ '3': D
902
+ splits:
903
+ - name: test
904
+ num_bytes: 107176.31733371314
905
+ num_examples: 216
906
+ - name: validation
907
+ num_bytes: 11469.713912475507
908
+ num_examples: 23
909
+ - name: dev
910
+ num_bytes: 2199.1754385964914
911
+ num_examples: 5
912
+ download_size: 74924
913
+ dataset_size: 120845.20668478514
914
+ - config_name: high_school_us_history
915
+ features:
916
+ - name: question
917
+ dtype: string
918
+ - name: subject
919
+ dtype: string
920
+ - name: choices
921
+ sequence: string
922
+ - name: answer
923
+ dtype:
924
+ class_label:
925
+ names:
926
+ '0': A
927
+ '1': B
928
+ '2': C
929
+ '3': D
930
+ splits:
931
+ - name: test
932
+ num_bytes: 101222.0774818402
933
+ num_examples: 204
934
+ - name: validation
935
+ num_bytes: 10971.030698889615
936
+ num_examples: 22
937
+ - name: dev
938
+ num_bytes: 2199.1754385964914
939
+ num_examples: 5
940
+ download_size: 200043
941
+ dataset_size: 114392.2836193263
942
+ - config_name: high_school_world_history
943
+ features:
944
+ - name: question
945
+ dtype: string
946
+ - name: subject
947
+ dtype: string
948
+ - name: choices
949
+ sequence: string
950
+ - name: answer
951
+ dtype:
952
+ class_label:
953
+ names:
954
+ '0': A
955
+ '1': B
956
+ '2': C
957
+ '3': D
958
+ splits:
959
+ - name: test
960
+ num_bytes: 117596.23707449081
961
+ num_examples: 237
962
+ - name: validation
963
+ num_bytes: 12965.76355323318
964
+ num_examples: 26
965
+ - name: dev
966
+ num_bytes: 2199.1754385964914
967
+ num_examples: 5
968
+ download_size: 250302
969
+ dataset_size: 132761.17606632048
970
+ - config_name: human_aging
971
+ features:
972
+ - name: question
973
+ dtype: string
974
+ - name: subject
975
+ dtype: string
976
+ - name: choices
977
+ sequence: string
978
+ - name: answer
979
+ dtype:
980
+ class_label:
981
+ names:
982
+ '0': A
983
+ '1': B
984
+ '2': C
985
+ '3': D
986
+ splits:
987
+ - name: test
988
+ num_bytes: 110649.62391397236
989
+ num_examples: 223
990
+ - name: validation
991
+ num_bytes: 11469.713912475507
992
+ num_examples: 23
993
+ - name: dev
994
+ num_bytes: 2199.1754385964914
995
+ num_examples: 5
996
+ download_size: 41196
997
+ dataset_size: 124318.51326504436
998
+ - config_name: human_sexuality
999
+ features:
1000
+ - name: question
1001
+ dtype: string
1002
+ - name: subject
1003
+ dtype: string
1004
+ - name: choices
1005
+ sequence: string
1006
+ - name: answer
1007
+ dtype:
1008
+ class_label:
1009
+ names:
1010
+ '0': A
1011
+ '1': B
1012
+ '2': C
1013
+ '3': D
1014
+ splits:
1015
+ - name: test
1016
+ num_bytes: 65000.451716279735
1017
+ num_examples: 131
1018
+ - name: validation
1019
+ num_bytes: 5984.198563030699
1020
+ num_examples: 12
1021
+ - name: dev
1022
+ num_bytes: 2199.1754385964914
1023
+ num_examples: 5
1024
+ download_size: 32533
1025
+ dataset_size: 73183.82571790692
1026
+ - config_name: international_law
1027
+ features:
1028
+ - name: question
1029
+ dtype: string
1030
+ - name: subject
1031
+ dtype: string
1032
+ - name: choices
1033
+ sequence: string
1034
+ - name: answer
1035
+ dtype:
1036
+ class_label:
1037
+ names:
1038
+ '0': A
1039
+ '1': B
1040
+ '2': C
1041
+ '3': D
1042
+ splits:
1043
+ - name: test
1044
+ num_bytes: 60038.58517305227
1045
+ num_examples: 121
1046
+ - name: validation
1047
+ num_bytes: 6482.88177661659
1048
+ num_examples: 13
1049
+ - name: dev
1050
+ num_bytes: 2199.1754385964914
1051
+ num_examples: 5
1052
+ download_size: 41592
1053
+ dataset_size: 68720.64238826535
1054
+ - config_name: jurisprudence
1055
+ features:
1056
+ - name: question
1057
+ dtype: string
1058
+ - name: subject
1059
+ dtype: string
1060
+ - name: choices
1061
+ sequence: string
1062
+ - name: answer
1063
+ dtype:
1064
+ class_label:
1065
+ names:
1066
+ '0': A
1067
+ '1': B
1068
+ '2': C
1069
+ '3': D
1070
+ splits:
1071
+ - name: test
1072
+ num_bytes: 53588.15866685657
1073
+ num_examples: 108
1074
+ - name: validation
1075
+ num_bytes: 5485.515349444808
1076
+ num_examples: 11
1077
+ - name: dev
1078
+ num_bytes: 2199.1754385964914
1079
+ num_examples: 5
1080
+ download_size: 33578
1081
+ dataset_size: 61272.84945489787
1082
+ - config_name: logical_fallacies
1083
+ features:
1084
+ - name: question
1085
+ dtype: string
1086
+ - name: subject
1087
+ dtype: string
1088
+ - name: choices
1089
+ sequence: string
1090
+ - name: answer
1091
+ dtype:
1092
+ class_label:
1093
+ names:
1094
+ '0': A
1095
+ '1': B
1096
+ '2': C
1097
+ '3': D
1098
+ splits:
1099
+ - name: test
1100
+ num_bytes: 80878.4246546076
1101
+ num_examples: 163
1102
+ - name: validation
1103
+ num_bytes: 8976.297844546049
1104
+ num_examples: 18
1105
+ - name: dev
1106
+ num_bytes: 2199.1754385964914
1107
+ num_examples: 5
1108
+ download_size: 33669
1109
+ dataset_size: 92053.89793775014
1110
+ - config_name: machine_learning
1111
+ features:
1112
+ - name: question
1113
+ dtype: string
1114
+ - name: subject
1115
+ dtype: string
1116
+ - name: choices
1117
+ sequence: string
1118
+ - name: answer
1119
+ dtype:
1120
+ class_label:
1121
+ names:
1122
+ '0': A
1123
+ '1': B
1124
+ '2': C
1125
+ '3': D
1126
+ splits:
1127
+ - name: test
1128
+ num_bytes: 55572.90528414756
1129
+ num_examples: 112
1130
+ - name: validation
1131
+ num_bytes: 5485.515349444808
1132
+ num_examples: 11
1133
+ - name: dev
1134
+ num_bytes: 2199.1754385964914
1135
+ num_examples: 5
1136
+ download_size: 31121
1137
+ dataset_size: 63257.596072188855
1138
+ - config_name: management
1139
+ features:
1140
+ - name: question
1141
+ dtype: string
1142
+ - name: subject
1143
+ dtype: string
1144
+ - name: choices
1145
+ sequence: string
1146
+ - name: answer
1147
+ dtype:
1148
+ class_label:
1149
+ names:
1150
+ '0': A
1151
+ '1': B
1152
+ '2': C
1153
+ '3': D
1154
+ splits:
1155
+ - name: test
1156
+ num_bytes: 51107.225395242844
1157
+ num_examples: 103
1158
+ - name: validation
1159
+ num_bytes: 5485.515349444808
1160
+ num_examples: 11
1161
+ - name: dev
1162
+ num_bytes: 2199.1754385964914
1163
+ num_examples: 5
1164
+ download_size: 22828
1165
+ dataset_size: 58791.91618328414
1166
+ - config_name: marketing
1167
+ features:
1168
+ - name: question
1169
+ dtype: string
1170
+ - name: subject
1171
+ dtype: string
1172
+ - name: choices
1173
+ sequence: string
1174
+ - name: answer
1175
+ dtype:
1176
+ class_label:
1177
+ names:
1178
+ '0': A
1179
+ '1': B
1180
+ '2': C
1181
+ '3': D
1182
+ splits:
1183
+ - name: test
1184
+ num_bytes: 116107.67711152257
1185
+ num_examples: 234
1186
+ - name: validation
1187
+ num_bytes: 12467.08033964729
1188
+ num_examples: 25
1189
+ - name: dev
1190
+ num_bytes: 2199.1754385964914
1191
+ num_examples: 5
1192
+ download_size: 49747
1193
+ dataset_size: 130773.93288976635
1194
+ - config_name: medical_genetics
1195
+ features:
1196
+ - name: question
1197
+ dtype: string
1198
+ - name: subject
1199
+ dtype: string
1200
+ - name: choices
1201
+ sequence: string
1202
+ - name: answer
1203
+ dtype:
1204
+ class_label:
1205
+ names:
1206
+ '0': A
1207
+ '1': B
1208
+ '2': C
1209
+ '3': D
1210
+ splits:
1211
+ - name: test
1212
+ num_bytes: 49618.6654322746
1213
+ num_examples: 100
1214
+ - name: validation
1215
+ num_bytes: 5485.515349444808
1216
+ num_examples: 11
1217
+ - name: dev
1218
+ num_bytes: 2199.1754385964914
1219
+ num_examples: 5
1220
+ download_size: 25775
1221
+ dataset_size: 57303.3562203159
1222
+ - config_name: miscellaneous
1223
+ features:
1224
+ - name: question
1225
+ dtype: string
1226
+ - name: subject
1227
+ dtype: string
1228
+ - name: choices
1229
+ sequence: string
1230
+ - name: answer
1231
+ dtype:
1232
+ class_label:
1233
+ names:
1234
+ '0': A
1235
+ '1': B
1236
+ '2': C
1237
+ '3': D
1238
+ splits:
1239
+ - name: test
1240
+ num_bytes: 388514.15033471014
1241
+ num_examples: 783
1242
+ - name: validation
1243
+ num_bytes: 42886.756368386676
1244
+ num_examples: 86
1245
+ - name: dev
1246
+ num_bytes: 2199.1754385964914
1247
+ num_examples: 5
1248
+ download_size: 115097
1249
+ dataset_size: 433600.08214169333
1250
+ - config_name: moral_disputes
1251
+ features:
1252
+ - name: question
1253
+ dtype: string
1254
+ - name: subject
1255
+ dtype: string
1256
+ - name: choices
1257
+ sequence: string
1258
+ - name: answer
1259
+ dtype:
1260
+ class_label:
1261
+ names:
1262
+ '0': A
1263
+ '1': B
1264
+ '2': C
1265
+ '3': D
1266
+ splits:
1267
+ - name: test
1268
+ num_bytes: 171680.58239567012
1269
+ num_examples: 346
1270
+ - name: validation
1271
+ num_bytes: 18949.96211626388
1272
+ num_examples: 38
1273
+ - name: dev
1274
+ num_bytes: 2199.1754385964914
1275
+ num_examples: 5
1276
+ download_size: 76043
1277
+ dataset_size: 192829.71995053047
1278
+ - config_name: moral_scenarios
1279
+ features:
1280
+ - name: question
1281
+ dtype: string
1282
+ - name: subject
1283
+ dtype: string
1284
+ - name: choices
1285
+ sequence: string
1286
+ - name: answer
1287
+ dtype:
1288
+ class_label:
1289
+ names:
1290
+ '0': A
1291
+ '1': B
1292
+ '2': C
1293
+ '3': D
1294
+ splits:
1295
+ - name: test
1296
+ num_bytes: 444087.05561885773
1297
+ num_examples: 895
1298
+ - name: validation
1299
+ num_bytes: 49868.32135858916
1300
+ num_examples: 100
1301
+ - name: dev
1302
+ num_bytes: 2199.1754385964914
1303
+ num_examples: 5
1304
+ download_size: 109869
1305
+ dataset_size: 496154.5524160434
1306
+ - config_name: nutrition
1307
+ features:
1308
+ - name: question
1309
+ dtype: string
1310
+ - name: subject
1311
+ dtype: string
1312
+ - name: choices
1313
+ sequence: string
1314
+ - name: answer
1315
+ dtype:
1316
+ class_label:
1317
+ names:
1318
+ '0': A
1319
+ '1': B
1320
+ '2': C
1321
+ '3': D
1322
+ splits:
1323
+ - name: test
1324
+ num_bytes: 151833.1162227603
1325
+ num_examples: 306
1326
+ - name: validation
1327
+ num_bytes: 16456.54604833442
1328
+ num_examples: 33
1329
+ - name: dev
1330
+ num_bytes: 2199.1754385964914
1331
+ num_examples: 5
1332
+ download_size: 69050
1333
+ dataset_size: 170488.8377096912
1334
+ - config_name: philosophy
1335
+ features:
1336
+ - name: question
1337
+ dtype: string
1338
+ - name: subject
1339
+ dtype: string
1340
+ - name: choices
1341
+ sequence: string
1342
+ - name: answer
1343
+ dtype:
1344
+ class_label:
1345
+ names:
1346
+ '0': A
1347
+ '1': B
1348
+ '2': C
1349
+ '3': D
1350
+ splits:
1351
+ - name: test
1352
+ num_bytes: 154314.04949437402
1353
+ num_examples: 311
1354
+ - name: validation
1355
+ num_bytes: 16955.229261920314
1356
+ num_examples: 34
1357
+ - name: dev
1358
+ num_bytes: 2199.1754385964914
1359
+ num_examples: 5
1360
+ download_size: 61912
1361
+ dataset_size: 173468.45419489083
1362
+ - config_name: prehistory
1363
+ features:
1364
+ - name: question
1365
+ dtype: string
1366
+ - name: subject
1367
+ dtype: string
1368
+ - name: choices
1369
+ sequence: string
1370
+ - name: answer
1371
+ dtype:
1372
+ class_label:
1373
+ names:
1374
+ '0': A
1375
+ '1': B
1376
+ '2': C
1377
+ '3': D
1378
+ splits:
1379
+ - name: test
1380
+ num_bytes: 160764.47600056973
1381
+ num_examples: 324
1382
+ - name: validation
1383
+ num_bytes: 17453.912475506204
1384
+ num_examples: 35
1385
+ - name: dev
1386
+ num_bytes: 2199.1754385964914
1387
+ num_examples: 5
1388
+ download_size: 68826
1389
+ dataset_size: 180417.5639146724
1390
+ - config_name: professional_accounting
1391
+ features:
1392
+ - name: question
1393
+ dtype: string
1394
+ - name: subject
1395
+ dtype: string
1396
+ - name: choices
1397
+ sequence: string
1398
+ - name: answer
1399
+ dtype:
1400
+ class_label:
1401
+ names:
1402
+ '0': A
1403
+ '1': B
1404
+ '2': C
1405
+ '3': D
1406
+ splits:
1407
+ - name: test
1408
+ num_bytes: 139924.6365190144
1409
+ num_examples: 282
1410
+ - name: validation
1411
+ num_bytes: 15459.179621162639
1412
+ num_examples: 31
1413
+ - name: dev
1414
+ num_bytes: 2199.1754385964914
1415
+ num_examples: 5
1416
+ download_size: 87297
1417
+ dataset_size: 157582.99157877354
1418
+ - config_name: professional_law
1419
+ features:
1420
+ - name: question
1421
+ dtype: string
1422
+ - name: subject
1423
+ dtype: string
1424
+ - name: choices
1425
+ sequence: string
1426
+ - name: answer
1427
+ dtype:
1428
+ class_label:
1429
+ names:
1430
+ '0': A
1431
+ '1': B
1432
+ '2': C
1433
+ '3': D
1434
+ splits:
1435
+ - name: test
1436
+ num_bytes: 761150.3277310925
1437
+ num_examples: 1534
1438
+ - name: validation
1439
+ num_bytes: 84776.14630960157
1440
+ num_examples: 170
1441
+ - name: dev
1442
+ num_bytes: 2199.1754385964914
1443
+ num_examples: 5
1444
+ download_size: 1167828
1445
+ dataset_size: 848125.6494792906
1446
+ - config_name: professional_medicine
1447
+ features:
1448
+ - name: question
1449
+ dtype: string
1450
+ - name: subject
1451
+ dtype: string
1452
+ - name: choices
1453
+ sequence: string
1454
+ - name: answer
1455
+ dtype:
1456
+ class_label:
1457
+ names:
1458
+ '0': A
1459
+ '1': B
1460
+ '2': C
1461
+ '3': D
1462
+ splits:
1463
+ - name: test
1464
+ num_bytes: 134962.7699757869
1465
+ num_examples: 272
1466
+ - name: validation
1467
+ num_bytes: 15459.179621162639
1468
+ num_examples: 31
1469
+ - name: dev
1470
+ num_bytes: 2199.1754385964914
1471
+ num_examples: 5
1472
+ download_size: 153242
1473
+ dataset_size: 152621.12503554605
1474
+ - config_name: professional_psychology
1475
+ features:
1476
+ - name: question
1477
+ dtype: string
1478
+ - name: subject
1479
+ dtype: string
1480
+ - name: choices
1481
+ sequence: string
1482
+ - name: answer
1483
+ dtype:
1484
+ class_label:
1485
+ names:
1486
+ '0': A
1487
+ '1': B
1488
+ '2': C
1489
+ '3': D
1490
+ splits:
1491
+ - name: test
1492
+ num_bytes: 303666.2324455206
1493
+ num_examples: 612
1494
+ - name: validation
1495
+ num_bytes: 34409.14173742652
1496
+ num_examples: 69
1497
+ - name: dev
1498
+ num_bytes: 2199.1754385964914
1499
+ num_examples: 5
1500
+ download_size: 159357
1501
+ dataset_size: 340274.5496215436
1502
+ - config_name: public_relations
1503
+ features:
1504
+ - name: question
1505
+ dtype: string
1506
+ - name: subject
1507
+ dtype: string
1508
+ - name: choices
1509
+ sequence: string
1510
+ - name: answer
1511
+ dtype:
1512
+ class_label:
1513
+ names:
1514
+ '0': A
1515
+ '1': B
1516
+ '2': C
1517
+ '3': D
1518
+ splits:
1519
+ - name: test
1520
+ num_bytes: 54580.53197550207
1521
+ num_examples: 110
1522
+ - name: validation
1523
+ num_bytes: 5984.198563030699
1524
+ num_examples: 12
1525
+ - name: dev
1526
+ num_bytes: 2199.1754385964914
1527
+ num_examples: 5
1528
+ download_size: 31500
1529
+ dataset_size: 62763.90597712925
1530
+ - config_name: security_studies
1531
+ features:
1532
+ - name: question
1533
+ dtype: string
1534
+ - name: subject
1535
+ dtype: string
1536
+ - name: choices
1537
+ sequence: string
1538
+ - name: answer
1539
+ dtype:
1540
+ class_label:
1541
+ names:
1542
+ '0': A
1543
+ '1': B
1544
+ '2': C
1545
+ '3': D
1546
+ splits:
1547
+ - name: test
1548
+ num_bytes: 121565.73030907278
1549
+ num_examples: 245
1550
+ - name: validation
1551
+ num_bytes: 13464.446766819072
1552
+ num_examples: 27
1553
+ - name: dev
1554
+ num_bytes: 2199.1754385964914
1555
+ num_examples: 5
1556
+ download_size: 140258
1557
+ dataset_size: 137229.35251448833
1558
+ - config_name: sociology
1559
+ features:
1560
+ - name: question
1561
+ dtype: string
1562
+ - name: subject
1563
+ dtype: string
1564
+ - name: choices
1565
+ sequence: string
1566
+ - name: answer
1567
+ dtype:
1568
+ class_label:
1569
+ names:
1570
+ '0': A
1571
+ '1': B
1572
+ '2': C
1573
+ '3': D
1574
+ splits:
1575
+ - name: test
1576
+ num_bytes: 99733.51751887196
1577
+ num_examples: 201
1578
+ - name: validation
1579
+ num_bytes: 10971.030698889615
1580
+ num_examples: 22
1581
+ - name: dev
1582
+ num_bytes: 2199.1754385964914
1583
+ num_examples: 5
1584
+ download_size: 56480
1585
+ dataset_size: 112903.72365635807
1586
+ - config_name: us_foreign_policy
1587
+ features:
1588
+ - name: question
1589
+ dtype: string
1590
+ - name: subject
1591
+ dtype: string
1592
+ - name: choices
1593
+ sequence: string
1594
+ - name: answer
1595
+ dtype:
1596
+ class_label:
1597
+ names:
1598
+ '0': A
1599
+ '1': B
1600
+ '2': C
1601
+ '3': D
1602
+ splits:
1603
+ - name: test
1604
+ num_bytes: 49618.6654322746
1605
+ num_examples: 100
1606
+ - name: validation
1607
+ num_bytes: 5485.515349444808
1608
+ num_examples: 11
1609
+ - name: dev
1610
+ num_bytes: 2199.1754385964914
1611
+ num_examples: 5
1612
+ download_size: 29027
1613
+ dataset_size: 57303.3562203159
1614
+ - config_name: virology
1615
+ features:
1616
+ - name: question
1617
+ dtype: string
1618
+ - name: subject
1619
+ dtype: string
1620
+ - name: choices
1621
+ sequence: string
1622
+ - name: answer
1623
+ dtype:
1624
+ class_label:
1625
+ names:
1626
+ '0': A
1627
+ '1': B
1628
+ '2': C
1629
+ '3': D
1630
+ splits:
1631
+ - name: test
1632
+ num_bytes: 82366.98461757584
1633
+ num_examples: 166
1634
+ - name: validation
1635
+ num_bytes: 8976.297844546049
1636
+ num_examples: 18
1637
+ - name: dev
1638
+ num_bytes: 2199.1754385964914
1639
+ num_examples: 5
1640
+ download_size: 38229
1641
+ dataset_size: 93542.45790071838
1642
+ - config_name: world_religions
1643
+ features:
1644
+ - name: question
1645
+ dtype: string
1646
+ - name: subject
1647
+ dtype: string
1648
+ - name: choices
1649
+ sequence: string
1650
+ - name: answer
1651
+ dtype:
1652
+ class_label:
1653
+ names:
1654
+ '0': A
1655
+ '1': B
1656
+ '2': C
1657
+ '3': D
1658
+ splits:
1659
+ - name: test
1660
+ num_bytes: 84847.91788918957
1661
+ num_examples: 171
1662
+ - name: validation
1663
+ num_bytes: 9474.98105813194
1664
+ num_examples: 19
1665
+ - name: dev
1666
+ num_bytes: 2199.1754385964914
1667
+ num_examples: 5
1668
+ download_size: 27165
1669
+ dataset_size: 96522.07438591801
1670
+ configs:
1671
+ - config_name: abstract_algebra
1672
+ data_files:
1673
+ - split: test
1674
+ path: abstract_algebra/test-*
1675
+ - split: validation
1676
+ path: abstract_algebra/validation-*
1677
+ - split: dev
1678
+ path: abstract_algebra/dev-*
1679
+ - config_name: all
1680
+ data_files:
1681
+ - split: test
1682
+ path: all/test-*
1683
+ - split: validation
1684
+ path: all/validation-*
1685
+ - split: dev
1686
+ path: all/dev-*
1687
+ - split: auxiliary_train
1688
+ path: all/auxiliary_train-*
1689
+ - config_name: anatomy
1690
+ data_files:
1691
+ - split: test
1692
+ path: anatomy/test-*
1693
+ - split: validation
1694
+ path: anatomy/validation-*
1695
+ - split: dev
1696
+ path: anatomy/dev-*
1697
+ - config_name: astronomy
1698
+ data_files:
1699
+ - split: test
1700
+ path: astronomy/test-*
1701
+ - split: validation
1702
+ path: astronomy/validation-*
1703
+ - split: dev
1704
+ path: astronomy/dev-*
1705
+ - config_name: auxiliary_train
1706
+ data_files:
1707
+ - split: train
1708
+ path: auxiliary_train/train-*
1709
+ - config_name: business_ethics
1710
+ data_files:
1711
+ - split: test
1712
+ path: business_ethics/test-*
1713
+ - split: validation
1714
+ path: business_ethics/validation-*
1715
+ - split: dev
1716
+ path: business_ethics/dev-*
1717
+ - config_name: clinical_knowledge
1718
+ data_files:
1719
+ - split: test
1720
+ path: clinical_knowledge/test-*
1721
+ - split: validation
1722
+ path: clinical_knowledge/validation-*
1723
+ - split: dev
1724
+ path: clinical_knowledge/dev-*
1725
+ - config_name: college_biology
1726
+ data_files:
1727
+ - split: test
1728
+ path: college_biology/test-*
1729
+ - split: validation
1730
+ path: college_biology/validation-*
1731
+ - split: dev
1732
+ path: college_biology/dev-*
1733
+ - config_name: college_chemistry
1734
+ data_files:
1735
+ - split: test
1736
+ path: college_chemistry/test-*
1737
+ - split: validation
1738
+ path: college_chemistry/validation-*
1739
+ - split: dev
1740
+ path: college_chemistry/dev-*
1741
+ - config_name: college_computer_science
1742
+ data_files:
1743
+ - split: test
1744
+ path: college_computer_science/test-*
1745
+ - split: validation
1746
+ path: college_computer_science/validation-*
1747
+ - split: dev
1748
+ path: college_computer_science/dev-*
1749
+ - config_name: college_mathematics
1750
+ data_files:
1751
+ - split: test
1752
+ path: college_mathematics/test-*
1753
+ - split: validation
1754
+ path: college_mathematics/validation-*
1755
+ - split: dev
1756
+ path: college_mathematics/dev-*
1757
+ - config_name: college_medicine
1758
+ data_files:
1759
+ - split: test
1760
+ path: college_medicine/test-*
1761
+ - split: validation
1762
+ path: college_medicine/validation-*
1763
+ - split: dev
1764
+ path: college_medicine/dev-*
1765
+ - config_name: college_physics
1766
+ data_files:
1767
+ - split: test
1768
+ path: college_physics/test-*
1769
+ - split: validation
1770
+ path: college_physics/validation-*
1771
+ - split: dev
1772
+ path: college_physics/dev-*
1773
+ - config_name: computer_security
1774
+ data_files:
1775
+ - split: test
1776
+ path: computer_security/test-*
1777
+ - split: validation
1778
+ path: computer_security/validation-*
1779
+ - split: dev
1780
+ path: computer_security/dev-*
1781
+ - config_name: conceptual_physics
1782
+ data_files:
1783
+ - split: test
1784
+ path: conceptual_physics/test-*
1785
+ - split: validation
1786
+ path: conceptual_physics/validation-*
1787
+ - split: dev
1788
+ path: conceptual_physics/dev-*
1789
+ - config_name: econometrics
1790
+ data_files:
1791
+ - split: test
1792
+ path: econometrics/test-*
1793
+ - split: validation
1794
+ path: econometrics/validation-*
1795
+ - split: dev
1796
+ path: econometrics/dev-*
1797
+ - config_name: electrical_engineering
1798
+ data_files:
1799
+ - split: test
1800
+ path: electrical_engineering/test-*
1801
+ - split: validation
1802
+ path: electrical_engineering/validation-*
1803
+ - split: dev
1804
+ path: electrical_engineering/dev-*
1805
+ - config_name: elementary_mathematics
1806
+ data_files:
1807
+ - split: test
1808
+ path: elementary_mathematics/test-*
1809
+ - split: validation
1810
+ path: elementary_mathematics/validation-*
1811
+ - split: dev
1812
+ path: elementary_mathematics/dev-*
1813
+ - config_name: formal_logic
1814
+ data_files:
1815
+ - split: test
1816
+ path: formal_logic/test-*
1817
+ - split: validation
1818
+ path: formal_logic/validation-*
1819
+ - split: dev
1820
+ path: formal_logic/dev-*
1821
+ - config_name: global_facts
1822
+ data_files:
1823
+ - split: test
1824
+ path: global_facts/test-*
1825
+ - split: validation
1826
+ path: global_facts/validation-*
1827
+ - split: dev
1828
+ path: global_facts/dev-*
1829
+ - config_name: high_school_biology
1830
+ data_files:
1831
+ - split: test
1832
+ path: high_school_biology/test-*
1833
+ - split: validation
1834
+ path: high_school_biology/validation-*
1835
+ - split: dev
1836
+ path: high_school_biology/dev-*
1837
+ - config_name: high_school_chemistry
1838
+ data_files:
1839
+ - split: test
1840
+ path: high_school_chemistry/test-*
1841
+ - split: validation
1842
+ path: high_school_chemistry/validation-*
1843
+ - split: dev
1844
+ path: high_school_chemistry/dev-*
1845
+ - config_name: high_school_computer_science
1846
+ data_files:
1847
+ - split: test
1848
+ path: high_school_computer_science/test-*
1849
+ - split: validation
1850
+ path: high_school_computer_science/validation-*
1851
+ - split: dev
1852
+ path: high_school_computer_science/dev-*
1853
+ - config_name: high_school_european_history
1854
+ data_files:
1855
+ - split: test
1856
+ path: high_school_european_history/test-*
1857
+ - split: validation
1858
+ path: high_school_european_history/validation-*
1859
+ - split: dev
1860
+ path: high_school_european_history/dev-*
1861
+ - config_name: high_school_geography
1862
+ data_files:
1863
+ - split: test
1864
+ path: high_school_geography/test-*
1865
+ - split: validation
1866
+ path: high_school_geography/validation-*
1867
+ - split: dev
1868
+ path: high_school_geography/dev-*
1869
+ - config_name: high_school_government_and_politics
1870
+ data_files:
1871
+ - split: test
1872
+ path: high_school_government_and_politics/test-*
1873
+ - split: validation
1874
+ path: high_school_government_and_politics/validation-*
1875
+ - split: dev
1876
+ path: high_school_government_and_politics/dev-*
1877
+ - config_name: high_school_macroeconomics
1878
+ data_files:
1879
+ - split: test
1880
+ path: high_school_macroeconomics/test-*
1881
+ - split: validation
1882
+ path: high_school_macroeconomics/validation-*
1883
+ - split: dev
1884
+ path: high_school_macroeconomics/dev-*
1885
+ - config_name: high_school_mathematics
1886
+ data_files:
1887
+ - split: test
1888
+ path: high_school_mathematics/test-*
1889
+ - split: validation
1890
+ path: high_school_mathematics/validation-*
1891
+ - split: dev
1892
+ path: high_school_mathematics/dev-*
1893
+ - config_name: high_school_microeconomics
1894
+ data_files:
1895
+ - split: test
1896
+ path: high_school_microeconomics/test-*
1897
+ - split: validation
1898
+ path: high_school_microeconomics/validation-*
1899
+ - split: dev
1900
+ path: high_school_microeconomics/dev-*
1901
+ - config_name: high_school_physics
1902
+ data_files:
1903
+ - split: test
1904
+ path: high_school_physics/test-*
1905
+ - split: validation
1906
+ path: high_school_physics/validation-*
1907
+ - split: dev
1908
+ path: high_school_physics/dev-*
1909
+ - config_name: high_school_psychology
1910
+ data_files:
1911
+ - split: test
1912
+ path: high_school_psychology/test-*
1913
+ - split: validation
1914
+ path: high_school_psychology/validation-*
1915
+ - split: dev
1916
+ path: high_school_psychology/dev-*
1917
+ - config_name: high_school_statistics
1918
+ data_files:
1919
+ - split: test
1920
+ path: high_school_statistics/test-*
1921
+ - split: validation
1922
+ path: high_school_statistics/validation-*
1923
+ - split: dev
1924
+ path: high_school_statistics/dev-*
1925
+ - config_name: high_school_us_history
1926
+ data_files:
1927
+ - split: test
1928
+ path: high_school_us_history/test-*
1929
+ - split: validation
1930
+ path: high_school_us_history/validation-*
1931
+ - split: dev
1932
+ path: high_school_us_history/dev-*
1933
+ - config_name: high_school_world_history
1934
+ data_files:
1935
+ - split: test
1936
+ path: high_school_world_history/test-*
1937
+ - split: validation
1938
+ path: high_school_world_history/validation-*
1939
+ - split: dev
1940
+ path: high_school_world_history/dev-*
1941
+ - config_name: human_aging
1942
+ data_files:
1943
+ - split: test
1944
+ path: human_aging/test-*
1945
+ - split: validation
1946
+ path: human_aging/validation-*
1947
+ - split: dev
1948
+ path: human_aging/dev-*
1949
+ - config_name: human_sexuality
1950
+ data_files:
1951
+ - split: test
1952
+ path: human_sexuality/test-*
1953
+ - split: validation
1954
+ path: human_sexuality/validation-*
1955
+ - split: dev
1956
+ path: human_sexuality/dev-*
1957
+ - config_name: international_law
1958
+ data_files:
1959
+ - split: test
1960
+ path: international_law/test-*
1961
+ - split: validation
1962
+ path: international_law/validation-*
1963
+ - split: dev
1964
+ path: international_law/dev-*
1965
+ - config_name: jurisprudence
1966
+ data_files:
1967
+ - split: test
1968
+ path: jurisprudence/test-*
1969
+ - split: validation
1970
+ path: jurisprudence/validation-*
1971
+ - split: dev
1972
+ path: jurisprudence/dev-*
1973
+ - config_name: logical_fallacies
1974
+ data_files:
1975
+ - split: test
1976
+ path: logical_fallacies/test-*
1977
+ - split: validation
1978
+ path: logical_fallacies/validation-*
1979
+ - split: dev
1980
+ path: logical_fallacies/dev-*
1981
+ - config_name: machine_learning
1982
+ data_files:
1983
+ - split: test
1984
+ path: machine_learning/test-*
1985
+ - split: validation
1986
+ path: machine_learning/validation-*
1987
+ - split: dev
1988
+ path: machine_learning/dev-*
1989
+ - config_name: management
1990
+ data_files:
1991
+ - split: test
1992
+ path: management/test-*
1993
+ - split: validation
1994
+ path: management/validation-*
1995
+ - split: dev
1996
+ path: management/dev-*
1997
+ - config_name: marketing
1998
+ data_files:
1999
+ - split: test
2000
+ path: marketing/test-*
2001
+ - split: validation
2002
+ path: marketing/validation-*
2003
+ - split: dev
2004
+ path: marketing/dev-*
2005
+ - config_name: medical_genetics
2006
+ data_files:
2007
+ - split: test
2008
+ path: medical_genetics/test-*
2009
+ - split: validation
2010
+ path: medical_genetics/validation-*
2011
+ - split: dev
2012
+ path: medical_genetics/dev-*
2013
+ - config_name: miscellaneous
2014
+ data_files:
2015
+ - split: test
2016
+ path: miscellaneous/test-*
2017
+ - split: validation
2018
+ path: miscellaneous/validation-*
2019
+ - split: dev
2020
+ path: miscellaneous/dev-*
2021
+ - config_name: moral_disputes
2022
+ data_files:
2023
+ - split: test
2024
+ path: moral_disputes/test-*
2025
+ - split: validation
2026
+ path: moral_disputes/validation-*
2027
+ - split: dev
2028
+ path: moral_disputes/dev-*
2029
+ - config_name: moral_scenarios
2030
+ data_files:
2031
+ - split: test
2032
+ path: moral_scenarios/test-*
2033
+ - split: validation
2034
+ path: moral_scenarios/validation-*
2035
+ - split: dev
2036
+ path: moral_scenarios/dev-*
2037
+ - config_name: nutrition
2038
+ data_files:
2039
+ - split: test
2040
+ path: nutrition/test-*
2041
+ - split: validation
2042
+ path: nutrition/validation-*
2043
+ - split: dev
2044
+ path: nutrition/dev-*
2045
+ - config_name: philosophy
2046
+ data_files:
2047
+ - split: test
2048
+ path: philosophy/test-*
2049
+ - split: validation
2050
+ path: philosophy/validation-*
2051
+ - split: dev
2052
+ path: philosophy/dev-*
2053
+ - config_name: prehistory
2054
+ data_files:
2055
+ - split: test
2056
+ path: prehistory/test-*
2057
+ - split: validation
2058
+ path: prehistory/validation-*
2059
+ - split: dev
2060
+ path: prehistory/dev-*
2061
+ - config_name: professional_accounting
2062
+ data_files:
2063
+ - split: test
2064
+ path: professional_accounting/test-*
2065
+ - split: validation
2066
+ path: professional_accounting/validation-*
2067
+ - split: dev
2068
+ path: professional_accounting/dev-*
2069
+ - config_name: professional_law
2070
+ data_files:
2071
+ - split: test
2072
+ path: professional_law/test-*
2073
+ - split: validation
2074
+ path: professional_law/validation-*
2075
+ - split: dev
2076
+ path: professional_law/dev-*
2077
+ - config_name: professional_medicine
2078
+ data_files:
2079
+ - split: test
2080
+ path: professional_medicine/test-*
2081
+ - split: validation
2082
+ path: professional_medicine/validation-*
2083
+ - split: dev
2084
+ path: professional_medicine/dev-*
2085
+ - config_name: professional_psychology
2086
+ data_files:
2087
+ - split: test
2088
+ path: professional_psychology/test-*
2089
+ - split: validation
2090
+ path: professional_psychology/validation-*
2091
+ - split: dev
2092
+ path: professional_psychology/dev-*
2093
+ - config_name: public_relations
2094
+ data_files:
2095
+ - split: test
2096
+ path: public_relations/test-*
2097
+ - split: validation
2098
+ path: public_relations/validation-*
2099
+ - split: dev
2100
+ path: public_relations/dev-*
2101
+ - config_name: security_studies
2102
+ data_files:
2103
+ - split: test
2104
+ path: security_studies/test-*
2105
+ - split: validation
2106
+ path: security_studies/validation-*
2107
+ - split: dev
2108
+ path: security_studies/dev-*
2109
+ - config_name: sociology
2110
+ data_files:
2111
+ - split: test
2112
+ path: sociology/test-*
2113
+ - split: validation
2114
+ path: sociology/validation-*
2115
+ - split: dev
2116
+ path: sociology/dev-*
2117
+ - config_name: us_foreign_policy
2118
+ data_files:
2119
+ - split: test
2120
+ path: us_foreign_policy/test-*
2121
+ - split: validation
2122
+ path: us_foreign_policy/validation-*
2123
+ - split: dev
2124
+ path: us_foreign_policy/dev-*
2125
+ - config_name: virology
2126
+ data_files:
2127
+ - split: test
2128
+ path: virology/test-*
2129
+ - split: validation
2130
+ path: virology/validation-*
2131
+ - split: dev
2132
+ path: virology/dev-*
2133
+ - config_name: world_religions
2134
+ data_files:
2135
+ - split: test
2136
+ path: world_religions/test-*
2137
+ - split: validation
2138
+ path: world_religions/validation-*
2139
+ - split: dev
2140
+ path: world_religions/dev-*
2141
+ ---
2142
+
2143
+ # Dataset Card for MMLU
2144
+
2145
+ ## Table of Contents
2146
+ - [Table of Contents](#table-of-contents)
2147
+ - [Dataset Description](#dataset-description)
2148
+ - [Dataset Summary](#dataset-summary)
2149
+ - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
2150
+ - [Languages](#languages)
2151
+ - [Dataset Structure](#dataset-structure)
2152
+ - [Data Instances](#data-instances)
2153
+ - [Data Fields](#data-fields)
2154
+ - [Data Splits](#data-splits)
2155
+ - [Dataset Creation](#dataset-creation)
2156
+ - [Curation Rationale](#curation-rationale)
2157
+ - [Source Data](#source-data)
2158
+ - [Annotations](#annotations)
2159
+ - [Personal and Sensitive Information](#personal-and-sensitive-information)
2160
+ - [Considerations for Using the Data](#considerations-for-using-the-data)
2161
+ - [Social Impact of Dataset](#social-impact-of-dataset)
2162
+ - [Discussion of Biases](#discussion-of-biases)
2163
+ - [Other Known Limitations](#other-known-limitations)
2164
+ - [Additional Information](#additional-information)
2165
+ - [Dataset Curators](#dataset-curators)
2166
+ - [Licensing Information](#licensing-information)
2167
+ - [Citation Information](#citation-information)
2168
+ - [Contributions](#contributions)
2169
+
2170
+ ## Dataset Description
2171
+
2172
+ - **Repository**: https://github.com/hendrycks/test
2173
+ - **Paper**: https://arxiv.org/abs/2009.03300
2174
+
2175
+ ### Dataset Summary
2176
+
2177
+ [Measuring Massive Multitask Language Understanding](https://arxiv.org/pdf/2009.03300) by [Dan Hendrycks](https://people.eecs.berkeley.edu/~hendrycks/), [Collin Burns](http://collinpburns.com), [Steven Basart](https://stevenbas.art), Andy Zou, Mantas Mazeika, [Dawn Song](https://people.eecs.berkeley.edu/~dawnsong/), and [Jacob Steinhardt](https://www.stat.berkeley.edu/~jsteinhardt/) (ICLR 2021).
2178
+
2179
+ This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn. This covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models must possess extensive world knowledge and problem solving ability.
2180
+
2181
+ A complete list of tasks: ['abstract_algebra', 'anatomy', 'astronomy', 'business_ethics', 'clinical_knowledge', 'college_biology', 'college_chemistry', 'college_computer_science', 'college_mathematics', 'college_medicine', 'college_physics', 'computer_security', 'conceptual_physics', 'econometrics', 'electrical_engineering', 'elementary_mathematics', 'formal_logic', 'global_facts', 'high_school_biology', 'high_school_chemistry', 'high_school_computer_science', 'high_school_european_history', 'high_school_geography', 'high_school_government_and_politics', 'high_school_macroeconomics', 'high_school_mathematics', 'high_school_microeconomics', 'high_school_physics', 'high_school_psychology', 'high_school_statistics', 'high_school_us_history', 'high_school_world_history', 'human_aging', 'human_sexuality', 'international_law', 'jurisprudence', 'logical_fallacies', 'machine_learning', 'management', 'marketing', 'medical_genetics', 'miscellaneous', 'moral_disputes', 'moral_scenarios', 'nutrition', 'philosophy', 'prehistory', 'professional_accounting', 'professional_law', 'professional_medicine', 'professional_psychology', 'public_relations', 'security_studies', 'sociology', 'us_foreign_policy', 'virology', 'world_religions']
2182
+
2183
+ ### Supported Tasks and Leaderboards
2184
+
2185
+ | Model | Authors | Humanities | Social Science | STEM | Other | Average |
2186
+ |------------------------------------|----------|:-------:|:-------:|:-------:|:-------:|:-------:|
2187
+ | [UnifiedQA](https://arxiv.org/abs/2005.00700) | Khashabi et al., 2020 | 45.6 | 56.6 | 40.2 | 54.6 | 48.9
2188
+ | [GPT-3](https://arxiv.org/abs/2005.14165) (few-shot) | Brown et al., 2020 | 40.8 | 50.4 | 36.7 | 48.8 | 43.9
2189
+ | [GPT-2](https://arxiv.org/abs/2005.14165) | Radford et al., 2019 | 32.8 | 33.3 | 30.2 | 33.1 | 32.4
2190
+ | Random Baseline | N/A | 25.0 | 25.0 | 25.0 | 25.0 | 25.0 | 25.0
2191
+
2192
+ ### Languages
2193
+
2194
+ English
2195
+
2196
+ ## Dataset Structure
2197
+
2198
+ ### Data Instances
2199
+
2200
+ An example from anatomy subtask looks as follows:
2201
+ ```
2202
+ {
2203
+ "question": "What is the embryological origin of the hyoid bone?",
2204
+ "choices": ["The first pharyngeal arch", "The first and second pharyngeal arches", "The second pharyngeal arch", "The second and third pharyngeal arches"],
2205
+ "answer": "D"
2206
+ }
2207
+ ```
2208
+
2209
+ ### Data Fields
2210
+
2211
+ - `question`: a string feature
2212
+ - `choices`: a list of 4 string features
2213
+ - `answer`: a ClassLabel feature
2214
+
2215
+ ### Data Splits
2216
+
2217
+ - `auxiliary_train`: auxiliary multiple-choice training questions from ARC, MC_TEST, OBQA, RACE, etc.
2218
+ - `dev`: 5 examples per subtask, meant for few-shot setting
2219
+ - `test`: there are at least 100 examples per subtask
2220
+
2221
+ | | auxiliary_train | dev | val | test |
2222
+ | ----- | :------: | :-----: | :-----: | :-----: |
2223
+ | TOTAL | 99842 | 285 | 1531 | 14042
2224
+
2225
+ ## Dataset Creation
2226
+
2227
+ ### Curation Rationale
2228
+
2229
+ Transformer models have driven this recent progress by pretraining on massive text corpora, including all of Wikipedia, thousands of books, and numerous websites. These models consequently see extensive information about specialized topics, most of which is not assessed by existing NLP benchmarks. To bridge the gap between the wide-ranging knowledge that models see during pretraining and the existing measures of success, we introduce a new benchmark for assessing models across a diverse set of subjects that humans learn.
2230
+
2231
+ ### Source Data
2232
+
2233
+ #### Initial Data Collection and Normalization
2234
+
2235
+ [More Information Needed]
2236
+
2237
+ #### Who are the source language producers?
2238
+
2239
+ [More Information Needed]
2240
+
2241
+ ### Annotations
2242
+
2243
+ #### Annotation process
2244
+
2245
+ [More Information Needed]
2246
+
2247
+ #### Who are the annotators?
2248
+
2249
+ [More Information Needed]
2250
+
2251
+ ### Personal and Sensitive Information
2252
+
2253
+ [More Information Needed]
2254
+
2255
+ ## Considerations for Using the Data
2256
+
2257
+ ### Social Impact of Dataset
2258
+
2259
+ [More Information Needed]
2260
+
2261
+ ### Discussion of Biases
2262
+
2263
+ [More Information Needed]
2264
+
2265
+ ### Other Known Limitations
2266
+
2267
+ [More Information Needed]
2268
+
2269
+ ## Additional Information
2270
+
2271
+ ### Dataset Curators
2272
+
2273
+ [More Information Needed]
2274
+
2275
+ ### Licensing Information
2276
+
2277
+ [MIT License](https://github.com/hendrycks/test/blob/master/LICENSE)
2278
+
2279
+ ### Citation Information
2280
+
2281
+ If you find this useful in your research, please consider citing the test and also the [ETHICS](https://arxiv.org/abs/2008.02275) dataset it draws from:
2282
+ ```
2283
+ @article{hendryckstest2021,
2284
+ title={Measuring Massive Multitask Language Understanding},
2285
+ author={Dan Hendrycks and Collin Burns and Steven Basart and Andy Zou and Mantas Mazeika and Dawn Song and Jacob Steinhardt},
2286
+ journal={Proceedings of the International Conference on Learning Representations (ICLR)},
2287
+ year={2021}
2288
+ }
2289
+
2290
+ @article{hendrycks2021ethics,
2291
+ title={Aligning AI With Shared Human Values},
2292
+ author={Dan Hendrycks and Collin Burns and Steven Basart and Andrew Critch and Jerry Li and Dawn Song and Jacob Steinhardt},
2293
+ journal={Proceedings of the International Conference on Learning Representations (ICLR)},
2294
+ year={2021}
2295
+ }
2296
+ ```
2297
+ ### Contributions
2298
+
2299
+ Thanks to [@andyzoujm](https://github.com/andyzoujm) for adding this dataset.
hf_cache/hub/datasets--cais--mmlu/blobs/e133c92a3269e646dfd034d5c94ad41a31c7b194 ADDED
The diff for this file is too large to render. See raw diff