Add files using upload-large-folder tool
Browse filesThis view is limited to 50 files because it contains too many changes. See raw diff
- hf_cache/hub/CACHEDIR.TAG +4 -0
- hf_cache/hub/datasets--EdinburghNLP--xsum/.no_exist/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/.huggingface.yaml +0 -0
- hf_cache/hub/datasets--EdinburghNLP--xsum/.no_exist/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/dataset_infos.json +0 -0
- hf_cache/hub/datasets--EdinburghNLP--xsum/.no_exist/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/xsum.py +0 -0
- hf_cache/hub/datasets--EdinburghNLP--xsum/blobs/fee00d9f711981e883d7a06af05d4ee18b7fe5d9 +222 -0
- hf_cache/hub/datasets--EdinburghNLP--xsum/refs/main +1 -0
- hf_cache/hub/datasets--EdinburghNLP--xsum/snapshots/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/README.md +222 -0
- hf_cache/hub/datasets--HuggingFaceFW--fineweb-2/blobs/67fa474443e4c2a3a8f1b9f19502898a1d86ef29 +0 -0
- hf_cache/hub/datasets--HuggingFaceFW--fineweb-2/refs/main +1 -0
- hf_cache/hub/datasets--HuggingFaceFW--fineweb-2/snapshots/af9c13333eb981300149d5ca60a8e9d659b276b9/README.md +0 -0
- hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/.no_exist/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/.huggingface.yaml +0 -0
- hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/.no_exist/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/dataset_infos.json +0 -0
- hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/.no_exist/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/fineweb-edu.py +0 -0
- hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/blobs/16e64d907529c7f742578a2fee4b72039d64211d +651 -0
- hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/refs/main +1 -0
- hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/snapshots/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/README.md +651 -0
- hf_cache/hub/datasets--Rowan--hellaswag/.no_exist/218ec52e09a7e7462a5400043bb9a69a41d06b76/.huggingface.yaml +0 -0
- hf_cache/hub/datasets--Rowan--hellaswag/.no_exist/218ec52e09a7e7462a5400043bb9a69a41d06b76/dataset_infos.json +0 -0
- hf_cache/hub/datasets--Rowan--hellaswag/.no_exist/218ec52e09a7e7462a5400043bb9a69a41d06b76/hellaswag.py +0 -0
- hf_cache/hub/datasets--Rowan--hellaswag/blobs/29f11d90eb3a5b319cfe8ce2a4e78d9f2a1aea3f +218 -0
- hf_cache/hub/datasets--Rowan--hellaswag/refs/main +1 -0
- hf_cache/hub/datasets--Rowan--hellaswag/snapshots/218ec52e09a7e7462a5400043bb9a69a41d06b76/README.md +218 -0
- hf_cache/hub/datasets--Salesforce--wikitext/.no_exist/b08601e04326c79dfdd32d625aee71d232d685c3/.huggingface.yaml +0 -0
- hf_cache/hub/datasets--Salesforce--wikitext/.no_exist/b08601e04326c79dfdd32d625aee71d232d685c3/dataset_infos.json +0 -0
- hf_cache/hub/datasets--Salesforce--wikitext/.no_exist/b08601e04326c79dfdd32d625aee71d232d685c3/wikitext.py +0 -0
- hf_cache/hub/datasets--Salesforce--wikitext/blobs/2a4fec2bc8df76c9d4da1c8e8865b625eb221c76 +344 -0
- hf_cache/hub/datasets--Salesforce--wikitext/refs/main +1 -0
- hf_cache/hub/datasets--Salesforce--wikitext/snapshots/b08601e04326c79dfdd32d625aee71d232d685c3/README.md +344 -0
- hf_cache/hub/datasets--abisee--cnn_dailymail/.no_exist/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/.huggingface.yaml +0 -0
- hf_cache/hub/datasets--abisee--cnn_dailymail/.no_exist/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/cnn_dailymail.py +0 -0
- hf_cache/hub/datasets--abisee--cnn_dailymail/.no_exist/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/dataset_infos.json +0 -0
- hf_cache/hub/datasets--abisee--cnn_dailymail/blobs/feadf7d245b4f6818e20b4cf65841d03b5703d47 +305 -0
- hf_cache/hub/datasets--abisee--cnn_dailymail/refs/main +1 -0
- hf_cache/hub/datasets--abisee--cnn_dailymail/snapshots/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/README.md +305 -0
- hf_cache/hub/datasets--allenai--openbookqa/.no_exist/388097ea7776314e93a529163e0fea805b8a6454/.huggingface.yaml +0 -0
- hf_cache/hub/datasets--allenai--openbookqa/.no_exist/388097ea7776314e93a529163e0fea805b8a6454/dataset_infos.json +0 -0
- hf_cache/hub/datasets--allenai--openbookqa/.no_exist/388097ea7776314e93a529163e0fea805b8a6454/openbookqa.py +0 -0
- hf_cache/hub/datasets--allenai--openbookqa/blobs/08128898cc0433b97f7a9c9ff09c5054c8587e3e +301 -0
- hf_cache/hub/datasets--allenai--openbookqa/refs/main +1 -0
- hf_cache/hub/datasets--allenai--openbookqa/snapshots/388097ea7776314e93a529163e0fea805b8a6454/README.md +301 -0
- hf_cache/hub/datasets--allenai--sciq/.no_exist/2c94ad3e1aafab77146f384e23536f97a4849815/.huggingface.yaml +0 -0
- hf_cache/hub/datasets--allenai--sciq/.no_exist/2c94ad3e1aafab77146f384e23536f97a4849815/dataset_infos.json +0 -0
- hf_cache/hub/datasets--allenai--sciq/.no_exist/2c94ad3e1aafab77146f384e23536f97a4849815/sciq.py +0 -0
- hf_cache/hub/datasets--allenai--sciq/blobs/c644057869cabcde87a2b5ab9665ec0d0bd1405b +216 -0
- hf_cache/hub/datasets--allenai--sciq/refs/main +1 -0
- hf_cache/hub/datasets--allenai--sciq/snapshots/2c94ad3e1aafab77146f384e23536f97a4849815/README.md +216 -0
- hf_cache/hub/datasets--cais--mmlu/.no_exist/c30699e8356da336a370243923dbaf21066bb9fe/.huggingface.yaml +0 -0
- hf_cache/hub/datasets--cais--mmlu/.no_exist/c30699e8356da336a370243923dbaf21066bb9fe/mmlu.py +0 -0
- hf_cache/hub/datasets--cais--mmlu/blobs/08de94c560ad7420252bff6e4729f1d1683def4f +2299 -0
- hf_cache/hub/datasets--cais--mmlu/blobs/e133c92a3269e646dfd034d5c94ad41a31c7b194 +0 -0
hf_cache/hub/CACHEDIR.TAG
ADDED
|
@@ -0,0 +1,4 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Signature: 8a477f597d28d172789f06886806bc55
|
| 2 |
+
# This file is a cache directory tag created by huggingface_hub.
|
| 3 |
+
# For information about cache directory tags, see:
|
| 4 |
+
# https://bford.info/cachedir/
|
hf_cache/hub/datasets--EdinburghNLP--xsum/.no_exist/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/.huggingface.yaml
ADDED
|
File without changes
|
hf_cache/hub/datasets--EdinburghNLP--xsum/.no_exist/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/dataset_infos.json
ADDED
|
File without changes
|
hf_cache/hub/datasets--EdinburghNLP--xsum/.no_exist/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/xsum.py
ADDED
|
File without changes
|
hf_cache/hub/datasets--EdinburghNLP--xsum/blobs/fee00d9f711981e883d7a06af05d4ee18b7fe5d9
ADDED
|
@@ -0,0 +1,222 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
annotations_creators:
|
| 3 |
+
- found
|
| 4 |
+
language_creators:
|
| 5 |
+
- found
|
| 6 |
+
language:
|
| 7 |
+
- en
|
| 8 |
+
license:
|
| 9 |
+
- unknown
|
| 10 |
+
multilinguality:
|
| 11 |
+
- monolingual
|
| 12 |
+
pretty_name: Extreme Summarization (XSum)
|
| 13 |
+
paperswithcode_id: xsum
|
| 14 |
+
size_categories:
|
| 15 |
+
- 100K<n<1M
|
| 16 |
+
source_datasets:
|
| 17 |
+
- original
|
| 18 |
+
task_categories:
|
| 19 |
+
- summarization
|
| 20 |
+
task_ids:
|
| 21 |
+
- news-articles-summarization
|
| 22 |
+
dataset_info:
|
| 23 |
+
features:
|
| 24 |
+
- name: document
|
| 25 |
+
dtype: string
|
| 26 |
+
- name: summary
|
| 27 |
+
dtype: string
|
| 28 |
+
- name: id
|
| 29 |
+
dtype: string
|
| 30 |
+
splits:
|
| 31 |
+
- name: train
|
| 32 |
+
num_bytes: 479206363
|
| 33 |
+
num_examples: 204045
|
| 34 |
+
- name: validation
|
| 35 |
+
num_bytes: 26292877
|
| 36 |
+
num_examples: 11332
|
| 37 |
+
- name: test
|
| 38 |
+
num_bytes: 26756141
|
| 39 |
+
num_examples: 11334
|
| 40 |
+
download_size: 332791351
|
| 41 |
+
dataset_size: 532255381
|
| 42 |
+
configs:
|
| 43 |
+
- config_name: default
|
| 44 |
+
data_files:
|
| 45 |
+
- split: train
|
| 46 |
+
path: data/train-*
|
| 47 |
+
- split: validation
|
| 48 |
+
path: data/validation-*
|
| 49 |
+
- split: test
|
| 50 |
+
path: data/test-*
|
| 51 |
+
train-eval-index:
|
| 52 |
+
- config: default
|
| 53 |
+
task: summarization
|
| 54 |
+
task_id: summarization
|
| 55 |
+
splits:
|
| 56 |
+
train_split: train
|
| 57 |
+
eval_split: test
|
| 58 |
+
col_mapping:
|
| 59 |
+
document: text
|
| 60 |
+
summary: target
|
| 61 |
+
metrics:
|
| 62 |
+
- type: rouge
|
| 63 |
+
name: Rouge
|
| 64 |
+
---
|
| 65 |
+
|
| 66 |
+
# Dataset Card for "xsum"
|
| 67 |
+
|
| 68 |
+
## Table of Contents
|
| 69 |
+
- [Dataset Description](#dataset-description)
|
| 70 |
+
- [Dataset Summary](#dataset-summary)
|
| 71 |
+
- [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
|
| 72 |
+
- [Languages](#languages)
|
| 73 |
+
- [Dataset Structure](#dataset-structure)
|
| 74 |
+
- [Data Instances](#data-instances)
|
| 75 |
+
- [Data Fields](#data-fields)
|
| 76 |
+
- [Data Splits](#data-splits)
|
| 77 |
+
- [Dataset Creation](#dataset-creation)
|
| 78 |
+
- [Curation Rationale](#curation-rationale)
|
| 79 |
+
- [Source Data](#source-data)
|
| 80 |
+
- [Annotations](#annotations)
|
| 81 |
+
- [Personal and Sensitive Information](#personal-and-sensitive-information)
|
| 82 |
+
- [Considerations for Using the Data](#considerations-for-using-the-data)
|
| 83 |
+
- [Social Impact of Dataset](#social-impact-of-dataset)
|
| 84 |
+
- [Discussion of Biases](#discussion-of-biases)
|
| 85 |
+
- [Other Known Limitations](#other-known-limitations)
|
| 86 |
+
- [Additional Information](#additional-information)
|
| 87 |
+
- [Dataset Curators](#dataset-curators)
|
| 88 |
+
- [Licensing Information](#licensing-information)
|
| 89 |
+
- [Citation Information](#citation-information)
|
| 90 |
+
- [Contributions](#contributions)
|
| 91 |
+
|
| 92 |
+
## Dataset Description
|
| 93 |
+
|
| 94 |
+
- **Homepage:**
|
| 95 |
+
- **Repository:** https://github.com/EdinburghNLP/XSum
|
| 96 |
+
- **Paper:** [Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization](https://arxiv.org/abs/1808.08745)
|
| 97 |
+
- **Point of Contact:** [Shashi Narayan](mailto:shashi.narayan@ed.ac.uk)
|
| 98 |
+
- **Size of downloaded dataset files:** 257.30 MB
|
| 99 |
+
- **Size of the generated dataset:** 532.26 MB
|
| 100 |
+
- **Total amount of disk used:** 789.56 MB
|
| 101 |
+
|
| 102 |
+
### Dataset Summary
|
| 103 |
+
|
| 104 |
+
Extreme Summarization (XSum) Dataset.
|
| 105 |
+
|
| 106 |
+
There are three features:
|
| 107 |
+
- document: Input news article.
|
| 108 |
+
- summary: One sentence summary of the article.
|
| 109 |
+
- id: BBC ID of the article.
|
| 110 |
+
|
| 111 |
+
### Supported Tasks and Leaderboards
|
| 112 |
+
|
| 113 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 114 |
+
|
| 115 |
+
### Languages
|
| 116 |
+
|
| 117 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 118 |
+
|
| 119 |
+
## Dataset Structure
|
| 120 |
+
|
| 121 |
+
### Data Instances
|
| 122 |
+
|
| 123 |
+
#### default
|
| 124 |
+
|
| 125 |
+
- **Size of downloaded dataset files:** 257.30 MB
|
| 126 |
+
- **Size of the generated dataset:** 532.26 MB
|
| 127 |
+
- **Total amount of disk used:** 789.56 MB
|
| 128 |
+
|
| 129 |
+
An example of 'validation' looks as follows.
|
| 130 |
+
```
|
| 131 |
+
{
|
| 132 |
+
"document": "some-body",
|
| 133 |
+
"id": "29750031",
|
| 134 |
+
"summary": "some-sentence"
|
| 135 |
+
}
|
| 136 |
+
```
|
| 137 |
+
|
| 138 |
+
### Data Fields
|
| 139 |
+
|
| 140 |
+
The data fields are the same among all splits.
|
| 141 |
+
|
| 142 |
+
#### default
|
| 143 |
+
- `document`: a `string` feature.
|
| 144 |
+
- `summary`: a `string` feature.
|
| 145 |
+
- `id`: a `string` feature.
|
| 146 |
+
|
| 147 |
+
### Data Splits
|
| 148 |
+
|
| 149 |
+
| name |train |validation|test |
|
| 150 |
+
|-------|-----:|---------:|----:|
|
| 151 |
+
|default|204045| 11332|11334|
|
| 152 |
+
|
| 153 |
+
## Dataset Creation
|
| 154 |
+
|
| 155 |
+
### Curation Rationale
|
| 156 |
+
|
| 157 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 158 |
+
|
| 159 |
+
### Source Data
|
| 160 |
+
|
| 161 |
+
#### Initial Data Collection and Normalization
|
| 162 |
+
|
| 163 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 164 |
+
|
| 165 |
+
#### Who are the source language producers?
|
| 166 |
+
|
| 167 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 168 |
+
|
| 169 |
+
### Annotations
|
| 170 |
+
|
| 171 |
+
#### Annotation process
|
| 172 |
+
|
| 173 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 174 |
+
|
| 175 |
+
#### Who are the annotators?
|
| 176 |
+
|
| 177 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 178 |
+
|
| 179 |
+
### Personal and Sensitive Information
|
| 180 |
+
|
| 181 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 182 |
+
|
| 183 |
+
## Considerations for Using the Data
|
| 184 |
+
|
| 185 |
+
### Social Impact of Dataset
|
| 186 |
+
|
| 187 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 188 |
+
|
| 189 |
+
### Discussion of Biases
|
| 190 |
+
|
| 191 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 192 |
+
|
| 193 |
+
### Other Known Limitations
|
| 194 |
+
|
| 195 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 196 |
+
|
| 197 |
+
## Additional Information
|
| 198 |
+
|
| 199 |
+
### Dataset Curators
|
| 200 |
+
|
| 201 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 202 |
+
|
| 203 |
+
### Licensing Information
|
| 204 |
+
|
| 205 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 206 |
+
|
| 207 |
+
### Citation Information
|
| 208 |
+
|
| 209 |
+
```
|
| 210 |
+
@article{Narayan2018DontGM,
|
| 211 |
+
title={Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization},
|
| 212 |
+
author={Shashi Narayan and Shay B. Cohen and Mirella Lapata},
|
| 213 |
+
journal={ArXiv},
|
| 214 |
+
year={2018},
|
| 215 |
+
volume={abs/1808.08745}
|
| 216 |
+
}
|
| 217 |
+
```
|
| 218 |
+
|
| 219 |
+
|
| 220 |
+
### Contributions
|
| 221 |
+
|
| 222 |
+
Thanks to [@thomwolf](https://github.com/thomwolf), [@lewtun](https://github.com/lewtun), [@mariamabarham](https://github.com/mariamabarham), [@jbragg](https://github.com/jbragg), [@lhoestq](https://github.com/lhoestq), [@patrickvonplaten](https://github.com/patrickvonplaten) for adding this dataset.
|
hf_cache/hub/datasets--EdinburghNLP--xsum/refs/main
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
7d4d486c2f8ef850b1a11aead99b894ff3dd7da9
|
hf_cache/hub/datasets--EdinburghNLP--xsum/snapshots/7d4d486c2f8ef850b1a11aead99b894ff3dd7da9/README.md
ADDED
|
@@ -0,0 +1,222 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
annotations_creators:
|
| 3 |
+
- found
|
| 4 |
+
language_creators:
|
| 5 |
+
- found
|
| 6 |
+
language:
|
| 7 |
+
- en
|
| 8 |
+
license:
|
| 9 |
+
- unknown
|
| 10 |
+
multilinguality:
|
| 11 |
+
- monolingual
|
| 12 |
+
pretty_name: Extreme Summarization (XSum)
|
| 13 |
+
paperswithcode_id: xsum
|
| 14 |
+
size_categories:
|
| 15 |
+
- 100K<n<1M
|
| 16 |
+
source_datasets:
|
| 17 |
+
- original
|
| 18 |
+
task_categories:
|
| 19 |
+
- summarization
|
| 20 |
+
task_ids:
|
| 21 |
+
- news-articles-summarization
|
| 22 |
+
dataset_info:
|
| 23 |
+
features:
|
| 24 |
+
- name: document
|
| 25 |
+
dtype: string
|
| 26 |
+
- name: summary
|
| 27 |
+
dtype: string
|
| 28 |
+
- name: id
|
| 29 |
+
dtype: string
|
| 30 |
+
splits:
|
| 31 |
+
- name: train
|
| 32 |
+
num_bytes: 479206363
|
| 33 |
+
num_examples: 204045
|
| 34 |
+
- name: validation
|
| 35 |
+
num_bytes: 26292877
|
| 36 |
+
num_examples: 11332
|
| 37 |
+
- name: test
|
| 38 |
+
num_bytes: 26756141
|
| 39 |
+
num_examples: 11334
|
| 40 |
+
download_size: 332791351
|
| 41 |
+
dataset_size: 532255381
|
| 42 |
+
configs:
|
| 43 |
+
- config_name: default
|
| 44 |
+
data_files:
|
| 45 |
+
- split: train
|
| 46 |
+
path: data/train-*
|
| 47 |
+
- split: validation
|
| 48 |
+
path: data/validation-*
|
| 49 |
+
- split: test
|
| 50 |
+
path: data/test-*
|
| 51 |
+
train-eval-index:
|
| 52 |
+
- config: default
|
| 53 |
+
task: summarization
|
| 54 |
+
task_id: summarization
|
| 55 |
+
splits:
|
| 56 |
+
train_split: train
|
| 57 |
+
eval_split: test
|
| 58 |
+
col_mapping:
|
| 59 |
+
document: text
|
| 60 |
+
summary: target
|
| 61 |
+
metrics:
|
| 62 |
+
- type: rouge
|
| 63 |
+
name: Rouge
|
| 64 |
+
---
|
| 65 |
+
|
| 66 |
+
# Dataset Card for "xsum"
|
| 67 |
+
|
| 68 |
+
## Table of Contents
|
| 69 |
+
- [Dataset Description](#dataset-description)
|
| 70 |
+
- [Dataset Summary](#dataset-summary)
|
| 71 |
+
- [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
|
| 72 |
+
- [Languages](#languages)
|
| 73 |
+
- [Dataset Structure](#dataset-structure)
|
| 74 |
+
- [Data Instances](#data-instances)
|
| 75 |
+
- [Data Fields](#data-fields)
|
| 76 |
+
- [Data Splits](#data-splits)
|
| 77 |
+
- [Dataset Creation](#dataset-creation)
|
| 78 |
+
- [Curation Rationale](#curation-rationale)
|
| 79 |
+
- [Source Data](#source-data)
|
| 80 |
+
- [Annotations](#annotations)
|
| 81 |
+
- [Personal and Sensitive Information](#personal-and-sensitive-information)
|
| 82 |
+
- [Considerations for Using the Data](#considerations-for-using-the-data)
|
| 83 |
+
- [Social Impact of Dataset](#social-impact-of-dataset)
|
| 84 |
+
- [Discussion of Biases](#discussion-of-biases)
|
| 85 |
+
- [Other Known Limitations](#other-known-limitations)
|
| 86 |
+
- [Additional Information](#additional-information)
|
| 87 |
+
- [Dataset Curators](#dataset-curators)
|
| 88 |
+
- [Licensing Information](#licensing-information)
|
| 89 |
+
- [Citation Information](#citation-information)
|
| 90 |
+
- [Contributions](#contributions)
|
| 91 |
+
|
| 92 |
+
## Dataset Description
|
| 93 |
+
|
| 94 |
+
- **Homepage:**
|
| 95 |
+
- **Repository:** https://github.com/EdinburghNLP/XSum
|
| 96 |
+
- **Paper:** [Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization](https://arxiv.org/abs/1808.08745)
|
| 97 |
+
- **Point of Contact:** [Shashi Narayan](mailto:shashi.narayan@ed.ac.uk)
|
| 98 |
+
- **Size of downloaded dataset files:** 257.30 MB
|
| 99 |
+
- **Size of the generated dataset:** 532.26 MB
|
| 100 |
+
- **Total amount of disk used:** 789.56 MB
|
| 101 |
+
|
| 102 |
+
### Dataset Summary
|
| 103 |
+
|
| 104 |
+
Extreme Summarization (XSum) Dataset.
|
| 105 |
+
|
| 106 |
+
There are three features:
|
| 107 |
+
- document: Input news article.
|
| 108 |
+
- summary: One sentence summary of the article.
|
| 109 |
+
- id: BBC ID of the article.
|
| 110 |
+
|
| 111 |
+
### Supported Tasks and Leaderboards
|
| 112 |
+
|
| 113 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 114 |
+
|
| 115 |
+
### Languages
|
| 116 |
+
|
| 117 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 118 |
+
|
| 119 |
+
## Dataset Structure
|
| 120 |
+
|
| 121 |
+
### Data Instances
|
| 122 |
+
|
| 123 |
+
#### default
|
| 124 |
+
|
| 125 |
+
- **Size of downloaded dataset files:** 257.30 MB
|
| 126 |
+
- **Size of the generated dataset:** 532.26 MB
|
| 127 |
+
- **Total amount of disk used:** 789.56 MB
|
| 128 |
+
|
| 129 |
+
An example of 'validation' looks as follows.
|
| 130 |
+
```
|
| 131 |
+
{
|
| 132 |
+
"document": "some-body",
|
| 133 |
+
"id": "29750031",
|
| 134 |
+
"summary": "some-sentence"
|
| 135 |
+
}
|
| 136 |
+
```
|
| 137 |
+
|
| 138 |
+
### Data Fields
|
| 139 |
+
|
| 140 |
+
The data fields are the same among all splits.
|
| 141 |
+
|
| 142 |
+
#### default
|
| 143 |
+
- `document`: a `string` feature.
|
| 144 |
+
- `summary`: a `string` feature.
|
| 145 |
+
- `id`: a `string` feature.
|
| 146 |
+
|
| 147 |
+
### Data Splits
|
| 148 |
+
|
| 149 |
+
| name |train |validation|test |
|
| 150 |
+
|-------|-----:|---------:|----:|
|
| 151 |
+
|default|204045| 11332|11334|
|
| 152 |
+
|
| 153 |
+
## Dataset Creation
|
| 154 |
+
|
| 155 |
+
### Curation Rationale
|
| 156 |
+
|
| 157 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 158 |
+
|
| 159 |
+
### Source Data
|
| 160 |
+
|
| 161 |
+
#### Initial Data Collection and Normalization
|
| 162 |
+
|
| 163 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 164 |
+
|
| 165 |
+
#### Who are the source language producers?
|
| 166 |
+
|
| 167 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 168 |
+
|
| 169 |
+
### Annotations
|
| 170 |
+
|
| 171 |
+
#### Annotation process
|
| 172 |
+
|
| 173 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 174 |
+
|
| 175 |
+
#### Who are the annotators?
|
| 176 |
+
|
| 177 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 178 |
+
|
| 179 |
+
### Personal and Sensitive Information
|
| 180 |
+
|
| 181 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 182 |
+
|
| 183 |
+
## Considerations for Using the Data
|
| 184 |
+
|
| 185 |
+
### Social Impact of Dataset
|
| 186 |
+
|
| 187 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 188 |
+
|
| 189 |
+
### Discussion of Biases
|
| 190 |
+
|
| 191 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 192 |
+
|
| 193 |
+
### Other Known Limitations
|
| 194 |
+
|
| 195 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 196 |
+
|
| 197 |
+
## Additional Information
|
| 198 |
+
|
| 199 |
+
### Dataset Curators
|
| 200 |
+
|
| 201 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 202 |
+
|
| 203 |
+
### Licensing Information
|
| 204 |
+
|
| 205 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 206 |
+
|
| 207 |
+
### Citation Information
|
| 208 |
+
|
| 209 |
+
```
|
| 210 |
+
@article{Narayan2018DontGM,
|
| 211 |
+
title={Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization},
|
| 212 |
+
author={Shashi Narayan and Shay B. Cohen and Mirella Lapata},
|
| 213 |
+
journal={ArXiv},
|
| 214 |
+
year={2018},
|
| 215 |
+
volume={abs/1808.08745}
|
| 216 |
+
}
|
| 217 |
+
```
|
| 218 |
+
|
| 219 |
+
|
| 220 |
+
### Contributions
|
| 221 |
+
|
| 222 |
+
Thanks to [@thomwolf](https://github.com/thomwolf), [@lewtun](https://github.com/lewtun), [@mariamabarham](https://github.com/mariamabarham), [@jbragg](https://github.com/jbragg), [@lhoestq](https://github.com/lhoestq), [@patrickvonplaten](https://github.com/patrickvonplaten) for adding this dataset.
|
hf_cache/hub/datasets--HuggingFaceFW--fineweb-2/blobs/67fa474443e4c2a3a8f1b9f19502898a1d86ef29
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
hf_cache/hub/datasets--HuggingFaceFW--fineweb-2/refs/main
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
af9c13333eb981300149d5ca60a8e9d659b276b9
|
hf_cache/hub/datasets--HuggingFaceFW--fineweb-2/snapshots/af9c13333eb981300149d5ca60a8e9d659b276b9/README.md
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/.no_exist/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/.huggingface.yaml
ADDED
|
File without changes
|
hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/.no_exist/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/dataset_infos.json
ADDED
|
File without changes
|
hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/.no_exist/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/fineweb-edu.py
ADDED
|
File without changes
|
hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/blobs/16e64d907529c7f742578a2fee4b72039d64211d
ADDED
|
@@ -0,0 +1,651 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: odc-by
|
| 3 |
+
task_categories:
|
| 4 |
+
- text-generation
|
| 5 |
+
language:
|
| 6 |
+
- en
|
| 7 |
+
pretty_name: FineWeb-Edu
|
| 8 |
+
size_categories:
|
| 9 |
+
- n>1T
|
| 10 |
+
configs:
|
| 11 |
+
- config_name: default
|
| 12 |
+
data_files:
|
| 13 |
+
- split: train
|
| 14 |
+
path: data/*/*
|
| 15 |
+
features:
|
| 16 |
+
- name: text
|
| 17 |
+
dtype: string
|
| 18 |
+
- name: id
|
| 19 |
+
dtype: string
|
| 20 |
+
- name: dump
|
| 21 |
+
dtype: string
|
| 22 |
+
- name: url
|
| 23 |
+
dtype: string
|
| 24 |
+
- name: date
|
| 25 |
+
dtype: string
|
| 26 |
+
- name: file_path
|
| 27 |
+
dtype: string
|
| 28 |
+
- name: language
|
| 29 |
+
dtype: string
|
| 30 |
+
- name: language_score
|
| 31 |
+
dtype: float64
|
| 32 |
+
- name: token_count
|
| 33 |
+
dtype: int64
|
| 34 |
+
- name: score
|
| 35 |
+
dtype: float64
|
| 36 |
+
- name: int_score
|
| 37 |
+
dtype: int64
|
| 38 |
+
- config_name: sample-10BT
|
| 39 |
+
data_files:
|
| 40 |
+
- split: train
|
| 41 |
+
path: sample/10BT/*
|
| 42 |
+
- config_name: sample-100BT
|
| 43 |
+
data_files:
|
| 44 |
+
- split: train
|
| 45 |
+
path: sample/100BT/*
|
| 46 |
+
- config_name: sample-350BT
|
| 47 |
+
data_files:
|
| 48 |
+
- split: train
|
| 49 |
+
path: sample/350BT/*
|
| 50 |
+
- config_name: CC-MAIN-2025-05
|
| 51 |
+
data_files:
|
| 52 |
+
- split: train
|
| 53 |
+
path: data/CC-MAIN-2025-05/*
|
| 54 |
+
- config_name: CC-MAIN-2025-08
|
| 55 |
+
data_files:
|
| 56 |
+
- split: train
|
| 57 |
+
path: data/CC-MAIN-2025-08/*
|
| 58 |
+
- config_name: CC-MAIN-2025-13
|
| 59 |
+
data_files:
|
| 60 |
+
- split: train
|
| 61 |
+
path: data/CC-MAIN-2025-13/*
|
| 62 |
+
- config_name: CC-MAIN-2025-18
|
| 63 |
+
data_files:
|
| 64 |
+
- split: train
|
| 65 |
+
path: data/CC-MAIN-2025-18/*
|
| 66 |
+
- config_name: CC-MAIN-2025-21
|
| 67 |
+
data_files:
|
| 68 |
+
- split: train
|
| 69 |
+
path: data/CC-MAIN-2025-21/*
|
| 70 |
+
- config_name: CC-MAIN-2025-26
|
| 71 |
+
data_files:
|
| 72 |
+
- split: train
|
| 73 |
+
path: data/CC-MAIN-2025-26/*
|
| 74 |
+
- config_name: CC-MAIN-2024-51
|
| 75 |
+
data_files:
|
| 76 |
+
- split: train
|
| 77 |
+
path: data/CC-MAIN-2024-51/*
|
| 78 |
+
- config_name: CC-MAIN-2024-46
|
| 79 |
+
data_files:
|
| 80 |
+
- split: train
|
| 81 |
+
path: data/CC-MAIN-2024-46/*
|
| 82 |
+
- config_name: CC-MAIN-2024-42
|
| 83 |
+
data_files:
|
| 84 |
+
- split: train
|
| 85 |
+
path: data/CC-MAIN-2024-42/*
|
| 86 |
+
- config_name: CC-MAIN-2024-38
|
| 87 |
+
data_files:
|
| 88 |
+
- split: train
|
| 89 |
+
path: data/CC-MAIN-2024-38/*
|
| 90 |
+
- config_name: CC-MAIN-2024-33
|
| 91 |
+
data_files:
|
| 92 |
+
- split: train
|
| 93 |
+
path: data/CC-MAIN-2024-33/*
|
| 94 |
+
- config_name: CC-MAIN-2024-30
|
| 95 |
+
data_files:
|
| 96 |
+
- split: train
|
| 97 |
+
path: data/CC-MAIN-2024-30/*
|
| 98 |
+
- config_name: CC-MAIN-2024-26
|
| 99 |
+
data_files:
|
| 100 |
+
- split: train
|
| 101 |
+
path: data/CC-MAIN-2024-26/*
|
| 102 |
+
- config_name: CC-MAIN-2024-22
|
| 103 |
+
data_files:
|
| 104 |
+
- split: train
|
| 105 |
+
path: data/CC-MAIN-2024-22/*
|
| 106 |
+
- config_name: CC-MAIN-2024-18
|
| 107 |
+
data_files:
|
| 108 |
+
- split: train
|
| 109 |
+
path: data/CC-MAIN-2024-18/*
|
| 110 |
+
- config_name: CC-MAIN-2024-10
|
| 111 |
+
data_files:
|
| 112 |
+
- split: train
|
| 113 |
+
path: data/CC-MAIN-2024-10/*
|
| 114 |
+
- config_name: CC-MAIN-2023-50
|
| 115 |
+
data_files:
|
| 116 |
+
- split: train
|
| 117 |
+
path: data/CC-MAIN-2023-50/*
|
| 118 |
+
- config_name: CC-MAIN-2023-40
|
| 119 |
+
data_files:
|
| 120 |
+
- split: train
|
| 121 |
+
path: data/CC-MAIN-2023-40/*
|
| 122 |
+
- config_name: CC-MAIN-2023-23
|
| 123 |
+
data_files:
|
| 124 |
+
- split: train
|
| 125 |
+
path: data/CC-MAIN-2023-23/*
|
| 126 |
+
- config_name: CC-MAIN-2023-14
|
| 127 |
+
data_files:
|
| 128 |
+
- split: train
|
| 129 |
+
path: data/CC-MAIN-2023-14/*
|
| 130 |
+
- config_name: CC-MAIN-2023-06
|
| 131 |
+
data_files:
|
| 132 |
+
- split: train
|
| 133 |
+
path: data/CC-MAIN-2023-06/*
|
| 134 |
+
- config_name: CC-MAIN-2022-49
|
| 135 |
+
data_files:
|
| 136 |
+
- split: train
|
| 137 |
+
path: data/CC-MAIN-2022-49/*
|
| 138 |
+
- config_name: CC-MAIN-2022-40
|
| 139 |
+
data_files:
|
| 140 |
+
- split: train
|
| 141 |
+
path: data/CC-MAIN-2022-40/*
|
| 142 |
+
- config_name: CC-MAIN-2022-33
|
| 143 |
+
data_files:
|
| 144 |
+
- split: train
|
| 145 |
+
path: data/CC-MAIN-2022-33/*
|
| 146 |
+
- config_name: CC-MAIN-2022-27
|
| 147 |
+
data_files:
|
| 148 |
+
- split: train
|
| 149 |
+
path: data/CC-MAIN-2022-27/*
|
| 150 |
+
- config_name: CC-MAIN-2022-21
|
| 151 |
+
data_files:
|
| 152 |
+
- split: train
|
| 153 |
+
path: data/CC-MAIN-2022-21/*
|
| 154 |
+
- config_name: CC-MAIN-2022-05
|
| 155 |
+
data_files:
|
| 156 |
+
- split: train
|
| 157 |
+
path: data/CC-MAIN-2022-05/*
|
| 158 |
+
- config_name: CC-MAIN-2021-49
|
| 159 |
+
data_files:
|
| 160 |
+
- split: train
|
| 161 |
+
path: data/CC-MAIN-2021-49/*
|
| 162 |
+
- config_name: CC-MAIN-2021-43
|
| 163 |
+
data_files:
|
| 164 |
+
- split: train
|
| 165 |
+
path: data/CC-MAIN-2021-43/*
|
| 166 |
+
- config_name: CC-MAIN-2021-39
|
| 167 |
+
data_files:
|
| 168 |
+
- split: train
|
| 169 |
+
path: data/CC-MAIN-2021-39/*
|
| 170 |
+
- config_name: CC-MAIN-2021-31
|
| 171 |
+
data_files:
|
| 172 |
+
- split: train
|
| 173 |
+
path: data/CC-MAIN-2021-31/*
|
| 174 |
+
- config_name: CC-MAIN-2021-25
|
| 175 |
+
data_files:
|
| 176 |
+
- split: train
|
| 177 |
+
path: data/CC-MAIN-2021-25/*
|
| 178 |
+
- config_name: CC-MAIN-2021-21
|
| 179 |
+
data_files:
|
| 180 |
+
- split: train
|
| 181 |
+
path: data/CC-MAIN-2021-21/*
|
| 182 |
+
- config_name: CC-MAIN-2021-17
|
| 183 |
+
data_files:
|
| 184 |
+
- split: train
|
| 185 |
+
path: data/CC-MAIN-2021-17/*
|
| 186 |
+
- config_name: CC-MAIN-2021-10
|
| 187 |
+
data_files:
|
| 188 |
+
- split: train
|
| 189 |
+
path: data/CC-MAIN-2021-10/*
|
| 190 |
+
- config_name: CC-MAIN-2021-04
|
| 191 |
+
data_files:
|
| 192 |
+
- split: train
|
| 193 |
+
path: data/CC-MAIN-2021-04/*
|
| 194 |
+
- config_name: CC-MAIN-2020-50
|
| 195 |
+
data_files:
|
| 196 |
+
- split: train
|
| 197 |
+
path: data/CC-MAIN-2020-50/*
|
| 198 |
+
- config_name: CC-MAIN-2020-45
|
| 199 |
+
data_files:
|
| 200 |
+
- split: train
|
| 201 |
+
path: data/CC-MAIN-2020-45/*
|
| 202 |
+
- config_name: CC-MAIN-2020-40
|
| 203 |
+
data_files:
|
| 204 |
+
- split: train
|
| 205 |
+
path: data/CC-MAIN-2020-40/*
|
| 206 |
+
- config_name: CC-MAIN-2020-34
|
| 207 |
+
data_files:
|
| 208 |
+
- split: train
|
| 209 |
+
path: data/CC-MAIN-2020-34/*
|
| 210 |
+
- config_name: CC-MAIN-2020-29
|
| 211 |
+
data_files:
|
| 212 |
+
- split: train
|
| 213 |
+
path: data/CC-MAIN-2020-29/*
|
| 214 |
+
- config_name: CC-MAIN-2020-24
|
| 215 |
+
data_files:
|
| 216 |
+
- split: train
|
| 217 |
+
path: data/CC-MAIN-2020-24/*
|
| 218 |
+
- config_name: CC-MAIN-2020-16
|
| 219 |
+
data_files:
|
| 220 |
+
- split: train
|
| 221 |
+
path: data/CC-MAIN-2020-16/*
|
| 222 |
+
- config_name: CC-MAIN-2020-10
|
| 223 |
+
data_files:
|
| 224 |
+
- split: train
|
| 225 |
+
path: data/CC-MAIN-2020-10/*
|
| 226 |
+
- config_name: CC-MAIN-2020-05
|
| 227 |
+
data_files:
|
| 228 |
+
- split: train
|
| 229 |
+
path: data/CC-MAIN-2020-05/*
|
| 230 |
+
- config_name: CC-MAIN-2019-51
|
| 231 |
+
data_files:
|
| 232 |
+
- split: train
|
| 233 |
+
path: data/CC-MAIN-2019-51/*
|
| 234 |
+
- config_name: CC-MAIN-2019-47
|
| 235 |
+
data_files:
|
| 236 |
+
- split: train
|
| 237 |
+
path: data/CC-MAIN-2019-47/*
|
| 238 |
+
- config_name: CC-MAIN-2019-43
|
| 239 |
+
data_files:
|
| 240 |
+
- split: train
|
| 241 |
+
path: data/CC-MAIN-2019-43/*
|
| 242 |
+
- config_name: CC-MAIN-2019-39
|
| 243 |
+
data_files:
|
| 244 |
+
- split: train
|
| 245 |
+
path: data/CC-MAIN-2019-39/*
|
| 246 |
+
- config_name: CC-MAIN-2019-35
|
| 247 |
+
data_files:
|
| 248 |
+
- split: train
|
| 249 |
+
path: data/CC-MAIN-2019-35/*
|
| 250 |
+
- config_name: CC-MAIN-2019-30
|
| 251 |
+
data_files:
|
| 252 |
+
- split: train
|
| 253 |
+
path: data/CC-MAIN-2019-30/*
|
| 254 |
+
- config_name: CC-MAIN-2019-26
|
| 255 |
+
data_files:
|
| 256 |
+
- split: train
|
| 257 |
+
path: data/CC-MAIN-2019-26/*
|
| 258 |
+
- config_name: CC-MAIN-2019-22
|
| 259 |
+
data_files:
|
| 260 |
+
- split: train
|
| 261 |
+
path: data/CC-MAIN-2019-22/*
|
| 262 |
+
- config_name: CC-MAIN-2019-18
|
| 263 |
+
data_files:
|
| 264 |
+
- split: train
|
| 265 |
+
path: data/CC-MAIN-2019-18/*
|
| 266 |
+
- config_name: CC-MAIN-2019-13
|
| 267 |
+
data_files:
|
| 268 |
+
- split: train
|
| 269 |
+
path: data/CC-MAIN-2019-13/*
|
| 270 |
+
- config_name: CC-MAIN-2019-09
|
| 271 |
+
data_files:
|
| 272 |
+
- split: train
|
| 273 |
+
path: data/CC-MAIN-2019-09/*
|
| 274 |
+
- config_name: CC-MAIN-2019-04
|
| 275 |
+
data_files:
|
| 276 |
+
- split: train
|
| 277 |
+
path: data/CC-MAIN-2019-04/*
|
| 278 |
+
- config_name: CC-MAIN-2018-51
|
| 279 |
+
data_files:
|
| 280 |
+
- split: train
|
| 281 |
+
path: data/CC-MAIN-2018-51/*
|
| 282 |
+
- config_name: CC-MAIN-2018-47
|
| 283 |
+
data_files:
|
| 284 |
+
- split: train
|
| 285 |
+
path: data/CC-MAIN-2018-47/*
|
| 286 |
+
- config_name: CC-MAIN-2018-43
|
| 287 |
+
data_files:
|
| 288 |
+
- split: train
|
| 289 |
+
path: data/CC-MAIN-2018-43/*
|
| 290 |
+
- config_name: CC-MAIN-2018-39
|
| 291 |
+
data_files:
|
| 292 |
+
- split: train
|
| 293 |
+
path: data/CC-MAIN-2018-39/*
|
| 294 |
+
- config_name: CC-MAIN-2018-34
|
| 295 |
+
data_files:
|
| 296 |
+
- split: train
|
| 297 |
+
path: data/CC-MAIN-2018-34/*
|
| 298 |
+
- config_name: CC-MAIN-2018-30
|
| 299 |
+
data_files:
|
| 300 |
+
- split: train
|
| 301 |
+
path: data/CC-MAIN-2018-30/*
|
| 302 |
+
- config_name: CC-MAIN-2018-26
|
| 303 |
+
data_files:
|
| 304 |
+
- split: train
|
| 305 |
+
path: data/CC-MAIN-2018-26/*
|
| 306 |
+
- config_name: CC-MAIN-2018-22
|
| 307 |
+
data_files:
|
| 308 |
+
- split: train
|
| 309 |
+
path: data/CC-MAIN-2018-22/*
|
| 310 |
+
- config_name: CC-MAIN-2018-17
|
| 311 |
+
data_files:
|
| 312 |
+
- split: train
|
| 313 |
+
path: data/CC-MAIN-2018-17/*
|
| 314 |
+
- config_name: CC-MAIN-2018-13
|
| 315 |
+
data_files:
|
| 316 |
+
- split: train
|
| 317 |
+
path: data/CC-MAIN-2018-13/*
|
| 318 |
+
- config_name: CC-MAIN-2018-09
|
| 319 |
+
data_files:
|
| 320 |
+
- split: train
|
| 321 |
+
path: data/CC-MAIN-2018-09/*
|
| 322 |
+
- config_name: CC-MAIN-2018-05
|
| 323 |
+
data_files:
|
| 324 |
+
- split: train
|
| 325 |
+
path: data/CC-MAIN-2018-05/*
|
| 326 |
+
- config_name: CC-MAIN-2017-51
|
| 327 |
+
data_files:
|
| 328 |
+
- split: train
|
| 329 |
+
path: data/CC-MAIN-2017-51/*
|
| 330 |
+
- config_name: CC-MAIN-2017-47
|
| 331 |
+
data_files:
|
| 332 |
+
- split: train
|
| 333 |
+
path: data/CC-MAIN-2017-47/*
|
| 334 |
+
- config_name: CC-MAIN-2017-43
|
| 335 |
+
data_files:
|
| 336 |
+
- split: train
|
| 337 |
+
path: data/CC-MAIN-2017-43/*
|
| 338 |
+
- config_name: CC-MAIN-2017-39
|
| 339 |
+
data_files:
|
| 340 |
+
- split: train
|
| 341 |
+
path: data/CC-MAIN-2017-39/*
|
| 342 |
+
- config_name: CC-MAIN-2017-34
|
| 343 |
+
data_files:
|
| 344 |
+
- split: train
|
| 345 |
+
path: data/CC-MAIN-2017-34/*
|
| 346 |
+
- config_name: CC-MAIN-2017-30
|
| 347 |
+
data_files:
|
| 348 |
+
- split: train
|
| 349 |
+
path: data/CC-MAIN-2017-30/*
|
| 350 |
+
- config_name: CC-MAIN-2017-26
|
| 351 |
+
data_files:
|
| 352 |
+
- split: train
|
| 353 |
+
path: data/CC-MAIN-2017-26/*
|
| 354 |
+
- config_name: CC-MAIN-2017-22
|
| 355 |
+
data_files:
|
| 356 |
+
- split: train
|
| 357 |
+
path: data/CC-MAIN-2017-22/*
|
| 358 |
+
- config_name: CC-MAIN-2017-17
|
| 359 |
+
data_files:
|
| 360 |
+
- split: train
|
| 361 |
+
path: data/CC-MAIN-2017-17/*
|
| 362 |
+
- config_name: CC-MAIN-2017-13
|
| 363 |
+
data_files:
|
| 364 |
+
- split: train
|
| 365 |
+
path: data/CC-MAIN-2017-13/*
|
| 366 |
+
- config_name: CC-MAIN-2017-09
|
| 367 |
+
data_files:
|
| 368 |
+
- split: train
|
| 369 |
+
path: data/CC-MAIN-2017-09/*
|
| 370 |
+
- config_name: CC-MAIN-2017-04
|
| 371 |
+
data_files:
|
| 372 |
+
- split: train
|
| 373 |
+
path: data/CC-MAIN-2017-04/*
|
| 374 |
+
- config_name: CC-MAIN-2016-50
|
| 375 |
+
data_files:
|
| 376 |
+
- split: train
|
| 377 |
+
path: data/CC-MAIN-2016-50/*
|
| 378 |
+
- config_name: CC-MAIN-2016-44
|
| 379 |
+
data_files:
|
| 380 |
+
- split: train
|
| 381 |
+
path: data/CC-MAIN-2016-44/*
|
| 382 |
+
- config_name: CC-MAIN-2016-40
|
| 383 |
+
data_files:
|
| 384 |
+
- split: train
|
| 385 |
+
path: data/CC-MAIN-2016-40/*
|
| 386 |
+
- config_name: CC-MAIN-2016-36
|
| 387 |
+
data_files:
|
| 388 |
+
- split: train
|
| 389 |
+
path: data/CC-MAIN-2016-36/*
|
| 390 |
+
- config_name: CC-MAIN-2016-30
|
| 391 |
+
data_files:
|
| 392 |
+
- split: train
|
| 393 |
+
path: data/CC-MAIN-2016-30/*
|
| 394 |
+
- config_name: CC-MAIN-2016-26
|
| 395 |
+
data_files:
|
| 396 |
+
- split: train
|
| 397 |
+
path: data/CC-MAIN-2016-26/*
|
| 398 |
+
- config_name: CC-MAIN-2016-22
|
| 399 |
+
data_files:
|
| 400 |
+
- split: train
|
| 401 |
+
path: data/CC-MAIN-2016-22/*
|
| 402 |
+
- config_name: CC-MAIN-2016-18
|
| 403 |
+
data_files:
|
| 404 |
+
- split: train
|
| 405 |
+
path: data/CC-MAIN-2016-18/*
|
| 406 |
+
- config_name: CC-MAIN-2016-07
|
| 407 |
+
data_files:
|
| 408 |
+
- split: train
|
| 409 |
+
path: data/CC-MAIN-2016-07/*
|
| 410 |
+
- config_name: CC-MAIN-2015-48
|
| 411 |
+
data_files:
|
| 412 |
+
- split: train
|
| 413 |
+
path: data/CC-MAIN-2015-48/*
|
| 414 |
+
- config_name: CC-MAIN-2015-40
|
| 415 |
+
data_files:
|
| 416 |
+
- split: train
|
| 417 |
+
path: data/CC-MAIN-2015-40/*
|
| 418 |
+
- config_name: CC-MAIN-2015-35
|
| 419 |
+
data_files:
|
| 420 |
+
- split: train
|
| 421 |
+
path: data/CC-MAIN-2015-35/*
|
| 422 |
+
- config_name: CC-MAIN-2015-32
|
| 423 |
+
data_files:
|
| 424 |
+
- split: train
|
| 425 |
+
path: data/CC-MAIN-2015-32/*
|
| 426 |
+
- config_name: CC-MAIN-2015-27
|
| 427 |
+
data_files:
|
| 428 |
+
- split: train
|
| 429 |
+
path: data/CC-MAIN-2015-27/*
|
| 430 |
+
- config_name: CC-MAIN-2015-22
|
| 431 |
+
data_files:
|
| 432 |
+
- split: train
|
| 433 |
+
path: data/CC-MAIN-2015-22/*
|
| 434 |
+
- config_name: CC-MAIN-2015-18
|
| 435 |
+
data_files:
|
| 436 |
+
- split: train
|
| 437 |
+
path: data/CC-MAIN-2015-18/*
|
| 438 |
+
- config_name: CC-MAIN-2015-14
|
| 439 |
+
data_files:
|
| 440 |
+
- split: train
|
| 441 |
+
path: data/CC-MAIN-2015-14/*
|
| 442 |
+
- config_name: CC-MAIN-2015-11
|
| 443 |
+
data_files:
|
| 444 |
+
- split: train
|
| 445 |
+
path: data/CC-MAIN-2015-11/*
|
| 446 |
+
- config_name: CC-MAIN-2015-06
|
| 447 |
+
data_files:
|
| 448 |
+
- split: train
|
| 449 |
+
path: data/CC-MAIN-2015-06/*
|
| 450 |
+
- config_name: CC-MAIN-2014-52
|
| 451 |
+
data_files:
|
| 452 |
+
- split: train
|
| 453 |
+
path: data/CC-MAIN-2014-52/*
|
| 454 |
+
- config_name: CC-MAIN-2014-49
|
| 455 |
+
data_files:
|
| 456 |
+
- split: train
|
| 457 |
+
path: data/CC-MAIN-2014-49/*
|
| 458 |
+
- config_name: CC-MAIN-2014-42
|
| 459 |
+
data_files:
|
| 460 |
+
- split: train
|
| 461 |
+
path: data/CC-MAIN-2014-42/*
|
| 462 |
+
- config_name: CC-MAIN-2014-41
|
| 463 |
+
data_files:
|
| 464 |
+
- split: train
|
| 465 |
+
path: data/CC-MAIN-2014-41/*
|
| 466 |
+
- config_name: CC-MAIN-2014-35
|
| 467 |
+
data_files:
|
| 468 |
+
- split: train
|
| 469 |
+
path: data/CC-MAIN-2014-35/*
|
| 470 |
+
- config_name: CC-MAIN-2014-23
|
| 471 |
+
data_files:
|
| 472 |
+
- split: train
|
| 473 |
+
path: data/CC-MAIN-2014-23/*
|
| 474 |
+
- config_name: CC-MAIN-2014-15
|
| 475 |
+
data_files:
|
| 476 |
+
- split: train
|
| 477 |
+
path: data/CC-MAIN-2014-15/*
|
| 478 |
+
- config_name: CC-MAIN-2014-10
|
| 479 |
+
data_files:
|
| 480 |
+
- split: train
|
| 481 |
+
path: data/CC-MAIN-2014-10/*
|
| 482 |
+
- config_name: CC-MAIN-2013-48
|
| 483 |
+
data_files:
|
| 484 |
+
- split: train
|
| 485 |
+
path: data/CC-MAIN-2013-48/*
|
| 486 |
+
- config_name: CC-MAIN-2013-20
|
| 487 |
+
data_files:
|
| 488 |
+
- split: train
|
| 489 |
+
path: data/CC-MAIN-2013-20/*
|
| 490 |
+
---
|
| 491 |
+
|
| 492 |
+
# 📚 FineWeb-Edu
|
| 493 |
+
<center>
|
| 494 |
+
<img src="https://cdn-uploads.huggingface.co/production/uploads/61c141342aac764ce1654e43/wwRnEQydH9qdRtFofIE-A.png" alt="FineWeb-Edu: The finest collection of educational content the web has to offer">
|
| 495 |
+
</center>
|
| 496 |
+
|
| 497 |
+
> 1.3 trillion tokens of the finest educational data the 🌐 web has to offer
|
| 498 |
+
|
| 499 |
+
**Paper:** https://arxiv.org/abs/2406.17557
|
| 500 |
+
|
| 501 |
+
## What is it?
|
| 502 |
+
|
| 503 |
+
📚 FineWeb-Edu dataset consists of **1.3T tokens** and **5.4T tokens** ([FineWeb-Edu-score-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu-score-2)) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version.
|
| 504 |
+
|
| 505 |
+
To enhance FineWeb's quality, we developed an [educational quality classifier](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier) using annotations generated by LLama3-70B-Instruct. We then used this classifier to retain only the most educational web pages. FineWeb-Edu outperforms FineWeb on popular benchmarks and shows the power of classifiers trained on synthetic data.
|
| 506 |
+
|
| 507 |
+
The [Dataset Curation](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu#dataset-curation) section details the process for creating the dataset.
|
| 508 |
+
|
| 509 |
+

|
| 510 |
+
|
| 511 |
+
You can find a deduplicated version of FineWeb-edu in [SmolLM-Corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus). We find that the deduplication of this dataset doesn't have any impact on model performance in our ablation setup (1.8B trained on 350B tokens).
|
| 512 |
+
|
| 513 |
+
## What is being released?
|
| 514 |
+
|
| 515 |
+
Along with the dataset, which includes all filtered CommonCrawl dumps since 2013, we also release the educational classifier used for the filtering as well as the code for training it and running inference at: https://github.com/huggingface/cosmopedia/tree/main/classification
|
| 516 |
+
|
| 517 |
+
## Changelog
|
| 518 |
+
_Previous versions remain available in the branch `version name`._
|
| 519 |
+
|
| 520 |
+
- **v1.4.0 (11-07-2025):** Added 6 new snapshots: `CC-MAIN-2025-05`, `CC-MAIN-2025-08`, `CC-MAIN-2025-13`, `CC-MAIN-2025-18`, `CC-MAIN-2025-21`, and `CC-MAIN-2025-26` (January to June 2025)
|
| 521 |
+
- **v1.3.0 (31-01-2025):** Fixed an issue with some dumps where some documents hadn't been processed: `CC-MAIN-2024-10`, `CC-MAIN-2024-18`, `CC-MAIN-2024-22`, `CC-MAIN-2024-26`, `CC-MAIN-2024-30`, `CC-MAIN-2024-33`, `CC-MAIN-2024-38`, `CC-MAIN-2024-42`, `CC-MAIN-2024-46` -- they now contain more data (~35B additional tokens).
|
| 522 |
+
- **v1.2.0 (03-01-2025):** Added 9 new snapshots: `CC-MAIN-2024-18`, `CC-MAIN-2024-22`, `CC-MAIN-2024-26`, `CC-MAIN-2024-30`, `CC-MAIN-2024-33`, `CC-MAIN-2024-38`, `CC-MAIN-2024-42`, `CC-MAIN-2024-46`, `CC-MAIN-2024-51`, covering April to December 2024.
|
| 523 |
+
- **v1.0.0 (02-06-2024):** Initial version
|
| 524 |
+
|
| 525 |
+
|
| 526 |
+
## How to load the dataset
|
| 527 |
+
Similarily to FineWeb, You can load the full dataset or a specific crawl/dump. Dumps have the format `CC-MAIN-(year)-(week number)`.
|
| 528 |
+
|
| 529 |
+
### (Smaller) sample versions
|
| 530 |
+
Along with config `default` (all the data), and the configs for each individual dump, you can also download the following configs:
|
| 531 |
+
- `sample-350BT`: a subset randomly sampled from the whole dataset of around 350B gpt2 tokens
|
| 532 |
+
- `sample-100BT`: a subset randomly sampled from the whole dataset of around 100B gpt2 tokens
|
| 533 |
+
- `sample-10BT`: a subset randomly sampled from the whole dataset of around 10B gpt2 tokens
|
| 534 |
+
|
| 535 |
+
`sample-10BT` was sampled from `sample-100BT` which in turn was sampled from `sample-350BT`.
|
| 536 |
+
|
| 537 |
+
### Using 🏭 [`datatrove`](https://github.com/huggingface/datatrove/)
|
| 538 |
+
|
| 539 |
+
```python
|
| 540 |
+
from datatrove.pipeline.readers import ParquetReader
|
| 541 |
+
|
| 542 |
+
# limit determines how many documents will be streamed (remove for all)
|
| 543 |
+
data_reader = ParquetReader("hf://datasets/HuggingFaceFW/fineweb-edu", glob_pattern="data/*/*.parquet", limit=1000)
|
| 544 |
+
# or to fetch a specific dump CC-MAIN-2024-10, eplace "CC-MAIN-2024-10" with "sample/100BT" to use the 100BT sample
|
| 545 |
+
data_reader = ParquetReader("hf://datasets/HuggingFaceFW/fineweb-edu/CC-MAIN-2024-10", limit=1000)
|
| 546 |
+
for document in data_reader():
|
| 547 |
+
# do something with document
|
| 548 |
+
print(document)
|
| 549 |
+
|
| 550 |
+
###############################
|
| 551 |
+
# OR for a processing pipeline:
|
| 552 |
+
###############################
|
| 553 |
+
|
| 554 |
+
from datatrove.executor import LocalPipelineExecutor
|
| 555 |
+
from datatrove.pipeline.readers import ParquetReader
|
| 556 |
+
from datatrove.pipeline.filters import LambdaFilter
|
| 557 |
+
from datatrove.pipeline.writers import JsonlWriter
|
| 558 |
+
|
| 559 |
+
pipeline_exec = LocalPipelineExecutor(
|
| 560 |
+
pipeline=[
|
| 561 |
+
# replace "CC-MAIN-2024-10" with "sample/100BT" to use the 100BT sample
|
| 562 |
+
ParquetReader("hf://datasets/HuggingFaceFW/fineweb-edu/CC-MAIN-2024-10", limit=1000),
|
| 563 |
+
LambdaFilter(lambda doc: "hugging" in doc.text),
|
| 564 |
+
JsonlWriter("some-output-path")
|
| 565 |
+
],
|
| 566 |
+
tasks=10
|
| 567 |
+
)
|
| 568 |
+
pipeline_exec.run()
|
| 569 |
+
```
|
| 570 |
+
|
| 571 |
+
### Using `datasets`
|
| 572 |
+
|
| 573 |
+
```python
|
| 574 |
+
from datasets import load_dataset
|
| 575 |
+
# use name="sample-10BT" to use the 10BT sample
|
| 576 |
+
fw = load_dataset("HuggingFaceFW/fineweb-edu", name="CC-MAIN-2024-10", split="train", streaming=True)
|
| 577 |
+
```
|
| 578 |
+
|
| 579 |
+
## Dataset curation
|
| 580 |
+
A new approach has recently emerged for filtering LLM training datasets: using synthetic data to develop classifiers for identifying educational content. This technique was used in the trainings of [LLama3](https://ai.meta.com/blog/meta-llama-3-meta-ai-responsibility/) and [Phi3](https://arxiv.org/abs/2404.14219), but its large-scale impact on web data filtering hasn't been fully explored or published.
|
| 581 |
+
|
| 582 |
+
The highly popular Phi3 models were trained on 3.3 and 4.8 trillion tokens, with the paper stating: “Our training data consists of heavily filtered publicly available web data (according to the 'educational level') from various open internet sources, as well as synthetic LLM-generated data". Similarly, the LLama3 blog post notes: “We found that previous generations of Llama are good at identifying high-quality data, so we used Llama 2 to help build the text-quality classifiers that are powering Llama 3.” However these classifiers and filtered datasets are not publicly available. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by [LLama3-70B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct) to create FineWeb-Edu.
|
| 583 |
+
|
| 584 |
+
### Annotation
|
| 585 |
+
We used [Llama3-70B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct) to score 500k FineWeb samples for their educational quality on a scale from 0 to 5.
|
| 586 |
+
|
| 587 |
+
We explored various prompts and found that the additive scale by [Yuan et al.](https://arxiv.org/pdf/2401.10020) worked best. To avoid the LLM favoring highly technical pages like arXiv abstracts and submissions, we focused on grade-school and middle-school level knowledge. By setting a threshold of 3 (on a scale of 0 to 5) during the filtering process, we were able to also retain some high-level educational pages. The final prompt can be found [here](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier/blob/main/utils/prompt.txt).
|
| 588 |
+
|
| 589 |
+
We also experimented with different LLMs: Llama3-70B-Instruct, Mixtral-8x-7B-Instruct, and Mixtral-8x22B-Instruct. Llama 3 and Mixtral-8x22B produced similar scores, while Mixtral-8x7B tended to be more generous, not fully adhering to the score scale. Verga et al. suggest using multiple LLMs as juries. We tried averaging the scores from the three models, but this shifted the distribution to the right due to the higher scores from Mixtral-8x7B. Training on a dataset filtered with a classifier using jury annotations performed worse than using a classifier based on Llama3 annotations. We hypothesize that the jury-based approach retains more low-quality samples.
|
| 590 |
+
|
| 591 |
+
### Classifier training
|
| 592 |
+
We fine-tuned a Bert-like regression model using these annotations, based on [Snowflake-arctic-embed](https://huggingface.co/Snowflake/snowflake-arctic-embed-m). When converted to a binary classification using a score of 3 as a threshold for keeping and removing files, the model achieved an F1 score of 82%. The classification of FineWeb 15T tokens took 6k H100 GPU hours.
|
| 593 |
+
|
| 594 |
+
The classifier is available at: [HuggingFaceFW/fineweb-edu-classifier/](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier/)
|
| 595 |
+
|
| 596 |
+
### Filtering and results
|
| 597 |
+
**Note**: You can find more details about the ablations and results in the FineWeb [blog post](https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1).
|
| 598 |
+
|
| 599 |
+
We investigated the impact of using different thresholds for the filtering and found that threshold 3 gave the best overall results. Although using a threshold higher than 3 improves performance on knowledge and reasoning intensive benchmarks, it significantly degrades performance on HellaSwag and PIQA.
|
| 600 |
+
|
| 601 |
+
We then built 📚 FineWeb-Edu by filtering out samples with scores lower than 3. This removed 92% of the dataset, leaving us with 1.3T educational tokens. Our ablation demonstrated that this refined dataset surpasses 🍷 FineWeb and all other open web datasets, with remarkable improvements on educational benchmarks such as MMLU, ARC, and OpenBookQA. The plot below compares FineWeb-Edu to other web datasets:
|
| 602 |
+
|
| 603 |
+

|
| 604 |
+
|
| 605 |
+
To retain more tokens, we also experimented with a less strict threshold of 2 instead of 3. While being less performant than using threshold 3, it still outperformed FineWeb and it preserved 5.4T tokens. We release these two dataset as [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) and [FineWeb-Edu-score-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu-score-2) along with the [classifier](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier).
|
| 606 |
+
|
| 607 |
+
You will find all the ablation models in [this collection](https://huggingface.co/collections/HuggingFaceFW/ablation-models-662457b0d213e8c14fe47f32). The FineWeb-Edu ablation model (trained on 350B tokens) is available at [https://huggingface.co/HuggingFaceFW/ablation-model-fineweb-edu](https://huggingface.co/HuggingFaceFW/ablation-model-fineweb-edu).
|
| 608 |
+
|
| 609 |
+
## Considerations for Using the Data
|
| 610 |
+
This section is copied from the parent dataset: [FineWeb](https://huggingface.co/datasets/HuggingFaceFW/fineweb).
|
| 611 |
+
|
| 612 |
+
### Social Impact of Dataset
|
| 613 |
+
|
| 614 |
+
With the release of this dataset we aim to make model training more accessible to the machine learning community at large.
|
| 615 |
+
|
| 616 |
+
While multiple open-weights models with strong performance have been publicly released in the past, more often than not these releases are not accompanied by the corresponding training dataset. This is unfortunate as the dataset specificities and characteristics have been demonstrated to have a very large impact and role in the performances of the models. As the creation of a high quality training dataset is a fundamental requirement to training an LLM capable of excelling at downstream tasks, with 🍷 FineWeb we (a) not only make the dataset creation process more transparent, by sharing our entire processing setup including the codebase used, we also (b) help alleviate the costs of dataset curation, both in time and in compute, for model creators by publicly releasing our dataset with the community.
|
| 617 |
+
|
| 618 |
+
### Discussion of Biases
|
| 619 |
+
|
| 620 |
+
Efforts were made to minimize the amount of NSFW and toxic content present in the dataset by employing filtering on the URL level. However, there are still a significant number of documents present in the final dataset that could be considered toxic or contain harmful content. As 🍷 FineWeb was sourced from the web as a whole, any harmful biases typically present in it may be reproduced on our dataset.
|
| 621 |
+
|
| 622 |
+
We deliberately avoided using machine learning filtering methods that define text quality based on the similarity to a “gold” source such as wikipedia or toxicity classifiers as these methods have been known to [disproportionately remove content in specific dialects](https://aclanthology.org/D16-1120/) and [overclassify as toxic text related to specific social identities](https://arxiv.org/pdf/2109.07445.pdf), respectively.
|
| 623 |
+
|
| 624 |
+
### Other Known Limitations
|
| 625 |
+
|
| 626 |
+
As a consequence of some of the filtering steps applied, it is likely that code content is not prevalent in our dataset. If you are training a model that should also perform code tasks, we recommend you use 🍷 FineWeb with a code dataset, such as [The Stack v2](https://huggingface.co/datasets/bigcode/the-stack-v2). You should also probably consider complementing 🍷 FineWeb with specialized curated sources (such as Wikipedia, for example) as they will likely have better formatting than the wikipedia content included in 🍷 FineWeb (we did not tailor the processing to individual websites).
|
| 627 |
+
|
| 628 |
+
## Additional Information
|
| 629 |
+
|
| 630 |
+
### Licensing Information
|
| 631 |
+
|
| 632 |
+
The dataset is released under the **Open Data Commons Attribution License (ODC-By) v1.0** [license](https://opendatacommons.org/licenses/by/1-0/). The use of this dataset is also subject to [CommonCrawl's Terms of Use](https://commoncrawl.org/terms-of-use).
|
| 633 |
+
|
| 634 |
+
### Future work
|
| 635 |
+
|
| 636 |
+
We plan to work on better educational classifier to improve the quality of FineWeb-Edu.
|
| 637 |
+
|
| 638 |
+
### Citation Information
|
| 639 |
+
|
| 640 |
+
You can cite our paper https://arxiv.org/abs/2406.17557 or this dataset:
|
| 641 |
+
|
| 642 |
+
```
|
| 643 |
+
@misc{lozhkov2024fineweb-edu,
|
| 644 |
+
author = { Lozhkov, Anton and Ben Allal, Loubna and von Werra, Leandro and Wolf, Thomas },
|
| 645 |
+
title = { FineWeb-Edu: the Finest Collection of Educational Content },
|
| 646 |
+
year = 2024,
|
| 647 |
+
url = { https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu },
|
| 648 |
+
doi = { 10.57967/hf/2497 },
|
| 649 |
+
publisher = { Hugging Face }
|
| 650 |
+
}
|
| 651 |
+
```
|
hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/refs/main
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
87f09149ef4734204d70ed1d046ddc9ca3f2b8f9
|
hf_cache/hub/datasets--HuggingFaceFW--fineweb-edu/snapshots/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/README.md
ADDED
|
@@ -0,0 +1,651 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: odc-by
|
| 3 |
+
task_categories:
|
| 4 |
+
- text-generation
|
| 5 |
+
language:
|
| 6 |
+
- en
|
| 7 |
+
pretty_name: FineWeb-Edu
|
| 8 |
+
size_categories:
|
| 9 |
+
- n>1T
|
| 10 |
+
configs:
|
| 11 |
+
- config_name: default
|
| 12 |
+
data_files:
|
| 13 |
+
- split: train
|
| 14 |
+
path: data/*/*
|
| 15 |
+
features:
|
| 16 |
+
- name: text
|
| 17 |
+
dtype: string
|
| 18 |
+
- name: id
|
| 19 |
+
dtype: string
|
| 20 |
+
- name: dump
|
| 21 |
+
dtype: string
|
| 22 |
+
- name: url
|
| 23 |
+
dtype: string
|
| 24 |
+
- name: date
|
| 25 |
+
dtype: string
|
| 26 |
+
- name: file_path
|
| 27 |
+
dtype: string
|
| 28 |
+
- name: language
|
| 29 |
+
dtype: string
|
| 30 |
+
- name: language_score
|
| 31 |
+
dtype: float64
|
| 32 |
+
- name: token_count
|
| 33 |
+
dtype: int64
|
| 34 |
+
- name: score
|
| 35 |
+
dtype: float64
|
| 36 |
+
- name: int_score
|
| 37 |
+
dtype: int64
|
| 38 |
+
- config_name: sample-10BT
|
| 39 |
+
data_files:
|
| 40 |
+
- split: train
|
| 41 |
+
path: sample/10BT/*
|
| 42 |
+
- config_name: sample-100BT
|
| 43 |
+
data_files:
|
| 44 |
+
- split: train
|
| 45 |
+
path: sample/100BT/*
|
| 46 |
+
- config_name: sample-350BT
|
| 47 |
+
data_files:
|
| 48 |
+
- split: train
|
| 49 |
+
path: sample/350BT/*
|
| 50 |
+
- config_name: CC-MAIN-2025-05
|
| 51 |
+
data_files:
|
| 52 |
+
- split: train
|
| 53 |
+
path: data/CC-MAIN-2025-05/*
|
| 54 |
+
- config_name: CC-MAIN-2025-08
|
| 55 |
+
data_files:
|
| 56 |
+
- split: train
|
| 57 |
+
path: data/CC-MAIN-2025-08/*
|
| 58 |
+
- config_name: CC-MAIN-2025-13
|
| 59 |
+
data_files:
|
| 60 |
+
- split: train
|
| 61 |
+
path: data/CC-MAIN-2025-13/*
|
| 62 |
+
- config_name: CC-MAIN-2025-18
|
| 63 |
+
data_files:
|
| 64 |
+
- split: train
|
| 65 |
+
path: data/CC-MAIN-2025-18/*
|
| 66 |
+
- config_name: CC-MAIN-2025-21
|
| 67 |
+
data_files:
|
| 68 |
+
- split: train
|
| 69 |
+
path: data/CC-MAIN-2025-21/*
|
| 70 |
+
- config_name: CC-MAIN-2025-26
|
| 71 |
+
data_files:
|
| 72 |
+
- split: train
|
| 73 |
+
path: data/CC-MAIN-2025-26/*
|
| 74 |
+
- config_name: CC-MAIN-2024-51
|
| 75 |
+
data_files:
|
| 76 |
+
- split: train
|
| 77 |
+
path: data/CC-MAIN-2024-51/*
|
| 78 |
+
- config_name: CC-MAIN-2024-46
|
| 79 |
+
data_files:
|
| 80 |
+
- split: train
|
| 81 |
+
path: data/CC-MAIN-2024-46/*
|
| 82 |
+
- config_name: CC-MAIN-2024-42
|
| 83 |
+
data_files:
|
| 84 |
+
- split: train
|
| 85 |
+
path: data/CC-MAIN-2024-42/*
|
| 86 |
+
- config_name: CC-MAIN-2024-38
|
| 87 |
+
data_files:
|
| 88 |
+
- split: train
|
| 89 |
+
path: data/CC-MAIN-2024-38/*
|
| 90 |
+
- config_name: CC-MAIN-2024-33
|
| 91 |
+
data_files:
|
| 92 |
+
- split: train
|
| 93 |
+
path: data/CC-MAIN-2024-33/*
|
| 94 |
+
- config_name: CC-MAIN-2024-30
|
| 95 |
+
data_files:
|
| 96 |
+
- split: train
|
| 97 |
+
path: data/CC-MAIN-2024-30/*
|
| 98 |
+
- config_name: CC-MAIN-2024-26
|
| 99 |
+
data_files:
|
| 100 |
+
- split: train
|
| 101 |
+
path: data/CC-MAIN-2024-26/*
|
| 102 |
+
- config_name: CC-MAIN-2024-22
|
| 103 |
+
data_files:
|
| 104 |
+
- split: train
|
| 105 |
+
path: data/CC-MAIN-2024-22/*
|
| 106 |
+
- config_name: CC-MAIN-2024-18
|
| 107 |
+
data_files:
|
| 108 |
+
- split: train
|
| 109 |
+
path: data/CC-MAIN-2024-18/*
|
| 110 |
+
- config_name: CC-MAIN-2024-10
|
| 111 |
+
data_files:
|
| 112 |
+
- split: train
|
| 113 |
+
path: data/CC-MAIN-2024-10/*
|
| 114 |
+
- config_name: CC-MAIN-2023-50
|
| 115 |
+
data_files:
|
| 116 |
+
- split: train
|
| 117 |
+
path: data/CC-MAIN-2023-50/*
|
| 118 |
+
- config_name: CC-MAIN-2023-40
|
| 119 |
+
data_files:
|
| 120 |
+
- split: train
|
| 121 |
+
path: data/CC-MAIN-2023-40/*
|
| 122 |
+
- config_name: CC-MAIN-2023-23
|
| 123 |
+
data_files:
|
| 124 |
+
- split: train
|
| 125 |
+
path: data/CC-MAIN-2023-23/*
|
| 126 |
+
- config_name: CC-MAIN-2023-14
|
| 127 |
+
data_files:
|
| 128 |
+
- split: train
|
| 129 |
+
path: data/CC-MAIN-2023-14/*
|
| 130 |
+
- config_name: CC-MAIN-2023-06
|
| 131 |
+
data_files:
|
| 132 |
+
- split: train
|
| 133 |
+
path: data/CC-MAIN-2023-06/*
|
| 134 |
+
- config_name: CC-MAIN-2022-49
|
| 135 |
+
data_files:
|
| 136 |
+
- split: train
|
| 137 |
+
path: data/CC-MAIN-2022-49/*
|
| 138 |
+
- config_name: CC-MAIN-2022-40
|
| 139 |
+
data_files:
|
| 140 |
+
- split: train
|
| 141 |
+
path: data/CC-MAIN-2022-40/*
|
| 142 |
+
- config_name: CC-MAIN-2022-33
|
| 143 |
+
data_files:
|
| 144 |
+
- split: train
|
| 145 |
+
path: data/CC-MAIN-2022-33/*
|
| 146 |
+
- config_name: CC-MAIN-2022-27
|
| 147 |
+
data_files:
|
| 148 |
+
- split: train
|
| 149 |
+
path: data/CC-MAIN-2022-27/*
|
| 150 |
+
- config_name: CC-MAIN-2022-21
|
| 151 |
+
data_files:
|
| 152 |
+
- split: train
|
| 153 |
+
path: data/CC-MAIN-2022-21/*
|
| 154 |
+
- config_name: CC-MAIN-2022-05
|
| 155 |
+
data_files:
|
| 156 |
+
- split: train
|
| 157 |
+
path: data/CC-MAIN-2022-05/*
|
| 158 |
+
- config_name: CC-MAIN-2021-49
|
| 159 |
+
data_files:
|
| 160 |
+
- split: train
|
| 161 |
+
path: data/CC-MAIN-2021-49/*
|
| 162 |
+
- config_name: CC-MAIN-2021-43
|
| 163 |
+
data_files:
|
| 164 |
+
- split: train
|
| 165 |
+
path: data/CC-MAIN-2021-43/*
|
| 166 |
+
- config_name: CC-MAIN-2021-39
|
| 167 |
+
data_files:
|
| 168 |
+
- split: train
|
| 169 |
+
path: data/CC-MAIN-2021-39/*
|
| 170 |
+
- config_name: CC-MAIN-2021-31
|
| 171 |
+
data_files:
|
| 172 |
+
- split: train
|
| 173 |
+
path: data/CC-MAIN-2021-31/*
|
| 174 |
+
- config_name: CC-MAIN-2021-25
|
| 175 |
+
data_files:
|
| 176 |
+
- split: train
|
| 177 |
+
path: data/CC-MAIN-2021-25/*
|
| 178 |
+
- config_name: CC-MAIN-2021-21
|
| 179 |
+
data_files:
|
| 180 |
+
- split: train
|
| 181 |
+
path: data/CC-MAIN-2021-21/*
|
| 182 |
+
- config_name: CC-MAIN-2021-17
|
| 183 |
+
data_files:
|
| 184 |
+
- split: train
|
| 185 |
+
path: data/CC-MAIN-2021-17/*
|
| 186 |
+
- config_name: CC-MAIN-2021-10
|
| 187 |
+
data_files:
|
| 188 |
+
- split: train
|
| 189 |
+
path: data/CC-MAIN-2021-10/*
|
| 190 |
+
- config_name: CC-MAIN-2021-04
|
| 191 |
+
data_files:
|
| 192 |
+
- split: train
|
| 193 |
+
path: data/CC-MAIN-2021-04/*
|
| 194 |
+
- config_name: CC-MAIN-2020-50
|
| 195 |
+
data_files:
|
| 196 |
+
- split: train
|
| 197 |
+
path: data/CC-MAIN-2020-50/*
|
| 198 |
+
- config_name: CC-MAIN-2020-45
|
| 199 |
+
data_files:
|
| 200 |
+
- split: train
|
| 201 |
+
path: data/CC-MAIN-2020-45/*
|
| 202 |
+
- config_name: CC-MAIN-2020-40
|
| 203 |
+
data_files:
|
| 204 |
+
- split: train
|
| 205 |
+
path: data/CC-MAIN-2020-40/*
|
| 206 |
+
- config_name: CC-MAIN-2020-34
|
| 207 |
+
data_files:
|
| 208 |
+
- split: train
|
| 209 |
+
path: data/CC-MAIN-2020-34/*
|
| 210 |
+
- config_name: CC-MAIN-2020-29
|
| 211 |
+
data_files:
|
| 212 |
+
- split: train
|
| 213 |
+
path: data/CC-MAIN-2020-29/*
|
| 214 |
+
- config_name: CC-MAIN-2020-24
|
| 215 |
+
data_files:
|
| 216 |
+
- split: train
|
| 217 |
+
path: data/CC-MAIN-2020-24/*
|
| 218 |
+
- config_name: CC-MAIN-2020-16
|
| 219 |
+
data_files:
|
| 220 |
+
- split: train
|
| 221 |
+
path: data/CC-MAIN-2020-16/*
|
| 222 |
+
- config_name: CC-MAIN-2020-10
|
| 223 |
+
data_files:
|
| 224 |
+
- split: train
|
| 225 |
+
path: data/CC-MAIN-2020-10/*
|
| 226 |
+
- config_name: CC-MAIN-2020-05
|
| 227 |
+
data_files:
|
| 228 |
+
- split: train
|
| 229 |
+
path: data/CC-MAIN-2020-05/*
|
| 230 |
+
- config_name: CC-MAIN-2019-51
|
| 231 |
+
data_files:
|
| 232 |
+
- split: train
|
| 233 |
+
path: data/CC-MAIN-2019-51/*
|
| 234 |
+
- config_name: CC-MAIN-2019-47
|
| 235 |
+
data_files:
|
| 236 |
+
- split: train
|
| 237 |
+
path: data/CC-MAIN-2019-47/*
|
| 238 |
+
- config_name: CC-MAIN-2019-43
|
| 239 |
+
data_files:
|
| 240 |
+
- split: train
|
| 241 |
+
path: data/CC-MAIN-2019-43/*
|
| 242 |
+
- config_name: CC-MAIN-2019-39
|
| 243 |
+
data_files:
|
| 244 |
+
- split: train
|
| 245 |
+
path: data/CC-MAIN-2019-39/*
|
| 246 |
+
- config_name: CC-MAIN-2019-35
|
| 247 |
+
data_files:
|
| 248 |
+
- split: train
|
| 249 |
+
path: data/CC-MAIN-2019-35/*
|
| 250 |
+
- config_name: CC-MAIN-2019-30
|
| 251 |
+
data_files:
|
| 252 |
+
- split: train
|
| 253 |
+
path: data/CC-MAIN-2019-30/*
|
| 254 |
+
- config_name: CC-MAIN-2019-26
|
| 255 |
+
data_files:
|
| 256 |
+
- split: train
|
| 257 |
+
path: data/CC-MAIN-2019-26/*
|
| 258 |
+
- config_name: CC-MAIN-2019-22
|
| 259 |
+
data_files:
|
| 260 |
+
- split: train
|
| 261 |
+
path: data/CC-MAIN-2019-22/*
|
| 262 |
+
- config_name: CC-MAIN-2019-18
|
| 263 |
+
data_files:
|
| 264 |
+
- split: train
|
| 265 |
+
path: data/CC-MAIN-2019-18/*
|
| 266 |
+
- config_name: CC-MAIN-2019-13
|
| 267 |
+
data_files:
|
| 268 |
+
- split: train
|
| 269 |
+
path: data/CC-MAIN-2019-13/*
|
| 270 |
+
- config_name: CC-MAIN-2019-09
|
| 271 |
+
data_files:
|
| 272 |
+
- split: train
|
| 273 |
+
path: data/CC-MAIN-2019-09/*
|
| 274 |
+
- config_name: CC-MAIN-2019-04
|
| 275 |
+
data_files:
|
| 276 |
+
- split: train
|
| 277 |
+
path: data/CC-MAIN-2019-04/*
|
| 278 |
+
- config_name: CC-MAIN-2018-51
|
| 279 |
+
data_files:
|
| 280 |
+
- split: train
|
| 281 |
+
path: data/CC-MAIN-2018-51/*
|
| 282 |
+
- config_name: CC-MAIN-2018-47
|
| 283 |
+
data_files:
|
| 284 |
+
- split: train
|
| 285 |
+
path: data/CC-MAIN-2018-47/*
|
| 286 |
+
- config_name: CC-MAIN-2018-43
|
| 287 |
+
data_files:
|
| 288 |
+
- split: train
|
| 289 |
+
path: data/CC-MAIN-2018-43/*
|
| 290 |
+
- config_name: CC-MAIN-2018-39
|
| 291 |
+
data_files:
|
| 292 |
+
- split: train
|
| 293 |
+
path: data/CC-MAIN-2018-39/*
|
| 294 |
+
- config_name: CC-MAIN-2018-34
|
| 295 |
+
data_files:
|
| 296 |
+
- split: train
|
| 297 |
+
path: data/CC-MAIN-2018-34/*
|
| 298 |
+
- config_name: CC-MAIN-2018-30
|
| 299 |
+
data_files:
|
| 300 |
+
- split: train
|
| 301 |
+
path: data/CC-MAIN-2018-30/*
|
| 302 |
+
- config_name: CC-MAIN-2018-26
|
| 303 |
+
data_files:
|
| 304 |
+
- split: train
|
| 305 |
+
path: data/CC-MAIN-2018-26/*
|
| 306 |
+
- config_name: CC-MAIN-2018-22
|
| 307 |
+
data_files:
|
| 308 |
+
- split: train
|
| 309 |
+
path: data/CC-MAIN-2018-22/*
|
| 310 |
+
- config_name: CC-MAIN-2018-17
|
| 311 |
+
data_files:
|
| 312 |
+
- split: train
|
| 313 |
+
path: data/CC-MAIN-2018-17/*
|
| 314 |
+
- config_name: CC-MAIN-2018-13
|
| 315 |
+
data_files:
|
| 316 |
+
- split: train
|
| 317 |
+
path: data/CC-MAIN-2018-13/*
|
| 318 |
+
- config_name: CC-MAIN-2018-09
|
| 319 |
+
data_files:
|
| 320 |
+
- split: train
|
| 321 |
+
path: data/CC-MAIN-2018-09/*
|
| 322 |
+
- config_name: CC-MAIN-2018-05
|
| 323 |
+
data_files:
|
| 324 |
+
- split: train
|
| 325 |
+
path: data/CC-MAIN-2018-05/*
|
| 326 |
+
- config_name: CC-MAIN-2017-51
|
| 327 |
+
data_files:
|
| 328 |
+
- split: train
|
| 329 |
+
path: data/CC-MAIN-2017-51/*
|
| 330 |
+
- config_name: CC-MAIN-2017-47
|
| 331 |
+
data_files:
|
| 332 |
+
- split: train
|
| 333 |
+
path: data/CC-MAIN-2017-47/*
|
| 334 |
+
- config_name: CC-MAIN-2017-43
|
| 335 |
+
data_files:
|
| 336 |
+
- split: train
|
| 337 |
+
path: data/CC-MAIN-2017-43/*
|
| 338 |
+
- config_name: CC-MAIN-2017-39
|
| 339 |
+
data_files:
|
| 340 |
+
- split: train
|
| 341 |
+
path: data/CC-MAIN-2017-39/*
|
| 342 |
+
- config_name: CC-MAIN-2017-34
|
| 343 |
+
data_files:
|
| 344 |
+
- split: train
|
| 345 |
+
path: data/CC-MAIN-2017-34/*
|
| 346 |
+
- config_name: CC-MAIN-2017-30
|
| 347 |
+
data_files:
|
| 348 |
+
- split: train
|
| 349 |
+
path: data/CC-MAIN-2017-30/*
|
| 350 |
+
- config_name: CC-MAIN-2017-26
|
| 351 |
+
data_files:
|
| 352 |
+
- split: train
|
| 353 |
+
path: data/CC-MAIN-2017-26/*
|
| 354 |
+
- config_name: CC-MAIN-2017-22
|
| 355 |
+
data_files:
|
| 356 |
+
- split: train
|
| 357 |
+
path: data/CC-MAIN-2017-22/*
|
| 358 |
+
- config_name: CC-MAIN-2017-17
|
| 359 |
+
data_files:
|
| 360 |
+
- split: train
|
| 361 |
+
path: data/CC-MAIN-2017-17/*
|
| 362 |
+
- config_name: CC-MAIN-2017-13
|
| 363 |
+
data_files:
|
| 364 |
+
- split: train
|
| 365 |
+
path: data/CC-MAIN-2017-13/*
|
| 366 |
+
- config_name: CC-MAIN-2017-09
|
| 367 |
+
data_files:
|
| 368 |
+
- split: train
|
| 369 |
+
path: data/CC-MAIN-2017-09/*
|
| 370 |
+
- config_name: CC-MAIN-2017-04
|
| 371 |
+
data_files:
|
| 372 |
+
- split: train
|
| 373 |
+
path: data/CC-MAIN-2017-04/*
|
| 374 |
+
- config_name: CC-MAIN-2016-50
|
| 375 |
+
data_files:
|
| 376 |
+
- split: train
|
| 377 |
+
path: data/CC-MAIN-2016-50/*
|
| 378 |
+
- config_name: CC-MAIN-2016-44
|
| 379 |
+
data_files:
|
| 380 |
+
- split: train
|
| 381 |
+
path: data/CC-MAIN-2016-44/*
|
| 382 |
+
- config_name: CC-MAIN-2016-40
|
| 383 |
+
data_files:
|
| 384 |
+
- split: train
|
| 385 |
+
path: data/CC-MAIN-2016-40/*
|
| 386 |
+
- config_name: CC-MAIN-2016-36
|
| 387 |
+
data_files:
|
| 388 |
+
- split: train
|
| 389 |
+
path: data/CC-MAIN-2016-36/*
|
| 390 |
+
- config_name: CC-MAIN-2016-30
|
| 391 |
+
data_files:
|
| 392 |
+
- split: train
|
| 393 |
+
path: data/CC-MAIN-2016-30/*
|
| 394 |
+
- config_name: CC-MAIN-2016-26
|
| 395 |
+
data_files:
|
| 396 |
+
- split: train
|
| 397 |
+
path: data/CC-MAIN-2016-26/*
|
| 398 |
+
- config_name: CC-MAIN-2016-22
|
| 399 |
+
data_files:
|
| 400 |
+
- split: train
|
| 401 |
+
path: data/CC-MAIN-2016-22/*
|
| 402 |
+
- config_name: CC-MAIN-2016-18
|
| 403 |
+
data_files:
|
| 404 |
+
- split: train
|
| 405 |
+
path: data/CC-MAIN-2016-18/*
|
| 406 |
+
- config_name: CC-MAIN-2016-07
|
| 407 |
+
data_files:
|
| 408 |
+
- split: train
|
| 409 |
+
path: data/CC-MAIN-2016-07/*
|
| 410 |
+
- config_name: CC-MAIN-2015-48
|
| 411 |
+
data_files:
|
| 412 |
+
- split: train
|
| 413 |
+
path: data/CC-MAIN-2015-48/*
|
| 414 |
+
- config_name: CC-MAIN-2015-40
|
| 415 |
+
data_files:
|
| 416 |
+
- split: train
|
| 417 |
+
path: data/CC-MAIN-2015-40/*
|
| 418 |
+
- config_name: CC-MAIN-2015-35
|
| 419 |
+
data_files:
|
| 420 |
+
- split: train
|
| 421 |
+
path: data/CC-MAIN-2015-35/*
|
| 422 |
+
- config_name: CC-MAIN-2015-32
|
| 423 |
+
data_files:
|
| 424 |
+
- split: train
|
| 425 |
+
path: data/CC-MAIN-2015-32/*
|
| 426 |
+
- config_name: CC-MAIN-2015-27
|
| 427 |
+
data_files:
|
| 428 |
+
- split: train
|
| 429 |
+
path: data/CC-MAIN-2015-27/*
|
| 430 |
+
- config_name: CC-MAIN-2015-22
|
| 431 |
+
data_files:
|
| 432 |
+
- split: train
|
| 433 |
+
path: data/CC-MAIN-2015-22/*
|
| 434 |
+
- config_name: CC-MAIN-2015-18
|
| 435 |
+
data_files:
|
| 436 |
+
- split: train
|
| 437 |
+
path: data/CC-MAIN-2015-18/*
|
| 438 |
+
- config_name: CC-MAIN-2015-14
|
| 439 |
+
data_files:
|
| 440 |
+
- split: train
|
| 441 |
+
path: data/CC-MAIN-2015-14/*
|
| 442 |
+
- config_name: CC-MAIN-2015-11
|
| 443 |
+
data_files:
|
| 444 |
+
- split: train
|
| 445 |
+
path: data/CC-MAIN-2015-11/*
|
| 446 |
+
- config_name: CC-MAIN-2015-06
|
| 447 |
+
data_files:
|
| 448 |
+
- split: train
|
| 449 |
+
path: data/CC-MAIN-2015-06/*
|
| 450 |
+
- config_name: CC-MAIN-2014-52
|
| 451 |
+
data_files:
|
| 452 |
+
- split: train
|
| 453 |
+
path: data/CC-MAIN-2014-52/*
|
| 454 |
+
- config_name: CC-MAIN-2014-49
|
| 455 |
+
data_files:
|
| 456 |
+
- split: train
|
| 457 |
+
path: data/CC-MAIN-2014-49/*
|
| 458 |
+
- config_name: CC-MAIN-2014-42
|
| 459 |
+
data_files:
|
| 460 |
+
- split: train
|
| 461 |
+
path: data/CC-MAIN-2014-42/*
|
| 462 |
+
- config_name: CC-MAIN-2014-41
|
| 463 |
+
data_files:
|
| 464 |
+
- split: train
|
| 465 |
+
path: data/CC-MAIN-2014-41/*
|
| 466 |
+
- config_name: CC-MAIN-2014-35
|
| 467 |
+
data_files:
|
| 468 |
+
- split: train
|
| 469 |
+
path: data/CC-MAIN-2014-35/*
|
| 470 |
+
- config_name: CC-MAIN-2014-23
|
| 471 |
+
data_files:
|
| 472 |
+
- split: train
|
| 473 |
+
path: data/CC-MAIN-2014-23/*
|
| 474 |
+
- config_name: CC-MAIN-2014-15
|
| 475 |
+
data_files:
|
| 476 |
+
- split: train
|
| 477 |
+
path: data/CC-MAIN-2014-15/*
|
| 478 |
+
- config_name: CC-MAIN-2014-10
|
| 479 |
+
data_files:
|
| 480 |
+
- split: train
|
| 481 |
+
path: data/CC-MAIN-2014-10/*
|
| 482 |
+
- config_name: CC-MAIN-2013-48
|
| 483 |
+
data_files:
|
| 484 |
+
- split: train
|
| 485 |
+
path: data/CC-MAIN-2013-48/*
|
| 486 |
+
- config_name: CC-MAIN-2013-20
|
| 487 |
+
data_files:
|
| 488 |
+
- split: train
|
| 489 |
+
path: data/CC-MAIN-2013-20/*
|
| 490 |
+
---
|
| 491 |
+
|
| 492 |
+
# 📚 FineWeb-Edu
|
| 493 |
+
<center>
|
| 494 |
+
<img src="https://cdn-uploads.huggingface.co/production/uploads/61c141342aac764ce1654e43/wwRnEQydH9qdRtFofIE-A.png" alt="FineWeb-Edu: The finest collection of educational content the web has to offer">
|
| 495 |
+
</center>
|
| 496 |
+
|
| 497 |
+
> 1.3 trillion tokens of the finest educational data the 🌐 web has to offer
|
| 498 |
+
|
| 499 |
+
**Paper:** https://arxiv.org/abs/2406.17557
|
| 500 |
+
|
| 501 |
+
## What is it?
|
| 502 |
+
|
| 503 |
+
📚 FineWeb-Edu dataset consists of **1.3T tokens** and **5.4T tokens** ([FineWeb-Edu-score-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu-score-2)) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version.
|
| 504 |
+
|
| 505 |
+
To enhance FineWeb's quality, we developed an [educational quality classifier](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier) using annotations generated by LLama3-70B-Instruct. We then used this classifier to retain only the most educational web pages. FineWeb-Edu outperforms FineWeb on popular benchmarks and shows the power of classifiers trained on synthetic data.
|
| 506 |
+
|
| 507 |
+
The [Dataset Curation](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu#dataset-curation) section details the process for creating the dataset.
|
| 508 |
+
|
| 509 |
+

|
| 510 |
+
|
| 511 |
+
You can find a deduplicated version of FineWeb-edu in [SmolLM-Corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus). We find that the deduplication of this dataset doesn't have any impact on model performance in our ablation setup (1.8B trained on 350B tokens).
|
| 512 |
+
|
| 513 |
+
## What is being released?
|
| 514 |
+
|
| 515 |
+
Along with the dataset, which includes all filtered CommonCrawl dumps since 2013, we also release the educational classifier used for the filtering as well as the code for training it and running inference at: https://github.com/huggingface/cosmopedia/tree/main/classification
|
| 516 |
+
|
| 517 |
+
## Changelog
|
| 518 |
+
_Previous versions remain available in the branch `version name`._
|
| 519 |
+
|
| 520 |
+
- **v1.4.0 (11-07-2025):** Added 6 new snapshots: `CC-MAIN-2025-05`, `CC-MAIN-2025-08`, `CC-MAIN-2025-13`, `CC-MAIN-2025-18`, `CC-MAIN-2025-21`, and `CC-MAIN-2025-26` (January to June 2025)
|
| 521 |
+
- **v1.3.0 (31-01-2025):** Fixed an issue with some dumps where some documents hadn't been processed: `CC-MAIN-2024-10`, `CC-MAIN-2024-18`, `CC-MAIN-2024-22`, `CC-MAIN-2024-26`, `CC-MAIN-2024-30`, `CC-MAIN-2024-33`, `CC-MAIN-2024-38`, `CC-MAIN-2024-42`, `CC-MAIN-2024-46` -- they now contain more data (~35B additional tokens).
|
| 522 |
+
- **v1.2.0 (03-01-2025):** Added 9 new snapshots: `CC-MAIN-2024-18`, `CC-MAIN-2024-22`, `CC-MAIN-2024-26`, `CC-MAIN-2024-30`, `CC-MAIN-2024-33`, `CC-MAIN-2024-38`, `CC-MAIN-2024-42`, `CC-MAIN-2024-46`, `CC-MAIN-2024-51`, covering April to December 2024.
|
| 523 |
+
- **v1.0.0 (02-06-2024):** Initial version
|
| 524 |
+
|
| 525 |
+
|
| 526 |
+
## How to load the dataset
|
| 527 |
+
Similarily to FineWeb, You can load the full dataset or a specific crawl/dump. Dumps have the format `CC-MAIN-(year)-(week number)`.
|
| 528 |
+
|
| 529 |
+
### (Smaller) sample versions
|
| 530 |
+
Along with config `default` (all the data), and the configs for each individual dump, you can also download the following configs:
|
| 531 |
+
- `sample-350BT`: a subset randomly sampled from the whole dataset of around 350B gpt2 tokens
|
| 532 |
+
- `sample-100BT`: a subset randomly sampled from the whole dataset of around 100B gpt2 tokens
|
| 533 |
+
- `sample-10BT`: a subset randomly sampled from the whole dataset of around 10B gpt2 tokens
|
| 534 |
+
|
| 535 |
+
`sample-10BT` was sampled from `sample-100BT` which in turn was sampled from `sample-350BT`.
|
| 536 |
+
|
| 537 |
+
### Using 🏭 [`datatrove`](https://github.com/huggingface/datatrove/)
|
| 538 |
+
|
| 539 |
+
```python
|
| 540 |
+
from datatrove.pipeline.readers import ParquetReader
|
| 541 |
+
|
| 542 |
+
# limit determines how many documents will be streamed (remove for all)
|
| 543 |
+
data_reader = ParquetReader("hf://datasets/HuggingFaceFW/fineweb-edu", glob_pattern="data/*/*.parquet", limit=1000)
|
| 544 |
+
# or to fetch a specific dump CC-MAIN-2024-10, eplace "CC-MAIN-2024-10" with "sample/100BT" to use the 100BT sample
|
| 545 |
+
data_reader = ParquetReader("hf://datasets/HuggingFaceFW/fineweb-edu/CC-MAIN-2024-10", limit=1000)
|
| 546 |
+
for document in data_reader():
|
| 547 |
+
# do something with document
|
| 548 |
+
print(document)
|
| 549 |
+
|
| 550 |
+
###############################
|
| 551 |
+
# OR for a processing pipeline:
|
| 552 |
+
###############################
|
| 553 |
+
|
| 554 |
+
from datatrove.executor import LocalPipelineExecutor
|
| 555 |
+
from datatrove.pipeline.readers import ParquetReader
|
| 556 |
+
from datatrove.pipeline.filters import LambdaFilter
|
| 557 |
+
from datatrove.pipeline.writers import JsonlWriter
|
| 558 |
+
|
| 559 |
+
pipeline_exec = LocalPipelineExecutor(
|
| 560 |
+
pipeline=[
|
| 561 |
+
# replace "CC-MAIN-2024-10" with "sample/100BT" to use the 100BT sample
|
| 562 |
+
ParquetReader("hf://datasets/HuggingFaceFW/fineweb-edu/CC-MAIN-2024-10", limit=1000),
|
| 563 |
+
LambdaFilter(lambda doc: "hugging" in doc.text),
|
| 564 |
+
JsonlWriter("some-output-path")
|
| 565 |
+
],
|
| 566 |
+
tasks=10
|
| 567 |
+
)
|
| 568 |
+
pipeline_exec.run()
|
| 569 |
+
```
|
| 570 |
+
|
| 571 |
+
### Using `datasets`
|
| 572 |
+
|
| 573 |
+
```python
|
| 574 |
+
from datasets import load_dataset
|
| 575 |
+
# use name="sample-10BT" to use the 10BT sample
|
| 576 |
+
fw = load_dataset("HuggingFaceFW/fineweb-edu", name="CC-MAIN-2024-10", split="train", streaming=True)
|
| 577 |
+
```
|
| 578 |
+
|
| 579 |
+
## Dataset curation
|
| 580 |
+
A new approach has recently emerged for filtering LLM training datasets: using synthetic data to develop classifiers for identifying educational content. This technique was used in the trainings of [LLama3](https://ai.meta.com/blog/meta-llama-3-meta-ai-responsibility/) and [Phi3](https://arxiv.org/abs/2404.14219), but its large-scale impact on web data filtering hasn't been fully explored or published.
|
| 581 |
+
|
| 582 |
+
The highly popular Phi3 models were trained on 3.3 and 4.8 trillion tokens, with the paper stating: “Our training data consists of heavily filtered publicly available web data (according to the 'educational level') from various open internet sources, as well as synthetic LLM-generated data". Similarly, the LLama3 blog post notes: “We found that previous generations of Llama are good at identifying high-quality data, so we used Llama 2 to help build the text-quality classifiers that are powering Llama 3.” However these classifiers and filtered datasets are not publicly available. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by [LLama3-70B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct) to create FineWeb-Edu.
|
| 583 |
+
|
| 584 |
+
### Annotation
|
| 585 |
+
We used [Llama3-70B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct) to score 500k FineWeb samples for their educational quality on a scale from 0 to 5.
|
| 586 |
+
|
| 587 |
+
We explored various prompts and found that the additive scale by [Yuan et al.](https://arxiv.org/pdf/2401.10020) worked best. To avoid the LLM favoring highly technical pages like arXiv abstracts and submissions, we focused on grade-school and middle-school level knowledge. By setting a threshold of 3 (on a scale of 0 to 5) during the filtering process, we were able to also retain some high-level educational pages. The final prompt can be found [here](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier/blob/main/utils/prompt.txt).
|
| 588 |
+
|
| 589 |
+
We also experimented with different LLMs: Llama3-70B-Instruct, Mixtral-8x-7B-Instruct, and Mixtral-8x22B-Instruct. Llama 3 and Mixtral-8x22B produced similar scores, while Mixtral-8x7B tended to be more generous, not fully adhering to the score scale. Verga et al. suggest using multiple LLMs as juries. We tried averaging the scores from the three models, but this shifted the distribution to the right due to the higher scores from Mixtral-8x7B. Training on a dataset filtered with a classifier using jury annotations performed worse than using a classifier based on Llama3 annotations. We hypothesize that the jury-based approach retains more low-quality samples.
|
| 590 |
+
|
| 591 |
+
### Classifier training
|
| 592 |
+
We fine-tuned a Bert-like regression model using these annotations, based on [Snowflake-arctic-embed](https://huggingface.co/Snowflake/snowflake-arctic-embed-m). When converted to a binary classification using a score of 3 as a threshold for keeping and removing files, the model achieved an F1 score of 82%. The classification of FineWeb 15T tokens took 6k H100 GPU hours.
|
| 593 |
+
|
| 594 |
+
The classifier is available at: [HuggingFaceFW/fineweb-edu-classifier/](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier/)
|
| 595 |
+
|
| 596 |
+
### Filtering and results
|
| 597 |
+
**Note**: You can find more details about the ablations and results in the FineWeb [blog post](https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1).
|
| 598 |
+
|
| 599 |
+
We investigated the impact of using different thresholds for the filtering and found that threshold 3 gave the best overall results. Although using a threshold higher than 3 improves performance on knowledge and reasoning intensive benchmarks, it significantly degrades performance on HellaSwag and PIQA.
|
| 600 |
+
|
| 601 |
+
We then built 📚 FineWeb-Edu by filtering out samples with scores lower than 3. This removed 92% of the dataset, leaving us with 1.3T educational tokens. Our ablation demonstrated that this refined dataset surpasses 🍷 FineWeb and all other open web datasets, with remarkable improvements on educational benchmarks such as MMLU, ARC, and OpenBookQA. The plot below compares FineWeb-Edu to other web datasets:
|
| 602 |
+
|
| 603 |
+

|
| 604 |
+
|
| 605 |
+
To retain more tokens, we also experimented with a less strict threshold of 2 instead of 3. While being less performant than using threshold 3, it still outperformed FineWeb and it preserved 5.4T tokens. We release these two dataset as [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) and [FineWeb-Edu-score-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu-score-2) along with the [classifier](https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier).
|
| 606 |
+
|
| 607 |
+
You will find all the ablation models in [this collection](https://huggingface.co/collections/HuggingFaceFW/ablation-models-662457b0d213e8c14fe47f32). The FineWeb-Edu ablation model (trained on 350B tokens) is available at [https://huggingface.co/HuggingFaceFW/ablation-model-fineweb-edu](https://huggingface.co/HuggingFaceFW/ablation-model-fineweb-edu).
|
| 608 |
+
|
| 609 |
+
## Considerations for Using the Data
|
| 610 |
+
This section is copied from the parent dataset: [FineWeb](https://huggingface.co/datasets/HuggingFaceFW/fineweb).
|
| 611 |
+
|
| 612 |
+
### Social Impact of Dataset
|
| 613 |
+
|
| 614 |
+
With the release of this dataset we aim to make model training more accessible to the machine learning community at large.
|
| 615 |
+
|
| 616 |
+
While multiple open-weights models with strong performance have been publicly released in the past, more often than not these releases are not accompanied by the corresponding training dataset. This is unfortunate as the dataset specificities and characteristics have been demonstrated to have a very large impact and role in the performances of the models. As the creation of a high quality training dataset is a fundamental requirement to training an LLM capable of excelling at downstream tasks, with 🍷 FineWeb we (a) not only make the dataset creation process more transparent, by sharing our entire processing setup including the codebase used, we also (b) help alleviate the costs of dataset curation, both in time and in compute, for model creators by publicly releasing our dataset with the community.
|
| 617 |
+
|
| 618 |
+
### Discussion of Biases
|
| 619 |
+
|
| 620 |
+
Efforts were made to minimize the amount of NSFW and toxic content present in the dataset by employing filtering on the URL level. However, there are still a significant number of documents present in the final dataset that could be considered toxic or contain harmful content. As 🍷 FineWeb was sourced from the web as a whole, any harmful biases typically present in it may be reproduced on our dataset.
|
| 621 |
+
|
| 622 |
+
We deliberately avoided using machine learning filtering methods that define text quality based on the similarity to a “gold” source such as wikipedia or toxicity classifiers as these methods have been known to [disproportionately remove content in specific dialects](https://aclanthology.org/D16-1120/) and [overclassify as toxic text related to specific social identities](https://arxiv.org/pdf/2109.07445.pdf), respectively.
|
| 623 |
+
|
| 624 |
+
### Other Known Limitations
|
| 625 |
+
|
| 626 |
+
As a consequence of some of the filtering steps applied, it is likely that code content is not prevalent in our dataset. If you are training a model that should also perform code tasks, we recommend you use 🍷 FineWeb with a code dataset, such as [The Stack v2](https://huggingface.co/datasets/bigcode/the-stack-v2). You should also probably consider complementing 🍷 FineWeb with specialized curated sources (such as Wikipedia, for example) as they will likely have better formatting than the wikipedia content included in 🍷 FineWeb (we did not tailor the processing to individual websites).
|
| 627 |
+
|
| 628 |
+
## Additional Information
|
| 629 |
+
|
| 630 |
+
### Licensing Information
|
| 631 |
+
|
| 632 |
+
The dataset is released under the **Open Data Commons Attribution License (ODC-By) v1.0** [license](https://opendatacommons.org/licenses/by/1-0/). The use of this dataset is also subject to [CommonCrawl's Terms of Use](https://commoncrawl.org/terms-of-use).
|
| 633 |
+
|
| 634 |
+
### Future work
|
| 635 |
+
|
| 636 |
+
We plan to work on better educational classifier to improve the quality of FineWeb-Edu.
|
| 637 |
+
|
| 638 |
+
### Citation Information
|
| 639 |
+
|
| 640 |
+
You can cite our paper https://arxiv.org/abs/2406.17557 or this dataset:
|
| 641 |
+
|
| 642 |
+
```
|
| 643 |
+
@misc{lozhkov2024fineweb-edu,
|
| 644 |
+
author = { Lozhkov, Anton and Ben Allal, Loubna and von Werra, Leandro and Wolf, Thomas },
|
| 645 |
+
title = { FineWeb-Edu: the Finest Collection of Educational Content },
|
| 646 |
+
year = 2024,
|
| 647 |
+
url = { https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu },
|
| 648 |
+
doi = { 10.57967/hf/2497 },
|
| 649 |
+
publisher = { Hugging Face }
|
| 650 |
+
}
|
| 651 |
+
```
|
hf_cache/hub/datasets--Rowan--hellaswag/.no_exist/218ec52e09a7e7462a5400043bb9a69a41d06b76/.huggingface.yaml
ADDED
|
File without changes
|
hf_cache/hub/datasets--Rowan--hellaswag/.no_exist/218ec52e09a7e7462a5400043bb9a69a41d06b76/dataset_infos.json
ADDED
|
File without changes
|
hf_cache/hub/datasets--Rowan--hellaswag/.no_exist/218ec52e09a7e7462a5400043bb9a69a41d06b76/hellaswag.py
ADDED
|
File without changes
|
hf_cache/hub/datasets--Rowan--hellaswag/blobs/29f11d90eb3a5b319cfe8ce2a4e78d9f2a1aea3f
ADDED
|
@@ -0,0 +1,218 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language:
|
| 3 |
+
- en
|
| 4 |
+
paperswithcode_id: hellaswag
|
| 5 |
+
pretty_name: HellaSwag
|
| 6 |
+
dataset_info:
|
| 7 |
+
features:
|
| 8 |
+
- name: ind
|
| 9 |
+
dtype: int32
|
| 10 |
+
- name: activity_label
|
| 11 |
+
dtype: string
|
| 12 |
+
- name: ctx_a
|
| 13 |
+
dtype: string
|
| 14 |
+
- name: ctx_b
|
| 15 |
+
dtype: string
|
| 16 |
+
- name: ctx
|
| 17 |
+
dtype: string
|
| 18 |
+
- name: endings
|
| 19 |
+
sequence: string
|
| 20 |
+
- name: source_id
|
| 21 |
+
dtype: string
|
| 22 |
+
- name: split
|
| 23 |
+
dtype: string
|
| 24 |
+
- name: split_type
|
| 25 |
+
dtype: string
|
| 26 |
+
- name: label
|
| 27 |
+
dtype: string
|
| 28 |
+
splits:
|
| 29 |
+
- name: train
|
| 30 |
+
num_bytes: 43232624
|
| 31 |
+
num_examples: 39905
|
| 32 |
+
- name: test
|
| 33 |
+
num_bytes: 10791853
|
| 34 |
+
num_examples: 10003
|
| 35 |
+
- name: validation
|
| 36 |
+
num_bytes: 11175717
|
| 37 |
+
num_examples: 10042
|
| 38 |
+
download_size: 36793872
|
| 39 |
+
dataset_size: 65200194
|
| 40 |
+
configs:
|
| 41 |
+
- config_name: default
|
| 42 |
+
data_files:
|
| 43 |
+
- split: train
|
| 44 |
+
path: data/train-*
|
| 45 |
+
- split: test
|
| 46 |
+
path: data/test-*
|
| 47 |
+
- split: validation
|
| 48 |
+
path: data/validation-*
|
| 49 |
+
---
|
| 50 |
+
|
| 51 |
+
# Dataset Card for "hellaswag"
|
| 52 |
+
|
| 53 |
+
## Table of Contents
|
| 54 |
+
- [Dataset Description](#dataset-description)
|
| 55 |
+
- [Dataset Summary](#dataset-summary)
|
| 56 |
+
- [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
|
| 57 |
+
- [Languages](#languages)
|
| 58 |
+
- [Dataset Structure](#dataset-structure)
|
| 59 |
+
- [Data Instances](#data-instances)
|
| 60 |
+
- [Data Fields](#data-fields)
|
| 61 |
+
- [Data Splits](#data-splits)
|
| 62 |
+
- [Dataset Creation](#dataset-creation)
|
| 63 |
+
- [Curation Rationale](#curation-rationale)
|
| 64 |
+
- [Source Data](#source-data)
|
| 65 |
+
- [Annotations](#annotations)
|
| 66 |
+
- [Personal and Sensitive Information](#personal-and-sensitive-information)
|
| 67 |
+
- [Considerations for Using the Data](#considerations-for-using-the-data)
|
| 68 |
+
- [Social Impact of Dataset](#social-impact-of-dataset)
|
| 69 |
+
- [Discussion of Biases](#discussion-of-biases)
|
| 70 |
+
- [Other Known Limitations](#other-known-limitations)
|
| 71 |
+
- [Additional Information](#additional-information)
|
| 72 |
+
- [Dataset Curators](#dataset-curators)
|
| 73 |
+
- [Licensing Information](#licensing-information)
|
| 74 |
+
- [Citation Information](#citation-information)
|
| 75 |
+
- [Contributions](#contributions)
|
| 76 |
+
|
| 77 |
+
## Dataset Description
|
| 78 |
+
|
| 79 |
+
- **Homepage:** [https://rowanzellers.com/hellaswag/](https://rowanzellers.com/hellaswag/)
|
| 80 |
+
- **Repository:** [https://github.com/rowanz/hellaswag/](https://github.com/rowanz/hellaswag/)
|
| 81 |
+
- **Paper:** [HellaSwag: Can a Machine Really Finish Your Sentence?](https://arxiv.org/abs/1905.07830)
|
| 82 |
+
- **Point of Contact:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 83 |
+
- **Size of downloaded dataset files:** 71.49 MB
|
| 84 |
+
- **Size of the generated dataset:** 65.32 MB
|
| 85 |
+
- **Total amount of disk used:** 136.81 MB
|
| 86 |
+
|
| 87 |
+
### Dataset Summary
|
| 88 |
+
|
| 89 |
+
HellaSwag: Can a Machine Really Finish Your Sentence? is a new dataset for commonsense NLI. A paper was published at ACL2019.
|
| 90 |
+
|
| 91 |
+
### Supported Tasks and Leaderboards
|
| 92 |
+
|
| 93 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 94 |
+
|
| 95 |
+
### Languages
|
| 96 |
+
|
| 97 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 98 |
+
|
| 99 |
+
## Dataset Structure
|
| 100 |
+
|
| 101 |
+
### Data Instances
|
| 102 |
+
|
| 103 |
+
#### default
|
| 104 |
+
|
| 105 |
+
- **Size of downloaded dataset files:** 71.49 MB
|
| 106 |
+
- **Size of the generated dataset:** 65.32 MB
|
| 107 |
+
- **Total amount of disk used:** 136.81 MB
|
| 108 |
+
|
| 109 |
+
An example of 'train' looks as follows.
|
| 110 |
+
```
|
| 111 |
+
This example was too long and was cropped:
|
| 112 |
+
|
| 113 |
+
{
|
| 114 |
+
"activity_label": "Removing ice from car",
|
| 115 |
+
"ctx": "Then, the man writes over the snow covering the window of a car, and a woman wearing winter clothes smiles. then",
|
| 116 |
+
"ctx_a": "Then, the man writes over the snow covering the window of a car, and a woman wearing winter clothes smiles.",
|
| 117 |
+
"ctx_b": "then",
|
| 118 |
+
"endings": "[\", the man adds wax to the windshield and cuts it.\", \", a person board a ski lift, while two men supporting the head of the per...",
|
| 119 |
+
"ind": 4,
|
| 120 |
+
"label": "3",
|
| 121 |
+
"source_id": "activitynet~v_-1IBHYS3L-Y",
|
| 122 |
+
"split": "train",
|
| 123 |
+
"split_type": "indomain"
|
| 124 |
+
}
|
| 125 |
+
```
|
| 126 |
+
|
| 127 |
+
### Data Fields
|
| 128 |
+
|
| 129 |
+
The data fields are the same among all splits.
|
| 130 |
+
|
| 131 |
+
#### default
|
| 132 |
+
- `ind`: a `int32` feature.
|
| 133 |
+
- `activity_label`: a `string` feature.
|
| 134 |
+
- `ctx_a`: a `string` feature.
|
| 135 |
+
- `ctx_b`: a `string` feature.
|
| 136 |
+
- `ctx`: a `string` feature.
|
| 137 |
+
- `endings`: a `list` of `string` features.
|
| 138 |
+
- `source_id`: a `string` feature.
|
| 139 |
+
- `split`: a `string` feature.
|
| 140 |
+
- `split_type`: a `string` feature.
|
| 141 |
+
- `label`: a `string` feature.
|
| 142 |
+
|
| 143 |
+
### Data Splits
|
| 144 |
+
|
| 145 |
+
| name |train|validation|test |
|
| 146 |
+
|-------|----:|---------:|----:|
|
| 147 |
+
|default|39905| 10042|10003|
|
| 148 |
+
|
| 149 |
+
## Dataset Creation
|
| 150 |
+
|
| 151 |
+
### Curation Rationale
|
| 152 |
+
|
| 153 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 154 |
+
|
| 155 |
+
### Source Data
|
| 156 |
+
|
| 157 |
+
#### Initial Data Collection and Normalization
|
| 158 |
+
|
| 159 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 160 |
+
|
| 161 |
+
#### Who are the source language producers?
|
| 162 |
+
|
| 163 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 164 |
+
|
| 165 |
+
### Annotations
|
| 166 |
+
|
| 167 |
+
#### Annotation process
|
| 168 |
+
|
| 169 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 170 |
+
|
| 171 |
+
#### Who are the annotators?
|
| 172 |
+
|
| 173 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 174 |
+
|
| 175 |
+
### Personal and Sensitive Information
|
| 176 |
+
|
| 177 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 178 |
+
|
| 179 |
+
## Considerations for Using the Data
|
| 180 |
+
|
| 181 |
+
### Social Impact of Dataset
|
| 182 |
+
|
| 183 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 184 |
+
|
| 185 |
+
### Discussion of Biases
|
| 186 |
+
|
| 187 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 188 |
+
|
| 189 |
+
### Other Known Limitations
|
| 190 |
+
|
| 191 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 192 |
+
|
| 193 |
+
## Additional Information
|
| 194 |
+
|
| 195 |
+
### Dataset Curators
|
| 196 |
+
|
| 197 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 198 |
+
|
| 199 |
+
### Licensing Information
|
| 200 |
+
|
| 201 |
+
MIT https://github.com/rowanz/hellaswag/blob/master/LICENSE
|
| 202 |
+
|
| 203 |
+
### Citation Information
|
| 204 |
+
|
| 205 |
+
```
|
| 206 |
+
@inproceedings{zellers2019hellaswag,
|
| 207 |
+
title={HellaSwag: Can a Machine Really Finish Your Sentence?},
|
| 208 |
+
author={Zellers, Rowan and Holtzman, Ari and Bisk, Yonatan and Farhadi, Ali and Choi, Yejin},
|
| 209 |
+
booktitle ={Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics},
|
| 210 |
+
year={2019}
|
| 211 |
+
}
|
| 212 |
+
|
| 213 |
+
```
|
| 214 |
+
|
| 215 |
+
|
| 216 |
+
### Contributions
|
| 217 |
+
|
| 218 |
+
Thanks to [@albertvillanova](https://github.com/albertvillanova), [@mariamabarham](https://github.com/mariamabarham), [@thomwolf](https://github.com/thomwolf), [@patrickvonplaten](https://github.com/patrickvonplaten), [@lewtun](https://github.com/lewtun) for adding this dataset.
|
hf_cache/hub/datasets--Rowan--hellaswag/refs/main
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
218ec52e09a7e7462a5400043bb9a69a41d06b76
|
hf_cache/hub/datasets--Rowan--hellaswag/snapshots/218ec52e09a7e7462a5400043bb9a69a41d06b76/README.md
ADDED
|
@@ -0,0 +1,218 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language:
|
| 3 |
+
- en
|
| 4 |
+
paperswithcode_id: hellaswag
|
| 5 |
+
pretty_name: HellaSwag
|
| 6 |
+
dataset_info:
|
| 7 |
+
features:
|
| 8 |
+
- name: ind
|
| 9 |
+
dtype: int32
|
| 10 |
+
- name: activity_label
|
| 11 |
+
dtype: string
|
| 12 |
+
- name: ctx_a
|
| 13 |
+
dtype: string
|
| 14 |
+
- name: ctx_b
|
| 15 |
+
dtype: string
|
| 16 |
+
- name: ctx
|
| 17 |
+
dtype: string
|
| 18 |
+
- name: endings
|
| 19 |
+
sequence: string
|
| 20 |
+
- name: source_id
|
| 21 |
+
dtype: string
|
| 22 |
+
- name: split
|
| 23 |
+
dtype: string
|
| 24 |
+
- name: split_type
|
| 25 |
+
dtype: string
|
| 26 |
+
- name: label
|
| 27 |
+
dtype: string
|
| 28 |
+
splits:
|
| 29 |
+
- name: train
|
| 30 |
+
num_bytes: 43232624
|
| 31 |
+
num_examples: 39905
|
| 32 |
+
- name: test
|
| 33 |
+
num_bytes: 10791853
|
| 34 |
+
num_examples: 10003
|
| 35 |
+
- name: validation
|
| 36 |
+
num_bytes: 11175717
|
| 37 |
+
num_examples: 10042
|
| 38 |
+
download_size: 36793872
|
| 39 |
+
dataset_size: 65200194
|
| 40 |
+
configs:
|
| 41 |
+
- config_name: default
|
| 42 |
+
data_files:
|
| 43 |
+
- split: train
|
| 44 |
+
path: data/train-*
|
| 45 |
+
- split: test
|
| 46 |
+
path: data/test-*
|
| 47 |
+
- split: validation
|
| 48 |
+
path: data/validation-*
|
| 49 |
+
---
|
| 50 |
+
|
| 51 |
+
# Dataset Card for "hellaswag"
|
| 52 |
+
|
| 53 |
+
## Table of Contents
|
| 54 |
+
- [Dataset Description](#dataset-description)
|
| 55 |
+
- [Dataset Summary](#dataset-summary)
|
| 56 |
+
- [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
|
| 57 |
+
- [Languages](#languages)
|
| 58 |
+
- [Dataset Structure](#dataset-structure)
|
| 59 |
+
- [Data Instances](#data-instances)
|
| 60 |
+
- [Data Fields](#data-fields)
|
| 61 |
+
- [Data Splits](#data-splits)
|
| 62 |
+
- [Dataset Creation](#dataset-creation)
|
| 63 |
+
- [Curation Rationale](#curation-rationale)
|
| 64 |
+
- [Source Data](#source-data)
|
| 65 |
+
- [Annotations](#annotations)
|
| 66 |
+
- [Personal and Sensitive Information](#personal-and-sensitive-information)
|
| 67 |
+
- [Considerations for Using the Data](#considerations-for-using-the-data)
|
| 68 |
+
- [Social Impact of Dataset](#social-impact-of-dataset)
|
| 69 |
+
- [Discussion of Biases](#discussion-of-biases)
|
| 70 |
+
- [Other Known Limitations](#other-known-limitations)
|
| 71 |
+
- [Additional Information](#additional-information)
|
| 72 |
+
- [Dataset Curators](#dataset-curators)
|
| 73 |
+
- [Licensing Information](#licensing-information)
|
| 74 |
+
- [Citation Information](#citation-information)
|
| 75 |
+
- [Contributions](#contributions)
|
| 76 |
+
|
| 77 |
+
## Dataset Description
|
| 78 |
+
|
| 79 |
+
- **Homepage:** [https://rowanzellers.com/hellaswag/](https://rowanzellers.com/hellaswag/)
|
| 80 |
+
- **Repository:** [https://github.com/rowanz/hellaswag/](https://github.com/rowanz/hellaswag/)
|
| 81 |
+
- **Paper:** [HellaSwag: Can a Machine Really Finish Your Sentence?](https://arxiv.org/abs/1905.07830)
|
| 82 |
+
- **Point of Contact:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 83 |
+
- **Size of downloaded dataset files:** 71.49 MB
|
| 84 |
+
- **Size of the generated dataset:** 65.32 MB
|
| 85 |
+
- **Total amount of disk used:** 136.81 MB
|
| 86 |
+
|
| 87 |
+
### Dataset Summary
|
| 88 |
+
|
| 89 |
+
HellaSwag: Can a Machine Really Finish Your Sentence? is a new dataset for commonsense NLI. A paper was published at ACL2019.
|
| 90 |
+
|
| 91 |
+
### Supported Tasks and Leaderboards
|
| 92 |
+
|
| 93 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 94 |
+
|
| 95 |
+
### Languages
|
| 96 |
+
|
| 97 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 98 |
+
|
| 99 |
+
## Dataset Structure
|
| 100 |
+
|
| 101 |
+
### Data Instances
|
| 102 |
+
|
| 103 |
+
#### default
|
| 104 |
+
|
| 105 |
+
- **Size of downloaded dataset files:** 71.49 MB
|
| 106 |
+
- **Size of the generated dataset:** 65.32 MB
|
| 107 |
+
- **Total amount of disk used:** 136.81 MB
|
| 108 |
+
|
| 109 |
+
An example of 'train' looks as follows.
|
| 110 |
+
```
|
| 111 |
+
This example was too long and was cropped:
|
| 112 |
+
|
| 113 |
+
{
|
| 114 |
+
"activity_label": "Removing ice from car",
|
| 115 |
+
"ctx": "Then, the man writes over the snow covering the window of a car, and a woman wearing winter clothes smiles. then",
|
| 116 |
+
"ctx_a": "Then, the man writes over the snow covering the window of a car, and a woman wearing winter clothes smiles.",
|
| 117 |
+
"ctx_b": "then",
|
| 118 |
+
"endings": "[\", the man adds wax to the windshield and cuts it.\", \", a person board a ski lift, while two men supporting the head of the per...",
|
| 119 |
+
"ind": 4,
|
| 120 |
+
"label": "3",
|
| 121 |
+
"source_id": "activitynet~v_-1IBHYS3L-Y",
|
| 122 |
+
"split": "train",
|
| 123 |
+
"split_type": "indomain"
|
| 124 |
+
}
|
| 125 |
+
```
|
| 126 |
+
|
| 127 |
+
### Data Fields
|
| 128 |
+
|
| 129 |
+
The data fields are the same among all splits.
|
| 130 |
+
|
| 131 |
+
#### default
|
| 132 |
+
- `ind`: a `int32` feature.
|
| 133 |
+
- `activity_label`: a `string` feature.
|
| 134 |
+
- `ctx_a`: a `string` feature.
|
| 135 |
+
- `ctx_b`: a `string` feature.
|
| 136 |
+
- `ctx`: a `string` feature.
|
| 137 |
+
- `endings`: a `list` of `string` features.
|
| 138 |
+
- `source_id`: a `string` feature.
|
| 139 |
+
- `split`: a `string` feature.
|
| 140 |
+
- `split_type`: a `string` feature.
|
| 141 |
+
- `label`: a `string` feature.
|
| 142 |
+
|
| 143 |
+
### Data Splits
|
| 144 |
+
|
| 145 |
+
| name |train|validation|test |
|
| 146 |
+
|-------|----:|---------:|----:|
|
| 147 |
+
|default|39905| 10042|10003|
|
| 148 |
+
|
| 149 |
+
## Dataset Creation
|
| 150 |
+
|
| 151 |
+
### Curation Rationale
|
| 152 |
+
|
| 153 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 154 |
+
|
| 155 |
+
### Source Data
|
| 156 |
+
|
| 157 |
+
#### Initial Data Collection and Normalization
|
| 158 |
+
|
| 159 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 160 |
+
|
| 161 |
+
#### Who are the source language producers?
|
| 162 |
+
|
| 163 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 164 |
+
|
| 165 |
+
### Annotations
|
| 166 |
+
|
| 167 |
+
#### Annotation process
|
| 168 |
+
|
| 169 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 170 |
+
|
| 171 |
+
#### Who are the annotators?
|
| 172 |
+
|
| 173 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 174 |
+
|
| 175 |
+
### Personal and Sensitive Information
|
| 176 |
+
|
| 177 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 178 |
+
|
| 179 |
+
## Considerations for Using the Data
|
| 180 |
+
|
| 181 |
+
### Social Impact of Dataset
|
| 182 |
+
|
| 183 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 184 |
+
|
| 185 |
+
### Discussion of Biases
|
| 186 |
+
|
| 187 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 188 |
+
|
| 189 |
+
### Other Known Limitations
|
| 190 |
+
|
| 191 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 192 |
+
|
| 193 |
+
## Additional Information
|
| 194 |
+
|
| 195 |
+
### Dataset Curators
|
| 196 |
+
|
| 197 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 198 |
+
|
| 199 |
+
### Licensing Information
|
| 200 |
+
|
| 201 |
+
MIT https://github.com/rowanz/hellaswag/blob/master/LICENSE
|
| 202 |
+
|
| 203 |
+
### Citation Information
|
| 204 |
+
|
| 205 |
+
```
|
| 206 |
+
@inproceedings{zellers2019hellaswag,
|
| 207 |
+
title={HellaSwag: Can a Machine Really Finish Your Sentence?},
|
| 208 |
+
author={Zellers, Rowan and Holtzman, Ari and Bisk, Yonatan and Farhadi, Ali and Choi, Yejin},
|
| 209 |
+
booktitle ={Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics},
|
| 210 |
+
year={2019}
|
| 211 |
+
}
|
| 212 |
+
|
| 213 |
+
```
|
| 214 |
+
|
| 215 |
+
|
| 216 |
+
### Contributions
|
| 217 |
+
|
| 218 |
+
Thanks to [@albertvillanova](https://github.com/albertvillanova), [@mariamabarham](https://github.com/mariamabarham), [@thomwolf](https://github.com/thomwolf), [@patrickvonplaten](https://github.com/patrickvonplaten), [@lewtun](https://github.com/lewtun) for adding this dataset.
|
hf_cache/hub/datasets--Salesforce--wikitext/.no_exist/b08601e04326c79dfdd32d625aee71d232d685c3/.huggingface.yaml
ADDED
|
File without changes
|
hf_cache/hub/datasets--Salesforce--wikitext/.no_exist/b08601e04326c79dfdd32d625aee71d232d685c3/dataset_infos.json
ADDED
|
File without changes
|
hf_cache/hub/datasets--Salesforce--wikitext/.no_exist/b08601e04326c79dfdd32d625aee71d232d685c3/wikitext.py
ADDED
|
File without changes
|
hf_cache/hub/datasets--Salesforce--wikitext/blobs/2a4fec2bc8df76c9d4da1c8e8865b625eb221c76
ADDED
|
@@ -0,0 +1,344 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
annotations_creators:
|
| 3 |
+
- no-annotation
|
| 4 |
+
language_creators:
|
| 5 |
+
- crowdsourced
|
| 6 |
+
language:
|
| 7 |
+
- en
|
| 8 |
+
license:
|
| 9 |
+
- cc-by-sa-3.0
|
| 10 |
+
- gfdl
|
| 11 |
+
multilinguality:
|
| 12 |
+
- monolingual
|
| 13 |
+
size_categories:
|
| 14 |
+
- 1M<n<10M
|
| 15 |
+
source_datasets:
|
| 16 |
+
- original
|
| 17 |
+
task_categories:
|
| 18 |
+
- text-generation
|
| 19 |
+
- fill-mask
|
| 20 |
+
task_ids:
|
| 21 |
+
- language-modeling
|
| 22 |
+
- masked-language-modeling
|
| 23 |
+
paperswithcode_id: wikitext-2
|
| 24 |
+
pretty_name: WikiText
|
| 25 |
+
dataset_info:
|
| 26 |
+
- config_name: wikitext-103-raw-v1
|
| 27 |
+
features:
|
| 28 |
+
- name: text
|
| 29 |
+
dtype: string
|
| 30 |
+
splits:
|
| 31 |
+
- name: test
|
| 32 |
+
num_bytes: 1305088
|
| 33 |
+
num_examples: 4358
|
| 34 |
+
- name: train
|
| 35 |
+
num_bytes: 546500949
|
| 36 |
+
num_examples: 1801350
|
| 37 |
+
- name: validation
|
| 38 |
+
num_bytes: 1159288
|
| 39 |
+
num_examples: 3760
|
| 40 |
+
download_size: 315466397
|
| 41 |
+
dataset_size: 548965325
|
| 42 |
+
- config_name: wikitext-103-v1
|
| 43 |
+
features:
|
| 44 |
+
- name: text
|
| 45 |
+
dtype: string
|
| 46 |
+
splits:
|
| 47 |
+
- name: test
|
| 48 |
+
num_bytes: 1295575
|
| 49 |
+
num_examples: 4358
|
| 50 |
+
- name: train
|
| 51 |
+
num_bytes: 545141915
|
| 52 |
+
num_examples: 1801350
|
| 53 |
+
- name: validation
|
| 54 |
+
num_bytes: 1154751
|
| 55 |
+
num_examples: 3760
|
| 56 |
+
download_size: 313093838
|
| 57 |
+
dataset_size: 547592241
|
| 58 |
+
- config_name: wikitext-2-raw-v1
|
| 59 |
+
features:
|
| 60 |
+
- name: text
|
| 61 |
+
dtype: string
|
| 62 |
+
splits:
|
| 63 |
+
- name: test
|
| 64 |
+
num_bytes: 1305088
|
| 65 |
+
num_examples: 4358
|
| 66 |
+
- name: train
|
| 67 |
+
num_bytes: 11061717
|
| 68 |
+
num_examples: 36718
|
| 69 |
+
- name: validation
|
| 70 |
+
num_bytes: 1159288
|
| 71 |
+
num_examples: 3760
|
| 72 |
+
download_size: 7747362
|
| 73 |
+
dataset_size: 13526093
|
| 74 |
+
- config_name: wikitext-2-v1
|
| 75 |
+
features:
|
| 76 |
+
- name: text
|
| 77 |
+
dtype: string
|
| 78 |
+
splits:
|
| 79 |
+
- name: test
|
| 80 |
+
num_bytes: 1270947
|
| 81 |
+
num_examples: 4358
|
| 82 |
+
- name: train
|
| 83 |
+
num_bytes: 10918118
|
| 84 |
+
num_examples: 36718
|
| 85 |
+
- name: validation
|
| 86 |
+
num_bytes: 1134123
|
| 87 |
+
num_examples: 3760
|
| 88 |
+
download_size: 7371282
|
| 89 |
+
dataset_size: 13323188
|
| 90 |
+
configs:
|
| 91 |
+
- config_name: wikitext-103-raw-v1
|
| 92 |
+
data_files:
|
| 93 |
+
- split: test
|
| 94 |
+
path: wikitext-103-raw-v1/test-*
|
| 95 |
+
- split: train
|
| 96 |
+
path: wikitext-103-raw-v1/train-*
|
| 97 |
+
- split: validation
|
| 98 |
+
path: wikitext-103-raw-v1/validation-*
|
| 99 |
+
- config_name: wikitext-103-v1
|
| 100 |
+
data_files:
|
| 101 |
+
- split: test
|
| 102 |
+
path: wikitext-103-v1/test-*
|
| 103 |
+
- split: train
|
| 104 |
+
path: wikitext-103-v1/train-*
|
| 105 |
+
- split: validation
|
| 106 |
+
path: wikitext-103-v1/validation-*
|
| 107 |
+
- config_name: wikitext-2-raw-v1
|
| 108 |
+
data_files:
|
| 109 |
+
- split: test
|
| 110 |
+
path: wikitext-2-raw-v1/test-*
|
| 111 |
+
- split: train
|
| 112 |
+
path: wikitext-2-raw-v1/train-*
|
| 113 |
+
- split: validation
|
| 114 |
+
path: wikitext-2-raw-v1/validation-*
|
| 115 |
+
- config_name: wikitext-2-v1
|
| 116 |
+
data_files:
|
| 117 |
+
- split: test
|
| 118 |
+
path: wikitext-2-v1/test-*
|
| 119 |
+
- split: train
|
| 120 |
+
path: wikitext-2-v1/train-*
|
| 121 |
+
- split: validation
|
| 122 |
+
path: wikitext-2-v1/validation-*
|
| 123 |
+
---
|
| 124 |
+
|
| 125 |
+
# Dataset Card for "wikitext"
|
| 126 |
+
|
| 127 |
+
## Table of Contents
|
| 128 |
+
- [Dataset Description](#dataset-description)
|
| 129 |
+
- [Dataset Summary](#dataset-summary)
|
| 130 |
+
- [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
|
| 131 |
+
- [Languages](#languages)
|
| 132 |
+
- [Dataset Structure](#dataset-structure)
|
| 133 |
+
- [Data Instances](#data-instances)
|
| 134 |
+
- [Data Fields](#data-fields)
|
| 135 |
+
- [Data Splits](#data-splits)
|
| 136 |
+
- [Dataset Creation](#dataset-creation)
|
| 137 |
+
- [Curation Rationale](#curation-rationale)
|
| 138 |
+
- [Source Data](#source-data)
|
| 139 |
+
- [Annotations](#annotations)
|
| 140 |
+
- [Personal and Sensitive Information](#personal-and-sensitive-information)
|
| 141 |
+
- [Considerations for Using the Data](#considerations-for-using-the-data)
|
| 142 |
+
- [Social Impact of Dataset](#social-impact-of-dataset)
|
| 143 |
+
- [Discussion of Biases](#discussion-of-biases)
|
| 144 |
+
- [Other Known Limitations](#other-known-limitations)
|
| 145 |
+
- [Additional Information](#additional-information)
|
| 146 |
+
- [Dataset Curators](#dataset-curators)
|
| 147 |
+
- [Licensing Information](#licensing-information)
|
| 148 |
+
- [Citation Information](#citation-information)
|
| 149 |
+
- [Contributions](#contributions)
|
| 150 |
+
|
| 151 |
+
## Dataset Description
|
| 152 |
+
|
| 153 |
+
- **Homepage:** [https://blog.einstein.ai/the-wikitext-long-term-dependency-language-modeling-dataset/](https://blog.einstein.ai/the-wikitext-long-term-dependency-language-modeling-dataset/)
|
| 154 |
+
- **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 155 |
+
- **Paper:** [Pointer Sentinel Mixture Models](https://arxiv.org/abs/1609.07843)
|
| 156 |
+
- **Point of Contact:** [Stephen Merity](mailto:smerity@salesforce.com)
|
| 157 |
+
- **Size of downloaded dataset files:** 391.41 MB
|
| 158 |
+
- **Size of the generated dataset:** 1.12 GB
|
| 159 |
+
- **Total amount of disk used:** 1.52 GB
|
| 160 |
+
|
| 161 |
+
### Dataset Summary
|
| 162 |
+
|
| 163 |
+
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
|
| 164 |
+
Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.
|
| 165 |
+
|
| 166 |
+
Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over
|
| 167 |
+
110 times larger. The WikiText dataset also features a far larger vocabulary and retains the original case, punctuation
|
| 168 |
+
and numbers - all of which are removed in PTB. As it is composed of full articles, the dataset is well suited for models
|
| 169 |
+
that can take advantage of long term dependencies.
|
| 170 |
+
|
| 171 |
+
Each subset comes in two different variants:
|
| 172 |
+
- Raw (for character level work) contain the raw tokens, before the addition of the <unk> (unknown) tokens.
|
| 173 |
+
- Non-raw (for word level work) contain only the tokens in their vocabulary (wiki.train.tokens, wiki.valid.tokens, and wiki.test.tokens).
|
| 174 |
+
The out-of-vocabulary tokens have been replaced with the the <unk> token.
|
| 175 |
+
|
| 176 |
+
|
| 177 |
+
### Supported Tasks and Leaderboards
|
| 178 |
+
|
| 179 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 180 |
+
|
| 181 |
+
### Languages
|
| 182 |
+
|
| 183 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 184 |
+
|
| 185 |
+
## Dataset Structure
|
| 186 |
+
|
| 187 |
+
### Data Instances
|
| 188 |
+
|
| 189 |
+
#### wikitext-103-raw-v1
|
| 190 |
+
|
| 191 |
+
- **Size of downloaded dataset files:** 191.98 MB
|
| 192 |
+
- **Size of the generated dataset:** 549.42 MB
|
| 193 |
+
- **Total amount of disk used:** 741.41 MB
|
| 194 |
+
|
| 195 |
+
An example of 'validation' looks as follows.
|
| 196 |
+
```
|
| 197 |
+
This example was too long and was cropped:
|
| 198 |
+
|
| 199 |
+
{
|
| 200 |
+
"text": "\" The gold dollar or gold one @-@ dollar piece was a coin struck as a regular issue by the United States Bureau of the Mint from..."
|
| 201 |
+
}
|
| 202 |
+
```
|
| 203 |
+
|
| 204 |
+
#### wikitext-103-v1
|
| 205 |
+
|
| 206 |
+
- **Size of downloaded dataset files:** 190.23 MB
|
| 207 |
+
- **Size of the generated dataset:** 548.05 MB
|
| 208 |
+
- **Total amount of disk used:** 738.27 MB
|
| 209 |
+
|
| 210 |
+
An example of 'train' looks as follows.
|
| 211 |
+
```
|
| 212 |
+
This example was too long and was cropped:
|
| 213 |
+
|
| 214 |
+
{
|
| 215 |
+
"text": "\" Senjō no Valkyria 3 : <unk> Chronicles ( Japanese : 戦場のヴァルキュリア3 , lit . Valkyria of the Battlefield 3 ) , commonly referred to..."
|
| 216 |
+
}
|
| 217 |
+
```
|
| 218 |
+
|
| 219 |
+
#### wikitext-2-raw-v1
|
| 220 |
+
|
| 221 |
+
- **Size of downloaded dataset files:** 4.72 MB
|
| 222 |
+
- **Size of the generated dataset:** 13.54 MB
|
| 223 |
+
- **Total amount of disk used:** 18.26 MB
|
| 224 |
+
|
| 225 |
+
An example of 'train' looks as follows.
|
| 226 |
+
```
|
| 227 |
+
This example was too long and was cropped:
|
| 228 |
+
|
| 229 |
+
{
|
| 230 |
+
"text": "\" The Sinclair Scientific Programmable was introduced in 1975 , with the same case as the Sinclair Oxford . It was larger than t..."
|
| 231 |
+
}
|
| 232 |
+
```
|
| 233 |
+
|
| 234 |
+
#### wikitext-2-v1
|
| 235 |
+
|
| 236 |
+
- **Size of downloaded dataset files:** 4.48 MB
|
| 237 |
+
- **Size of the generated dataset:** 13.34 MB
|
| 238 |
+
- **Total amount of disk used:** 17.82 MB
|
| 239 |
+
|
| 240 |
+
An example of 'train' looks as follows.
|
| 241 |
+
```
|
| 242 |
+
This example was too long and was cropped:
|
| 243 |
+
|
| 244 |
+
{
|
| 245 |
+
"text": "\" Senjō no Valkyria 3 : <unk> Chronicles ( Japanese : 戦場のヴァルキュリア3 , lit . Valkyria of the Battlefield 3 ) , commonly referred to..."
|
| 246 |
+
}
|
| 247 |
+
```
|
| 248 |
+
|
| 249 |
+
### Data Fields
|
| 250 |
+
|
| 251 |
+
The data fields are the same among all splits.
|
| 252 |
+
|
| 253 |
+
#### wikitext-103-raw-v1
|
| 254 |
+
- `text`: a `string` feature.
|
| 255 |
+
|
| 256 |
+
#### wikitext-103-v1
|
| 257 |
+
- `text`: a `string` feature.
|
| 258 |
+
|
| 259 |
+
#### wikitext-2-raw-v1
|
| 260 |
+
- `text`: a `string` feature.
|
| 261 |
+
|
| 262 |
+
#### wikitext-2-v1
|
| 263 |
+
- `text`: a `string` feature.
|
| 264 |
+
|
| 265 |
+
### Data Splits
|
| 266 |
+
|
| 267 |
+
| name | train |validation|test|
|
| 268 |
+
|-------------------|------:|---------:|---:|
|
| 269 |
+
|wikitext-103-raw-v1|1801350| 3760|4358|
|
| 270 |
+
|wikitext-103-v1 |1801350| 3760|4358|
|
| 271 |
+
|wikitext-2-raw-v1 | 36718| 3760|4358|
|
| 272 |
+
|wikitext-2-v1 | 36718| 3760|4358|
|
| 273 |
+
|
| 274 |
+
## Dataset Creation
|
| 275 |
+
|
| 276 |
+
### Curation Rationale
|
| 277 |
+
|
| 278 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 279 |
+
|
| 280 |
+
### Source Data
|
| 281 |
+
|
| 282 |
+
#### Initial Data Collection and Normalization
|
| 283 |
+
|
| 284 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 285 |
+
|
| 286 |
+
#### Who are the source language producers?
|
| 287 |
+
|
| 288 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 289 |
+
|
| 290 |
+
### Annotations
|
| 291 |
+
|
| 292 |
+
#### Annotation process
|
| 293 |
+
|
| 294 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 295 |
+
|
| 296 |
+
#### Who are the annotators?
|
| 297 |
+
|
| 298 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 299 |
+
|
| 300 |
+
### Personal and Sensitive Information
|
| 301 |
+
|
| 302 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 303 |
+
|
| 304 |
+
## Considerations for Using the Data
|
| 305 |
+
|
| 306 |
+
### Social Impact of Dataset
|
| 307 |
+
|
| 308 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 309 |
+
|
| 310 |
+
### Discussion of Biases
|
| 311 |
+
|
| 312 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 313 |
+
|
| 314 |
+
### Other Known Limitations
|
| 315 |
+
|
| 316 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 317 |
+
|
| 318 |
+
## Additional Information
|
| 319 |
+
|
| 320 |
+
### Dataset Curators
|
| 321 |
+
|
| 322 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 323 |
+
|
| 324 |
+
### Licensing Information
|
| 325 |
+
|
| 326 |
+
The dataset is available under the [Creative Commons Attribution-ShareAlike License (CC BY-SA 4.0)](https://creativecommons.org/licenses/by-sa/4.0/).
|
| 327 |
+
|
| 328 |
+
### Citation Information
|
| 329 |
+
|
| 330 |
+
```
|
| 331 |
+
@misc{merity2016pointer,
|
| 332 |
+
title={Pointer Sentinel Mixture Models},
|
| 333 |
+
author={Stephen Merity and Caiming Xiong and James Bradbury and Richard Socher},
|
| 334 |
+
year={2016},
|
| 335 |
+
eprint={1609.07843},
|
| 336 |
+
archivePrefix={arXiv},
|
| 337 |
+
primaryClass={cs.CL}
|
| 338 |
+
}
|
| 339 |
+
```
|
| 340 |
+
|
| 341 |
+
|
| 342 |
+
### Contributions
|
| 343 |
+
|
| 344 |
+
Thanks to [@thomwolf](https://github.com/thomwolf), [@lewtun](https://github.com/lewtun), [@patrickvonplaten](https://github.com/patrickvonplaten), [@mariamabarham](https://github.com/mariamabarham) for adding this dataset.
|
hf_cache/hub/datasets--Salesforce--wikitext/refs/main
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
b08601e04326c79dfdd32d625aee71d232d685c3
|
hf_cache/hub/datasets--Salesforce--wikitext/snapshots/b08601e04326c79dfdd32d625aee71d232d685c3/README.md
ADDED
|
@@ -0,0 +1,344 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
annotations_creators:
|
| 3 |
+
- no-annotation
|
| 4 |
+
language_creators:
|
| 5 |
+
- crowdsourced
|
| 6 |
+
language:
|
| 7 |
+
- en
|
| 8 |
+
license:
|
| 9 |
+
- cc-by-sa-3.0
|
| 10 |
+
- gfdl
|
| 11 |
+
multilinguality:
|
| 12 |
+
- monolingual
|
| 13 |
+
size_categories:
|
| 14 |
+
- 1M<n<10M
|
| 15 |
+
source_datasets:
|
| 16 |
+
- original
|
| 17 |
+
task_categories:
|
| 18 |
+
- text-generation
|
| 19 |
+
- fill-mask
|
| 20 |
+
task_ids:
|
| 21 |
+
- language-modeling
|
| 22 |
+
- masked-language-modeling
|
| 23 |
+
paperswithcode_id: wikitext-2
|
| 24 |
+
pretty_name: WikiText
|
| 25 |
+
dataset_info:
|
| 26 |
+
- config_name: wikitext-103-raw-v1
|
| 27 |
+
features:
|
| 28 |
+
- name: text
|
| 29 |
+
dtype: string
|
| 30 |
+
splits:
|
| 31 |
+
- name: test
|
| 32 |
+
num_bytes: 1305088
|
| 33 |
+
num_examples: 4358
|
| 34 |
+
- name: train
|
| 35 |
+
num_bytes: 546500949
|
| 36 |
+
num_examples: 1801350
|
| 37 |
+
- name: validation
|
| 38 |
+
num_bytes: 1159288
|
| 39 |
+
num_examples: 3760
|
| 40 |
+
download_size: 315466397
|
| 41 |
+
dataset_size: 548965325
|
| 42 |
+
- config_name: wikitext-103-v1
|
| 43 |
+
features:
|
| 44 |
+
- name: text
|
| 45 |
+
dtype: string
|
| 46 |
+
splits:
|
| 47 |
+
- name: test
|
| 48 |
+
num_bytes: 1295575
|
| 49 |
+
num_examples: 4358
|
| 50 |
+
- name: train
|
| 51 |
+
num_bytes: 545141915
|
| 52 |
+
num_examples: 1801350
|
| 53 |
+
- name: validation
|
| 54 |
+
num_bytes: 1154751
|
| 55 |
+
num_examples: 3760
|
| 56 |
+
download_size: 313093838
|
| 57 |
+
dataset_size: 547592241
|
| 58 |
+
- config_name: wikitext-2-raw-v1
|
| 59 |
+
features:
|
| 60 |
+
- name: text
|
| 61 |
+
dtype: string
|
| 62 |
+
splits:
|
| 63 |
+
- name: test
|
| 64 |
+
num_bytes: 1305088
|
| 65 |
+
num_examples: 4358
|
| 66 |
+
- name: train
|
| 67 |
+
num_bytes: 11061717
|
| 68 |
+
num_examples: 36718
|
| 69 |
+
- name: validation
|
| 70 |
+
num_bytes: 1159288
|
| 71 |
+
num_examples: 3760
|
| 72 |
+
download_size: 7747362
|
| 73 |
+
dataset_size: 13526093
|
| 74 |
+
- config_name: wikitext-2-v1
|
| 75 |
+
features:
|
| 76 |
+
- name: text
|
| 77 |
+
dtype: string
|
| 78 |
+
splits:
|
| 79 |
+
- name: test
|
| 80 |
+
num_bytes: 1270947
|
| 81 |
+
num_examples: 4358
|
| 82 |
+
- name: train
|
| 83 |
+
num_bytes: 10918118
|
| 84 |
+
num_examples: 36718
|
| 85 |
+
- name: validation
|
| 86 |
+
num_bytes: 1134123
|
| 87 |
+
num_examples: 3760
|
| 88 |
+
download_size: 7371282
|
| 89 |
+
dataset_size: 13323188
|
| 90 |
+
configs:
|
| 91 |
+
- config_name: wikitext-103-raw-v1
|
| 92 |
+
data_files:
|
| 93 |
+
- split: test
|
| 94 |
+
path: wikitext-103-raw-v1/test-*
|
| 95 |
+
- split: train
|
| 96 |
+
path: wikitext-103-raw-v1/train-*
|
| 97 |
+
- split: validation
|
| 98 |
+
path: wikitext-103-raw-v1/validation-*
|
| 99 |
+
- config_name: wikitext-103-v1
|
| 100 |
+
data_files:
|
| 101 |
+
- split: test
|
| 102 |
+
path: wikitext-103-v1/test-*
|
| 103 |
+
- split: train
|
| 104 |
+
path: wikitext-103-v1/train-*
|
| 105 |
+
- split: validation
|
| 106 |
+
path: wikitext-103-v1/validation-*
|
| 107 |
+
- config_name: wikitext-2-raw-v1
|
| 108 |
+
data_files:
|
| 109 |
+
- split: test
|
| 110 |
+
path: wikitext-2-raw-v1/test-*
|
| 111 |
+
- split: train
|
| 112 |
+
path: wikitext-2-raw-v1/train-*
|
| 113 |
+
- split: validation
|
| 114 |
+
path: wikitext-2-raw-v1/validation-*
|
| 115 |
+
- config_name: wikitext-2-v1
|
| 116 |
+
data_files:
|
| 117 |
+
- split: test
|
| 118 |
+
path: wikitext-2-v1/test-*
|
| 119 |
+
- split: train
|
| 120 |
+
path: wikitext-2-v1/train-*
|
| 121 |
+
- split: validation
|
| 122 |
+
path: wikitext-2-v1/validation-*
|
| 123 |
+
---
|
| 124 |
+
|
| 125 |
+
# Dataset Card for "wikitext"
|
| 126 |
+
|
| 127 |
+
## Table of Contents
|
| 128 |
+
- [Dataset Description](#dataset-description)
|
| 129 |
+
- [Dataset Summary](#dataset-summary)
|
| 130 |
+
- [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
|
| 131 |
+
- [Languages](#languages)
|
| 132 |
+
- [Dataset Structure](#dataset-structure)
|
| 133 |
+
- [Data Instances](#data-instances)
|
| 134 |
+
- [Data Fields](#data-fields)
|
| 135 |
+
- [Data Splits](#data-splits)
|
| 136 |
+
- [Dataset Creation](#dataset-creation)
|
| 137 |
+
- [Curation Rationale](#curation-rationale)
|
| 138 |
+
- [Source Data](#source-data)
|
| 139 |
+
- [Annotations](#annotations)
|
| 140 |
+
- [Personal and Sensitive Information](#personal-and-sensitive-information)
|
| 141 |
+
- [Considerations for Using the Data](#considerations-for-using-the-data)
|
| 142 |
+
- [Social Impact of Dataset](#social-impact-of-dataset)
|
| 143 |
+
- [Discussion of Biases](#discussion-of-biases)
|
| 144 |
+
- [Other Known Limitations](#other-known-limitations)
|
| 145 |
+
- [Additional Information](#additional-information)
|
| 146 |
+
- [Dataset Curators](#dataset-curators)
|
| 147 |
+
- [Licensing Information](#licensing-information)
|
| 148 |
+
- [Citation Information](#citation-information)
|
| 149 |
+
- [Contributions](#contributions)
|
| 150 |
+
|
| 151 |
+
## Dataset Description
|
| 152 |
+
|
| 153 |
+
- **Homepage:** [https://blog.einstein.ai/the-wikitext-long-term-dependency-language-modeling-dataset/](https://blog.einstein.ai/the-wikitext-long-term-dependency-language-modeling-dataset/)
|
| 154 |
+
- **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 155 |
+
- **Paper:** [Pointer Sentinel Mixture Models](https://arxiv.org/abs/1609.07843)
|
| 156 |
+
- **Point of Contact:** [Stephen Merity](mailto:smerity@salesforce.com)
|
| 157 |
+
- **Size of downloaded dataset files:** 391.41 MB
|
| 158 |
+
- **Size of the generated dataset:** 1.12 GB
|
| 159 |
+
- **Total amount of disk used:** 1.52 GB
|
| 160 |
+
|
| 161 |
+
### Dataset Summary
|
| 162 |
+
|
| 163 |
+
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
|
| 164 |
+
Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.
|
| 165 |
+
|
| 166 |
+
Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over
|
| 167 |
+
110 times larger. The WikiText dataset also features a far larger vocabulary and retains the original case, punctuation
|
| 168 |
+
and numbers - all of which are removed in PTB. As it is composed of full articles, the dataset is well suited for models
|
| 169 |
+
that can take advantage of long term dependencies.
|
| 170 |
+
|
| 171 |
+
Each subset comes in two different variants:
|
| 172 |
+
- Raw (for character level work) contain the raw tokens, before the addition of the <unk> (unknown) tokens.
|
| 173 |
+
- Non-raw (for word level work) contain only the tokens in their vocabulary (wiki.train.tokens, wiki.valid.tokens, and wiki.test.tokens).
|
| 174 |
+
The out-of-vocabulary tokens have been replaced with the the <unk> token.
|
| 175 |
+
|
| 176 |
+
|
| 177 |
+
### Supported Tasks and Leaderboards
|
| 178 |
+
|
| 179 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 180 |
+
|
| 181 |
+
### Languages
|
| 182 |
+
|
| 183 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 184 |
+
|
| 185 |
+
## Dataset Structure
|
| 186 |
+
|
| 187 |
+
### Data Instances
|
| 188 |
+
|
| 189 |
+
#### wikitext-103-raw-v1
|
| 190 |
+
|
| 191 |
+
- **Size of downloaded dataset files:** 191.98 MB
|
| 192 |
+
- **Size of the generated dataset:** 549.42 MB
|
| 193 |
+
- **Total amount of disk used:** 741.41 MB
|
| 194 |
+
|
| 195 |
+
An example of 'validation' looks as follows.
|
| 196 |
+
```
|
| 197 |
+
This example was too long and was cropped:
|
| 198 |
+
|
| 199 |
+
{
|
| 200 |
+
"text": "\" The gold dollar or gold one @-@ dollar piece was a coin struck as a regular issue by the United States Bureau of the Mint from..."
|
| 201 |
+
}
|
| 202 |
+
```
|
| 203 |
+
|
| 204 |
+
#### wikitext-103-v1
|
| 205 |
+
|
| 206 |
+
- **Size of downloaded dataset files:** 190.23 MB
|
| 207 |
+
- **Size of the generated dataset:** 548.05 MB
|
| 208 |
+
- **Total amount of disk used:** 738.27 MB
|
| 209 |
+
|
| 210 |
+
An example of 'train' looks as follows.
|
| 211 |
+
```
|
| 212 |
+
This example was too long and was cropped:
|
| 213 |
+
|
| 214 |
+
{
|
| 215 |
+
"text": "\" Senjō no Valkyria 3 : <unk> Chronicles ( Japanese : 戦場のヴァルキュリア3 , lit . Valkyria of the Battlefield 3 ) , commonly referred to..."
|
| 216 |
+
}
|
| 217 |
+
```
|
| 218 |
+
|
| 219 |
+
#### wikitext-2-raw-v1
|
| 220 |
+
|
| 221 |
+
- **Size of downloaded dataset files:** 4.72 MB
|
| 222 |
+
- **Size of the generated dataset:** 13.54 MB
|
| 223 |
+
- **Total amount of disk used:** 18.26 MB
|
| 224 |
+
|
| 225 |
+
An example of 'train' looks as follows.
|
| 226 |
+
```
|
| 227 |
+
This example was too long and was cropped:
|
| 228 |
+
|
| 229 |
+
{
|
| 230 |
+
"text": "\" The Sinclair Scientific Programmable was introduced in 1975 , with the same case as the Sinclair Oxford . It was larger than t..."
|
| 231 |
+
}
|
| 232 |
+
```
|
| 233 |
+
|
| 234 |
+
#### wikitext-2-v1
|
| 235 |
+
|
| 236 |
+
- **Size of downloaded dataset files:** 4.48 MB
|
| 237 |
+
- **Size of the generated dataset:** 13.34 MB
|
| 238 |
+
- **Total amount of disk used:** 17.82 MB
|
| 239 |
+
|
| 240 |
+
An example of 'train' looks as follows.
|
| 241 |
+
```
|
| 242 |
+
This example was too long and was cropped:
|
| 243 |
+
|
| 244 |
+
{
|
| 245 |
+
"text": "\" Senjō no Valkyria 3 : <unk> Chronicles ( Japanese : 戦場のヴァルキュリア3 , lit . Valkyria of the Battlefield 3 ) , commonly referred to..."
|
| 246 |
+
}
|
| 247 |
+
```
|
| 248 |
+
|
| 249 |
+
### Data Fields
|
| 250 |
+
|
| 251 |
+
The data fields are the same among all splits.
|
| 252 |
+
|
| 253 |
+
#### wikitext-103-raw-v1
|
| 254 |
+
- `text`: a `string` feature.
|
| 255 |
+
|
| 256 |
+
#### wikitext-103-v1
|
| 257 |
+
- `text`: a `string` feature.
|
| 258 |
+
|
| 259 |
+
#### wikitext-2-raw-v1
|
| 260 |
+
- `text`: a `string` feature.
|
| 261 |
+
|
| 262 |
+
#### wikitext-2-v1
|
| 263 |
+
- `text`: a `string` feature.
|
| 264 |
+
|
| 265 |
+
### Data Splits
|
| 266 |
+
|
| 267 |
+
| name | train |validation|test|
|
| 268 |
+
|-------------------|------:|---------:|---:|
|
| 269 |
+
|wikitext-103-raw-v1|1801350| 3760|4358|
|
| 270 |
+
|wikitext-103-v1 |1801350| 3760|4358|
|
| 271 |
+
|wikitext-2-raw-v1 | 36718| 3760|4358|
|
| 272 |
+
|wikitext-2-v1 | 36718| 3760|4358|
|
| 273 |
+
|
| 274 |
+
## Dataset Creation
|
| 275 |
+
|
| 276 |
+
### Curation Rationale
|
| 277 |
+
|
| 278 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 279 |
+
|
| 280 |
+
### Source Data
|
| 281 |
+
|
| 282 |
+
#### Initial Data Collection and Normalization
|
| 283 |
+
|
| 284 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 285 |
+
|
| 286 |
+
#### Who are the source language producers?
|
| 287 |
+
|
| 288 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 289 |
+
|
| 290 |
+
### Annotations
|
| 291 |
+
|
| 292 |
+
#### Annotation process
|
| 293 |
+
|
| 294 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 295 |
+
|
| 296 |
+
#### Who are the annotators?
|
| 297 |
+
|
| 298 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 299 |
+
|
| 300 |
+
### Personal and Sensitive Information
|
| 301 |
+
|
| 302 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 303 |
+
|
| 304 |
+
## Considerations for Using the Data
|
| 305 |
+
|
| 306 |
+
### Social Impact of Dataset
|
| 307 |
+
|
| 308 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 309 |
+
|
| 310 |
+
### Discussion of Biases
|
| 311 |
+
|
| 312 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 313 |
+
|
| 314 |
+
### Other Known Limitations
|
| 315 |
+
|
| 316 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 317 |
+
|
| 318 |
+
## Additional Information
|
| 319 |
+
|
| 320 |
+
### Dataset Curators
|
| 321 |
+
|
| 322 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 323 |
+
|
| 324 |
+
### Licensing Information
|
| 325 |
+
|
| 326 |
+
The dataset is available under the [Creative Commons Attribution-ShareAlike License (CC BY-SA 4.0)](https://creativecommons.org/licenses/by-sa/4.0/).
|
| 327 |
+
|
| 328 |
+
### Citation Information
|
| 329 |
+
|
| 330 |
+
```
|
| 331 |
+
@misc{merity2016pointer,
|
| 332 |
+
title={Pointer Sentinel Mixture Models},
|
| 333 |
+
author={Stephen Merity and Caiming Xiong and James Bradbury and Richard Socher},
|
| 334 |
+
year={2016},
|
| 335 |
+
eprint={1609.07843},
|
| 336 |
+
archivePrefix={arXiv},
|
| 337 |
+
primaryClass={cs.CL}
|
| 338 |
+
}
|
| 339 |
+
```
|
| 340 |
+
|
| 341 |
+
|
| 342 |
+
### Contributions
|
| 343 |
+
|
| 344 |
+
Thanks to [@thomwolf](https://github.com/thomwolf), [@lewtun](https://github.com/lewtun), [@patrickvonplaten](https://github.com/patrickvonplaten), [@mariamabarham](https://github.com/mariamabarham) for adding this dataset.
|
hf_cache/hub/datasets--abisee--cnn_dailymail/.no_exist/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/.huggingface.yaml
ADDED
|
File without changes
|
hf_cache/hub/datasets--abisee--cnn_dailymail/.no_exist/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/cnn_dailymail.py
ADDED
|
File without changes
|
hf_cache/hub/datasets--abisee--cnn_dailymail/.no_exist/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/dataset_infos.json
ADDED
|
File without changes
|
hf_cache/hub/datasets--abisee--cnn_dailymail/blobs/feadf7d245b4f6818e20b4cf65841d03b5703d47
ADDED
|
@@ -0,0 +1,305 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
annotations_creators:
|
| 3 |
+
- no-annotation
|
| 4 |
+
language_creators:
|
| 5 |
+
- found
|
| 6 |
+
language:
|
| 7 |
+
- en
|
| 8 |
+
license:
|
| 9 |
+
- apache-2.0
|
| 10 |
+
multilinguality:
|
| 11 |
+
- monolingual
|
| 12 |
+
size_categories:
|
| 13 |
+
- 100K<n<1M
|
| 14 |
+
source_datasets:
|
| 15 |
+
- original
|
| 16 |
+
task_categories:
|
| 17 |
+
- summarization
|
| 18 |
+
task_ids:
|
| 19 |
+
- news-articles-summarization
|
| 20 |
+
paperswithcode_id: cnn-daily-mail-1
|
| 21 |
+
pretty_name: CNN / Daily Mail
|
| 22 |
+
dataset_info:
|
| 23 |
+
- config_name: 1.0.0
|
| 24 |
+
features:
|
| 25 |
+
- name: article
|
| 26 |
+
dtype: string
|
| 27 |
+
- name: highlights
|
| 28 |
+
dtype: string
|
| 29 |
+
- name: id
|
| 30 |
+
dtype: string
|
| 31 |
+
splits:
|
| 32 |
+
- name: train
|
| 33 |
+
num_bytes: 1261703785
|
| 34 |
+
num_examples: 287113
|
| 35 |
+
- name: validation
|
| 36 |
+
num_bytes: 57732412
|
| 37 |
+
num_examples: 13368
|
| 38 |
+
- name: test
|
| 39 |
+
num_bytes: 49925732
|
| 40 |
+
num_examples: 11490
|
| 41 |
+
download_size: 836927248
|
| 42 |
+
dataset_size: 1369361929
|
| 43 |
+
- config_name: 2.0.0
|
| 44 |
+
features:
|
| 45 |
+
- name: article
|
| 46 |
+
dtype: string
|
| 47 |
+
- name: highlights
|
| 48 |
+
dtype: string
|
| 49 |
+
- name: id
|
| 50 |
+
dtype: string
|
| 51 |
+
splits:
|
| 52 |
+
- name: train
|
| 53 |
+
num_bytes: 1261703785
|
| 54 |
+
num_examples: 287113
|
| 55 |
+
- name: validation
|
| 56 |
+
num_bytes: 57732412
|
| 57 |
+
num_examples: 13368
|
| 58 |
+
- name: test
|
| 59 |
+
num_bytes: 49925732
|
| 60 |
+
num_examples: 11490
|
| 61 |
+
download_size: 837094602
|
| 62 |
+
dataset_size: 1369361929
|
| 63 |
+
- config_name: 3.0.0
|
| 64 |
+
features:
|
| 65 |
+
- name: article
|
| 66 |
+
dtype: string
|
| 67 |
+
- name: highlights
|
| 68 |
+
dtype: string
|
| 69 |
+
- name: id
|
| 70 |
+
dtype: string
|
| 71 |
+
splits:
|
| 72 |
+
- name: train
|
| 73 |
+
num_bytes: 1261703785
|
| 74 |
+
num_examples: 287113
|
| 75 |
+
- name: validation
|
| 76 |
+
num_bytes: 57732412
|
| 77 |
+
num_examples: 13368
|
| 78 |
+
- name: test
|
| 79 |
+
num_bytes: 49925732
|
| 80 |
+
num_examples: 11490
|
| 81 |
+
download_size: 837094602
|
| 82 |
+
dataset_size: 1369361929
|
| 83 |
+
configs:
|
| 84 |
+
- config_name: 1.0.0
|
| 85 |
+
data_files:
|
| 86 |
+
- split: train
|
| 87 |
+
path: 1.0.0/train-*
|
| 88 |
+
- split: validation
|
| 89 |
+
path: 1.0.0/validation-*
|
| 90 |
+
- split: test
|
| 91 |
+
path: 1.0.0/test-*
|
| 92 |
+
- config_name: 2.0.0
|
| 93 |
+
data_files:
|
| 94 |
+
- split: train
|
| 95 |
+
path: 2.0.0/train-*
|
| 96 |
+
- split: validation
|
| 97 |
+
path: 2.0.0/validation-*
|
| 98 |
+
- split: test
|
| 99 |
+
path: 2.0.0/test-*
|
| 100 |
+
- config_name: 3.0.0
|
| 101 |
+
data_files:
|
| 102 |
+
- split: train
|
| 103 |
+
path: 3.0.0/train-*
|
| 104 |
+
- split: validation
|
| 105 |
+
path: 3.0.0/validation-*
|
| 106 |
+
- split: test
|
| 107 |
+
path: 3.0.0/test-*
|
| 108 |
+
train-eval-index:
|
| 109 |
+
- config: 3.0.0
|
| 110 |
+
task: summarization
|
| 111 |
+
task_id: summarization
|
| 112 |
+
splits:
|
| 113 |
+
eval_split: test
|
| 114 |
+
col_mapping:
|
| 115 |
+
article: text
|
| 116 |
+
highlights: target
|
| 117 |
+
---
|
| 118 |
+
# Dataset Card for CNN Dailymail Dataset
|
| 119 |
+
|
| 120 |
+
## Table of Contents
|
| 121 |
+
- [Dataset Description](#dataset-description)
|
| 122 |
+
- [Dataset Summary](#dataset-summary)
|
| 123 |
+
- [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
|
| 124 |
+
- [Languages](#languages)
|
| 125 |
+
- [Dataset Structure](#dataset-structure)
|
| 126 |
+
- [Data Instances](#data-instances)
|
| 127 |
+
- [Data Fields](#data-fields)
|
| 128 |
+
- [Data Splits](#data-splits)
|
| 129 |
+
- [Dataset Creation](#dataset-creation)
|
| 130 |
+
- [Curation Rationale](#curation-rationale)
|
| 131 |
+
- [Source Data](#source-data)
|
| 132 |
+
- [Annotations](#annotations)
|
| 133 |
+
- [Personal and Sensitive Information](#personal-and-sensitive-information)
|
| 134 |
+
- [Considerations for Using the Data](#considerations-for-using-the-data)
|
| 135 |
+
- [Social Impact of Dataset](#social-impact-of-dataset)
|
| 136 |
+
- [Discussion of Biases](#discussion-of-biases)
|
| 137 |
+
- [Other Known Limitations](#other-known-limitations)
|
| 138 |
+
- [Additional Information](#additional-information)
|
| 139 |
+
- [Dataset Curators](#dataset-curators)
|
| 140 |
+
- [Licensing Information](#licensing-information)
|
| 141 |
+
- [Citation Information](#citation-information)
|
| 142 |
+
- [Contributions](#contributions)
|
| 143 |
+
|
| 144 |
+
## Dataset Description
|
| 145 |
+
|
| 146 |
+
- **Homepage:**
|
| 147 |
+
- **Repository:** [CNN / DailyMail Dataset repository](https://github.com/abisee/cnn-dailymail)
|
| 148 |
+
- **Paper:** [Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond](https://papers.nips.cc/paper/5945-teaching-machines-to-read-and-comprehend.pdf), [Get To The Point: Summarization with Pointer-Generator Networks](https://www.aclweb.org/anthology/K16-1028.pdf)
|
| 149 |
+
- **Leaderboard:** [Papers with Code leaderboard for CNN / Dailymail Dataset](https://paperswithcode.com/sota/document-summarization-on-cnn-daily-mail)
|
| 150 |
+
- **Point of Contact:** [Abigail See](mailto:abisee@stanford.edu)
|
| 151 |
+
|
| 152 |
+
### Dataset Summary
|
| 153 |
+
|
| 154 |
+
The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering.
|
| 155 |
+
|
| 156 |
+
### Supported Tasks and Leaderboards
|
| 157 |
+
|
| 158 |
+
- 'summarization': [Versions 2.0.0 and 3.0.0 of the CNN / DailyMail Dataset](https://www.aclweb.org/anthology/K16-1028.pdf) can be used to train a model for abstractive and extractive summarization ([Version 1.0.0](https://papers.nips.cc/paper/5945-teaching-machines-to-read-and-comprehend.pdf) was developed for machine reading and comprehension and abstractive question answering). The model performance is measured by how high the output summary's [ROUGE](https://huggingface.co/metrics/rouge) score for a given article is when compared to the highlight as written by the original article author. [Zhong et al (2020)](https://www.aclweb.org/anthology/2020.acl-main.552.pdf) report a ROUGE-1 score of 44.41 when testing a model trained for extractive summarization. See the [Papers With Code leaderboard](https://paperswithcode.com/sota/document-summarization-on-cnn-daily-mail) for more models.
|
| 159 |
+
|
| 160 |
+
### Languages
|
| 161 |
+
|
| 162 |
+
The BCP-47 code for English as generally spoken in the United States is en-US and the BCP-47 code for English as generally spoken in the United Kingdom is en-GB. It is unknown if other varieties of English are represented in the data.
|
| 163 |
+
|
| 164 |
+
## Dataset Structure
|
| 165 |
+
|
| 166 |
+
### Data Instances
|
| 167 |
+
|
| 168 |
+
For each instance, there is a string for the article, a string for the highlights, and a string for the id. See the [CNN / Daily Mail dataset viewer](https://huggingface.co/datasets/viewer/?dataset=cnn_dailymail&config=3.0.0) to explore more examples.
|
| 169 |
+
|
| 170 |
+
```
|
| 171 |
+
{'id': '0054d6d30dbcad772e20b22771153a2a9cbeaf62',
|
| 172 |
+
'article': '(CNN) -- An American woman died aboard a cruise ship that docked at Rio de Janeiro on Tuesday, the same ship on which 86 passengers previously fell ill, according to the state-run Brazilian news agency, Agencia Brasil. The American tourist died aboard the MS Veendam, owned by cruise operator Holland America. Federal Police told Agencia Brasil that forensic doctors were investigating her death. The ship's doctors told police that the woman was elderly and suffered from diabetes and hypertension, according the agency. The other passengers came down with diarrhea prior to her death during an earlier part of the trip, the ship's doctors said. The Veendam left New York 36 days ago for a South America tour.'
|
| 173 |
+
'highlights': 'The elderly woman suffered from diabetes and hypertension, ship's doctors say .\nPreviously, 86 passengers had fallen ill on the ship, Agencia Brasil says .'}
|
| 174 |
+
```
|
| 175 |
+
|
| 176 |
+
The average token count for the articles and the highlights are provided below:
|
| 177 |
+
|
| 178 |
+
| Feature | Mean Token Count |
|
| 179 |
+
| ---------- | ---------------- |
|
| 180 |
+
| Article | 781 |
|
| 181 |
+
| Highlights | 56 |
|
| 182 |
+
|
| 183 |
+
### Data Fields
|
| 184 |
+
|
| 185 |
+
- `id`: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from
|
| 186 |
+
- `article`: a string containing the body of the news article
|
| 187 |
+
- `highlights`: a string containing the highlight of the article as written by the article author
|
| 188 |
+
|
| 189 |
+
### Data Splits
|
| 190 |
+
|
| 191 |
+
The CNN/DailyMail dataset has 3 splits: _train_, _validation_, and _test_. Below are the statistics for Version 3.0.0 of the dataset.
|
| 192 |
+
|
| 193 |
+
| Dataset Split | Number of Instances in Split |
|
| 194 |
+
| ------------- | ------------------------------------------- |
|
| 195 |
+
| Train | 287,113 |
|
| 196 |
+
| Validation | 13,368 |
|
| 197 |
+
| Test | 11,490 |
|
| 198 |
+
|
| 199 |
+
## Dataset Creation
|
| 200 |
+
|
| 201 |
+
### Curation Rationale
|
| 202 |
+
|
| 203 |
+
Version 1.0.0 aimed to support supervised neural methodologies for machine reading and question answering with a large amount of real natural language training data and released about 313k unique articles and nearly 1M Cloze style questions to go with the articles. Versions 2.0.0 and 3.0.0 changed the structure of the dataset to support summarization rather than question answering. Version 3.0.0 provided a non-anonymized version of the data, whereas both the previous versions were preprocessed to replace named entities with unique identifier labels.
|
| 204 |
+
|
| 205 |
+
### Source Data
|
| 206 |
+
|
| 207 |
+
#### Initial Data Collection and Normalization
|
| 208 |
+
|
| 209 |
+
The data consists of news articles and highlight sentences. In the question answering setting of the data, the articles are used as the context and entities are hidden one at a time in the highlight sentences, producing Cloze style questions where the goal of the model is to correctly guess which entity in the context has been hidden in the highlight. In the summarization setting, the highlight sentences are concatenated to form a summary of the article. The CNN articles were written between April 2007 and April 2015. The Daily Mail articles were written between June 2010 and April 2015.
|
| 210 |
+
|
| 211 |
+
The code for the original data collection is available at <https://github.com/deepmind/rc-data>. The articles were downloaded using archives of <www.cnn.com> and <www.dailymail.co.uk> on the Wayback Machine. Articles were not included in the Version 1.0.0 collection if they exceeded 2000 tokens. Due to accessibility issues with the Wayback Machine, Kyunghyun Cho has made the datasets available at <https://cs.nyu.edu/~kcho/DMQA/>. An updated version of the code that does not anonymize the data is available at <https://github.com/abisee/cnn-dailymail>.
|
| 212 |
+
|
| 213 |
+
Hermann et al provided their own tokenization script. The script provided by See uses the PTBTokenizer. It also lowercases the text and adds periods to lines missing them.
|
| 214 |
+
|
| 215 |
+
#### Who are the source language producers?
|
| 216 |
+
|
| 217 |
+
The text was written by journalists at CNN and the Daily Mail.
|
| 218 |
+
|
| 219 |
+
### Annotations
|
| 220 |
+
|
| 221 |
+
The dataset does not contain any additional annotations.
|
| 222 |
+
|
| 223 |
+
#### Annotation process
|
| 224 |
+
|
| 225 |
+
[N/A]
|
| 226 |
+
|
| 227 |
+
#### Who are the annotators?
|
| 228 |
+
|
| 229 |
+
[N/A]
|
| 230 |
+
|
| 231 |
+
### Personal and Sensitive Information
|
| 232 |
+
|
| 233 |
+
Version 3.0 is not anonymized, so individuals' names can be found in the dataset. Information about the original author is not included in the dataset.
|
| 234 |
+
|
| 235 |
+
## Considerations for Using the Data
|
| 236 |
+
|
| 237 |
+
### Social Impact of Dataset
|
| 238 |
+
|
| 239 |
+
The purpose of this dataset is to help develop models that can summarize long paragraphs of text in one or two sentences.
|
| 240 |
+
|
| 241 |
+
This task is useful for efficiently presenting information given a large quantity of text. It should be made clear that any summarizations produced by models trained on this dataset are reflective of the language used in the articles, but are in fact automatically generated.
|
| 242 |
+
|
| 243 |
+
### Discussion of Biases
|
| 244 |
+
|
| 245 |
+
[Bordia and Bowman (2019)](https://www.aclweb.org/anthology/N19-3002.pdf) explore measuring gender bias and debiasing techniques in the CNN / Dailymail dataset, the Penn Treebank, and WikiText-2. They find the CNN / Dailymail dataset to have a slightly lower gender bias based on their metric compared to the other datasets, but still show evidence of gender bias when looking at words such as 'fragile'.
|
| 246 |
+
|
| 247 |
+
Because the articles were written by and for people in the US and the UK, they will likely present specifically US and UK perspectives and feature events that are considered relevant to those populations during the time that the articles were published.
|
| 248 |
+
|
| 249 |
+
### Other Known Limitations
|
| 250 |
+
|
| 251 |
+
News articles have been shown to conform to writing conventions in which important information is primarily presented in the first third of the article [(Kryściński et al, 2019)](https://www.aclweb.org/anthology/D19-1051.pdf). [Chen et al (2016)](https://www.aclweb.org/anthology/P16-1223.pdf) conducted a manual study of 100 random instances of the first version of the dataset and found 25% of the samples to be difficult even for humans to answer correctly due to ambiguity and coreference errors.
|
| 252 |
+
|
| 253 |
+
It should also be noted that machine-generated summarizations, even when extractive, may differ in truth values when compared to the original articles.
|
| 254 |
+
|
| 255 |
+
## Additional Information
|
| 256 |
+
|
| 257 |
+
### Dataset Curators
|
| 258 |
+
|
| 259 |
+
The data was originally collected by Karl Moritz Hermann, Tomáš Kočiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom of Google DeepMind. Tomáš Kočiský and Phil Blunsom are also affiliated with the University of Oxford. They released scripts to collect and process the data into the question answering format.
|
| 260 |
+
|
| 261 |
+
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, and Bing Xiang of IMB Watson and Çağlar Gu̇lçehre of Université de Montréal modified Hermann et al's collection scripts to restore the data to a summary format. They also produced both anonymized and non-anonymized versions.
|
| 262 |
+
|
| 263 |
+
The code for the non-anonymized version is made publicly available by Abigail See of Stanford University, Peter J. Liu of Google Brain and Christopher D. Manning of Stanford University at <https://github.com/abisee/cnn-dailymail>. The work at Stanford University was supported by the DARPA DEFT ProgramAFRL contract no. FA8750-13-2-0040.
|
| 264 |
+
|
| 265 |
+
### Licensing Information
|
| 266 |
+
|
| 267 |
+
The CNN / Daily Mail dataset version 1.0.0 is released under the [Apache-2.0 License](http://www.apache.org/licenses/LICENSE-2.0).
|
| 268 |
+
|
| 269 |
+
### Citation Information
|
| 270 |
+
|
| 271 |
+
```
|
| 272 |
+
@inproceedings{see-etal-2017-get,
|
| 273 |
+
title = "Get To The Point: Summarization with Pointer-Generator Networks",
|
| 274 |
+
author = "See, Abigail and
|
| 275 |
+
Liu, Peter J. and
|
| 276 |
+
Manning, Christopher D.",
|
| 277 |
+
booktitle = "Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
|
| 278 |
+
month = jul,
|
| 279 |
+
year = "2017",
|
| 280 |
+
address = "Vancouver, Canada",
|
| 281 |
+
publisher = "Association for Computational Linguistics",
|
| 282 |
+
url = "https://www.aclweb.org/anthology/P17-1099",
|
| 283 |
+
doi = "10.18653/v1/P17-1099",
|
| 284 |
+
pages = "1073--1083",
|
| 285 |
+
abstract = "Neural sequence-to-sequence models have provided a viable new approach for abstractive text summarization (meaning they are not restricted to simply selecting and rearranging passages from the original text). However, these models have two shortcomings: they are liable to reproduce factual details inaccurately, and they tend to repeat themselves. In this work we propose a novel architecture that augments the standard sequence-to-sequence attentional model in two orthogonal ways. First, we use a hybrid pointer-generator network that can copy words from the source text via pointing, which aids accurate reproduction of information, while retaining the ability to produce novel words through the generator. Second, we use coverage to keep track of what has been summarized, which discourages repetition. We apply our model to the CNN / Daily Mail summarization task, outperforming the current abstractive state-of-the-art by at least 2 ROUGE points.",
|
| 286 |
+
}
|
| 287 |
+
```
|
| 288 |
+
|
| 289 |
+
```
|
| 290 |
+
@inproceedings{DBLP:conf/nips/HermannKGEKSB15,
|
| 291 |
+
author={Karl Moritz Hermann and Tomás Kociský and Edward Grefenstette and Lasse Espeholt and Will Kay and Mustafa Suleyman and Phil Blunsom},
|
| 292 |
+
title={Teaching Machines to Read and Comprehend},
|
| 293 |
+
year={2015},
|
| 294 |
+
cdate={1420070400000},
|
| 295 |
+
pages={1693-1701},
|
| 296 |
+
url={http://papers.nips.cc/paper/5945-teaching-machines-to-read-and-comprehend},
|
| 297 |
+
booktitle={NIPS},
|
| 298 |
+
crossref={conf/nips/2015}
|
| 299 |
+
}
|
| 300 |
+
|
| 301 |
+
```
|
| 302 |
+
|
| 303 |
+
### Contributions
|
| 304 |
+
|
| 305 |
+
Thanks to [@thomwolf](https://github.com/thomwolf), [@lewtun](https://github.com/lewtun), [@jplu](https://github.com/jplu), [@jbragg](https://github.com/jbragg), [@patrickvonplaten](https://github.com/patrickvonplaten) and [@mcmillanmajora](https://github.com/mcmillanmajora) for adding this dataset.
|
hf_cache/hub/datasets--abisee--cnn_dailymail/refs/main
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
96df5e686bee6baa90b8bee7c28b81fa3fa6223d
|
hf_cache/hub/datasets--abisee--cnn_dailymail/snapshots/96df5e686bee6baa90b8bee7c28b81fa3fa6223d/README.md
ADDED
|
@@ -0,0 +1,305 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
annotations_creators:
|
| 3 |
+
- no-annotation
|
| 4 |
+
language_creators:
|
| 5 |
+
- found
|
| 6 |
+
language:
|
| 7 |
+
- en
|
| 8 |
+
license:
|
| 9 |
+
- apache-2.0
|
| 10 |
+
multilinguality:
|
| 11 |
+
- monolingual
|
| 12 |
+
size_categories:
|
| 13 |
+
- 100K<n<1M
|
| 14 |
+
source_datasets:
|
| 15 |
+
- original
|
| 16 |
+
task_categories:
|
| 17 |
+
- summarization
|
| 18 |
+
task_ids:
|
| 19 |
+
- news-articles-summarization
|
| 20 |
+
paperswithcode_id: cnn-daily-mail-1
|
| 21 |
+
pretty_name: CNN / Daily Mail
|
| 22 |
+
dataset_info:
|
| 23 |
+
- config_name: 1.0.0
|
| 24 |
+
features:
|
| 25 |
+
- name: article
|
| 26 |
+
dtype: string
|
| 27 |
+
- name: highlights
|
| 28 |
+
dtype: string
|
| 29 |
+
- name: id
|
| 30 |
+
dtype: string
|
| 31 |
+
splits:
|
| 32 |
+
- name: train
|
| 33 |
+
num_bytes: 1261703785
|
| 34 |
+
num_examples: 287113
|
| 35 |
+
- name: validation
|
| 36 |
+
num_bytes: 57732412
|
| 37 |
+
num_examples: 13368
|
| 38 |
+
- name: test
|
| 39 |
+
num_bytes: 49925732
|
| 40 |
+
num_examples: 11490
|
| 41 |
+
download_size: 836927248
|
| 42 |
+
dataset_size: 1369361929
|
| 43 |
+
- config_name: 2.0.0
|
| 44 |
+
features:
|
| 45 |
+
- name: article
|
| 46 |
+
dtype: string
|
| 47 |
+
- name: highlights
|
| 48 |
+
dtype: string
|
| 49 |
+
- name: id
|
| 50 |
+
dtype: string
|
| 51 |
+
splits:
|
| 52 |
+
- name: train
|
| 53 |
+
num_bytes: 1261703785
|
| 54 |
+
num_examples: 287113
|
| 55 |
+
- name: validation
|
| 56 |
+
num_bytes: 57732412
|
| 57 |
+
num_examples: 13368
|
| 58 |
+
- name: test
|
| 59 |
+
num_bytes: 49925732
|
| 60 |
+
num_examples: 11490
|
| 61 |
+
download_size: 837094602
|
| 62 |
+
dataset_size: 1369361929
|
| 63 |
+
- config_name: 3.0.0
|
| 64 |
+
features:
|
| 65 |
+
- name: article
|
| 66 |
+
dtype: string
|
| 67 |
+
- name: highlights
|
| 68 |
+
dtype: string
|
| 69 |
+
- name: id
|
| 70 |
+
dtype: string
|
| 71 |
+
splits:
|
| 72 |
+
- name: train
|
| 73 |
+
num_bytes: 1261703785
|
| 74 |
+
num_examples: 287113
|
| 75 |
+
- name: validation
|
| 76 |
+
num_bytes: 57732412
|
| 77 |
+
num_examples: 13368
|
| 78 |
+
- name: test
|
| 79 |
+
num_bytes: 49925732
|
| 80 |
+
num_examples: 11490
|
| 81 |
+
download_size: 837094602
|
| 82 |
+
dataset_size: 1369361929
|
| 83 |
+
configs:
|
| 84 |
+
- config_name: 1.0.0
|
| 85 |
+
data_files:
|
| 86 |
+
- split: train
|
| 87 |
+
path: 1.0.0/train-*
|
| 88 |
+
- split: validation
|
| 89 |
+
path: 1.0.0/validation-*
|
| 90 |
+
- split: test
|
| 91 |
+
path: 1.0.0/test-*
|
| 92 |
+
- config_name: 2.0.0
|
| 93 |
+
data_files:
|
| 94 |
+
- split: train
|
| 95 |
+
path: 2.0.0/train-*
|
| 96 |
+
- split: validation
|
| 97 |
+
path: 2.0.0/validation-*
|
| 98 |
+
- split: test
|
| 99 |
+
path: 2.0.0/test-*
|
| 100 |
+
- config_name: 3.0.0
|
| 101 |
+
data_files:
|
| 102 |
+
- split: train
|
| 103 |
+
path: 3.0.0/train-*
|
| 104 |
+
- split: validation
|
| 105 |
+
path: 3.0.0/validation-*
|
| 106 |
+
- split: test
|
| 107 |
+
path: 3.0.0/test-*
|
| 108 |
+
train-eval-index:
|
| 109 |
+
- config: 3.0.0
|
| 110 |
+
task: summarization
|
| 111 |
+
task_id: summarization
|
| 112 |
+
splits:
|
| 113 |
+
eval_split: test
|
| 114 |
+
col_mapping:
|
| 115 |
+
article: text
|
| 116 |
+
highlights: target
|
| 117 |
+
---
|
| 118 |
+
# Dataset Card for CNN Dailymail Dataset
|
| 119 |
+
|
| 120 |
+
## Table of Contents
|
| 121 |
+
- [Dataset Description](#dataset-description)
|
| 122 |
+
- [Dataset Summary](#dataset-summary)
|
| 123 |
+
- [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
|
| 124 |
+
- [Languages](#languages)
|
| 125 |
+
- [Dataset Structure](#dataset-structure)
|
| 126 |
+
- [Data Instances](#data-instances)
|
| 127 |
+
- [Data Fields](#data-fields)
|
| 128 |
+
- [Data Splits](#data-splits)
|
| 129 |
+
- [Dataset Creation](#dataset-creation)
|
| 130 |
+
- [Curation Rationale](#curation-rationale)
|
| 131 |
+
- [Source Data](#source-data)
|
| 132 |
+
- [Annotations](#annotations)
|
| 133 |
+
- [Personal and Sensitive Information](#personal-and-sensitive-information)
|
| 134 |
+
- [Considerations for Using the Data](#considerations-for-using-the-data)
|
| 135 |
+
- [Social Impact of Dataset](#social-impact-of-dataset)
|
| 136 |
+
- [Discussion of Biases](#discussion-of-biases)
|
| 137 |
+
- [Other Known Limitations](#other-known-limitations)
|
| 138 |
+
- [Additional Information](#additional-information)
|
| 139 |
+
- [Dataset Curators](#dataset-curators)
|
| 140 |
+
- [Licensing Information](#licensing-information)
|
| 141 |
+
- [Citation Information](#citation-information)
|
| 142 |
+
- [Contributions](#contributions)
|
| 143 |
+
|
| 144 |
+
## Dataset Description
|
| 145 |
+
|
| 146 |
+
- **Homepage:**
|
| 147 |
+
- **Repository:** [CNN / DailyMail Dataset repository](https://github.com/abisee/cnn-dailymail)
|
| 148 |
+
- **Paper:** [Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond](https://papers.nips.cc/paper/5945-teaching-machines-to-read-and-comprehend.pdf), [Get To The Point: Summarization with Pointer-Generator Networks](https://www.aclweb.org/anthology/K16-1028.pdf)
|
| 149 |
+
- **Leaderboard:** [Papers with Code leaderboard for CNN / Dailymail Dataset](https://paperswithcode.com/sota/document-summarization-on-cnn-daily-mail)
|
| 150 |
+
- **Point of Contact:** [Abigail See](mailto:abisee@stanford.edu)
|
| 151 |
+
|
| 152 |
+
### Dataset Summary
|
| 153 |
+
|
| 154 |
+
The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering.
|
| 155 |
+
|
| 156 |
+
### Supported Tasks and Leaderboards
|
| 157 |
+
|
| 158 |
+
- 'summarization': [Versions 2.0.0 and 3.0.0 of the CNN / DailyMail Dataset](https://www.aclweb.org/anthology/K16-1028.pdf) can be used to train a model for abstractive and extractive summarization ([Version 1.0.0](https://papers.nips.cc/paper/5945-teaching-machines-to-read-and-comprehend.pdf) was developed for machine reading and comprehension and abstractive question answering). The model performance is measured by how high the output summary's [ROUGE](https://huggingface.co/metrics/rouge) score for a given article is when compared to the highlight as written by the original article author. [Zhong et al (2020)](https://www.aclweb.org/anthology/2020.acl-main.552.pdf) report a ROUGE-1 score of 44.41 when testing a model trained for extractive summarization. See the [Papers With Code leaderboard](https://paperswithcode.com/sota/document-summarization-on-cnn-daily-mail) for more models.
|
| 159 |
+
|
| 160 |
+
### Languages
|
| 161 |
+
|
| 162 |
+
The BCP-47 code for English as generally spoken in the United States is en-US and the BCP-47 code for English as generally spoken in the United Kingdom is en-GB. It is unknown if other varieties of English are represented in the data.
|
| 163 |
+
|
| 164 |
+
## Dataset Structure
|
| 165 |
+
|
| 166 |
+
### Data Instances
|
| 167 |
+
|
| 168 |
+
For each instance, there is a string for the article, a string for the highlights, and a string for the id. See the [CNN / Daily Mail dataset viewer](https://huggingface.co/datasets/viewer/?dataset=cnn_dailymail&config=3.0.0) to explore more examples.
|
| 169 |
+
|
| 170 |
+
```
|
| 171 |
+
{'id': '0054d6d30dbcad772e20b22771153a2a9cbeaf62',
|
| 172 |
+
'article': '(CNN) -- An American woman died aboard a cruise ship that docked at Rio de Janeiro on Tuesday, the same ship on which 86 passengers previously fell ill, according to the state-run Brazilian news agency, Agencia Brasil. The American tourist died aboard the MS Veendam, owned by cruise operator Holland America. Federal Police told Agencia Brasil that forensic doctors were investigating her death. The ship's doctors told police that the woman was elderly and suffered from diabetes and hypertension, according the agency. The other passengers came down with diarrhea prior to her death during an earlier part of the trip, the ship's doctors said. The Veendam left New York 36 days ago for a South America tour.'
|
| 173 |
+
'highlights': 'The elderly woman suffered from diabetes and hypertension, ship's doctors say .\nPreviously, 86 passengers had fallen ill on the ship, Agencia Brasil says .'}
|
| 174 |
+
```
|
| 175 |
+
|
| 176 |
+
The average token count for the articles and the highlights are provided below:
|
| 177 |
+
|
| 178 |
+
| Feature | Mean Token Count |
|
| 179 |
+
| ---------- | ---------------- |
|
| 180 |
+
| Article | 781 |
|
| 181 |
+
| Highlights | 56 |
|
| 182 |
+
|
| 183 |
+
### Data Fields
|
| 184 |
+
|
| 185 |
+
- `id`: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from
|
| 186 |
+
- `article`: a string containing the body of the news article
|
| 187 |
+
- `highlights`: a string containing the highlight of the article as written by the article author
|
| 188 |
+
|
| 189 |
+
### Data Splits
|
| 190 |
+
|
| 191 |
+
The CNN/DailyMail dataset has 3 splits: _train_, _validation_, and _test_. Below are the statistics for Version 3.0.0 of the dataset.
|
| 192 |
+
|
| 193 |
+
| Dataset Split | Number of Instances in Split |
|
| 194 |
+
| ------------- | ------------------------------------------- |
|
| 195 |
+
| Train | 287,113 |
|
| 196 |
+
| Validation | 13,368 |
|
| 197 |
+
| Test | 11,490 |
|
| 198 |
+
|
| 199 |
+
## Dataset Creation
|
| 200 |
+
|
| 201 |
+
### Curation Rationale
|
| 202 |
+
|
| 203 |
+
Version 1.0.0 aimed to support supervised neural methodologies for machine reading and question answering with a large amount of real natural language training data and released about 313k unique articles and nearly 1M Cloze style questions to go with the articles. Versions 2.0.0 and 3.0.0 changed the structure of the dataset to support summarization rather than question answering. Version 3.0.0 provided a non-anonymized version of the data, whereas both the previous versions were preprocessed to replace named entities with unique identifier labels.
|
| 204 |
+
|
| 205 |
+
### Source Data
|
| 206 |
+
|
| 207 |
+
#### Initial Data Collection and Normalization
|
| 208 |
+
|
| 209 |
+
The data consists of news articles and highlight sentences. In the question answering setting of the data, the articles are used as the context and entities are hidden one at a time in the highlight sentences, producing Cloze style questions where the goal of the model is to correctly guess which entity in the context has been hidden in the highlight. In the summarization setting, the highlight sentences are concatenated to form a summary of the article. The CNN articles were written between April 2007 and April 2015. The Daily Mail articles were written between June 2010 and April 2015.
|
| 210 |
+
|
| 211 |
+
The code for the original data collection is available at <https://github.com/deepmind/rc-data>. The articles were downloaded using archives of <www.cnn.com> and <www.dailymail.co.uk> on the Wayback Machine. Articles were not included in the Version 1.0.0 collection if they exceeded 2000 tokens. Due to accessibility issues with the Wayback Machine, Kyunghyun Cho has made the datasets available at <https://cs.nyu.edu/~kcho/DMQA/>. An updated version of the code that does not anonymize the data is available at <https://github.com/abisee/cnn-dailymail>.
|
| 212 |
+
|
| 213 |
+
Hermann et al provided their own tokenization script. The script provided by See uses the PTBTokenizer. It also lowercases the text and adds periods to lines missing them.
|
| 214 |
+
|
| 215 |
+
#### Who are the source language producers?
|
| 216 |
+
|
| 217 |
+
The text was written by journalists at CNN and the Daily Mail.
|
| 218 |
+
|
| 219 |
+
### Annotations
|
| 220 |
+
|
| 221 |
+
The dataset does not contain any additional annotations.
|
| 222 |
+
|
| 223 |
+
#### Annotation process
|
| 224 |
+
|
| 225 |
+
[N/A]
|
| 226 |
+
|
| 227 |
+
#### Who are the annotators?
|
| 228 |
+
|
| 229 |
+
[N/A]
|
| 230 |
+
|
| 231 |
+
### Personal and Sensitive Information
|
| 232 |
+
|
| 233 |
+
Version 3.0 is not anonymized, so individuals' names can be found in the dataset. Information about the original author is not included in the dataset.
|
| 234 |
+
|
| 235 |
+
## Considerations for Using the Data
|
| 236 |
+
|
| 237 |
+
### Social Impact of Dataset
|
| 238 |
+
|
| 239 |
+
The purpose of this dataset is to help develop models that can summarize long paragraphs of text in one or two sentences.
|
| 240 |
+
|
| 241 |
+
This task is useful for efficiently presenting information given a large quantity of text. It should be made clear that any summarizations produced by models trained on this dataset are reflective of the language used in the articles, but are in fact automatically generated.
|
| 242 |
+
|
| 243 |
+
### Discussion of Biases
|
| 244 |
+
|
| 245 |
+
[Bordia and Bowman (2019)](https://www.aclweb.org/anthology/N19-3002.pdf) explore measuring gender bias and debiasing techniques in the CNN / Dailymail dataset, the Penn Treebank, and WikiText-2. They find the CNN / Dailymail dataset to have a slightly lower gender bias based on their metric compared to the other datasets, but still show evidence of gender bias when looking at words such as 'fragile'.
|
| 246 |
+
|
| 247 |
+
Because the articles were written by and for people in the US and the UK, they will likely present specifically US and UK perspectives and feature events that are considered relevant to those populations during the time that the articles were published.
|
| 248 |
+
|
| 249 |
+
### Other Known Limitations
|
| 250 |
+
|
| 251 |
+
News articles have been shown to conform to writing conventions in which important information is primarily presented in the first third of the article [(Kryściński et al, 2019)](https://www.aclweb.org/anthology/D19-1051.pdf). [Chen et al (2016)](https://www.aclweb.org/anthology/P16-1223.pdf) conducted a manual study of 100 random instances of the first version of the dataset and found 25% of the samples to be difficult even for humans to answer correctly due to ambiguity and coreference errors.
|
| 252 |
+
|
| 253 |
+
It should also be noted that machine-generated summarizations, even when extractive, may differ in truth values when compared to the original articles.
|
| 254 |
+
|
| 255 |
+
## Additional Information
|
| 256 |
+
|
| 257 |
+
### Dataset Curators
|
| 258 |
+
|
| 259 |
+
The data was originally collected by Karl Moritz Hermann, Tomáš Kočiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom of Google DeepMind. Tomáš Kočiský and Phil Blunsom are also affiliated with the University of Oxford. They released scripts to collect and process the data into the question answering format.
|
| 260 |
+
|
| 261 |
+
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, and Bing Xiang of IMB Watson and Çağlar Gu̇lçehre of Université de Montréal modified Hermann et al's collection scripts to restore the data to a summary format. They also produced both anonymized and non-anonymized versions.
|
| 262 |
+
|
| 263 |
+
The code for the non-anonymized version is made publicly available by Abigail See of Stanford University, Peter J. Liu of Google Brain and Christopher D. Manning of Stanford University at <https://github.com/abisee/cnn-dailymail>. The work at Stanford University was supported by the DARPA DEFT ProgramAFRL contract no. FA8750-13-2-0040.
|
| 264 |
+
|
| 265 |
+
### Licensing Information
|
| 266 |
+
|
| 267 |
+
The CNN / Daily Mail dataset version 1.0.0 is released under the [Apache-2.0 License](http://www.apache.org/licenses/LICENSE-2.0).
|
| 268 |
+
|
| 269 |
+
### Citation Information
|
| 270 |
+
|
| 271 |
+
```
|
| 272 |
+
@inproceedings{see-etal-2017-get,
|
| 273 |
+
title = "Get To The Point: Summarization with Pointer-Generator Networks",
|
| 274 |
+
author = "See, Abigail and
|
| 275 |
+
Liu, Peter J. and
|
| 276 |
+
Manning, Christopher D.",
|
| 277 |
+
booktitle = "Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
|
| 278 |
+
month = jul,
|
| 279 |
+
year = "2017",
|
| 280 |
+
address = "Vancouver, Canada",
|
| 281 |
+
publisher = "Association for Computational Linguistics",
|
| 282 |
+
url = "https://www.aclweb.org/anthology/P17-1099",
|
| 283 |
+
doi = "10.18653/v1/P17-1099",
|
| 284 |
+
pages = "1073--1083",
|
| 285 |
+
abstract = "Neural sequence-to-sequence models have provided a viable new approach for abstractive text summarization (meaning they are not restricted to simply selecting and rearranging passages from the original text). However, these models have two shortcomings: they are liable to reproduce factual details inaccurately, and they tend to repeat themselves. In this work we propose a novel architecture that augments the standard sequence-to-sequence attentional model in two orthogonal ways. First, we use a hybrid pointer-generator network that can copy words from the source text via pointing, which aids accurate reproduction of information, while retaining the ability to produce novel words through the generator. Second, we use coverage to keep track of what has been summarized, which discourages repetition. We apply our model to the CNN / Daily Mail summarization task, outperforming the current abstractive state-of-the-art by at least 2 ROUGE points.",
|
| 286 |
+
}
|
| 287 |
+
```
|
| 288 |
+
|
| 289 |
+
```
|
| 290 |
+
@inproceedings{DBLP:conf/nips/HermannKGEKSB15,
|
| 291 |
+
author={Karl Moritz Hermann and Tomás Kociský and Edward Grefenstette and Lasse Espeholt and Will Kay and Mustafa Suleyman and Phil Blunsom},
|
| 292 |
+
title={Teaching Machines to Read and Comprehend},
|
| 293 |
+
year={2015},
|
| 294 |
+
cdate={1420070400000},
|
| 295 |
+
pages={1693-1701},
|
| 296 |
+
url={http://papers.nips.cc/paper/5945-teaching-machines-to-read-and-comprehend},
|
| 297 |
+
booktitle={NIPS},
|
| 298 |
+
crossref={conf/nips/2015}
|
| 299 |
+
}
|
| 300 |
+
|
| 301 |
+
```
|
| 302 |
+
|
| 303 |
+
### Contributions
|
| 304 |
+
|
| 305 |
+
Thanks to [@thomwolf](https://github.com/thomwolf), [@lewtun](https://github.com/lewtun), [@jplu](https://github.com/jplu), [@jbragg](https://github.com/jbragg), [@patrickvonplaten](https://github.com/patrickvonplaten) and [@mcmillanmajora](https://github.com/mcmillanmajora) for adding this dataset.
|
hf_cache/hub/datasets--allenai--openbookqa/.no_exist/388097ea7776314e93a529163e0fea805b8a6454/.huggingface.yaml
ADDED
|
File without changes
|
hf_cache/hub/datasets--allenai--openbookqa/.no_exist/388097ea7776314e93a529163e0fea805b8a6454/dataset_infos.json
ADDED
|
File without changes
|
hf_cache/hub/datasets--allenai--openbookqa/.no_exist/388097ea7776314e93a529163e0fea805b8a6454/openbookqa.py
ADDED
|
File without changes
|
hf_cache/hub/datasets--allenai--openbookqa/blobs/08128898cc0433b97f7a9c9ff09c5054c8587e3e
ADDED
|
@@ -0,0 +1,301 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
annotations_creators:
|
| 3 |
+
- crowdsourced
|
| 4 |
+
- expert-generated
|
| 5 |
+
language_creators:
|
| 6 |
+
- expert-generated
|
| 7 |
+
language:
|
| 8 |
+
- en
|
| 9 |
+
license:
|
| 10 |
+
- unknown
|
| 11 |
+
multilinguality:
|
| 12 |
+
- monolingual
|
| 13 |
+
size_categories:
|
| 14 |
+
- 1K<n<10K
|
| 15 |
+
source_datasets:
|
| 16 |
+
- original
|
| 17 |
+
task_categories:
|
| 18 |
+
- question-answering
|
| 19 |
+
task_ids:
|
| 20 |
+
- open-domain-qa
|
| 21 |
+
paperswithcode_id: openbookqa
|
| 22 |
+
pretty_name: OpenBookQA
|
| 23 |
+
dataset_info:
|
| 24 |
+
- config_name: additional
|
| 25 |
+
features:
|
| 26 |
+
- name: id
|
| 27 |
+
dtype: string
|
| 28 |
+
- name: question_stem
|
| 29 |
+
dtype: string
|
| 30 |
+
- name: choices
|
| 31 |
+
sequence:
|
| 32 |
+
- name: text
|
| 33 |
+
dtype: string
|
| 34 |
+
- name: label
|
| 35 |
+
dtype: string
|
| 36 |
+
- name: answerKey
|
| 37 |
+
dtype: string
|
| 38 |
+
- name: fact1
|
| 39 |
+
dtype: string
|
| 40 |
+
- name: humanScore
|
| 41 |
+
dtype: float32
|
| 42 |
+
- name: clarity
|
| 43 |
+
dtype: float32
|
| 44 |
+
- name: turkIdAnonymized
|
| 45 |
+
dtype: string
|
| 46 |
+
splits:
|
| 47 |
+
- name: train
|
| 48 |
+
num_bytes: 1288577
|
| 49 |
+
num_examples: 4957
|
| 50 |
+
- name: validation
|
| 51 |
+
num_bytes: 135916
|
| 52 |
+
num_examples: 500
|
| 53 |
+
- name: test
|
| 54 |
+
num_bytes: 130701
|
| 55 |
+
num_examples: 500
|
| 56 |
+
download_size: 783789
|
| 57 |
+
dataset_size: 1555194
|
| 58 |
+
- config_name: main
|
| 59 |
+
features:
|
| 60 |
+
- name: id
|
| 61 |
+
dtype: string
|
| 62 |
+
- name: question_stem
|
| 63 |
+
dtype: string
|
| 64 |
+
- name: choices
|
| 65 |
+
sequence:
|
| 66 |
+
- name: text
|
| 67 |
+
dtype: string
|
| 68 |
+
- name: label
|
| 69 |
+
dtype: string
|
| 70 |
+
- name: answerKey
|
| 71 |
+
dtype: string
|
| 72 |
+
splits:
|
| 73 |
+
- name: train
|
| 74 |
+
num_bytes: 895386
|
| 75 |
+
num_examples: 4957
|
| 76 |
+
- name: validation
|
| 77 |
+
num_bytes: 95428
|
| 78 |
+
num_examples: 500
|
| 79 |
+
- name: test
|
| 80 |
+
num_bytes: 91759
|
| 81 |
+
num_examples: 500
|
| 82 |
+
download_size: 609613
|
| 83 |
+
dataset_size: 1082573
|
| 84 |
+
configs:
|
| 85 |
+
- config_name: additional
|
| 86 |
+
data_files:
|
| 87 |
+
- split: train
|
| 88 |
+
path: additional/train-*
|
| 89 |
+
- split: validation
|
| 90 |
+
path: additional/validation-*
|
| 91 |
+
- split: test
|
| 92 |
+
path: additional/test-*
|
| 93 |
+
- config_name: main
|
| 94 |
+
data_files:
|
| 95 |
+
- split: train
|
| 96 |
+
path: main/train-*
|
| 97 |
+
- split: validation
|
| 98 |
+
path: main/validation-*
|
| 99 |
+
- split: test
|
| 100 |
+
path: main/test-*
|
| 101 |
+
default: true
|
| 102 |
+
---
|
| 103 |
+
|
| 104 |
+
# Dataset Card for OpenBookQA
|
| 105 |
+
|
| 106 |
+
## Table of Contents
|
| 107 |
+
- [Dataset Description](#dataset-description)
|
| 108 |
+
- [Dataset Summary](#dataset-summary)
|
| 109 |
+
- [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
|
| 110 |
+
- [Languages](#languages)
|
| 111 |
+
- [Dataset Structure](#dataset-structure)
|
| 112 |
+
- [Data Instances](#data-instances)
|
| 113 |
+
- [Data Fields](#data-fields)
|
| 114 |
+
- [Data Splits](#data-splits)
|
| 115 |
+
- [Dataset Creation](#dataset-creation)
|
| 116 |
+
- [Curation Rationale](#curation-rationale)
|
| 117 |
+
- [Source Data](#source-data)
|
| 118 |
+
- [Annotations](#annotations)
|
| 119 |
+
- [Personal and Sensitive Information](#personal-and-sensitive-information)
|
| 120 |
+
- [Considerations for Using the Data](#considerations-for-using-the-data)
|
| 121 |
+
- [Social Impact of Dataset](#social-impact-of-dataset)
|
| 122 |
+
- [Discussion of Biases](#discussion-of-biases)
|
| 123 |
+
- [Other Known Limitations](#other-known-limitations)
|
| 124 |
+
- [Additional Information](#additional-information)
|
| 125 |
+
- [Dataset Curators](#dataset-curators)
|
| 126 |
+
- [Licensing Information](#licensing-information)
|
| 127 |
+
- [Citation Information](#citation-information)
|
| 128 |
+
- [Contributions](#contributions)
|
| 129 |
+
|
| 130 |
+
## Dataset Description
|
| 131 |
+
|
| 132 |
+
- **Homepage:** [https://allenai.org/data/open-book-qa](https://allenai.org/data/open-book-qa)
|
| 133 |
+
- **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 134 |
+
- **Paper:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 135 |
+
- **Point of Contact:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 136 |
+
- **Size of downloaded dataset files:** 2.89 MB
|
| 137 |
+
- **Size of the generated dataset:** 2.88 MB
|
| 138 |
+
- **Total amount of disk used:** 5.78 MB
|
| 139 |
+
|
| 140 |
+
### Dataset Summary
|
| 141 |
+
|
| 142 |
+
OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic
|
| 143 |
+
(with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In
|
| 144 |
+
particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge,
|
| 145 |
+
and rich text comprehension.
|
| 146 |
+
OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of
|
| 147 |
+
a subject.
|
| 148 |
+
|
| 149 |
+
### Supported Tasks and Leaderboards
|
| 150 |
+
|
| 151 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 152 |
+
|
| 153 |
+
### Languages
|
| 154 |
+
|
| 155 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 156 |
+
|
| 157 |
+
## Dataset Structure
|
| 158 |
+
|
| 159 |
+
### Data Instances
|
| 160 |
+
|
| 161 |
+
#### main
|
| 162 |
+
|
| 163 |
+
- **Size of downloaded dataset files:** 1.45 MB
|
| 164 |
+
- **Size of the generated dataset:** 1.45 MB
|
| 165 |
+
- **Total amount of disk used:** 2.88 MB
|
| 166 |
+
|
| 167 |
+
An example of 'train' looks as follows:
|
| 168 |
+
```
|
| 169 |
+
{'id': '7-980',
|
| 170 |
+
'question_stem': 'The sun is responsible for',
|
| 171 |
+
'choices': {'text': ['puppies learning new tricks',
|
| 172 |
+
'children growing up and getting old',
|
| 173 |
+
'flowers wilting in a vase',
|
| 174 |
+
'plants sprouting, blooming and wilting'],
|
| 175 |
+
'label': ['A', 'B', 'C', 'D']},
|
| 176 |
+
'answerKey': 'D'}
|
| 177 |
+
```
|
| 178 |
+
|
| 179 |
+
#### additional
|
| 180 |
+
|
| 181 |
+
- **Size of downloaded dataset files:** 1.45 MB
|
| 182 |
+
- **Size of the generated dataset:** 1.45 MB
|
| 183 |
+
- **Total amount of disk used:** 2.88 MB
|
| 184 |
+
|
| 185 |
+
An example of 'train' looks as follows:
|
| 186 |
+
```
|
| 187 |
+
{'id': '7-980',
|
| 188 |
+
'question_stem': 'The sun is responsible for',
|
| 189 |
+
'choices': {'text': ['puppies learning new tricks',
|
| 190 |
+
'children growing up and getting old',
|
| 191 |
+
'flowers wilting in a vase',
|
| 192 |
+
'plants sprouting, blooming and wilting'],
|
| 193 |
+
'label': ['A', 'B', 'C', 'D']},
|
| 194 |
+
'answerKey': 'D',
|
| 195 |
+
'fact1': 'the sun is the source of energy for physical cycles on Earth',
|
| 196 |
+
'humanScore': 1.0,
|
| 197 |
+
'clarity': 2.0,
|
| 198 |
+
'turkIdAnonymized': 'b356d338b7'}
|
| 199 |
+
```
|
| 200 |
+
|
| 201 |
+
### Data Fields
|
| 202 |
+
|
| 203 |
+
The data fields are the same among all splits.
|
| 204 |
+
|
| 205 |
+
#### main
|
| 206 |
+
- `id`: a `string` feature.
|
| 207 |
+
- `question_stem`: a `string` feature.
|
| 208 |
+
- `choices`: a dictionary feature containing:
|
| 209 |
+
- `text`: a `string` feature.
|
| 210 |
+
- `label`: a `string` feature.
|
| 211 |
+
- `answerKey`: a `string` feature.
|
| 212 |
+
|
| 213 |
+
#### additional
|
| 214 |
+
- `id`: a `string` feature.
|
| 215 |
+
- `question_stem`: a `string` feature.
|
| 216 |
+
- `choices`: a dictionary feature containing:
|
| 217 |
+
- `text`: a `string` feature.
|
| 218 |
+
- `label`: a `string` feature.
|
| 219 |
+
- `answerKey`: a `string` feature.
|
| 220 |
+
- `fact1` (`str`): oOriginating common knowledge core fact associated to the question.
|
| 221 |
+
- `humanScore` (`float`): Human accuracy score.
|
| 222 |
+
- `clarity` (`float`): Clarity score.
|
| 223 |
+
- `turkIdAnonymized` (`str`): Anonymized crowd-worker ID.
|
| 224 |
+
|
| 225 |
+
### Data Splits
|
| 226 |
+
|
| 227 |
+
| name | train | validation | test |
|
| 228 |
+
|------------|------:|-----------:|-----:|
|
| 229 |
+
| main | 4957 | 500 | 500 |
|
| 230 |
+
| additional | 4957 | 500 | 500 |
|
| 231 |
+
|
| 232 |
+
## Dataset Creation
|
| 233 |
+
|
| 234 |
+
### Curation Rationale
|
| 235 |
+
|
| 236 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 237 |
+
|
| 238 |
+
### Source Data
|
| 239 |
+
|
| 240 |
+
#### Initial Data Collection and Normalization
|
| 241 |
+
|
| 242 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 243 |
+
|
| 244 |
+
#### Who are the source language producers?
|
| 245 |
+
|
| 246 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 247 |
+
|
| 248 |
+
### Annotations
|
| 249 |
+
|
| 250 |
+
#### Annotation process
|
| 251 |
+
|
| 252 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 253 |
+
|
| 254 |
+
#### Who are the annotators?
|
| 255 |
+
|
| 256 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 257 |
+
|
| 258 |
+
### Personal and Sensitive Information
|
| 259 |
+
|
| 260 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 261 |
+
|
| 262 |
+
## Considerations for Using the Data
|
| 263 |
+
|
| 264 |
+
### Social Impact of Dataset
|
| 265 |
+
|
| 266 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 267 |
+
|
| 268 |
+
### Discussion of Biases
|
| 269 |
+
|
| 270 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 271 |
+
|
| 272 |
+
### Other Known Limitations
|
| 273 |
+
|
| 274 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 275 |
+
|
| 276 |
+
## Additional Information
|
| 277 |
+
|
| 278 |
+
### Dataset Curators
|
| 279 |
+
|
| 280 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 281 |
+
|
| 282 |
+
### Licensing Information
|
| 283 |
+
|
| 284 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 285 |
+
|
| 286 |
+
### Citation Information
|
| 287 |
+
|
| 288 |
+
```
|
| 289 |
+
@inproceedings{OpenBookQA2018,
|
| 290 |
+
title={Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering},
|
| 291 |
+
author={Todor Mihaylov and Peter Clark and Tushar Khot and Ashish Sabharwal},
|
| 292 |
+
booktitle={EMNLP},
|
| 293 |
+
year={2018}
|
| 294 |
+
}
|
| 295 |
+
|
| 296 |
+
```
|
| 297 |
+
|
| 298 |
+
|
| 299 |
+
### Contributions
|
| 300 |
+
|
| 301 |
+
Thanks to [@thomwolf](https://github.com/thomwolf), [@patrickvonplaten](https://github.com/patrickvonplaten), [@lewtun](https://github.com/lewtun) for adding this dataset.
|
hf_cache/hub/datasets--allenai--openbookqa/refs/main
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
388097ea7776314e93a529163e0fea805b8a6454
|
hf_cache/hub/datasets--allenai--openbookqa/snapshots/388097ea7776314e93a529163e0fea805b8a6454/README.md
ADDED
|
@@ -0,0 +1,301 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
annotations_creators:
|
| 3 |
+
- crowdsourced
|
| 4 |
+
- expert-generated
|
| 5 |
+
language_creators:
|
| 6 |
+
- expert-generated
|
| 7 |
+
language:
|
| 8 |
+
- en
|
| 9 |
+
license:
|
| 10 |
+
- unknown
|
| 11 |
+
multilinguality:
|
| 12 |
+
- monolingual
|
| 13 |
+
size_categories:
|
| 14 |
+
- 1K<n<10K
|
| 15 |
+
source_datasets:
|
| 16 |
+
- original
|
| 17 |
+
task_categories:
|
| 18 |
+
- question-answering
|
| 19 |
+
task_ids:
|
| 20 |
+
- open-domain-qa
|
| 21 |
+
paperswithcode_id: openbookqa
|
| 22 |
+
pretty_name: OpenBookQA
|
| 23 |
+
dataset_info:
|
| 24 |
+
- config_name: additional
|
| 25 |
+
features:
|
| 26 |
+
- name: id
|
| 27 |
+
dtype: string
|
| 28 |
+
- name: question_stem
|
| 29 |
+
dtype: string
|
| 30 |
+
- name: choices
|
| 31 |
+
sequence:
|
| 32 |
+
- name: text
|
| 33 |
+
dtype: string
|
| 34 |
+
- name: label
|
| 35 |
+
dtype: string
|
| 36 |
+
- name: answerKey
|
| 37 |
+
dtype: string
|
| 38 |
+
- name: fact1
|
| 39 |
+
dtype: string
|
| 40 |
+
- name: humanScore
|
| 41 |
+
dtype: float32
|
| 42 |
+
- name: clarity
|
| 43 |
+
dtype: float32
|
| 44 |
+
- name: turkIdAnonymized
|
| 45 |
+
dtype: string
|
| 46 |
+
splits:
|
| 47 |
+
- name: train
|
| 48 |
+
num_bytes: 1288577
|
| 49 |
+
num_examples: 4957
|
| 50 |
+
- name: validation
|
| 51 |
+
num_bytes: 135916
|
| 52 |
+
num_examples: 500
|
| 53 |
+
- name: test
|
| 54 |
+
num_bytes: 130701
|
| 55 |
+
num_examples: 500
|
| 56 |
+
download_size: 783789
|
| 57 |
+
dataset_size: 1555194
|
| 58 |
+
- config_name: main
|
| 59 |
+
features:
|
| 60 |
+
- name: id
|
| 61 |
+
dtype: string
|
| 62 |
+
- name: question_stem
|
| 63 |
+
dtype: string
|
| 64 |
+
- name: choices
|
| 65 |
+
sequence:
|
| 66 |
+
- name: text
|
| 67 |
+
dtype: string
|
| 68 |
+
- name: label
|
| 69 |
+
dtype: string
|
| 70 |
+
- name: answerKey
|
| 71 |
+
dtype: string
|
| 72 |
+
splits:
|
| 73 |
+
- name: train
|
| 74 |
+
num_bytes: 895386
|
| 75 |
+
num_examples: 4957
|
| 76 |
+
- name: validation
|
| 77 |
+
num_bytes: 95428
|
| 78 |
+
num_examples: 500
|
| 79 |
+
- name: test
|
| 80 |
+
num_bytes: 91759
|
| 81 |
+
num_examples: 500
|
| 82 |
+
download_size: 609613
|
| 83 |
+
dataset_size: 1082573
|
| 84 |
+
configs:
|
| 85 |
+
- config_name: additional
|
| 86 |
+
data_files:
|
| 87 |
+
- split: train
|
| 88 |
+
path: additional/train-*
|
| 89 |
+
- split: validation
|
| 90 |
+
path: additional/validation-*
|
| 91 |
+
- split: test
|
| 92 |
+
path: additional/test-*
|
| 93 |
+
- config_name: main
|
| 94 |
+
data_files:
|
| 95 |
+
- split: train
|
| 96 |
+
path: main/train-*
|
| 97 |
+
- split: validation
|
| 98 |
+
path: main/validation-*
|
| 99 |
+
- split: test
|
| 100 |
+
path: main/test-*
|
| 101 |
+
default: true
|
| 102 |
+
---
|
| 103 |
+
|
| 104 |
+
# Dataset Card for OpenBookQA
|
| 105 |
+
|
| 106 |
+
## Table of Contents
|
| 107 |
+
- [Dataset Description](#dataset-description)
|
| 108 |
+
- [Dataset Summary](#dataset-summary)
|
| 109 |
+
- [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
|
| 110 |
+
- [Languages](#languages)
|
| 111 |
+
- [Dataset Structure](#dataset-structure)
|
| 112 |
+
- [Data Instances](#data-instances)
|
| 113 |
+
- [Data Fields](#data-fields)
|
| 114 |
+
- [Data Splits](#data-splits)
|
| 115 |
+
- [Dataset Creation](#dataset-creation)
|
| 116 |
+
- [Curation Rationale](#curation-rationale)
|
| 117 |
+
- [Source Data](#source-data)
|
| 118 |
+
- [Annotations](#annotations)
|
| 119 |
+
- [Personal and Sensitive Information](#personal-and-sensitive-information)
|
| 120 |
+
- [Considerations for Using the Data](#considerations-for-using-the-data)
|
| 121 |
+
- [Social Impact of Dataset](#social-impact-of-dataset)
|
| 122 |
+
- [Discussion of Biases](#discussion-of-biases)
|
| 123 |
+
- [Other Known Limitations](#other-known-limitations)
|
| 124 |
+
- [Additional Information](#additional-information)
|
| 125 |
+
- [Dataset Curators](#dataset-curators)
|
| 126 |
+
- [Licensing Information](#licensing-information)
|
| 127 |
+
- [Citation Information](#citation-information)
|
| 128 |
+
- [Contributions](#contributions)
|
| 129 |
+
|
| 130 |
+
## Dataset Description
|
| 131 |
+
|
| 132 |
+
- **Homepage:** [https://allenai.org/data/open-book-qa](https://allenai.org/data/open-book-qa)
|
| 133 |
+
- **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 134 |
+
- **Paper:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 135 |
+
- **Point of Contact:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 136 |
+
- **Size of downloaded dataset files:** 2.89 MB
|
| 137 |
+
- **Size of the generated dataset:** 2.88 MB
|
| 138 |
+
- **Total amount of disk used:** 5.78 MB
|
| 139 |
+
|
| 140 |
+
### Dataset Summary
|
| 141 |
+
|
| 142 |
+
OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic
|
| 143 |
+
(with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In
|
| 144 |
+
particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge,
|
| 145 |
+
and rich text comprehension.
|
| 146 |
+
OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of
|
| 147 |
+
a subject.
|
| 148 |
+
|
| 149 |
+
### Supported Tasks and Leaderboards
|
| 150 |
+
|
| 151 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 152 |
+
|
| 153 |
+
### Languages
|
| 154 |
+
|
| 155 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 156 |
+
|
| 157 |
+
## Dataset Structure
|
| 158 |
+
|
| 159 |
+
### Data Instances
|
| 160 |
+
|
| 161 |
+
#### main
|
| 162 |
+
|
| 163 |
+
- **Size of downloaded dataset files:** 1.45 MB
|
| 164 |
+
- **Size of the generated dataset:** 1.45 MB
|
| 165 |
+
- **Total amount of disk used:** 2.88 MB
|
| 166 |
+
|
| 167 |
+
An example of 'train' looks as follows:
|
| 168 |
+
```
|
| 169 |
+
{'id': '7-980',
|
| 170 |
+
'question_stem': 'The sun is responsible for',
|
| 171 |
+
'choices': {'text': ['puppies learning new tricks',
|
| 172 |
+
'children growing up and getting old',
|
| 173 |
+
'flowers wilting in a vase',
|
| 174 |
+
'plants sprouting, blooming and wilting'],
|
| 175 |
+
'label': ['A', 'B', 'C', 'D']},
|
| 176 |
+
'answerKey': 'D'}
|
| 177 |
+
```
|
| 178 |
+
|
| 179 |
+
#### additional
|
| 180 |
+
|
| 181 |
+
- **Size of downloaded dataset files:** 1.45 MB
|
| 182 |
+
- **Size of the generated dataset:** 1.45 MB
|
| 183 |
+
- **Total amount of disk used:** 2.88 MB
|
| 184 |
+
|
| 185 |
+
An example of 'train' looks as follows:
|
| 186 |
+
```
|
| 187 |
+
{'id': '7-980',
|
| 188 |
+
'question_stem': 'The sun is responsible for',
|
| 189 |
+
'choices': {'text': ['puppies learning new tricks',
|
| 190 |
+
'children growing up and getting old',
|
| 191 |
+
'flowers wilting in a vase',
|
| 192 |
+
'plants sprouting, blooming and wilting'],
|
| 193 |
+
'label': ['A', 'B', 'C', 'D']},
|
| 194 |
+
'answerKey': 'D',
|
| 195 |
+
'fact1': 'the sun is the source of energy for physical cycles on Earth',
|
| 196 |
+
'humanScore': 1.0,
|
| 197 |
+
'clarity': 2.0,
|
| 198 |
+
'turkIdAnonymized': 'b356d338b7'}
|
| 199 |
+
```
|
| 200 |
+
|
| 201 |
+
### Data Fields
|
| 202 |
+
|
| 203 |
+
The data fields are the same among all splits.
|
| 204 |
+
|
| 205 |
+
#### main
|
| 206 |
+
- `id`: a `string` feature.
|
| 207 |
+
- `question_stem`: a `string` feature.
|
| 208 |
+
- `choices`: a dictionary feature containing:
|
| 209 |
+
- `text`: a `string` feature.
|
| 210 |
+
- `label`: a `string` feature.
|
| 211 |
+
- `answerKey`: a `string` feature.
|
| 212 |
+
|
| 213 |
+
#### additional
|
| 214 |
+
- `id`: a `string` feature.
|
| 215 |
+
- `question_stem`: a `string` feature.
|
| 216 |
+
- `choices`: a dictionary feature containing:
|
| 217 |
+
- `text`: a `string` feature.
|
| 218 |
+
- `label`: a `string` feature.
|
| 219 |
+
- `answerKey`: a `string` feature.
|
| 220 |
+
- `fact1` (`str`): oOriginating common knowledge core fact associated to the question.
|
| 221 |
+
- `humanScore` (`float`): Human accuracy score.
|
| 222 |
+
- `clarity` (`float`): Clarity score.
|
| 223 |
+
- `turkIdAnonymized` (`str`): Anonymized crowd-worker ID.
|
| 224 |
+
|
| 225 |
+
### Data Splits
|
| 226 |
+
|
| 227 |
+
| name | train | validation | test |
|
| 228 |
+
|------------|------:|-----------:|-----:|
|
| 229 |
+
| main | 4957 | 500 | 500 |
|
| 230 |
+
| additional | 4957 | 500 | 500 |
|
| 231 |
+
|
| 232 |
+
## Dataset Creation
|
| 233 |
+
|
| 234 |
+
### Curation Rationale
|
| 235 |
+
|
| 236 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 237 |
+
|
| 238 |
+
### Source Data
|
| 239 |
+
|
| 240 |
+
#### Initial Data Collection and Normalization
|
| 241 |
+
|
| 242 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 243 |
+
|
| 244 |
+
#### Who are the source language producers?
|
| 245 |
+
|
| 246 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 247 |
+
|
| 248 |
+
### Annotations
|
| 249 |
+
|
| 250 |
+
#### Annotation process
|
| 251 |
+
|
| 252 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 253 |
+
|
| 254 |
+
#### Who are the annotators?
|
| 255 |
+
|
| 256 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 257 |
+
|
| 258 |
+
### Personal and Sensitive Information
|
| 259 |
+
|
| 260 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 261 |
+
|
| 262 |
+
## Considerations for Using the Data
|
| 263 |
+
|
| 264 |
+
### Social Impact of Dataset
|
| 265 |
+
|
| 266 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 267 |
+
|
| 268 |
+
### Discussion of Biases
|
| 269 |
+
|
| 270 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 271 |
+
|
| 272 |
+
### Other Known Limitations
|
| 273 |
+
|
| 274 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 275 |
+
|
| 276 |
+
## Additional Information
|
| 277 |
+
|
| 278 |
+
### Dataset Curators
|
| 279 |
+
|
| 280 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 281 |
+
|
| 282 |
+
### Licensing Information
|
| 283 |
+
|
| 284 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 285 |
+
|
| 286 |
+
### Citation Information
|
| 287 |
+
|
| 288 |
+
```
|
| 289 |
+
@inproceedings{OpenBookQA2018,
|
| 290 |
+
title={Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering},
|
| 291 |
+
author={Todor Mihaylov and Peter Clark and Tushar Khot and Ashish Sabharwal},
|
| 292 |
+
booktitle={EMNLP},
|
| 293 |
+
year={2018}
|
| 294 |
+
}
|
| 295 |
+
|
| 296 |
+
```
|
| 297 |
+
|
| 298 |
+
|
| 299 |
+
### Contributions
|
| 300 |
+
|
| 301 |
+
Thanks to [@thomwolf](https://github.com/thomwolf), [@patrickvonplaten](https://github.com/patrickvonplaten), [@lewtun](https://github.com/lewtun) for adding this dataset.
|
hf_cache/hub/datasets--allenai--sciq/.no_exist/2c94ad3e1aafab77146f384e23536f97a4849815/.huggingface.yaml
ADDED
|
File without changes
|
hf_cache/hub/datasets--allenai--sciq/.no_exist/2c94ad3e1aafab77146f384e23536f97a4849815/dataset_infos.json
ADDED
|
File without changes
|
hf_cache/hub/datasets--allenai--sciq/.no_exist/2c94ad3e1aafab77146f384e23536f97a4849815/sciq.py
ADDED
|
File without changes
|
hf_cache/hub/datasets--allenai--sciq/blobs/c644057869cabcde87a2b5ab9665ec0d0bd1405b
ADDED
|
@@ -0,0 +1,216 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
annotations_creators:
|
| 3 |
+
- no-annotation
|
| 4 |
+
language_creators:
|
| 5 |
+
- crowdsourced
|
| 6 |
+
language:
|
| 7 |
+
- en
|
| 8 |
+
license:
|
| 9 |
+
- cc-by-nc-3.0
|
| 10 |
+
multilinguality:
|
| 11 |
+
- monolingual
|
| 12 |
+
size_categories:
|
| 13 |
+
- 10K<n<100K
|
| 14 |
+
source_datasets:
|
| 15 |
+
- original
|
| 16 |
+
task_categories:
|
| 17 |
+
- question-answering
|
| 18 |
+
task_ids:
|
| 19 |
+
- closed-domain-qa
|
| 20 |
+
paperswithcode_id: sciq
|
| 21 |
+
pretty_name: SciQ
|
| 22 |
+
dataset_info:
|
| 23 |
+
features:
|
| 24 |
+
- name: question
|
| 25 |
+
dtype: string
|
| 26 |
+
- name: distractor3
|
| 27 |
+
dtype: string
|
| 28 |
+
- name: distractor1
|
| 29 |
+
dtype: string
|
| 30 |
+
- name: distractor2
|
| 31 |
+
dtype: string
|
| 32 |
+
- name: correct_answer
|
| 33 |
+
dtype: string
|
| 34 |
+
- name: support
|
| 35 |
+
dtype: string
|
| 36 |
+
splits:
|
| 37 |
+
- name: train
|
| 38 |
+
num_bytes: 6546183
|
| 39 |
+
num_examples: 11679
|
| 40 |
+
- name: validation
|
| 41 |
+
num_bytes: 554120
|
| 42 |
+
num_examples: 1000
|
| 43 |
+
- name: test
|
| 44 |
+
num_bytes: 563927
|
| 45 |
+
num_examples: 1000
|
| 46 |
+
download_size: 4674410
|
| 47 |
+
dataset_size: 7664230
|
| 48 |
+
configs:
|
| 49 |
+
- config_name: default
|
| 50 |
+
data_files:
|
| 51 |
+
- split: train
|
| 52 |
+
path: data/train-*
|
| 53 |
+
- split: validation
|
| 54 |
+
path: data/validation-*
|
| 55 |
+
- split: test
|
| 56 |
+
path: data/test-*
|
| 57 |
+
---
|
| 58 |
+
|
| 59 |
+
# Dataset Card for "sciq"
|
| 60 |
+
|
| 61 |
+
## Table of Contents
|
| 62 |
+
- [Dataset Description](#dataset-description)
|
| 63 |
+
- [Dataset Summary](#dataset-summary)
|
| 64 |
+
- [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
|
| 65 |
+
- [Languages](#languages)
|
| 66 |
+
- [Dataset Structure](#dataset-structure)
|
| 67 |
+
- [Data Instances](#data-instances)
|
| 68 |
+
- [Data Fields](#data-fields)
|
| 69 |
+
- [Data Splits](#data-splits)
|
| 70 |
+
- [Dataset Creation](#dataset-creation)
|
| 71 |
+
- [Curation Rationale](#curation-rationale)
|
| 72 |
+
- [Source Data](#source-data)
|
| 73 |
+
- [Annotations](#annotations)
|
| 74 |
+
- [Personal and Sensitive Information](#personal-and-sensitive-information)
|
| 75 |
+
- [Considerations for Using the Data](#considerations-for-using-the-data)
|
| 76 |
+
- [Social Impact of Dataset](#social-impact-of-dataset)
|
| 77 |
+
- [Discussion of Biases](#discussion-of-biases)
|
| 78 |
+
- [Other Known Limitations](#other-known-limitations)
|
| 79 |
+
- [Additional Information](#additional-information)
|
| 80 |
+
- [Dataset Curators](#dataset-curators)
|
| 81 |
+
- [Licensing Information](#licensing-information)
|
| 82 |
+
- [Citation Information](#citation-information)
|
| 83 |
+
- [Contributions](#contributions)
|
| 84 |
+
|
| 85 |
+
## Dataset Description
|
| 86 |
+
|
| 87 |
+
- **Homepage:** [https://allenai.org/data/sciq](https://allenai.org/data/sciq)
|
| 88 |
+
- **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 89 |
+
- **Paper:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 90 |
+
- **Point of Contact:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 91 |
+
- **Size of downloaded dataset files:** 2.82 MB
|
| 92 |
+
- **Size of the generated dataset:** 7.68 MB
|
| 93 |
+
- **Total amount of disk used:** 10.50 MB
|
| 94 |
+
|
| 95 |
+
### Dataset Summary
|
| 96 |
+
|
| 97 |
+
The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others. The questions are in multiple-choice format with 4 answer options each. For the majority of the questions, an additional paragraph with supporting evidence for the correct answer is provided.
|
| 98 |
+
|
| 99 |
+
### Supported Tasks and Leaderboards
|
| 100 |
+
|
| 101 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 102 |
+
|
| 103 |
+
### Languages
|
| 104 |
+
|
| 105 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 106 |
+
|
| 107 |
+
## Dataset Structure
|
| 108 |
+
|
| 109 |
+
### Data Instances
|
| 110 |
+
|
| 111 |
+
#### default
|
| 112 |
+
|
| 113 |
+
- **Size of downloaded dataset files:** 2.82 MB
|
| 114 |
+
- **Size of the generated dataset:** 7.68 MB
|
| 115 |
+
- **Total amount of disk used:** 10.50 MB
|
| 116 |
+
|
| 117 |
+
An example of 'train' looks as follows.
|
| 118 |
+
```
|
| 119 |
+
This example was too long and was cropped:
|
| 120 |
+
|
| 121 |
+
{
|
| 122 |
+
"correct_answer": "coriolis effect",
|
| 123 |
+
"distractor1": "muon effect",
|
| 124 |
+
"distractor2": "centrifugal effect",
|
| 125 |
+
"distractor3": "tropical effect",
|
| 126 |
+
"question": "What phenomenon makes global winds blow northeast to southwest or the reverse in the northern hemisphere and northwest to southeast or the reverse in the southern hemisphere?",
|
| 127 |
+
"support": "\"Without Coriolis Effect the global winds would blow north to south or south to north. But Coriolis makes them blow northeast to..."
|
| 128 |
+
}
|
| 129 |
+
```
|
| 130 |
+
|
| 131 |
+
### Data Fields
|
| 132 |
+
|
| 133 |
+
The data fields are the same among all splits.
|
| 134 |
+
|
| 135 |
+
#### default
|
| 136 |
+
- `question`: a `string` feature.
|
| 137 |
+
- `distractor3`: a `string` feature.
|
| 138 |
+
- `distractor1`: a `string` feature.
|
| 139 |
+
- `distractor2`: a `string` feature.
|
| 140 |
+
- `correct_answer`: a `string` feature.
|
| 141 |
+
- `support`: a `string` feature.
|
| 142 |
+
|
| 143 |
+
### Data Splits
|
| 144 |
+
|
| 145 |
+
| name |train|validation|test|
|
| 146 |
+
|-------|----:|---------:|---:|
|
| 147 |
+
|default|11679| 1000|1000|
|
| 148 |
+
|
| 149 |
+
## Dataset Creation
|
| 150 |
+
|
| 151 |
+
### Curation Rationale
|
| 152 |
+
|
| 153 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 154 |
+
|
| 155 |
+
### Source Data
|
| 156 |
+
|
| 157 |
+
#### Initial Data Collection and Normalization
|
| 158 |
+
|
| 159 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 160 |
+
|
| 161 |
+
#### Who are the source language producers?
|
| 162 |
+
|
| 163 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 164 |
+
|
| 165 |
+
### Annotations
|
| 166 |
+
|
| 167 |
+
#### Annotation process
|
| 168 |
+
|
| 169 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 170 |
+
|
| 171 |
+
#### Who are the annotators?
|
| 172 |
+
|
| 173 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 174 |
+
|
| 175 |
+
### Personal and Sensitive Information
|
| 176 |
+
|
| 177 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 178 |
+
|
| 179 |
+
## Considerations for Using the Data
|
| 180 |
+
|
| 181 |
+
### Social Impact of Dataset
|
| 182 |
+
|
| 183 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 184 |
+
|
| 185 |
+
### Discussion of Biases
|
| 186 |
+
|
| 187 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 188 |
+
|
| 189 |
+
### Other Known Limitations
|
| 190 |
+
|
| 191 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 192 |
+
|
| 193 |
+
## Additional Information
|
| 194 |
+
|
| 195 |
+
### Dataset Curators
|
| 196 |
+
|
| 197 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 198 |
+
|
| 199 |
+
### Licensing Information
|
| 200 |
+
|
| 201 |
+
The dataset is licensed under the [Creative Commons Attribution-NonCommercial 3.0 Unported License](http://creativecommons.org/licenses/by-nc/3.0/).
|
| 202 |
+
|
| 203 |
+
### Citation Information
|
| 204 |
+
|
| 205 |
+
```
|
| 206 |
+
@inproceedings{SciQ,
|
| 207 |
+
title={Crowdsourcing Multiple Choice Science Questions},
|
| 208 |
+
author={Johannes Welbl, Nelson F. Liu, Matt Gardner},
|
| 209 |
+
year={2017},
|
| 210 |
+
journal={arXiv:1707.06209v1}
|
| 211 |
+
}
|
| 212 |
+
```
|
| 213 |
+
|
| 214 |
+
### Contributions
|
| 215 |
+
|
| 216 |
+
Thanks to [@patrickvonplaten](https://github.com/patrickvonplaten), [@lewtun](https://github.com/lewtun), [@thomwolf](https://github.com/thomwolf) for adding this dataset.
|
hf_cache/hub/datasets--allenai--sciq/refs/main
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
2c94ad3e1aafab77146f384e23536f97a4849815
|
hf_cache/hub/datasets--allenai--sciq/snapshots/2c94ad3e1aafab77146f384e23536f97a4849815/README.md
ADDED
|
@@ -0,0 +1,216 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
annotations_creators:
|
| 3 |
+
- no-annotation
|
| 4 |
+
language_creators:
|
| 5 |
+
- crowdsourced
|
| 6 |
+
language:
|
| 7 |
+
- en
|
| 8 |
+
license:
|
| 9 |
+
- cc-by-nc-3.0
|
| 10 |
+
multilinguality:
|
| 11 |
+
- monolingual
|
| 12 |
+
size_categories:
|
| 13 |
+
- 10K<n<100K
|
| 14 |
+
source_datasets:
|
| 15 |
+
- original
|
| 16 |
+
task_categories:
|
| 17 |
+
- question-answering
|
| 18 |
+
task_ids:
|
| 19 |
+
- closed-domain-qa
|
| 20 |
+
paperswithcode_id: sciq
|
| 21 |
+
pretty_name: SciQ
|
| 22 |
+
dataset_info:
|
| 23 |
+
features:
|
| 24 |
+
- name: question
|
| 25 |
+
dtype: string
|
| 26 |
+
- name: distractor3
|
| 27 |
+
dtype: string
|
| 28 |
+
- name: distractor1
|
| 29 |
+
dtype: string
|
| 30 |
+
- name: distractor2
|
| 31 |
+
dtype: string
|
| 32 |
+
- name: correct_answer
|
| 33 |
+
dtype: string
|
| 34 |
+
- name: support
|
| 35 |
+
dtype: string
|
| 36 |
+
splits:
|
| 37 |
+
- name: train
|
| 38 |
+
num_bytes: 6546183
|
| 39 |
+
num_examples: 11679
|
| 40 |
+
- name: validation
|
| 41 |
+
num_bytes: 554120
|
| 42 |
+
num_examples: 1000
|
| 43 |
+
- name: test
|
| 44 |
+
num_bytes: 563927
|
| 45 |
+
num_examples: 1000
|
| 46 |
+
download_size: 4674410
|
| 47 |
+
dataset_size: 7664230
|
| 48 |
+
configs:
|
| 49 |
+
- config_name: default
|
| 50 |
+
data_files:
|
| 51 |
+
- split: train
|
| 52 |
+
path: data/train-*
|
| 53 |
+
- split: validation
|
| 54 |
+
path: data/validation-*
|
| 55 |
+
- split: test
|
| 56 |
+
path: data/test-*
|
| 57 |
+
---
|
| 58 |
+
|
| 59 |
+
# Dataset Card for "sciq"
|
| 60 |
+
|
| 61 |
+
## Table of Contents
|
| 62 |
+
- [Dataset Description](#dataset-description)
|
| 63 |
+
- [Dataset Summary](#dataset-summary)
|
| 64 |
+
- [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
|
| 65 |
+
- [Languages](#languages)
|
| 66 |
+
- [Dataset Structure](#dataset-structure)
|
| 67 |
+
- [Data Instances](#data-instances)
|
| 68 |
+
- [Data Fields](#data-fields)
|
| 69 |
+
- [Data Splits](#data-splits)
|
| 70 |
+
- [Dataset Creation](#dataset-creation)
|
| 71 |
+
- [Curation Rationale](#curation-rationale)
|
| 72 |
+
- [Source Data](#source-data)
|
| 73 |
+
- [Annotations](#annotations)
|
| 74 |
+
- [Personal and Sensitive Information](#personal-and-sensitive-information)
|
| 75 |
+
- [Considerations for Using the Data](#considerations-for-using-the-data)
|
| 76 |
+
- [Social Impact of Dataset](#social-impact-of-dataset)
|
| 77 |
+
- [Discussion of Biases](#discussion-of-biases)
|
| 78 |
+
- [Other Known Limitations](#other-known-limitations)
|
| 79 |
+
- [Additional Information](#additional-information)
|
| 80 |
+
- [Dataset Curators](#dataset-curators)
|
| 81 |
+
- [Licensing Information](#licensing-information)
|
| 82 |
+
- [Citation Information](#citation-information)
|
| 83 |
+
- [Contributions](#contributions)
|
| 84 |
+
|
| 85 |
+
## Dataset Description
|
| 86 |
+
|
| 87 |
+
- **Homepage:** [https://allenai.org/data/sciq](https://allenai.org/data/sciq)
|
| 88 |
+
- **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 89 |
+
- **Paper:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 90 |
+
- **Point of Contact:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 91 |
+
- **Size of downloaded dataset files:** 2.82 MB
|
| 92 |
+
- **Size of the generated dataset:** 7.68 MB
|
| 93 |
+
- **Total amount of disk used:** 10.50 MB
|
| 94 |
+
|
| 95 |
+
### Dataset Summary
|
| 96 |
+
|
| 97 |
+
The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others. The questions are in multiple-choice format with 4 answer options each. For the majority of the questions, an additional paragraph with supporting evidence for the correct answer is provided.
|
| 98 |
+
|
| 99 |
+
### Supported Tasks and Leaderboards
|
| 100 |
+
|
| 101 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 102 |
+
|
| 103 |
+
### Languages
|
| 104 |
+
|
| 105 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 106 |
+
|
| 107 |
+
## Dataset Structure
|
| 108 |
+
|
| 109 |
+
### Data Instances
|
| 110 |
+
|
| 111 |
+
#### default
|
| 112 |
+
|
| 113 |
+
- **Size of downloaded dataset files:** 2.82 MB
|
| 114 |
+
- **Size of the generated dataset:** 7.68 MB
|
| 115 |
+
- **Total amount of disk used:** 10.50 MB
|
| 116 |
+
|
| 117 |
+
An example of 'train' looks as follows.
|
| 118 |
+
```
|
| 119 |
+
This example was too long and was cropped:
|
| 120 |
+
|
| 121 |
+
{
|
| 122 |
+
"correct_answer": "coriolis effect",
|
| 123 |
+
"distractor1": "muon effect",
|
| 124 |
+
"distractor2": "centrifugal effect",
|
| 125 |
+
"distractor3": "tropical effect",
|
| 126 |
+
"question": "What phenomenon makes global winds blow northeast to southwest or the reverse in the northern hemisphere and northwest to southeast or the reverse in the southern hemisphere?",
|
| 127 |
+
"support": "\"Without Coriolis Effect the global winds would blow north to south or south to north. But Coriolis makes them blow northeast to..."
|
| 128 |
+
}
|
| 129 |
+
```
|
| 130 |
+
|
| 131 |
+
### Data Fields
|
| 132 |
+
|
| 133 |
+
The data fields are the same among all splits.
|
| 134 |
+
|
| 135 |
+
#### default
|
| 136 |
+
- `question`: a `string` feature.
|
| 137 |
+
- `distractor3`: a `string` feature.
|
| 138 |
+
- `distractor1`: a `string` feature.
|
| 139 |
+
- `distractor2`: a `string` feature.
|
| 140 |
+
- `correct_answer`: a `string` feature.
|
| 141 |
+
- `support`: a `string` feature.
|
| 142 |
+
|
| 143 |
+
### Data Splits
|
| 144 |
+
|
| 145 |
+
| name |train|validation|test|
|
| 146 |
+
|-------|----:|---------:|---:|
|
| 147 |
+
|default|11679| 1000|1000|
|
| 148 |
+
|
| 149 |
+
## Dataset Creation
|
| 150 |
+
|
| 151 |
+
### Curation Rationale
|
| 152 |
+
|
| 153 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 154 |
+
|
| 155 |
+
### Source Data
|
| 156 |
+
|
| 157 |
+
#### Initial Data Collection and Normalization
|
| 158 |
+
|
| 159 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 160 |
+
|
| 161 |
+
#### Who are the source language producers?
|
| 162 |
+
|
| 163 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 164 |
+
|
| 165 |
+
### Annotations
|
| 166 |
+
|
| 167 |
+
#### Annotation process
|
| 168 |
+
|
| 169 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 170 |
+
|
| 171 |
+
#### Who are the annotators?
|
| 172 |
+
|
| 173 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 174 |
+
|
| 175 |
+
### Personal and Sensitive Information
|
| 176 |
+
|
| 177 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 178 |
+
|
| 179 |
+
## Considerations for Using the Data
|
| 180 |
+
|
| 181 |
+
### Social Impact of Dataset
|
| 182 |
+
|
| 183 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 184 |
+
|
| 185 |
+
### Discussion of Biases
|
| 186 |
+
|
| 187 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 188 |
+
|
| 189 |
+
### Other Known Limitations
|
| 190 |
+
|
| 191 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 192 |
+
|
| 193 |
+
## Additional Information
|
| 194 |
+
|
| 195 |
+
### Dataset Curators
|
| 196 |
+
|
| 197 |
+
[More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
|
| 198 |
+
|
| 199 |
+
### Licensing Information
|
| 200 |
+
|
| 201 |
+
The dataset is licensed under the [Creative Commons Attribution-NonCommercial 3.0 Unported License](http://creativecommons.org/licenses/by-nc/3.0/).
|
| 202 |
+
|
| 203 |
+
### Citation Information
|
| 204 |
+
|
| 205 |
+
```
|
| 206 |
+
@inproceedings{SciQ,
|
| 207 |
+
title={Crowdsourcing Multiple Choice Science Questions},
|
| 208 |
+
author={Johannes Welbl, Nelson F. Liu, Matt Gardner},
|
| 209 |
+
year={2017},
|
| 210 |
+
journal={arXiv:1707.06209v1}
|
| 211 |
+
}
|
| 212 |
+
```
|
| 213 |
+
|
| 214 |
+
### Contributions
|
| 215 |
+
|
| 216 |
+
Thanks to [@patrickvonplaten](https://github.com/patrickvonplaten), [@lewtun](https://github.com/lewtun), [@thomwolf](https://github.com/thomwolf) for adding this dataset.
|
hf_cache/hub/datasets--cais--mmlu/.no_exist/c30699e8356da336a370243923dbaf21066bb9fe/.huggingface.yaml
ADDED
|
File without changes
|
hf_cache/hub/datasets--cais--mmlu/.no_exist/c30699e8356da336a370243923dbaf21066bb9fe/mmlu.py
ADDED
|
File without changes
|
hf_cache/hub/datasets--cais--mmlu/blobs/08de94c560ad7420252bff6e4729f1d1683def4f
ADDED
|
@@ -0,0 +1,2299 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
annotations_creators:
|
| 3 |
+
- no-annotation
|
| 4 |
+
language_creators:
|
| 5 |
+
- expert-generated
|
| 6 |
+
language:
|
| 7 |
+
- en
|
| 8 |
+
license:
|
| 9 |
+
- mit
|
| 10 |
+
multilinguality:
|
| 11 |
+
- monolingual
|
| 12 |
+
size_categories:
|
| 13 |
+
- 10K<n<100K
|
| 14 |
+
source_datasets:
|
| 15 |
+
- original
|
| 16 |
+
task_categories:
|
| 17 |
+
- question-answering
|
| 18 |
+
task_ids:
|
| 19 |
+
- multiple-choice-qa
|
| 20 |
+
paperswithcode_id: mmlu
|
| 21 |
+
pretty_name: Measuring Massive Multitask Language Understanding
|
| 22 |
+
language_bcp47:
|
| 23 |
+
- en-US
|
| 24 |
+
dataset_info:
|
| 25 |
+
- config_name: abstract_algebra
|
| 26 |
+
features:
|
| 27 |
+
- name: question
|
| 28 |
+
dtype: string
|
| 29 |
+
- name: subject
|
| 30 |
+
dtype: string
|
| 31 |
+
- name: choices
|
| 32 |
+
sequence: string
|
| 33 |
+
- name: answer
|
| 34 |
+
dtype:
|
| 35 |
+
class_label:
|
| 36 |
+
names:
|
| 37 |
+
'0': A
|
| 38 |
+
'1': B
|
| 39 |
+
'2': C
|
| 40 |
+
'3': D
|
| 41 |
+
splits:
|
| 42 |
+
- name: test
|
| 43 |
+
num_bytes: 49618.6654322746
|
| 44 |
+
num_examples: 100
|
| 45 |
+
- name: validation
|
| 46 |
+
num_bytes: 5485.515349444808
|
| 47 |
+
num_examples: 11
|
| 48 |
+
- name: dev
|
| 49 |
+
num_bytes: 2199.1754385964914
|
| 50 |
+
num_examples: 5
|
| 51 |
+
download_size: 17143
|
| 52 |
+
dataset_size: 57303.3562203159
|
| 53 |
+
- config_name: all
|
| 54 |
+
features:
|
| 55 |
+
- name: question
|
| 56 |
+
dtype: string
|
| 57 |
+
- name: subject
|
| 58 |
+
dtype: string
|
| 59 |
+
- name: choices
|
| 60 |
+
sequence: string
|
| 61 |
+
- name: answer
|
| 62 |
+
dtype:
|
| 63 |
+
class_label:
|
| 64 |
+
names:
|
| 65 |
+
'0': A
|
| 66 |
+
'1': B
|
| 67 |
+
'2': C
|
| 68 |
+
'3': D
|
| 69 |
+
splits:
|
| 70 |
+
- name: test
|
| 71 |
+
num_bytes: 6967453
|
| 72 |
+
num_examples: 14042
|
| 73 |
+
- name: validation
|
| 74 |
+
num_bytes: 763484
|
| 75 |
+
num_examples: 1531
|
| 76 |
+
- name: dev
|
| 77 |
+
num_bytes: 125353
|
| 78 |
+
num_examples: 285
|
| 79 |
+
- name: auxiliary_train
|
| 80 |
+
num_bytes: 161000625
|
| 81 |
+
num_examples: 99842
|
| 82 |
+
download_size: 51503402
|
| 83 |
+
dataset_size: 168856915
|
| 84 |
+
- config_name: anatomy
|
| 85 |
+
features:
|
| 86 |
+
- name: question
|
| 87 |
+
dtype: string
|
| 88 |
+
- name: subject
|
| 89 |
+
dtype: string
|
| 90 |
+
- name: choices
|
| 91 |
+
sequence: string
|
| 92 |
+
- name: answer
|
| 93 |
+
dtype:
|
| 94 |
+
class_label:
|
| 95 |
+
names:
|
| 96 |
+
'0': A
|
| 97 |
+
'1': B
|
| 98 |
+
'2': C
|
| 99 |
+
'3': D
|
| 100 |
+
splits:
|
| 101 |
+
- name: test
|
| 102 |
+
num_bytes: 66985.19833357072
|
| 103 |
+
num_examples: 135
|
| 104 |
+
- name: validation
|
| 105 |
+
num_bytes: 6981.5649902024825
|
| 106 |
+
num_examples: 14
|
| 107 |
+
- name: dev
|
| 108 |
+
num_bytes: 2199.1754385964914
|
| 109 |
+
num_examples: 5
|
| 110 |
+
download_size: 28864
|
| 111 |
+
dataset_size: 76165.9387623697
|
| 112 |
+
- config_name: astronomy
|
| 113 |
+
features:
|
| 114 |
+
- name: question
|
| 115 |
+
dtype: string
|
| 116 |
+
- name: subject
|
| 117 |
+
dtype: string
|
| 118 |
+
- name: choices
|
| 119 |
+
sequence: string
|
| 120 |
+
- name: answer
|
| 121 |
+
dtype:
|
| 122 |
+
class_label:
|
| 123 |
+
names:
|
| 124 |
+
'0': A
|
| 125 |
+
'1': B
|
| 126 |
+
'2': C
|
| 127 |
+
'3': D
|
| 128 |
+
splits:
|
| 129 |
+
- name: test
|
| 130 |
+
num_bytes: 75420.3714570574
|
| 131 |
+
num_examples: 152
|
| 132 |
+
- name: validation
|
| 133 |
+
num_bytes: 7978.931417374265
|
| 134 |
+
num_examples: 16
|
| 135 |
+
- name: dev
|
| 136 |
+
num_bytes: 2199.1754385964914
|
| 137 |
+
num_examples: 5
|
| 138 |
+
download_size: 39316
|
| 139 |
+
dataset_size: 85598.47831302814
|
| 140 |
+
- config_name: auxiliary_train
|
| 141 |
+
features:
|
| 142 |
+
- name: train
|
| 143 |
+
struct:
|
| 144 |
+
- name: answer
|
| 145 |
+
dtype: int64
|
| 146 |
+
- name: choices
|
| 147 |
+
sequence: string
|
| 148 |
+
- name: question
|
| 149 |
+
dtype: string
|
| 150 |
+
- name: subject
|
| 151 |
+
dtype: string
|
| 152 |
+
splits:
|
| 153 |
+
- name: train
|
| 154 |
+
num_bytes: 161000625
|
| 155 |
+
num_examples: 99842
|
| 156 |
+
download_size: 47518592
|
| 157 |
+
dataset_size: 161000625
|
| 158 |
+
- config_name: business_ethics
|
| 159 |
+
features:
|
| 160 |
+
- name: question
|
| 161 |
+
dtype: string
|
| 162 |
+
- name: subject
|
| 163 |
+
dtype: string
|
| 164 |
+
- name: choices
|
| 165 |
+
sequence: string
|
| 166 |
+
- name: answer
|
| 167 |
+
dtype:
|
| 168 |
+
class_label:
|
| 169 |
+
names:
|
| 170 |
+
'0': A
|
| 171 |
+
'1': B
|
| 172 |
+
'2': C
|
| 173 |
+
'3': D
|
| 174 |
+
splits:
|
| 175 |
+
- name: test
|
| 176 |
+
num_bytes: 49618.6654322746
|
| 177 |
+
num_examples: 100
|
| 178 |
+
- name: validation
|
| 179 |
+
num_bytes: 5485.515349444808
|
| 180 |
+
num_examples: 11
|
| 181 |
+
- name: dev
|
| 182 |
+
num_bytes: 2199.1754385964914
|
| 183 |
+
num_examples: 5
|
| 184 |
+
download_size: 31619
|
| 185 |
+
dataset_size: 57303.3562203159
|
| 186 |
+
- config_name: clinical_knowledge
|
| 187 |
+
features:
|
| 188 |
+
- name: question
|
| 189 |
+
dtype: string
|
| 190 |
+
- name: subject
|
| 191 |
+
dtype: string
|
| 192 |
+
- name: choices
|
| 193 |
+
sequence: string
|
| 194 |
+
- name: answer
|
| 195 |
+
dtype:
|
| 196 |
+
class_label:
|
| 197 |
+
names:
|
| 198 |
+
'0': A
|
| 199 |
+
'1': B
|
| 200 |
+
'2': C
|
| 201 |
+
'3': D
|
| 202 |
+
splits:
|
| 203 |
+
- name: test
|
| 204 |
+
num_bytes: 131489.4633955277
|
| 205 |
+
num_examples: 265
|
| 206 |
+
- name: validation
|
| 207 |
+
num_bytes: 14461.813193990856
|
| 208 |
+
num_examples: 29
|
| 209 |
+
- name: dev
|
| 210 |
+
num_bytes: 2199.1754385964914
|
| 211 |
+
num_examples: 5
|
| 212 |
+
download_size: 51655
|
| 213 |
+
dataset_size: 148150.45202811505
|
| 214 |
+
- config_name: college_biology
|
| 215 |
+
features:
|
| 216 |
+
- name: question
|
| 217 |
+
dtype: string
|
| 218 |
+
- name: subject
|
| 219 |
+
dtype: string
|
| 220 |
+
- name: choices
|
| 221 |
+
sequence: string
|
| 222 |
+
- name: answer
|
| 223 |
+
dtype:
|
| 224 |
+
class_label:
|
| 225 |
+
names:
|
| 226 |
+
'0': A
|
| 227 |
+
'1': B
|
| 228 |
+
'2': C
|
| 229 |
+
'3': D
|
| 230 |
+
splits:
|
| 231 |
+
- name: test
|
| 232 |
+
num_bytes: 71450.87822247542
|
| 233 |
+
num_examples: 144
|
| 234 |
+
- name: validation
|
| 235 |
+
num_bytes: 7978.931417374265
|
| 236 |
+
num_examples: 16
|
| 237 |
+
- name: dev
|
| 238 |
+
num_bytes: 2199.1754385964914
|
| 239 |
+
num_examples: 5
|
| 240 |
+
download_size: 43017
|
| 241 |
+
dataset_size: 81628.98507844617
|
| 242 |
+
- config_name: college_chemistry
|
| 243 |
+
features:
|
| 244 |
+
- name: question
|
| 245 |
+
dtype: string
|
| 246 |
+
- name: subject
|
| 247 |
+
dtype: string
|
| 248 |
+
- name: choices
|
| 249 |
+
sequence: string
|
| 250 |
+
- name: answer
|
| 251 |
+
dtype:
|
| 252 |
+
class_label:
|
| 253 |
+
names:
|
| 254 |
+
'0': A
|
| 255 |
+
'1': B
|
| 256 |
+
'2': C
|
| 257 |
+
'3': D
|
| 258 |
+
splits:
|
| 259 |
+
- name: test
|
| 260 |
+
num_bytes: 49618.6654322746
|
| 261 |
+
num_examples: 100
|
| 262 |
+
- name: validation
|
| 263 |
+
num_bytes: 3989.4657086871325
|
| 264 |
+
num_examples: 8
|
| 265 |
+
- name: dev
|
| 266 |
+
num_bytes: 2199.1754385964914
|
| 267 |
+
num_examples: 5
|
| 268 |
+
download_size: 26781
|
| 269 |
+
dataset_size: 55807.30657955822
|
| 270 |
+
- config_name: college_computer_science
|
| 271 |
+
features:
|
| 272 |
+
- name: question
|
| 273 |
+
dtype: string
|
| 274 |
+
- name: subject
|
| 275 |
+
dtype: string
|
| 276 |
+
- name: choices
|
| 277 |
+
sequence: string
|
| 278 |
+
- name: answer
|
| 279 |
+
dtype:
|
| 280 |
+
class_label:
|
| 281 |
+
names:
|
| 282 |
+
'0': A
|
| 283 |
+
'1': B
|
| 284 |
+
'2': C
|
| 285 |
+
'3': D
|
| 286 |
+
splits:
|
| 287 |
+
- name: test
|
| 288 |
+
num_bytes: 49618.6654322746
|
| 289 |
+
num_examples: 100
|
| 290 |
+
- name: validation
|
| 291 |
+
num_bytes: 5485.515349444808
|
| 292 |
+
num_examples: 11
|
| 293 |
+
- name: dev
|
| 294 |
+
num_bytes: 2199.1754385964914
|
| 295 |
+
num_examples: 5
|
| 296 |
+
download_size: 41132
|
| 297 |
+
dataset_size: 57303.3562203159
|
| 298 |
+
- config_name: college_mathematics
|
| 299 |
+
features:
|
| 300 |
+
- name: question
|
| 301 |
+
dtype: string
|
| 302 |
+
- name: subject
|
| 303 |
+
dtype: string
|
| 304 |
+
- name: choices
|
| 305 |
+
sequence: string
|
| 306 |
+
- name: answer
|
| 307 |
+
dtype:
|
| 308 |
+
class_label:
|
| 309 |
+
names:
|
| 310 |
+
'0': A
|
| 311 |
+
'1': B
|
| 312 |
+
'2': C
|
| 313 |
+
'3': D
|
| 314 |
+
splits:
|
| 315 |
+
- name: test
|
| 316 |
+
num_bytes: 49618.6654322746
|
| 317 |
+
num_examples: 100
|
| 318 |
+
- name: validation
|
| 319 |
+
num_bytes: 5485.515349444808
|
| 320 |
+
num_examples: 11
|
| 321 |
+
- name: dev
|
| 322 |
+
num_bytes: 2199.1754385964914
|
| 323 |
+
num_examples: 5
|
| 324 |
+
download_size: 26779
|
| 325 |
+
dataset_size: 57303.3562203159
|
| 326 |
+
- config_name: college_medicine
|
| 327 |
+
features:
|
| 328 |
+
- name: question
|
| 329 |
+
dtype: string
|
| 330 |
+
- name: subject
|
| 331 |
+
dtype: string
|
| 332 |
+
- name: choices
|
| 333 |
+
sequence: string
|
| 334 |
+
- name: answer
|
| 335 |
+
dtype:
|
| 336 |
+
class_label:
|
| 337 |
+
names:
|
| 338 |
+
'0': A
|
| 339 |
+
'1': B
|
| 340 |
+
'2': C
|
| 341 |
+
'3': D
|
| 342 |
+
splits:
|
| 343 |
+
- name: test
|
| 344 |
+
num_bytes: 85840.29119783506
|
| 345 |
+
num_examples: 173
|
| 346 |
+
- name: validation
|
| 347 |
+
num_bytes: 10971.030698889615
|
| 348 |
+
num_examples: 22
|
| 349 |
+
- name: dev
|
| 350 |
+
num_bytes: 2199.1754385964914
|
| 351 |
+
num_examples: 5
|
| 352 |
+
download_size: 56303
|
| 353 |
+
dataset_size: 99010.49733532117
|
| 354 |
+
- config_name: college_physics
|
| 355 |
+
features:
|
| 356 |
+
- name: question
|
| 357 |
+
dtype: string
|
| 358 |
+
- name: subject
|
| 359 |
+
dtype: string
|
| 360 |
+
- name: choices
|
| 361 |
+
sequence: string
|
| 362 |
+
- name: answer
|
| 363 |
+
dtype:
|
| 364 |
+
class_label:
|
| 365 |
+
names:
|
| 366 |
+
'0': A
|
| 367 |
+
'1': B
|
| 368 |
+
'2': C
|
| 369 |
+
'3': D
|
| 370 |
+
splits:
|
| 371 |
+
- name: test
|
| 372 |
+
num_bytes: 50611.0387409201
|
| 373 |
+
num_examples: 102
|
| 374 |
+
- name: validation
|
| 375 |
+
num_bytes: 5485.515349444808
|
| 376 |
+
num_examples: 11
|
| 377 |
+
- name: dev
|
| 378 |
+
num_bytes: 2199.1754385964914
|
| 379 |
+
num_examples: 5
|
| 380 |
+
download_size: 29539
|
| 381 |
+
dataset_size: 58295.7295289614
|
| 382 |
+
- config_name: computer_security
|
| 383 |
+
features:
|
| 384 |
+
- name: question
|
| 385 |
+
dtype: string
|
| 386 |
+
- name: subject
|
| 387 |
+
dtype: string
|
| 388 |
+
- name: choices
|
| 389 |
+
sequence: string
|
| 390 |
+
- name: answer
|
| 391 |
+
dtype:
|
| 392 |
+
class_label:
|
| 393 |
+
names:
|
| 394 |
+
'0': A
|
| 395 |
+
'1': B
|
| 396 |
+
'2': C
|
| 397 |
+
'3': D
|
| 398 |
+
splits:
|
| 399 |
+
- name: test
|
| 400 |
+
num_bytes: 49618.6654322746
|
| 401 |
+
num_examples: 100
|
| 402 |
+
- name: validation
|
| 403 |
+
num_bytes: 5485.515349444808
|
| 404 |
+
num_examples: 11
|
| 405 |
+
- name: dev
|
| 406 |
+
num_bytes: 2199.1754385964914
|
| 407 |
+
num_examples: 5
|
| 408 |
+
download_size: 30150
|
| 409 |
+
dataset_size: 57303.3562203159
|
| 410 |
+
- config_name: conceptual_physics
|
| 411 |
+
features:
|
| 412 |
+
- name: question
|
| 413 |
+
dtype: string
|
| 414 |
+
- name: subject
|
| 415 |
+
dtype: string
|
| 416 |
+
- name: choices
|
| 417 |
+
sequence: string
|
| 418 |
+
- name: answer
|
| 419 |
+
dtype:
|
| 420 |
+
class_label:
|
| 421 |
+
names:
|
| 422 |
+
'0': A
|
| 423 |
+
'1': B
|
| 424 |
+
'2': C
|
| 425 |
+
'3': D
|
| 426 |
+
splits:
|
| 427 |
+
- name: test
|
| 428 |
+
num_bytes: 116603.86376584532
|
| 429 |
+
num_examples: 235
|
| 430 |
+
- name: validation
|
| 431 |
+
num_bytes: 12965.76355323318
|
| 432 |
+
num_examples: 26
|
| 433 |
+
- name: dev
|
| 434 |
+
num_bytes: 2199.1754385964914
|
| 435 |
+
num_examples: 5
|
| 436 |
+
download_size: 34968
|
| 437 |
+
dataset_size: 131768.802757675
|
| 438 |
+
- config_name: econometrics
|
| 439 |
+
features:
|
| 440 |
+
- name: question
|
| 441 |
+
dtype: string
|
| 442 |
+
- name: subject
|
| 443 |
+
dtype: string
|
| 444 |
+
- name: choices
|
| 445 |
+
sequence: string
|
| 446 |
+
- name: answer
|
| 447 |
+
dtype:
|
| 448 |
+
class_label:
|
| 449 |
+
names:
|
| 450 |
+
'0': A
|
| 451 |
+
'1': B
|
| 452 |
+
'2': C
|
| 453 |
+
'3': D
|
| 454 |
+
splits:
|
| 455 |
+
- name: test
|
| 456 |
+
num_bytes: 56565.27859279305
|
| 457 |
+
num_examples: 114
|
| 458 |
+
- name: validation
|
| 459 |
+
num_bytes: 5984.198563030699
|
| 460 |
+
num_examples: 12
|
| 461 |
+
- name: dev
|
| 462 |
+
num_bytes: 2199.1754385964914
|
| 463 |
+
num_examples: 5
|
| 464 |
+
download_size: 36040
|
| 465 |
+
dataset_size: 64748.652594420244
|
| 466 |
+
- config_name: electrical_engineering
|
| 467 |
+
features:
|
| 468 |
+
- name: question
|
| 469 |
+
dtype: string
|
| 470 |
+
- name: subject
|
| 471 |
+
dtype: string
|
| 472 |
+
- name: choices
|
| 473 |
+
sequence: string
|
| 474 |
+
- name: answer
|
| 475 |
+
dtype:
|
| 476 |
+
class_label:
|
| 477 |
+
names:
|
| 478 |
+
'0': A
|
| 479 |
+
'1': B
|
| 480 |
+
'2': C
|
| 481 |
+
'3': D
|
| 482 |
+
splits:
|
| 483 |
+
- name: test
|
| 484 |
+
num_bytes: 71947.06487679818
|
| 485 |
+
num_examples: 145
|
| 486 |
+
- name: validation
|
| 487 |
+
num_bytes: 7978.931417374265
|
| 488 |
+
num_examples: 16
|
| 489 |
+
- name: dev
|
| 490 |
+
num_bytes: 2199.1754385964914
|
| 491 |
+
num_examples: 5
|
| 492 |
+
download_size: 26746
|
| 493 |
+
dataset_size: 82125.17173276893
|
| 494 |
+
- config_name: elementary_mathematics
|
| 495 |
+
features:
|
| 496 |
+
- name: question
|
| 497 |
+
dtype: string
|
| 498 |
+
- name: subject
|
| 499 |
+
dtype: string
|
| 500 |
+
- name: choices
|
| 501 |
+
sequence: string
|
| 502 |
+
- name: answer
|
| 503 |
+
dtype:
|
| 504 |
+
class_label:
|
| 505 |
+
names:
|
| 506 |
+
'0': A
|
| 507 |
+
'1': B
|
| 508 |
+
'2': C
|
| 509 |
+
'3': D
|
| 510 |
+
splits:
|
| 511 |
+
- name: test
|
| 512 |
+
num_bytes: 187558.555333998
|
| 513 |
+
num_examples: 378
|
| 514 |
+
- name: validation
|
| 515 |
+
num_bytes: 20446.011757021555
|
| 516 |
+
num_examples: 41
|
| 517 |
+
- name: dev
|
| 518 |
+
num_bytes: 2199.1754385964914
|
| 519 |
+
num_examples: 5
|
| 520 |
+
download_size: 54987
|
| 521 |
+
dataset_size: 210203.74252961605
|
| 522 |
+
- config_name: formal_logic
|
| 523 |
+
features:
|
| 524 |
+
- name: question
|
| 525 |
+
dtype: string
|
| 526 |
+
- name: subject
|
| 527 |
+
dtype: string
|
| 528 |
+
- name: choices
|
| 529 |
+
sequence: string
|
| 530 |
+
- name: answer
|
| 531 |
+
dtype:
|
| 532 |
+
class_label:
|
| 533 |
+
names:
|
| 534 |
+
'0': A
|
| 535 |
+
'1': B
|
| 536 |
+
'2': C
|
| 537 |
+
'3': D
|
| 538 |
+
splits:
|
| 539 |
+
- name: test
|
| 540 |
+
num_bytes: 62519.518444666
|
| 541 |
+
num_examples: 126
|
| 542 |
+
- name: validation
|
| 543 |
+
num_bytes: 6981.5649902024825
|
| 544 |
+
num_examples: 14
|
| 545 |
+
- name: dev
|
| 546 |
+
num_bytes: 2199.1754385964914
|
| 547 |
+
num_examples: 5
|
| 548 |
+
download_size: 32884
|
| 549 |
+
dataset_size: 71700.25887346498
|
| 550 |
+
- config_name: global_facts
|
| 551 |
+
features:
|
| 552 |
+
- name: question
|
| 553 |
+
dtype: string
|
| 554 |
+
- name: subject
|
| 555 |
+
dtype: string
|
| 556 |
+
- name: choices
|
| 557 |
+
sequence: string
|
| 558 |
+
- name: answer
|
| 559 |
+
dtype:
|
| 560 |
+
class_label:
|
| 561 |
+
names:
|
| 562 |
+
'0': A
|
| 563 |
+
'1': B
|
| 564 |
+
'2': C
|
| 565 |
+
'3': D
|
| 566 |
+
splits:
|
| 567 |
+
- name: test
|
| 568 |
+
num_bytes: 49618.6654322746
|
| 569 |
+
num_examples: 100
|
| 570 |
+
- name: validation
|
| 571 |
+
num_bytes: 4986.8321358589155
|
| 572 |
+
num_examples: 10
|
| 573 |
+
- name: dev
|
| 574 |
+
num_bytes: 2199.1754385964914
|
| 575 |
+
num_examples: 5
|
| 576 |
+
download_size: 19258
|
| 577 |
+
dataset_size: 56804.67300673001
|
| 578 |
+
- config_name: high_school_biology
|
| 579 |
+
features:
|
| 580 |
+
- name: question
|
| 581 |
+
dtype: string
|
| 582 |
+
- name: subject
|
| 583 |
+
dtype: string
|
| 584 |
+
- name: choices
|
| 585 |
+
sequence: string
|
| 586 |
+
- name: answer
|
| 587 |
+
dtype:
|
| 588 |
+
class_label:
|
| 589 |
+
names:
|
| 590 |
+
'0': A
|
| 591 |
+
'1': B
|
| 592 |
+
'2': C
|
| 593 |
+
'3': D
|
| 594 |
+
splits:
|
| 595 |
+
- name: test
|
| 596 |
+
num_bytes: 153817.86284005127
|
| 597 |
+
num_examples: 310
|
| 598 |
+
- name: validation
|
| 599 |
+
num_bytes: 15957.86283474853
|
| 600 |
+
num_examples: 32
|
| 601 |
+
- name: dev
|
| 602 |
+
num_bytes: 2199.1754385964914
|
| 603 |
+
num_examples: 5
|
| 604 |
+
download_size: 78216
|
| 605 |
+
dataset_size: 171974.90111339628
|
| 606 |
+
- config_name: high_school_chemistry
|
| 607 |
+
features:
|
| 608 |
+
- name: question
|
| 609 |
+
dtype: string
|
| 610 |
+
- name: subject
|
| 611 |
+
dtype: string
|
| 612 |
+
- name: choices
|
| 613 |
+
sequence: string
|
| 614 |
+
- name: answer
|
| 615 |
+
dtype:
|
| 616 |
+
class_label:
|
| 617 |
+
names:
|
| 618 |
+
'0': A
|
| 619 |
+
'1': B
|
| 620 |
+
'2': C
|
| 621 |
+
'3': D
|
| 622 |
+
splits:
|
| 623 |
+
- name: test
|
| 624 |
+
num_bytes: 100725.89082751745
|
| 625 |
+
num_examples: 203
|
| 626 |
+
- name: validation
|
| 627 |
+
num_bytes: 10971.030698889615
|
| 628 |
+
num_examples: 22
|
| 629 |
+
- name: dev
|
| 630 |
+
num_bytes: 2199.1754385964914
|
| 631 |
+
num_examples: 5
|
| 632 |
+
download_size: 45799
|
| 633 |
+
dataset_size: 113896.09696500355
|
| 634 |
+
- config_name: high_school_computer_science
|
| 635 |
+
features:
|
| 636 |
+
- name: question
|
| 637 |
+
dtype: string
|
| 638 |
+
- name: subject
|
| 639 |
+
dtype: string
|
| 640 |
+
- name: choices
|
| 641 |
+
sequence: string
|
| 642 |
+
- name: answer
|
| 643 |
+
dtype:
|
| 644 |
+
class_label:
|
| 645 |
+
names:
|
| 646 |
+
'0': A
|
| 647 |
+
'1': B
|
| 648 |
+
'2': C
|
| 649 |
+
'3': D
|
| 650 |
+
splits:
|
| 651 |
+
- name: test
|
| 652 |
+
num_bytes: 49618.6654322746
|
| 653 |
+
num_examples: 100
|
| 654 |
+
- name: validation
|
| 655 |
+
num_bytes: 4488.148922273024
|
| 656 |
+
num_examples: 9
|
| 657 |
+
- name: dev
|
| 658 |
+
num_bytes: 2199.1754385964914
|
| 659 |
+
num_examples: 5
|
| 660 |
+
download_size: 39072
|
| 661 |
+
dataset_size: 56305.989793144116
|
| 662 |
+
- config_name: high_school_european_history
|
| 663 |
+
features:
|
| 664 |
+
- name: question
|
| 665 |
+
dtype: string
|
| 666 |
+
- name: subject
|
| 667 |
+
dtype: string
|
| 668 |
+
- name: choices
|
| 669 |
+
sequence: string
|
| 670 |
+
- name: answer
|
| 671 |
+
dtype:
|
| 672 |
+
class_label:
|
| 673 |
+
names:
|
| 674 |
+
'0': A
|
| 675 |
+
'1': B
|
| 676 |
+
'2': C
|
| 677 |
+
'3': D
|
| 678 |
+
splits:
|
| 679 |
+
- name: test
|
| 680 |
+
num_bytes: 81870.79796325309
|
| 681 |
+
num_examples: 165
|
| 682 |
+
- name: validation
|
| 683 |
+
num_bytes: 8976.297844546049
|
| 684 |
+
num_examples: 18
|
| 685 |
+
- name: dev
|
| 686 |
+
num_bytes: 2199.1754385964914
|
| 687 |
+
num_examples: 5
|
| 688 |
+
download_size: 196270
|
| 689 |
+
dataset_size: 93046.27124639563
|
| 690 |
+
- config_name: high_school_geography
|
| 691 |
+
features:
|
| 692 |
+
- name: question
|
| 693 |
+
dtype: string
|
| 694 |
+
- name: subject
|
| 695 |
+
dtype: string
|
| 696 |
+
- name: choices
|
| 697 |
+
sequence: string
|
| 698 |
+
- name: answer
|
| 699 |
+
dtype:
|
| 700 |
+
class_label:
|
| 701 |
+
names:
|
| 702 |
+
'0': A
|
| 703 |
+
'1': B
|
| 704 |
+
'2': C
|
| 705 |
+
'3': D
|
| 706 |
+
splits:
|
| 707 |
+
- name: test
|
| 708 |
+
num_bytes: 98244.95755590372
|
| 709 |
+
num_examples: 198
|
| 710 |
+
- name: validation
|
| 711 |
+
num_bytes: 10971.030698889615
|
| 712 |
+
num_examples: 22
|
| 713 |
+
- name: dev
|
| 714 |
+
num_bytes: 2199.1754385964914
|
| 715 |
+
num_examples: 5
|
| 716 |
+
download_size: 38255
|
| 717 |
+
dataset_size: 111415.16369338983
|
| 718 |
+
- config_name: high_school_government_and_politics
|
| 719 |
+
features:
|
| 720 |
+
- name: question
|
| 721 |
+
dtype: string
|
| 722 |
+
- name: subject
|
| 723 |
+
dtype: string
|
| 724 |
+
- name: choices
|
| 725 |
+
sequence: string
|
| 726 |
+
- name: answer
|
| 727 |
+
dtype:
|
| 728 |
+
class_label:
|
| 729 |
+
names:
|
| 730 |
+
'0': A
|
| 731 |
+
'1': B
|
| 732 |
+
'2': C
|
| 733 |
+
'3': D
|
| 734 |
+
splits:
|
| 735 |
+
- name: test
|
| 736 |
+
num_bytes: 95764.02428428999
|
| 737 |
+
num_examples: 193
|
| 738 |
+
- name: validation
|
| 739 |
+
num_bytes: 10472.347485303722
|
| 740 |
+
num_examples: 21
|
| 741 |
+
- name: dev
|
| 742 |
+
num_bytes: 2199.1754385964914
|
| 743 |
+
num_examples: 5
|
| 744 |
+
download_size: 52963
|
| 745 |
+
dataset_size: 108435.5472081902
|
| 746 |
+
- config_name: high_school_macroeconomics
|
| 747 |
+
features:
|
| 748 |
+
- name: question
|
| 749 |
+
dtype: string
|
| 750 |
+
- name: subject
|
| 751 |
+
dtype: string
|
| 752 |
+
- name: choices
|
| 753 |
+
sequence: string
|
| 754 |
+
- name: answer
|
| 755 |
+
dtype:
|
| 756 |
+
class_label:
|
| 757 |
+
names:
|
| 758 |
+
'0': A
|
| 759 |
+
'1': B
|
| 760 |
+
'2': C
|
| 761 |
+
'3': D
|
| 762 |
+
splits:
|
| 763 |
+
- name: test
|
| 764 |
+
num_bytes: 193512.79518587096
|
| 765 |
+
num_examples: 390
|
| 766 |
+
- name: validation
|
| 767 |
+
num_bytes: 21443.378184193338
|
| 768 |
+
num_examples: 43
|
| 769 |
+
- name: dev
|
| 770 |
+
num_bytes: 2199.1754385964914
|
| 771 |
+
num_examples: 5
|
| 772 |
+
download_size: 68758
|
| 773 |
+
dataset_size: 217155.34880866078
|
| 774 |
+
- config_name: high_school_mathematics
|
| 775 |
+
features:
|
| 776 |
+
- name: question
|
| 777 |
+
dtype: string
|
| 778 |
+
- name: subject
|
| 779 |
+
dtype: string
|
| 780 |
+
- name: choices
|
| 781 |
+
sequence: string
|
| 782 |
+
- name: answer
|
| 783 |
+
dtype:
|
| 784 |
+
class_label:
|
| 785 |
+
names:
|
| 786 |
+
'0': A
|
| 787 |
+
'1': B
|
| 788 |
+
'2': C
|
| 789 |
+
'3': D
|
| 790 |
+
splits:
|
| 791 |
+
- name: test
|
| 792 |
+
num_bytes: 133970.39666714144
|
| 793 |
+
num_examples: 270
|
| 794 |
+
- name: validation
|
| 795 |
+
num_bytes: 14461.813193990856
|
| 796 |
+
num_examples: 29
|
| 797 |
+
- name: dev
|
| 798 |
+
num_bytes: 2199.1754385964914
|
| 799 |
+
num_examples: 5
|
| 800 |
+
download_size: 45210
|
| 801 |
+
dataset_size: 150631.38529972878
|
| 802 |
+
- config_name: high_school_microeconomics
|
| 803 |
+
features:
|
| 804 |
+
- name: question
|
| 805 |
+
dtype: string
|
| 806 |
+
- name: subject
|
| 807 |
+
dtype: string
|
| 808 |
+
- name: choices
|
| 809 |
+
sequence: string
|
| 810 |
+
- name: answer
|
| 811 |
+
dtype:
|
| 812 |
+
class_label:
|
| 813 |
+
names:
|
| 814 |
+
'0': A
|
| 815 |
+
'1': B
|
| 816 |
+
'2': C
|
| 817 |
+
'3': D
|
| 818 |
+
splits:
|
| 819 |
+
- name: test
|
| 820 |
+
num_bytes: 118092.42372881356
|
| 821 |
+
num_examples: 238
|
| 822 |
+
- name: validation
|
| 823 |
+
num_bytes: 12965.76355323318
|
| 824 |
+
num_examples: 26
|
| 825 |
+
- name: dev
|
| 826 |
+
num_bytes: 2199.1754385964914
|
| 827 |
+
num_examples: 5
|
| 828 |
+
download_size: 49885
|
| 829 |
+
dataset_size: 133257.36272064323
|
| 830 |
+
- config_name: high_school_physics
|
| 831 |
+
features:
|
| 832 |
+
- name: question
|
| 833 |
+
dtype: string
|
| 834 |
+
- name: subject
|
| 835 |
+
dtype: string
|
| 836 |
+
- name: choices
|
| 837 |
+
sequence: string
|
| 838 |
+
- name: answer
|
| 839 |
+
dtype:
|
| 840 |
+
class_label:
|
| 841 |
+
names:
|
| 842 |
+
'0': A
|
| 843 |
+
'1': B
|
| 844 |
+
'2': C
|
| 845 |
+
'3': D
|
| 846 |
+
splits:
|
| 847 |
+
- name: test
|
| 848 |
+
num_bytes: 74924.18480273466
|
| 849 |
+
num_examples: 151
|
| 850 |
+
- name: validation
|
| 851 |
+
num_bytes: 8477.614630960157
|
| 852 |
+
num_examples: 17
|
| 853 |
+
- name: dev
|
| 854 |
+
num_bytes: 2199.1754385964914
|
| 855 |
+
num_examples: 5
|
| 856 |
+
download_size: 45483
|
| 857 |
+
dataset_size: 85600.9748722913
|
| 858 |
+
- config_name: high_school_psychology
|
| 859 |
+
features:
|
| 860 |
+
- name: question
|
| 861 |
+
dtype: string
|
| 862 |
+
- name: subject
|
| 863 |
+
dtype: string
|
| 864 |
+
- name: choices
|
| 865 |
+
sequence: string
|
| 866 |
+
- name: answer
|
| 867 |
+
dtype:
|
| 868 |
+
class_label:
|
| 869 |
+
names:
|
| 870 |
+
'0': A
|
| 871 |
+
'1': B
|
| 872 |
+
'2': C
|
| 873 |
+
'3': D
|
| 874 |
+
splits:
|
| 875 |
+
- name: test
|
| 876 |
+
num_bytes: 270421.7266058966
|
| 877 |
+
num_examples: 545
|
| 878 |
+
- name: validation
|
| 879 |
+
num_bytes: 29920.992815153495
|
| 880 |
+
num_examples: 60
|
| 881 |
+
- name: dev
|
| 882 |
+
num_bytes: 2199.1754385964914
|
| 883 |
+
num_examples: 5
|
| 884 |
+
download_size: 113158
|
| 885 |
+
dataset_size: 302541.8948596466
|
| 886 |
+
- config_name: high_school_statistics
|
| 887 |
+
features:
|
| 888 |
+
- name: question
|
| 889 |
+
dtype: string
|
| 890 |
+
- name: subject
|
| 891 |
+
dtype: string
|
| 892 |
+
- name: choices
|
| 893 |
+
sequence: string
|
| 894 |
+
- name: answer
|
| 895 |
+
dtype:
|
| 896 |
+
class_label:
|
| 897 |
+
names:
|
| 898 |
+
'0': A
|
| 899 |
+
'1': B
|
| 900 |
+
'2': C
|
| 901 |
+
'3': D
|
| 902 |
+
splits:
|
| 903 |
+
- name: test
|
| 904 |
+
num_bytes: 107176.31733371314
|
| 905 |
+
num_examples: 216
|
| 906 |
+
- name: validation
|
| 907 |
+
num_bytes: 11469.713912475507
|
| 908 |
+
num_examples: 23
|
| 909 |
+
- name: dev
|
| 910 |
+
num_bytes: 2199.1754385964914
|
| 911 |
+
num_examples: 5
|
| 912 |
+
download_size: 74924
|
| 913 |
+
dataset_size: 120845.20668478514
|
| 914 |
+
- config_name: high_school_us_history
|
| 915 |
+
features:
|
| 916 |
+
- name: question
|
| 917 |
+
dtype: string
|
| 918 |
+
- name: subject
|
| 919 |
+
dtype: string
|
| 920 |
+
- name: choices
|
| 921 |
+
sequence: string
|
| 922 |
+
- name: answer
|
| 923 |
+
dtype:
|
| 924 |
+
class_label:
|
| 925 |
+
names:
|
| 926 |
+
'0': A
|
| 927 |
+
'1': B
|
| 928 |
+
'2': C
|
| 929 |
+
'3': D
|
| 930 |
+
splits:
|
| 931 |
+
- name: test
|
| 932 |
+
num_bytes: 101222.0774818402
|
| 933 |
+
num_examples: 204
|
| 934 |
+
- name: validation
|
| 935 |
+
num_bytes: 10971.030698889615
|
| 936 |
+
num_examples: 22
|
| 937 |
+
- name: dev
|
| 938 |
+
num_bytes: 2199.1754385964914
|
| 939 |
+
num_examples: 5
|
| 940 |
+
download_size: 200043
|
| 941 |
+
dataset_size: 114392.2836193263
|
| 942 |
+
- config_name: high_school_world_history
|
| 943 |
+
features:
|
| 944 |
+
- name: question
|
| 945 |
+
dtype: string
|
| 946 |
+
- name: subject
|
| 947 |
+
dtype: string
|
| 948 |
+
- name: choices
|
| 949 |
+
sequence: string
|
| 950 |
+
- name: answer
|
| 951 |
+
dtype:
|
| 952 |
+
class_label:
|
| 953 |
+
names:
|
| 954 |
+
'0': A
|
| 955 |
+
'1': B
|
| 956 |
+
'2': C
|
| 957 |
+
'3': D
|
| 958 |
+
splits:
|
| 959 |
+
- name: test
|
| 960 |
+
num_bytes: 117596.23707449081
|
| 961 |
+
num_examples: 237
|
| 962 |
+
- name: validation
|
| 963 |
+
num_bytes: 12965.76355323318
|
| 964 |
+
num_examples: 26
|
| 965 |
+
- name: dev
|
| 966 |
+
num_bytes: 2199.1754385964914
|
| 967 |
+
num_examples: 5
|
| 968 |
+
download_size: 250302
|
| 969 |
+
dataset_size: 132761.17606632048
|
| 970 |
+
- config_name: human_aging
|
| 971 |
+
features:
|
| 972 |
+
- name: question
|
| 973 |
+
dtype: string
|
| 974 |
+
- name: subject
|
| 975 |
+
dtype: string
|
| 976 |
+
- name: choices
|
| 977 |
+
sequence: string
|
| 978 |
+
- name: answer
|
| 979 |
+
dtype:
|
| 980 |
+
class_label:
|
| 981 |
+
names:
|
| 982 |
+
'0': A
|
| 983 |
+
'1': B
|
| 984 |
+
'2': C
|
| 985 |
+
'3': D
|
| 986 |
+
splits:
|
| 987 |
+
- name: test
|
| 988 |
+
num_bytes: 110649.62391397236
|
| 989 |
+
num_examples: 223
|
| 990 |
+
- name: validation
|
| 991 |
+
num_bytes: 11469.713912475507
|
| 992 |
+
num_examples: 23
|
| 993 |
+
- name: dev
|
| 994 |
+
num_bytes: 2199.1754385964914
|
| 995 |
+
num_examples: 5
|
| 996 |
+
download_size: 41196
|
| 997 |
+
dataset_size: 124318.51326504436
|
| 998 |
+
- config_name: human_sexuality
|
| 999 |
+
features:
|
| 1000 |
+
- name: question
|
| 1001 |
+
dtype: string
|
| 1002 |
+
- name: subject
|
| 1003 |
+
dtype: string
|
| 1004 |
+
- name: choices
|
| 1005 |
+
sequence: string
|
| 1006 |
+
- name: answer
|
| 1007 |
+
dtype:
|
| 1008 |
+
class_label:
|
| 1009 |
+
names:
|
| 1010 |
+
'0': A
|
| 1011 |
+
'1': B
|
| 1012 |
+
'2': C
|
| 1013 |
+
'3': D
|
| 1014 |
+
splits:
|
| 1015 |
+
- name: test
|
| 1016 |
+
num_bytes: 65000.451716279735
|
| 1017 |
+
num_examples: 131
|
| 1018 |
+
- name: validation
|
| 1019 |
+
num_bytes: 5984.198563030699
|
| 1020 |
+
num_examples: 12
|
| 1021 |
+
- name: dev
|
| 1022 |
+
num_bytes: 2199.1754385964914
|
| 1023 |
+
num_examples: 5
|
| 1024 |
+
download_size: 32533
|
| 1025 |
+
dataset_size: 73183.82571790692
|
| 1026 |
+
- config_name: international_law
|
| 1027 |
+
features:
|
| 1028 |
+
- name: question
|
| 1029 |
+
dtype: string
|
| 1030 |
+
- name: subject
|
| 1031 |
+
dtype: string
|
| 1032 |
+
- name: choices
|
| 1033 |
+
sequence: string
|
| 1034 |
+
- name: answer
|
| 1035 |
+
dtype:
|
| 1036 |
+
class_label:
|
| 1037 |
+
names:
|
| 1038 |
+
'0': A
|
| 1039 |
+
'1': B
|
| 1040 |
+
'2': C
|
| 1041 |
+
'3': D
|
| 1042 |
+
splits:
|
| 1043 |
+
- name: test
|
| 1044 |
+
num_bytes: 60038.58517305227
|
| 1045 |
+
num_examples: 121
|
| 1046 |
+
- name: validation
|
| 1047 |
+
num_bytes: 6482.88177661659
|
| 1048 |
+
num_examples: 13
|
| 1049 |
+
- name: dev
|
| 1050 |
+
num_bytes: 2199.1754385964914
|
| 1051 |
+
num_examples: 5
|
| 1052 |
+
download_size: 41592
|
| 1053 |
+
dataset_size: 68720.64238826535
|
| 1054 |
+
- config_name: jurisprudence
|
| 1055 |
+
features:
|
| 1056 |
+
- name: question
|
| 1057 |
+
dtype: string
|
| 1058 |
+
- name: subject
|
| 1059 |
+
dtype: string
|
| 1060 |
+
- name: choices
|
| 1061 |
+
sequence: string
|
| 1062 |
+
- name: answer
|
| 1063 |
+
dtype:
|
| 1064 |
+
class_label:
|
| 1065 |
+
names:
|
| 1066 |
+
'0': A
|
| 1067 |
+
'1': B
|
| 1068 |
+
'2': C
|
| 1069 |
+
'3': D
|
| 1070 |
+
splits:
|
| 1071 |
+
- name: test
|
| 1072 |
+
num_bytes: 53588.15866685657
|
| 1073 |
+
num_examples: 108
|
| 1074 |
+
- name: validation
|
| 1075 |
+
num_bytes: 5485.515349444808
|
| 1076 |
+
num_examples: 11
|
| 1077 |
+
- name: dev
|
| 1078 |
+
num_bytes: 2199.1754385964914
|
| 1079 |
+
num_examples: 5
|
| 1080 |
+
download_size: 33578
|
| 1081 |
+
dataset_size: 61272.84945489787
|
| 1082 |
+
- config_name: logical_fallacies
|
| 1083 |
+
features:
|
| 1084 |
+
- name: question
|
| 1085 |
+
dtype: string
|
| 1086 |
+
- name: subject
|
| 1087 |
+
dtype: string
|
| 1088 |
+
- name: choices
|
| 1089 |
+
sequence: string
|
| 1090 |
+
- name: answer
|
| 1091 |
+
dtype:
|
| 1092 |
+
class_label:
|
| 1093 |
+
names:
|
| 1094 |
+
'0': A
|
| 1095 |
+
'1': B
|
| 1096 |
+
'2': C
|
| 1097 |
+
'3': D
|
| 1098 |
+
splits:
|
| 1099 |
+
- name: test
|
| 1100 |
+
num_bytes: 80878.4246546076
|
| 1101 |
+
num_examples: 163
|
| 1102 |
+
- name: validation
|
| 1103 |
+
num_bytes: 8976.297844546049
|
| 1104 |
+
num_examples: 18
|
| 1105 |
+
- name: dev
|
| 1106 |
+
num_bytes: 2199.1754385964914
|
| 1107 |
+
num_examples: 5
|
| 1108 |
+
download_size: 33669
|
| 1109 |
+
dataset_size: 92053.89793775014
|
| 1110 |
+
- config_name: machine_learning
|
| 1111 |
+
features:
|
| 1112 |
+
- name: question
|
| 1113 |
+
dtype: string
|
| 1114 |
+
- name: subject
|
| 1115 |
+
dtype: string
|
| 1116 |
+
- name: choices
|
| 1117 |
+
sequence: string
|
| 1118 |
+
- name: answer
|
| 1119 |
+
dtype:
|
| 1120 |
+
class_label:
|
| 1121 |
+
names:
|
| 1122 |
+
'0': A
|
| 1123 |
+
'1': B
|
| 1124 |
+
'2': C
|
| 1125 |
+
'3': D
|
| 1126 |
+
splits:
|
| 1127 |
+
- name: test
|
| 1128 |
+
num_bytes: 55572.90528414756
|
| 1129 |
+
num_examples: 112
|
| 1130 |
+
- name: validation
|
| 1131 |
+
num_bytes: 5485.515349444808
|
| 1132 |
+
num_examples: 11
|
| 1133 |
+
- name: dev
|
| 1134 |
+
num_bytes: 2199.1754385964914
|
| 1135 |
+
num_examples: 5
|
| 1136 |
+
download_size: 31121
|
| 1137 |
+
dataset_size: 63257.596072188855
|
| 1138 |
+
- config_name: management
|
| 1139 |
+
features:
|
| 1140 |
+
- name: question
|
| 1141 |
+
dtype: string
|
| 1142 |
+
- name: subject
|
| 1143 |
+
dtype: string
|
| 1144 |
+
- name: choices
|
| 1145 |
+
sequence: string
|
| 1146 |
+
- name: answer
|
| 1147 |
+
dtype:
|
| 1148 |
+
class_label:
|
| 1149 |
+
names:
|
| 1150 |
+
'0': A
|
| 1151 |
+
'1': B
|
| 1152 |
+
'2': C
|
| 1153 |
+
'3': D
|
| 1154 |
+
splits:
|
| 1155 |
+
- name: test
|
| 1156 |
+
num_bytes: 51107.225395242844
|
| 1157 |
+
num_examples: 103
|
| 1158 |
+
- name: validation
|
| 1159 |
+
num_bytes: 5485.515349444808
|
| 1160 |
+
num_examples: 11
|
| 1161 |
+
- name: dev
|
| 1162 |
+
num_bytes: 2199.1754385964914
|
| 1163 |
+
num_examples: 5
|
| 1164 |
+
download_size: 22828
|
| 1165 |
+
dataset_size: 58791.91618328414
|
| 1166 |
+
- config_name: marketing
|
| 1167 |
+
features:
|
| 1168 |
+
- name: question
|
| 1169 |
+
dtype: string
|
| 1170 |
+
- name: subject
|
| 1171 |
+
dtype: string
|
| 1172 |
+
- name: choices
|
| 1173 |
+
sequence: string
|
| 1174 |
+
- name: answer
|
| 1175 |
+
dtype:
|
| 1176 |
+
class_label:
|
| 1177 |
+
names:
|
| 1178 |
+
'0': A
|
| 1179 |
+
'1': B
|
| 1180 |
+
'2': C
|
| 1181 |
+
'3': D
|
| 1182 |
+
splits:
|
| 1183 |
+
- name: test
|
| 1184 |
+
num_bytes: 116107.67711152257
|
| 1185 |
+
num_examples: 234
|
| 1186 |
+
- name: validation
|
| 1187 |
+
num_bytes: 12467.08033964729
|
| 1188 |
+
num_examples: 25
|
| 1189 |
+
- name: dev
|
| 1190 |
+
num_bytes: 2199.1754385964914
|
| 1191 |
+
num_examples: 5
|
| 1192 |
+
download_size: 49747
|
| 1193 |
+
dataset_size: 130773.93288976635
|
| 1194 |
+
- config_name: medical_genetics
|
| 1195 |
+
features:
|
| 1196 |
+
- name: question
|
| 1197 |
+
dtype: string
|
| 1198 |
+
- name: subject
|
| 1199 |
+
dtype: string
|
| 1200 |
+
- name: choices
|
| 1201 |
+
sequence: string
|
| 1202 |
+
- name: answer
|
| 1203 |
+
dtype:
|
| 1204 |
+
class_label:
|
| 1205 |
+
names:
|
| 1206 |
+
'0': A
|
| 1207 |
+
'1': B
|
| 1208 |
+
'2': C
|
| 1209 |
+
'3': D
|
| 1210 |
+
splits:
|
| 1211 |
+
- name: test
|
| 1212 |
+
num_bytes: 49618.6654322746
|
| 1213 |
+
num_examples: 100
|
| 1214 |
+
- name: validation
|
| 1215 |
+
num_bytes: 5485.515349444808
|
| 1216 |
+
num_examples: 11
|
| 1217 |
+
- name: dev
|
| 1218 |
+
num_bytes: 2199.1754385964914
|
| 1219 |
+
num_examples: 5
|
| 1220 |
+
download_size: 25775
|
| 1221 |
+
dataset_size: 57303.3562203159
|
| 1222 |
+
- config_name: miscellaneous
|
| 1223 |
+
features:
|
| 1224 |
+
- name: question
|
| 1225 |
+
dtype: string
|
| 1226 |
+
- name: subject
|
| 1227 |
+
dtype: string
|
| 1228 |
+
- name: choices
|
| 1229 |
+
sequence: string
|
| 1230 |
+
- name: answer
|
| 1231 |
+
dtype:
|
| 1232 |
+
class_label:
|
| 1233 |
+
names:
|
| 1234 |
+
'0': A
|
| 1235 |
+
'1': B
|
| 1236 |
+
'2': C
|
| 1237 |
+
'3': D
|
| 1238 |
+
splits:
|
| 1239 |
+
- name: test
|
| 1240 |
+
num_bytes: 388514.15033471014
|
| 1241 |
+
num_examples: 783
|
| 1242 |
+
- name: validation
|
| 1243 |
+
num_bytes: 42886.756368386676
|
| 1244 |
+
num_examples: 86
|
| 1245 |
+
- name: dev
|
| 1246 |
+
num_bytes: 2199.1754385964914
|
| 1247 |
+
num_examples: 5
|
| 1248 |
+
download_size: 115097
|
| 1249 |
+
dataset_size: 433600.08214169333
|
| 1250 |
+
- config_name: moral_disputes
|
| 1251 |
+
features:
|
| 1252 |
+
- name: question
|
| 1253 |
+
dtype: string
|
| 1254 |
+
- name: subject
|
| 1255 |
+
dtype: string
|
| 1256 |
+
- name: choices
|
| 1257 |
+
sequence: string
|
| 1258 |
+
- name: answer
|
| 1259 |
+
dtype:
|
| 1260 |
+
class_label:
|
| 1261 |
+
names:
|
| 1262 |
+
'0': A
|
| 1263 |
+
'1': B
|
| 1264 |
+
'2': C
|
| 1265 |
+
'3': D
|
| 1266 |
+
splits:
|
| 1267 |
+
- name: test
|
| 1268 |
+
num_bytes: 171680.58239567012
|
| 1269 |
+
num_examples: 346
|
| 1270 |
+
- name: validation
|
| 1271 |
+
num_bytes: 18949.96211626388
|
| 1272 |
+
num_examples: 38
|
| 1273 |
+
- name: dev
|
| 1274 |
+
num_bytes: 2199.1754385964914
|
| 1275 |
+
num_examples: 5
|
| 1276 |
+
download_size: 76043
|
| 1277 |
+
dataset_size: 192829.71995053047
|
| 1278 |
+
- config_name: moral_scenarios
|
| 1279 |
+
features:
|
| 1280 |
+
- name: question
|
| 1281 |
+
dtype: string
|
| 1282 |
+
- name: subject
|
| 1283 |
+
dtype: string
|
| 1284 |
+
- name: choices
|
| 1285 |
+
sequence: string
|
| 1286 |
+
- name: answer
|
| 1287 |
+
dtype:
|
| 1288 |
+
class_label:
|
| 1289 |
+
names:
|
| 1290 |
+
'0': A
|
| 1291 |
+
'1': B
|
| 1292 |
+
'2': C
|
| 1293 |
+
'3': D
|
| 1294 |
+
splits:
|
| 1295 |
+
- name: test
|
| 1296 |
+
num_bytes: 444087.05561885773
|
| 1297 |
+
num_examples: 895
|
| 1298 |
+
- name: validation
|
| 1299 |
+
num_bytes: 49868.32135858916
|
| 1300 |
+
num_examples: 100
|
| 1301 |
+
- name: dev
|
| 1302 |
+
num_bytes: 2199.1754385964914
|
| 1303 |
+
num_examples: 5
|
| 1304 |
+
download_size: 109869
|
| 1305 |
+
dataset_size: 496154.5524160434
|
| 1306 |
+
- config_name: nutrition
|
| 1307 |
+
features:
|
| 1308 |
+
- name: question
|
| 1309 |
+
dtype: string
|
| 1310 |
+
- name: subject
|
| 1311 |
+
dtype: string
|
| 1312 |
+
- name: choices
|
| 1313 |
+
sequence: string
|
| 1314 |
+
- name: answer
|
| 1315 |
+
dtype:
|
| 1316 |
+
class_label:
|
| 1317 |
+
names:
|
| 1318 |
+
'0': A
|
| 1319 |
+
'1': B
|
| 1320 |
+
'2': C
|
| 1321 |
+
'3': D
|
| 1322 |
+
splits:
|
| 1323 |
+
- name: test
|
| 1324 |
+
num_bytes: 151833.1162227603
|
| 1325 |
+
num_examples: 306
|
| 1326 |
+
- name: validation
|
| 1327 |
+
num_bytes: 16456.54604833442
|
| 1328 |
+
num_examples: 33
|
| 1329 |
+
- name: dev
|
| 1330 |
+
num_bytes: 2199.1754385964914
|
| 1331 |
+
num_examples: 5
|
| 1332 |
+
download_size: 69050
|
| 1333 |
+
dataset_size: 170488.8377096912
|
| 1334 |
+
- config_name: philosophy
|
| 1335 |
+
features:
|
| 1336 |
+
- name: question
|
| 1337 |
+
dtype: string
|
| 1338 |
+
- name: subject
|
| 1339 |
+
dtype: string
|
| 1340 |
+
- name: choices
|
| 1341 |
+
sequence: string
|
| 1342 |
+
- name: answer
|
| 1343 |
+
dtype:
|
| 1344 |
+
class_label:
|
| 1345 |
+
names:
|
| 1346 |
+
'0': A
|
| 1347 |
+
'1': B
|
| 1348 |
+
'2': C
|
| 1349 |
+
'3': D
|
| 1350 |
+
splits:
|
| 1351 |
+
- name: test
|
| 1352 |
+
num_bytes: 154314.04949437402
|
| 1353 |
+
num_examples: 311
|
| 1354 |
+
- name: validation
|
| 1355 |
+
num_bytes: 16955.229261920314
|
| 1356 |
+
num_examples: 34
|
| 1357 |
+
- name: dev
|
| 1358 |
+
num_bytes: 2199.1754385964914
|
| 1359 |
+
num_examples: 5
|
| 1360 |
+
download_size: 61912
|
| 1361 |
+
dataset_size: 173468.45419489083
|
| 1362 |
+
- config_name: prehistory
|
| 1363 |
+
features:
|
| 1364 |
+
- name: question
|
| 1365 |
+
dtype: string
|
| 1366 |
+
- name: subject
|
| 1367 |
+
dtype: string
|
| 1368 |
+
- name: choices
|
| 1369 |
+
sequence: string
|
| 1370 |
+
- name: answer
|
| 1371 |
+
dtype:
|
| 1372 |
+
class_label:
|
| 1373 |
+
names:
|
| 1374 |
+
'0': A
|
| 1375 |
+
'1': B
|
| 1376 |
+
'2': C
|
| 1377 |
+
'3': D
|
| 1378 |
+
splits:
|
| 1379 |
+
- name: test
|
| 1380 |
+
num_bytes: 160764.47600056973
|
| 1381 |
+
num_examples: 324
|
| 1382 |
+
- name: validation
|
| 1383 |
+
num_bytes: 17453.912475506204
|
| 1384 |
+
num_examples: 35
|
| 1385 |
+
- name: dev
|
| 1386 |
+
num_bytes: 2199.1754385964914
|
| 1387 |
+
num_examples: 5
|
| 1388 |
+
download_size: 68826
|
| 1389 |
+
dataset_size: 180417.5639146724
|
| 1390 |
+
- config_name: professional_accounting
|
| 1391 |
+
features:
|
| 1392 |
+
- name: question
|
| 1393 |
+
dtype: string
|
| 1394 |
+
- name: subject
|
| 1395 |
+
dtype: string
|
| 1396 |
+
- name: choices
|
| 1397 |
+
sequence: string
|
| 1398 |
+
- name: answer
|
| 1399 |
+
dtype:
|
| 1400 |
+
class_label:
|
| 1401 |
+
names:
|
| 1402 |
+
'0': A
|
| 1403 |
+
'1': B
|
| 1404 |
+
'2': C
|
| 1405 |
+
'3': D
|
| 1406 |
+
splits:
|
| 1407 |
+
- name: test
|
| 1408 |
+
num_bytes: 139924.6365190144
|
| 1409 |
+
num_examples: 282
|
| 1410 |
+
- name: validation
|
| 1411 |
+
num_bytes: 15459.179621162639
|
| 1412 |
+
num_examples: 31
|
| 1413 |
+
- name: dev
|
| 1414 |
+
num_bytes: 2199.1754385964914
|
| 1415 |
+
num_examples: 5
|
| 1416 |
+
download_size: 87297
|
| 1417 |
+
dataset_size: 157582.99157877354
|
| 1418 |
+
- config_name: professional_law
|
| 1419 |
+
features:
|
| 1420 |
+
- name: question
|
| 1421 |
+
dtype: string
|
| 1422 |
+
- name: subject
|
| 1423 |
+
dtype: string
|
| 1424 |
+
- name: choices
|
| 1425 |
+
sequence: string
|
| 1426 |
+
- name: answer
|
| 1427 |
+
dtype:
|
| 1428 |
+
class_label:
|
| 1429 |
+
names:
|
| 1430 |
+
'0': A
|
| 1431 |
+
'1': B
|
| 1432 |
+
'2': C
|
| 1433 |
+
'3': D
|
| 1434 |
+
splits:
|
| 1435 |
+
- name: test
|
| 1436 |
+
num_bytes: 761150.3277310925
|
| 1437 |
+
num_examples: 1534
|
| 1438 |
+
- name: validation
|
| 1439 |
+
num_bytes: 84776.14630960157
|
| 1440 |
+
num_examples: 170
|
| 1441 |
+
- name: dev
|
| 1442 |
+
num_bytes: 2199.1754385964914
|
| 1443 |
+
num_examples: 5
|
| 1444 |
+
download_size: 1167828
|
| 1445 |
+
dataset_size: 848125.6494792906
|
| 1446 |
+
- config_name: professional_medicine
|
| 1447 |
+
features:
|
| 1448 |
+
- name: question
|
| 1449 |
+
dtype: string
|
| 1450 |
+
- name: subject
|
| 1451 |
+
dtype: string
|
| 1452 |
+
- name: choices
|
| 1453 |
+
sequence: string
|
| 1454 |
+
- name: answer
|
| 1455 |
+
dtype:
|
| 1456 |
+
class_label:
|
| 1457 |
+
names:
|
| 1458 |
+
'0': A
|
| 1459 |
+
'1': B
|
| 1460 |
+
'2': C
|
| 1461 |
+
'3': D
|
| 1462 |
+
splits:
|
| 1463 |
+
- name: test
|
| 1464 |
+
num_bytes: 134962.7699757869
|
| 1465 |
+
num_examples: 272
|
| 1466 |
+
- name: validation
|
| 1467 |
+
num_bytes: 15459.179621162639
|
| 1468 |
+
num_examples: 31
|
| 1469 |
+
- name: dev
|
| 1470 |
+
num_bytes: 2199.1754385964914
|
| 1471 |
+
num_examples: 5
|
| 1472 |
+
download_size: 153242
|
| 1473 |
+
dataset_size: 152621.12503554605
|
| 1474 |
+
- config_name: professional_psychology
|
| 1475 |
+
features:
|
| 1476 |
+
- name: question
|
| 1477 |
+
dtype: string
|
| 1478 |
+
- name: subject
|
| 1479 |
+
dtype: string
|
| 1480 |
+
- name: choices
|
| 1481 |
+
sequence: string
|
| 1482 |
+
- name: answer
|
| 1483 |
+
dtype:
|
| 1484 |
+
class_label:
|
| 1485 |
+
names:
|
| 1486 |
+
'0': A
|
| 1487 |
+
'1': B
|
| 1488 |
+
'2': C
|
| 1489 |
+
'3': D
|
| 1490 |
+
splits:
|
| 1491 |
+
- name: test
|
| 1492 |
+
num_bytes: 303666.2324455206
|
| 1493 |
+
num_examples: 612
|
| 1494 |
+
- name: validation
|
| 1495 |
+
num_bytes: 34409.14173742652
|
| 1496 |
+
num_examples: 69
|
| 1497 |
+
- name: dev
|
| 1498 |
+
num_bytes: 2199.1754385964914
|
| 1499 |
+
num_examples: 5
|
| 1500 |
+
download_size: 159357
|
| 1501 |
+
dataset_size: 340274.5496215436
|
| 1502 |
+
- config_name: public_relations
|
| 1503 |
+
features:
|
| 1504 |
+
- name: question
|
| 1505 |
+
dtype: string
|
| 1506 |
+
- name: subject
|
| 1507 |
+
dtype: string
|
| 1508 |
+
- name: choices
|
| 1509 |
+
sequence: string
|
| 1510 |
+
- name: answer
|
| 1511 |
+
dtype:
|
| 1512 |
+
class_label:
|
| 1513 |
+
names:
|
| 1514 |
+
'0': A
|
| 1515 |
+
'1': B
|
| 1516 |
+
'2': C
|
| 1517 |
+
'3': D
|
| 1518 |
+
splits:
|
| 1519 |
+
- name: test
|
| 1520 |
+
num_bytes: 54580.53197550207
|
| 1521 |
+
num_examples: 110
|
| 1522 |
+
- name: validation
|
| 1523 |
+
num_bytes: 5984.198563030699
|
| 1524 |
+
num_examples: 12
|
| 1525 |
+
- name: dev
|
| 1526 |
+
num_bytes: 2199.1754385964914
|
| 1527 |
+
num_examples: 5
|
| 1528 |
+
download_size: 31500
|
| 1529 |
+
dataset_size: 62763.90597712925
|
| 1530 |
+
- config_name: security_studies
|
| 1531 |
+
features:
|
| 1532 |
+
- name: question
|
| 1533 |
+
dtype: string
|
| 1534 |
+
- name: subject
|
| 1535 |
+
dtype: string
|
| 1536 |
+
- name: choices
|
| 1537 |
+
sequence: string
|
| 1538 |
+
- name: answer
|
| 1539 |
+
dtype:
|
| 1540 |
+
class_label:
|
| 1541 |
+
names:
|
| 1542 |
+
'0': A
|
| 1543 |
+
'1': B
|
| 1544 |
+
'2': C
|
| 1545 |
+
'3': D
|
| 1546 |
+
splits:
|
| 1547 |
+
- name: test
|
| 1548 |
+
num_bytes: 121565.73030907278
|
| 1549 |
+
num_examples: 245
|
| 1550 |
+
- name: validation
|
| 1551 |
+
num_bytes: 13464.446766819072
|
| 1552 |
+
num_examples: 27
|
| 1553 |
+
- name: dev
|
| 1554 |
+
num_bytes: 2199.1754385964914
|
| 1555 |
+
num_examples: 5
|
| 1556 |
+
download_size: 140258
|
| 1557 |
+
dataset_size: 137229.35251448833
|
| 1558 |
+
- config_name: sociology
|
| 1559 |
+
features:
|
| 1560 |
+
- name: question
|
| 1561 |
+
dtype: string
|
| 1562 |
+
- name: subject
|
| 1563 |
+
dtype: string
|
| 1564 |
+
- name: choices
|
| 1565 |
+
sequence: string
|
| 1566 |
+
- name: answer
|
| 1567 |
+
dtype:
|
| 1568 |
+
class_label:
|
| 1569 |
+
names:
|
| 1570 |
+
'0': A
|
| 1571 |
+
'1': B
|
| 1572 |
+
'2': C
|
| 1573 |
+
'3': D
|
| 1574 |
+
splits:
|
| 1575 |
+
- name: test
|
| 1576 |
+
num_bytes: 99733.51751887196
|
| 1577 |
+
num_examples: 201
|
| 1578 |
+
- name: validation
|
| 1579 |
+
num_bytes: 10971.030698889615
|
| 1580 |
+
num_examples: 22
|
| 1581 |
+
- name: dev
|
| 1582 |
+
num_bytes: 2199.1754385964914
|
| 1583 |
+
num_examples: 5
|
| 1584 |
+
download_size: 56480
|
| 1585 |
+
dataset_size: 112903.72365635807
|
| 1586 |
+
- config_name: us_foreign_policy
|
| 1587 |
+
features:
|
| 1588 |
+
- name: question
|
| 1589 |
+
dtype: string
|
| 1590 |
+
- name: subject
|
| 1591 |
+
dtype: string
|
| 1592 |
+
- name: choices
|
| 1593 |
+
sequence: string
|
| 1594 |
+
- name: answer
|
| 1595 |
+
dtype:
|
| 1596 |
+
class_label:
|
| 1597 |
+
names:
|
| 1598 |
+
'0': A
|
| 1599 |
+
'1': B
|
| 1600 |
+
'2': C
|
| 1601 |
+
'3': D
|
| 1602 |
+
splits:
|
| 1603 |
+
- name: test
|
| 1604 |
+
num_bytes: 49618.6654322746
|
| 1605 |
+
num_examples: 100
|
| 1606 |
+
- name: validation
|
| 1607 |
+
num_bytes: 5485.515349444808
|
| 1608 |
+
num_examples: 11
|
| 1609 |
+
- name: dev
|
| 1610 |
+
num_bytes: 2199.1754385964914
|
| 1611 |
+
num_examples: 5
|
| 1612 |
+
download_size: 29027
|
| 1613 |
+
dataset_size: 57303.3562203159
|
| 1614 |
+
- config_name: virology
|
| 1615 |
+
features:
|
| 1616 |
+
- name: question
|
| 1617 |
+
dtype: string
|
| 1618 |
+
- name: subject
|
| 1619 |
+
dtype: string
|
| 1620 |
+
- name: choices
|
| 1621 |
+
sequence: string
|
| 1622 |
+
- name: answer
|
| 1623 |
+
dtype:
|
| 1624 |
+
class_label:
|
| 1625 |
+
names:
|
| 1626 |
+
'0': A
|
| 1627 |
+
'1': B
|
| 1628 |
+
'2': C
|
| 1629 |
+
'3': D
|
| 1630 |
+
splits:
|
| 1631 |
+
- name: test
|
| 1632 |
+
num_bytes: 82366.98461757584
|
| 1633 |
+
num_examples: 166
|
| 1634 |
+
- name: validation
|
| 1635 |
+
num_bytes: 8976.297844546049
|
| 1636 |
+
num_examples: 18
|
| 1637 |
+
- name: dev
|
| 1638 |
+
num_bytes: 2199.1754385964914
|
| 1639 |
+
num_examples: 5
|
| 1640 |
+
download_size: 38229
|
| 1641 |
+
dataset_size: 93542.45790071838
|
| 1642 |
+
- config_name: world_religions
|
| 1643 |
+
features:
|
| 1644 |
+
- name: question
|
| 1645 |
+
dtype: string
|
| 1646 |
+
- name: subject
|
| 1647 |
+
dtype: string
|
| 1648 |
+
- name: choices
|
| 1649 |
+
sequence: string
|
| 1650 |
+
- name: answer
|
| 1651 |
+
dtype:
|
| 1652 |
+
class_label:
|
| 1653 |
+
names:
|
| 1654 |
+
'0': A
|
| 1655 |
+
'1': B
|
| 1656 |
+
'2': C
|
| 1657 |
+
'3': D
|
| 1658 |
+
splits:
|
| 1659 |
+
- name: test
|
| 1660 |
+
num_bytes: 84847.91788918957
|
| 1661 |
+
num_examples: 171
|
| 1662 |
+
- name: validation
|
| 1663 |
+
num_bytes: 9474.98105813194
|
| 1664 |
+
num_examples: 19
|
| 1665 |
+
- name: dev
|
| 1666 |
+
num_bytes: 2199.1754385964914
|
| 1667 |
+
num_examples: 5
|
| 1668 |
+
download_size: 27165
|
| 1669 |
+
dataset_size: 96522.07438591801
|
| 1670 |
+
configs:
|
| 1671 |
+
- config_name: abstract_algebra
|
| 1672 |
+
data_files:
|
| 1673 |
+
- split: test
|
| 1674 |
+
path: abstract_algebra/test-*
|
| 1675 |
+
- split: validation
|
| 1676 |
+
path: abstract_algebra/validation-*
|
| 1677 |
+
- split: dev
|
| 1678 |
+
path: abstract_algebra/dev-*
|
| 1679 |
+
- config_name: all
|
| 1680 |
+
data_files:
|
| 1681 |
+
- split: test
|
| 1682 |
+
path: all/test-*
|
| 1683 |
+
- split: validation
|
| 1684 |
+
path: all/validation-*
|
| 1685 |
+
- split: dev
|
| 1686 |
+
path: all/dev-*
|
| 1687 |
+
- split: auxiliary_train
|
| 1688 |
+
path: all/auxiliary_train-*
|
| 1689 |
+
- config_name: anatomy
|
| 1690 |
+
data_files:
|
| 1691 |
+
- split: test
|
| 1692 |
+
path: anatomy/test-*
|
| 1693 |
+
- split: validation
|
| 1694 |
+
path: anatomy/validation-*
|
| 1695 |
+
- split: dev
|
| 1696 |
+
path: anatomy/dev-*
|
| 1697 |
+
- config_name: astronomy
|
| 1698 |
+
data_files:
|
| 1699 |
+
- split: test
|
| 1700 |
+
path: astronomy/test-*
|
| 1701 |
+
- split: validation
|
| 1702 |
+
path: astronomy/validation-*
|
| 1703 |
+
- split: dev
|
| 1704 |
+
path: astronomy/dev-*
|
| 1705 |
+
- config_name: auxiliary_train
|
| 1706 |
+
data_files:
|
| 1707 |
+
- split: train
|
| 1708 |
+
path: auxiliary_train/train-*
|
| 1709 |
+
- config_name: business_ethics
|
| 1710 |
+
data_files:
|
| 1711 |
+
- split: test
|
| 1712 |
+
path: business_ethics/test-*
|
| 1713 |
+
- split: validation
|
| 1714 |
+
path: business_ethics/validation-*
|
| 1715 |
+
- split: dev
|
| 1716 |
+
path: business_ethics/dev-*
|
| 1717 |
+
- config_name: clinical_knowledge
|
| 1718 |
+
data_files:
|
| 1719 |
+
- split: test
|
| 1720 |
+
path: clinical_knowledge/test-*
|
| 1721 |
+
- split: validation
|
| 1722 |
+
path: clinical_knowledge/validation-*
|
| 1723 |
+
- split: dev
|
| 1724 |
+
path: clinical_knowledge/dev-*
|
| 1725 |
+
- config_name: college_biology
|
| 1726 |
+
data_files:
|
| 1727 |
+
- split: test
|
| 1728 |
+
path: college_biology/test-*
|
| 1729 |
+
- split: validation
|
| 1730 |
+
path: college_biology/validation-*
|
| 1731 |
+
- split: dev
|
| 1732 |
+
path: college_biology/dev-*
|
| 1733 |
+
- config_name: college_chemistry
|
| 1734 |
+
data_files:
|
| 1735 |
+
- split: test
|
| 1736 |
+
path: college_chemistry/test-*
|
| 1737 |
+
- split: validation
|
| 1738 |
+
path: college_chemistry/validation-*
|
| 1739 |
+
- split: dev
|
| 1740 |
+
path: college_chemistry/dev-*
|
| 1741 |
+
- config_name: college_computer_science
|
| 1742 |
+
data_files:
|
| 1743 |
+
- split: test
|
| 1744 |
+
path: college_computer_science/test-*
|
| 1745 |
+
- split: validation
|
| 1746 |
+
path: college_computer_science/validation-*
|
| 1747 |
+
- split: dev
|
| 1748 |
+
path: college_computer_science/dev-*
|
| 1749 |
+
- config_name: college_mathematics
|
| 1750 |
+
data_files:
|
| 1751 |
+
- split: test
|
| 1752 |
+
path: college_mathematics/test-*
|
| 1753 |
+
- split: validation
|
| 1754 |
+
path: college_mathematics/validation-*
|
| 1755 |
+
- split: dev
|
| 1756 |
+
path: college_mathematics/dev-*
|
| 1757 |
+
- config_name: college_medicine
|
| 1758 |
+
data_files:
|
| 1759 |
+
- split: test
|
| 1760 |
+
path: college_medicine/test-*
|
| 1761 |
+
- split: validation
|
| 1762 |
+
path: college_medicine/validation-*
|
| 1763 |
+
- split: dev
|
| 1764 |
+
path: college_medicine/dev-*
|
| 1765 |
+
- config_name: college_physics
|
| 1766 |
+
data_files:
|
| 1767 |
+
- split: test
|
| 1768 |
+
path: college_physics/test-*
|
| 1769 |
+
- split: validation
|
| 1770 |
+
path: college_physics/validation-*
|
| 1771 |
+
- split: dev
|
| 1772 |
+
path: college_physics/dev-*
|
| 1773 |
+
- config_name: computer_security
|
| 1774 |
+
data_files:
|
| 1775 |
+
- split: test
|
| 1776 |
+
path: computer_security/test-*
|
| 1777 |
+
- split: validation
|
| 1778 |
+
path: computer_security/validation-*
|
| 1779 |
+
- split: dev
|
| 1780 |
+
path: computer_security/dev-*
|
| 1781 |
+
- config_name: conceptual_physics
|
| 1782 |
+
data_files:
|
| 1783 |
+
- split: test
|
| 1784 |
+
path: conceptual_physics/test-*
|
| 1785 |
+
- split: validation
|
| 1786 |
+
path: conceptual_physics/validation-*
|
| 1787 |
+
- split: dev
|
| 1788 |
+
path: conceptual_physics/dev-*
|
| 1789 |
+
- config_name: econometrics
|
| 1790 |
+
data_files:
|
| 1791 |
+
- split: test
|
| 1792 |
+
path: econometrics/test-*
|
| 1793 |
+
- split: validation
|
| 1794 |
+
path: econometrics/validation-*
|
| 1795 |
+
- split: dev
|
| 1796 |
+
path: econometrics/dev-*
|
| 1797 |
+
- config_name: electrical_engineering
|
| 1798 |
+
data_files:
|
| 1799 |
+
- split: test
|
| 1800 |
+
path: electrical_engineering/test-*
|
| 1801 |
+
- split: validation
|
| 1802 |
+
path: electrical_engineering/validation-*
|
| 1803 |
+
- split: dev
|
| 1804 |
+
path: electrical_engineering/dev-*
|
| 1805 |
+
- config_name: elementary_mathematics
|
| 1806 |
+
data_files:
|
| 1807 |
+
- split: test
|
| 1808 |
+
path: elementary_mathematics/test-*
|
| 1809 |
+
- split: validation
|
| 1810 |
+
path: elementary_mathematics/validation-*
|
| 1811 |
+
- split: dev
|
| 1812 |
+
path: elementary_mathematics/dev-*
|
| 1813 |
+
- config_name: formal_logic
|
| 1814 |
+
data_files:
|
| 1815 |
+
- split: test
|
| 1816 |
+
path: formal_logic/test-*
|
| 1817 |
+
- split: validation
|
| 1818 |
+
path: formal_logic/validation-*
|
| 1819 |
+
- split: dev
|
| 1820 |
+
path: formal_logic/dev-*
|
| 1821 |
+
- config_name: global_facts
|
| 1822 |
+
data_files:
|
| 1823 |
+
- split: test
|
| 1824 |
+
path: global_facts/test-*
|
| 1825 |
+
- split: validation
|
| 1826 |
+
path: global_facts/validation-*
|
| 1827 |
+
- split: dev
|
| 1828 |
+
path: global_facts/dev-*
|
| 1829 |
+
- config_name: high_school_biology
|
| 1830 |
+
data_files:
|
| 1831 |
+
- split: test
|
| 1832 |
+
path: high_school_biology/test-*
|
| 1833 |
+
- split: validation
|
| 1834 |
+
path: high_school_biology/validation-*
|
| 1835 |
+
- split: dev
|
| 1836 |
+
path: high_school_biology/dev-*
|
| 1837 |
+
- config_name: high_school_chemistry
|
| 1838 |
+
data_files:
|
| 1839 |
+
- split: test
|
| 1840 |
+
path: high_school_chemistry/test-*
|
| 1841 |
+
- split: validation
|
| 1842 |
+
path: high_school_chemistry/validation-*
|
| 1843 |
+
- split: dev
|
| 1844 |
+
path: high_school_chemistry/dev-*
|
| 1845 |
+
- config_name: high_school_computer_science
|
| 1846 |
+
data_files:
|
| 1847 |
+
- split: test
|
| 1848 |
+
path: high_school_computer_science/test-*
|
| 1849 |
+
- split: validation
|
| 1850 |
+
path: high_school_computer_science/validation-*
|
| 1851 |
+
- split: dev
|
| 1852 |
+
path: high_school_computer_science/dev-*
|
| 1853 |
+
- config_name: high_school_european_history
|
| 1854 |
+
data_files:
|
| 1855 |
+
- split: test
|
| 1856 |
+
path: high_school_european_history/test-*
|
| 1857 |
+
- split: validation
|
| 1858 |
+
path: high_school_european_history/validation-*
|
| 1859 |
+
- split: dev
|
| 1860 |
+
path: high_school_european_history/dev-*
|
| 1861 |
+
- config_name: high_school_geography
|
| 1862 |
+
data_files:
|
| 1863 |
+
- split: test
|
| 1864 |
+
path: high_school_geography/test-*
|
| 1865 |
+
- split: validation
|
| 1866 |
+
path: high_school_geography/validation-*
|
| 1867 |
+
- split: dev
|
| 1868 |
+
path: high_school_geography/dev-*
|
| 1869 |
+
- config_name: high_school_government_and_politics
|
| 1870 |
+
data_files:
|
| 1871 |
+
- split: test
|
| 1872 |
+
path: high_school_government_and_politics/test-*
|
| 1873 |
+
- split: validation
|
| 1874 |
+
path: high_school_government_and_politics/validation-*
|
| 1875 |
+
- split: dev
|
| 1876 |
+
path: high_school_government_and_politics/dev-*
|
| 1877 |
+
- config_name: high_school_macroeconomics
|
| 1878 |
+
data_files:
|
| 1879 |
+
- split: test
|
| 1880 |
+
path: high_school_macroeconomics/test-*
|
| 1881 |
+
- split: validation
|
| 1882 |
+
path: high_school_macroeconomics/validation-*
|
| 1883 |
+
- split: dev
|
| 1884 |
+
path: high_school_macroeconomics/dev-*
|
| 1885 |
+
- config_name: high_school_mathematics
|
| 1886 |
+
data_files:
|
| 1887 |
+
- split: test
|
| 1888 |
+
path: high_school_mathematics/test-*
|
| 1889 |
+
- split: validation
|
| 1890 |
+
path: high_school_mathematics/validation-*
|
| 1891 |
+
- split: dev
|
| 1892 |
+
path: high_school_mathematics/dev-*
|
| 1893 |
+
- config_name: high_school_microeconomics
|
| 1894 |
+
data_files:
|
| 1895 |
+
- split: test
|
| 1896 |
+
path: high_school_microeconomics/test-*
|
| 1897 |
+
- split: validation
|
| 1898 |
+
path: high_school_microeconomics/validation-*
|
| 1899 |
+
- split: dev
|
| 1900 |
+
path: high_school_microeconomics/dev-*
|
| 1901 |
+
- config_name: high_school_physics
|
| 1902 |
+
data_files:
|
| 1903 |
+
- split: test
|
| 1904 |
+
path: high_school_physics/test-*
|
| 1905 |
+
- split: validation
|
| 1906 |
+
path: high_school_physics/validation-*
|
| 1907 |
+
- split: dev
|
| 1908 |
+
path: high_school_physics/dev-*
|
| 1909 |
+
- config_name: high_school_psychology
|
| 1910 |
+
data_files:
|
| 1911 |
+
- split: test
|
| 1912 |
+
path: high_school_psychology/test-*
|
| 1913 |
+
- split: validation
|
| 1914 |
+
path: high_school_psychology/validation-*
|
| 1915 |
+
- split: dev
|
| 1916 |
+
path: high_school_psychology/dev-*
|
| 1917 |
+
- config_name: high_school_statistics
|
| 1918 |
+
data_files:
|
| 1919 |
+
- split: test
|
| 1920 |
+
path: high_school_statistics/test-*
|
| 1921 |
+
- split: validation
|
| 1922 |
+
path: high_school_statistics/validation-*
|
| 1923 |
+
- split: dev
|
| 1924 |
+
path: high_school_statistics/dev-*
|
| 1925 |
+
- config_name: high_school_us_history
|
| 1926 |
+
data_files:
|
| 1927 |
+
- split: test
|
| 1928 |
+
path: high_school_us_history/test-*
|
| 1929 |
+
- split: validation
|
| 1930 |
+
path: high_school_us_history/validation-*
|
| 1931 |
+
- split: dev
|
| 1932 |
+
path: high_school_us_history/dev-*
|
| 1933 |
+
- config_name: high_school_world_history
|
| 1934 |
+
data_files:
|
| 1935 |
+
- split: test
|
| 1936 |
+
path: high_school_world_history/test-*
|
| 1937 |
+
- split: validation
|
| 1938 |
+
path: high_school_world_history/validation-*
|
| 1939 |
+
- split: dev
|
| 1940 |
+
path: high_school_world_history/dev-*
|
| 1941 |
+
- config_name: human_aging
|
| 1942 |
+
data_files:
|
| 1943 |
+
- split: test
|
| 1944 |
+
path: human_aging/test-*
|
| 1945 |
+
- split: validation
|
| 1946 |
+
path: human_aging/validation-*
|
| 1947 |
+
- split: dev
|
| 1948 |
+
path: human_aging/dev-*
|
| 1949 |
+
- config_name: human_sexuality
|
| 1950 |
+
data_files:
|
| 1951 |
+
- split: test
|
| 1952 |
+
path: human_sexuality/test-*
|
| 1953 |
+
- split: validation
|
| 1954 |
+
path: human_sexuality/validation-*
|
| 1955 |
+
- split: dev
|
| 1956 |
+
path: human_sexuality/dev-*
|
| 1957 |
+
- config_name: international_law
|
| 1958 |
+
data_files:
|
| 1959 |
+
- split: test
|
| 1960 |
+
path: international_law/test-*
|
| 1961 |
+
- split: validation
|
| 1962 |
+
path: international_law/validation-*
|
| 1963 |
+
- split: dev
|
| 1964 |
+
path: international_law/dev-*
|
| 1965 |
+
- config_name: jurisprudence
|
| 1966 |
+
data_files:
|
| 1967 |
+
- split: test
|
| 1968 |
+
path: jurisprudence/test-*
|
| 1969 |
+
- split: validation
|
| 1970 |
+
path: jurisprudence/validation-*
|
| 1971 |
+
- split: dev
|
| 1972 |
+
path: jurisprudence/dev-*
|
| 1973 |
+
- config_name: logical_fallacies
|
| 1974 |
+
data_files:
|
| 1975 |
+
- split: test
|
| 1976 |
+
path: logical_fallacies/test-*
|
| 1977 |
+
- split: validation
|
| 1978 |
+
path: logical_fallacies/validation-*
|
| 1979 |
+
- split: dev
|
| 1980 |
+
path: logical_fallacies/dev-*
|
| 1981 |
+
- config_name: machine_learning
|
| 1982 |
+
data_files:
|
| 1983 |
+
- split: test
|
| 1984 |
+
path: machine_learning/test-*
|
| 1985 |
+
- split: validation
|
| 1986 |
+
path: machine_learning/validation-*
|
| 1987 |
+
- split: dev
|
| 1988 |
+
path: machine_learning/dev-*
|
| 1989 |
+
- config_name: management
|
| 1990 |
+
data_files:
|
| 1991 |
+
- split: test
|
| 1992 |
+
path: management/test-*
|
| 1993 |
+
- split: validation
|
| 1994 |
+
path: management/validation-*
|
| 1995 |
+
- split: dev
|
| 1996 |
+
path: management/dev-*
|
| 1997 |
+
- config_name: marketing
|
| 1998 |
+
data_files:
|
| 1999 |
+
- split: test
|
| 2000 |
+
path: marketing/test-*
|
| 2001 |
+
- split: validation
|
| 2002 |
+
path: marketing/validation-*
|
| 2003 |
+
- split: dev
|
| 2004 |
+
path: marketing/dev-*
|
| 2005 |
+
- config_name: medical_genetics
|
| 2006 |
+
data_files:
|
| 2007 |
+
- split: test
|
| 2008 |
+
path: medical_genetics/test-*
|
| 2009 |
+
- split: validation
|
| 2010 |
+
path: medical_genetics/validation-*
|
| 2011 |
+
- split: dev
|
| 2012 |
+
path: medical_genetics/dev-*
|
| 2013 |
+
- config_name: miscellaneous
|
| 2014 |
+
data_files:
|
| 2015 |
+
- split: test
|
| 2016 |
+
path: miscellaneous/test-*
|
| 2017 |
+
- split: validation
|
| 2018 |
+
path: miscellaneous/validation-*
|
| 2019 |
+
- split: dev
|
| 2020 |
+
path: miscellaneous/dev-*
|
| 2021 |
+
- config_name: moral_disputes
|
| 2022 |
+
data_files:
|
| 2023 |
+
- split: test
|
| 2024 |
+
path: moral_disputes/test-*
|
| 2025 |
+
- split: validation
|
| 2026 |
+
path: moral_disputes/validation-*
|
| 2027 |
+
- split: dev
|
| 2028 |
+
path: moral_disputes/dev-*
|
| 2029 |
+
- config_name: moral_scenarios
|
| 2030 |
+
data_files:
|
| 2031 |
+
- split: test
|
| 2032 |
+
path: moral_scenarios/test-*
|
| 2033 |
+
- split: validation
|
| 2034 |
+
path: moral_scenarios/validation-*
|
| 2035 |
+
- split: dev
|
| 2036 |
+
path: moral_scenarios/dev-*
|
| 2037 |
+
- config_name: nutrition
|
| 2038 |
+
data_files:
|
| 2039 |
+
- split: test
|
| 2040 |
+
path: nutrition/test-*
|
| 2041 |
+
- split: validation
|
| 2042 |
+
path: nutrition/validation-*
|
| 2043 |
+
- split: dev
|
| 2044 |
+
path: nutrition/dev-*
|
| 2045 |
+
- config_name: philosophy
|
| 2046 |
+
data_files:
|
| 2047 |
+
- split: test
|
| 2048 |
+
path: philosophy/test-*
|
| 2049 |
+
- split: validation
|
| 2050 |
+
path: philosophy/validation-*
|
| 2051 |
+
- split: dev
|
| 2052 |
+
path: philosophy/dev-*
|
| 2053 |
+
- config_name: prehistory
|
| 2054 |
+
data_files:
|
| 2055 |
+
- split: test
|
| 2056 |
+
path: prehistory/test-*
|
| 2057 |
+
- split: validation
|
| 2058 |
+
path: prehistory/validation-*
|
| 2059 |
+
- split: dev
|
| 2060 |
+
path: prehistory/dev-*
|
| 2061 |
+
- config_name: professional_accounting
|
| 2062 |
+
data_files:
|
| 2063 |
+
- split: test
|
| 2064 |
+
path: professional_accounting/test-*
|
| 2065 |
+
- split: validation
|
| 2066 |
+
path: professional_accounting/validation-*
|
| 2067 |
+
- split: dev
|
| 2068 |
+
path: professional_accounting/dev-*
|
| 2069 |
+
- config_name: professional_law
|
| 2070 |
+
data_files:
|
| 2071 |
+
- split: test
|
| 2072 |
+
path: professional_law/test-*
|
| 2073 |
+
- split: validation
|
| 2074 |
+
path: professional_law/validation-*
|
| 2075 |
+
- split: dev
|
| 2076 |
+
path: professional_law/dev-*
|
| 2077 |
+
- config_name: professional_medicine
|
| 2078 |
+
data_files:
|
| 2079 |
+
- split: test
|
| 2080 |
+
path: professional_medicine/test-*
|
| 2081 |
+
- split: validation
|
| 2082 |
+
path: professional_medicine/validation-*
|
| 2083 |
+
- split: dev
|
| 2084 |
+
path: professional_medicine/dev-*
|
| 2085 |
+
- config_name: professional_psychology
|
| 2086 |
+
data_files:
|
| 2087 |
+
- split: test
|
| 2088 |
+
path: professional_psychology/test-*
|
| 2089 |
+
- split: validation
|
| 2090 |
+
path: professional_psychology/validation-*
|
| 2091 |
+
- split: dev
|
| 2092 |
+
path: professional_psychology/dev-*
|
| 2093 |
+
- config_name: public_relations
|
| 2094 |
+
data_files:
|
| 2095 |
+
- split: test
|
| 2096 |
+
path: public_relations/test-*
|
| 2097 |
+
- split: validation
|
| 2098 |
+
path: public_relations/validation-*
|
| 2099 |
+
- split: dev
|
| 2100 |
+
path: public_relations/dev-*
|
| 2101 |
+
- config_name: security_studies
|
| 2102 |
+
data_files:
|
| 2103 |
+
- split: test
|
| 2104 |
+
path: security_studies/test-*
|
| 2105 |
+
- split: validation
|
| 2106 |
+
path: security_studies/validation-*
|
| 2107 |
+
- split: dev
|
| 2108 |
+
path: security_studies/dev-*
|
| 2109 |
+
- config_name: sociology
|
| 2110 |
+
data_files:
|
| 2111 |
+
- split: test
|
| 2112 |
+
path: sociology/test-*
|
| 2113 |
+
- split: validation
|
| 2114 |
+
path: sociology/validation-*
|
| 2115 |
+
- split: dev
|
| 2116 |
+
path: sociology/dev-*
|
| 2117 |
+
- config_name: us_foreign_policy
|
| 2118 |
+
data_files:
|
| 2119 |
+
- split: test
|
| 2120 |
+
path: us_foreign_policy/test-*
|
| 2121 |
+
- split: validation
|
| 2122 |
+
path: us_foreign_policy/validation-*
|
| 2123 |
+
- split: dev
|
| 2124 |
+
path: us_foreign_policy/dev-*
|
| 2125 |
+
- config_name: virology
|
| 2126 |
+
data_files:
|
| 2127 |
+
- split: test
|
| 2128 |
+
path: virology/test-*
|
| 2129 |
+
- split: validation
|
| 2130 |
+
path: virology/validation-*
|
| 2131 |
+
- split: dev
|
| 2132 |
+
path: virology/dev-*
|
| 2133 |
+
- config_name: world_religions
|
| 2134 |
+
data_files:
|
| 2135 |
+
- split: test
|
| 2136 |
+
path: world_religions/test-*
|
| 2137 |
+
- split: validation
|
| 2138 |
+
path: world_religions/validation-*
|
| 2139 |
+
- split: dev
|
| 2140 |
+
path: world_religions/dev-*
|
| 2141 |
+
---
|
| 2142 |
+
|
| 2143 |
+
# Dataset Card for MMLU
|
| 2144 |
+
|
| 2145 |
+
## Table of Contents
|
| 2146 |
+
- [Table of Contents](#table-of-contents)
|
| 2147 |
+
- [Dataset Description](#dataset-description)
|
| 2148 |
+
- [Dataset Summary](#dataset-summary)
|
| 2149 |
+
- [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)
|
| 2150 |
+
- [Languages](#languages)
|
| 2151 |
+
- [Dataset Structure](#dataset-structure)
|
| 2152 |
+
- [Data Instances](#data-instances)
|
| 2153 |
+
- [Data Fields](#data-fields)
|
| 2154 |
+
- [Data Splits](#data-splits)
|
| 2155 |
+
- [Dataset Creation](#dataset-creation)
|
| 2156 |
+
- [Curation Rationale](#curation-rationale)
|
| 2157 |
+
- [Source Data](#source-data)
|
| 2158 |
+
- [Annotations](#annotations)
|
| 2159 |
+
- [Personal and Sensitive Information](#personal-and-sensitive-information)
|
| 2160 |
+
- [Considerations for Using the Data](#considerations-for-using-the-data)
|
| 2161 |
+
- [Social Impact of Dataset](#social-impact-of-dataset)
|
| 2162 |
+
- [Discussion of Biases](#discussion-of-biases)
|
| 2163 |
+
- [Other Known Limitations](#other-known-limitations)
|
| 2164 |
+
- [Additional Information](#additional-information)
|
| 2165 |
+
- [Dataset Curators](#dataset-curators)
|
| 2166 |
+
- [Licensing Information](#licensing-information)
|
| 2167 |
+
- [Citation Information](#citation-information)
|
| 2168 |
+
- [Contributions](#contributions)
|
| 2169 |
+
|
| 2170 |
+
## Dataset Description
|
| 2171 |
+
|
| 2172 |
+
- **Repository**: https://github.com/hendrycks/test
|
| 2173 |
+
- **Paper**: https://arxiv.org/abs/2009.03300
|
| 2174 |
+
|
| 2175 |
+
### Dataset Summary
|
| 2176 |
+
|
| 2177 |
+
[Measuring Massive Multitask Language Understanding](https://arxiv.org/pdf/2009.03300) by [Dan Hendrycks](https://people.eecs.berkeley.edu/~hendrycks/), [Collin Burns](http://collinpburns.com), [Steven Basart](https://stevenbas.art), Andy Zou, Mantas Mazeika, [Dawn Song](https://people.eecs.berkeley.edu/~dawnsong/), and [Jacob Steinhardt](https://www.stat.berkeley.edu/~jsteinhardt/) (ICLR 2021).
|
| 2178 |
+
|
| 2179 |
+
This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn. This covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models must possess extensive world knowledge and problem solving ability.
|
| 2180 |
+
|
| 2181 |
+
A complete list of tasks: ['abstract_algebra', 'anatomy', 'astronomy', 'business_ethics', 'clinical_knowledge', 'college_biology', 'college_chemistry', 'college_computer_science', 'college_mathematics', 'college_medicine', 'college_physics', 'computer_security', 'conceptual_physics', 'econometrics', 'electrical_engineering', 'elementary_mathematics', 'formal_logic', 'global_facts', 'high_school_biology', 'high_school_chemistry', 'high_school_computer_science', 'high_school_european_history', 'high_school_geography', 'high_school_government_and_politics', 'high_school_macroeconomics', 'high_school_mathematics', 'high_school_microeconomics', 'high_school_physics', 'high_school_psychology', 'high_school_statistics', 'high_school_us_history', 'high_school_world_history', 'human_aging', 'human_sexuality', 'international_law', 'jurisprudence', 'logical_fallacies', 'machine_learning', 'management', 'marketing', 'medical_genetics', 'miscellaneous', 'moral_disputes', 'moral_scenarios', 'nutrition', 'philosophy', 'prehistory', 'professional_accounting', 'professional_law', 'professional_medicine', 'professional_psychology', 'public_relations', 'security_studies', 'sociology', 'us_foreign_policy', 'virology', 'world_religions']
|
| 2182 |
+
|
| 2183 |
+
### Supported Tasks and Leaderboards
|
| 2184 |
+
|
| 2185 |
+
| Model | Authors | Humanities | Social Science | STEM | Other | Average |
|
| 2186 |
+
|------------------------------------|----------|:-------:|:-------:|:-------:|:-------:|:-------:|
|
| 2187 |
+
| [UnifiedQA](https://arxiv.org/abs/2005.00700) | Khashabi et al., 2020 | 45.6 | 56.6 | 40.2 | 54.6 | 48.9
|
| 2188 |
+
| [GPT-3](https://arxiv.org/abs/2005.14165) (few-shot) | Brown et al., 2020 | 40.8 | 50.4 | 36.7 | 48.8 | 43.9
|
| 2189 |
+
| [GPT-2](https://arxiv.org/abs/2005.14165) | Radford et al., 2019 | 32.8 | 33.3 | 30.2 | 33.1 | 32.4
|
| 2190 |
+
| Random Baseline | N/A | 25.0 | 25.0 | 25.0 | 25.0 | 25.0 | 25.0
|
| 2191 |
+
|
| 2192 |
+
### Languages
|
| 2193 |
+
|
| 2194 |
+
English
|
| 2195 |
+
|
| 2196 |
+
## Dataset Structure
|
| 2197 |
+
|
| 2198 |
+
### Data Instances
|
| 2199 |
+
|
| 2200 |
+
An example from anatomy subtask looks as follows:
|
| 2201 |
+
```
|
| 2202 |
+
{
|
| 2203 |
+
"question": "What is the embryological origin of the hyoid bone?",
|
| 2204 |
+
"choices": ["The first pharyngeal arch", "The first and second pharyngeal arches", "The second pharyngeal arch", "The second and third pharyngeal arches"],
|
| 2205 |
+
"answer": "D"
|
| 2206 |
+
}
|
| 2207 |
+
```
|
| 2208 |
+
|
| 2209 |
+
### Data Fields
|
| 2210 |
+
|
| 2211 |
+
- `question`: a string feature
|
| 2212 |
+
- `choices`: a list of 4 string features
|
| 2213 |
+
- `answer`: a ClassLabel feature
|
| 2214 |
+
|
| 2215 |
+
### Data Splits
|
| 2216 |
+
|
| 2217 |
+
- `auxiliary_train`: auxiliary multiple-choice training questions from ARC, MC_TEST, OBQA, RACE, etc.
|
| 2218 |
+
- `dev`: 5 examples per subtask, meant for few-shot setting
|
| 2219 |
+
- `test`: there are at least 100 examples per subtask
|
| 2220 |
+
|
| 2221 |
+
| | auxiliary_train | dev | val | test |
|
| 2222 |
+
| ----- | :------: | :-----: | :-----: | :-----: |
|
| 2223 |
+
| TOTAL | 99842 | 285 | 1531 | 14042
|
| 2224 |
+
|
| 2225 |
+
## Dataset Creation
|
| 2226 |
+
|
| 2227 |
+
### Curation Rationale
|
| 2228 |
+
|
| 2229 |
+
Transformer models have driven this recent progress by pretraining on massive text corpora, including all of Wikipedia, thousands of books, and numerous websites. These models consequently see extensive information about specialized topics, most of which is not assessed by existing NLP benchmarks. To bridge the gap between the wide-ranging knowledge that models see during pretraining and the existing measures of success, we introduce a new benchmark for assessing models across a diverse set of subjects that humans learn.
|
| 2230 |
+
|
| 2231 |
+
### Source Data
|
| 2232 |
+
|
| 2233 |
+
#### Initial Data Collection and Normalization
|
| 2234 |
+
|
| 2235 |
+
[More Information Needed]
|
| 2236 |
+
|
| 2237 |
+
#### Who are the source language producers?
|
| 2238 |
+
|
| 2239 |
+
[More Information Needed]
|
| 2240 |
+
|
| 2241 |
+
### Annotations
|
| 2242 |
+
|
| 2243 |
+
#### Annotation process
|
| 2244 |
+
|
| 2245 |
+
[More Information Needed]
|
| 2246 |
+
|
| 2247 |
+
#### Who are the annotators?
|
| 2248 |
+
|
| 2249 |
+
[More Information Needed]
|
| 2250 |
+
|
| 2251 |
+
### Personal and Sensitive Information
|
| 2252 |
+
|
| 2253 |
+
[More Information Needed]
|
| 2254 |
+
|
| 2255 |
+
## Considerations for Using the Data
|
| 2256 |
+
|
| 2257 |
+
### Social Impact of Dataset
|
| 2258 |
+
|
| 2259 |
+
[More Information Needed]
|
| 2260 |
+
|
| 2261 |
+
### Discussion of Biases
|
| 2262 |
+
|
| 2263 |
+
[More Information Needed]
|
| 2264 |
+
|
| 2265 |
+
### Other Known Limitations
|
| 2266 |
+
|
| 2267 |
+
[More Information Needed]
|
| 2268 |
+
|
| 2269 |
+
## Additional Information
|
| 2270 |
+
|
| 2271 |
+
### Dataset Curators
|
| 2272 |
+
|
| 2273 |
+
[More Information Needed]
|
| 2274 |
+
|
| 2275 |
+
### Licensing Information
|
| 2276 |
+
|
| 2277 |
+
[MIT License](https://github.com/hendrycks/test/blob/master/LICENSE)
|
| 2278 |
+
|
| 2279 |
+
### Citation Information
|
| 2280 |
+
|
| 2281 |
+
If you find this useful in your research, please consider citing the test and also the [ETHICS](https://arxiv.org/abs/2008.02275) dataset it draws from:
|
| 2282 |
+
```
|
| 2283 |
+
@article{hendryckstest2021,
|
| 2284 |
+
title={Measuring Massive Multitask Language Understanding},
|
| 2285 |
+
author={Dan Hendrycks and Collin Burns and Steven Basart and Andy Zou and Mantas Mazeika and Dawn Song and Jacob Steinhardt},
|
| 2286 |
+
journal={Proceedings of the International Conference on Learning Representations (ICLR)},
|
| 2287 |
+
year={2021}
|
| 2288 |
+
}
|
| 2289 |
+
|
| 2290 |
+
@article{hendrycks2021ethics,
|
| 2291 |
+
title={Aligning AI With Shared Human Values},
|
| 2292 |
+
author={Dan Hendrycks and Collin Burns and Steven Basart and Andrew Critch and Jerry Li and Dawn Song and Jacob Steinhardt},
|
| 2293 |
+
journal={Proceedings of the International Conference on Learning Representations (ICLR)},
|
| 2294 |
+
year={2021}
|
| 2295 |
+
}
|
| 2296 |
+
```
|
| 2297 |
+
### Contributions
|
| 2298 |
+
|
| 2299 |
+
Thanks to [@andyzoujm](https://github.com/andyzoujm) for adding this dataset.
|
hf_cache/hub/datasets--cais--mmlu/blobs/e133c92a3269e646dfd034d5c94ad41a31c7b194
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|