Buckets:
| license: mit | |
| tags: | |
| - education | |
| - evaluation | |
| - benchamark | |
| - portuguese | |
| - brazil | |
| pretty_name: Alvorada Bench | |
| size_categories: | |
| - 100K<n<1M | |
| dataset_info: | |
| - config_name: questions | |
| features: | |
| - name: question_id | |
| dtype: string | |
| - name: question_number | |
| dtype: string | |
| - name: subject | |
| dtype: string | |
| - name: question_statement | |
| dtype: string | |
| - name: correct_answer | |
| dtype: string | |
| - name: exam_name | |
| dtype: string | |
| - name: exam_year | |
| dtype: int64 | |
| - name: exam_type | |
| dtype: string | |
| - name: alternative_a | |
| dtype: string | |
| - name: alternative_b | |
| dtype: string | |
| - name: alternative_c | |
| dtype: string | |
| - name: alternative_d | |
| dtype: string | |
| - name: alternative_e | |
| dtype: string | |
| splits: | |
| - name: train | |
| num_bytes: 5000000 | |
| num_examples: 4515 | |
| download_size: 2500000 | |
| dataset_size: 5000000 | |
| - config_name: responses | |
| features: | |
| - name: model | |
| dtype: string | |
| - name: prompt_template | |
| dtype: string | |
| - name: question_id | |
| dtype: string | |
| - name: question_number | |
| dtype: string | |
| - name: subject | |
| dtype: string | |
| - name: chosen_answer | |
| dtype: string | |
| - name: difficulty_level | |
| dtype: string | |
| - name: uncertainty_level | |
| dtype: string | |
| - name: bloom_taxonomy | |
| dtype: string | |
| - name: is_correct | |
| dtype: string | |
| - name: exam_name | |
| dtype: string | |
| - name: provider | |
| dtype: string | |
| - name: exam_year | |
| dtype: string | |
| - name: exam_type | |
| dtype: string | |
| splits: | |
| - name: train | |
| num_bytes: 50000000 | |
| num_examples: 270840 | |
| download_size: 25000000 | |
| dataset_size: 50000000 | |
| configs: | |
| - config_name: questions | |
| data_files: | |
| - split: train | |
| path: questions_data.csv | |
| - config_name: responses | |
| data_files: | |
| - split: train | |
| path: responses_data.csv | |
| This dataset contains 4,515 multiple-choice questions from five major Brazilian university entrance exams (ENEM, FUVEST, UNICAMP, ITA, IME) spanning 32 years (1981-2025), along with model responses from 20 LLMs. | |
| ## Files | |
| ### 📄 `questions_data.csv` (4,515 rows) | |
| Contains the exam questions with: | |
| - `question_id`: Unique identifier | |
| - `question_statement`: Question text in Portuguese | |
| - `correct_answer`: Correct option (A-E) | |
| - `alternative_a` to `alternative_e`: Answer choices | |
| - `subject`: Academic subject | |
| - `exam_name`, `exam_year`, `exam_type`: Exam metadata | |
| ### 📄 `responses_data.csv` | |
| Contains model responses with: | |
| - `model`: Model name (o3, deepseek-reasoner, claude-opus-4-20250514) | |
| - `prompt_template`: Prompting strategy used (zero-shot, role-playing, chain-of-thought) | |
| - `chosen_answer`: Model's selected answer | |
| - `is_correct`: Whether the answer was correct | |
| - `difficulty_level`, `uncertainty_level`: Model's self-reported metrics (1-10 scale) | |
| - `bloom_taxonomy`: Cognitive complexity classification | |
| - Additional metadata matching questions_data | |
| ## Cite | |
| ``` | |
| @misc{godoy2025alvoradabenchlanguagemodelssolve, | |
| title={Alvorada-Bench: Can Language Models Solve Brazilian University Entrance Exams?}, | |
| author={Henrique Godoy}, | |
| year={2025}, | |
| eprint={2508.15835}, | |
| archivePrefix={arXiv}, | |
| primaryClass={cs.CL}, | |
| url={https://arxiv.org/abs/2508.15835}, | |
| } | |
| ``` |
Xet Storage Details
- Size:
- 3.17 kB
- Xet hash:
- 091ef5c005dfb4a9417f0deeb0d3e6293c5d4e5352f8797cd331ef496f06738d
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.