runawaydevil's picture
|
download
raw
3.17 kB
---
license: mit
tags:
- education
- evaluation
- benchamark
- portuguese
- brazil
pretty_name: Alvorada Bench
size_categories:
- 100K<n<1M
dataset_info:
- config_name: questions
features:
- name: question_id
dtype: string
- name: question_number
dtype: string
- name: subject
dtype: string
- name: question_statement
dtype: string
- name: correct_answer
dtype: string
- name: exam_name
dtype: string
- name: exam_year
dtype: int64
- name: exam_type
dtype: string
- name: alternative_a
dtype: string
- name: alternative_b
dtype: string
- name: alternative_c
dtype: string
- name: alternative_d
dtype: string
- name: alternative_e
dtype: string
splits:
- name: train
num_bytes: 5000000
num_examples: 4515
download_size: 2500000
dataset_size: 5000000
- config_name: responses
features:
- name: model
dtype: string
- name: prompt_template
dtype: string
- name: question_id
dtype: string
- name: question_number
dtype: string
- name: subject
dtype: string
- name: chosen_answer
dtype: string
- name: difficulty_level
dtype: string
- name: uncertainty_level
dtype: string
- name: bloom_taxonomy
dtype: string
- name: is_correct
dtype: string
- name: exam_name
dtype: string
- name: provider
dtype: string
- name: exam_year
dtype: string
- name: exam_type
dtype: string
splits:
- name: train
num_bytes: 50000000
num_examples: 270840
download_size: 25000000
dataset_size: 50000000
configs:
- config_name: questions
data_files:
- split: train
path: questions_data.csv
- config_name: responses
data_files:
- split: train
path: responses_data.csv
---
This dataset contains 4,515 multiple-choice questions from five major Brazilian university entrance exams (ENEM, FUVEST, UNICAMP, ITA, IME) spanning 32 years (1981-2025), along with model responses from 20 LLMs.
## Files
### 📄 `questions_data.csv` (4,515 rows)
Contains the exam questions with:
- `question_id`: Unique identifier
- `question_statement`: Question text in Portuguese
- `correct_answer`: Correct option (A-E)
- `alternative_a` to `alternative_e`: Answer choices
- `subject`: Academic subject
- `exam_name`, `exam_year`, `exam_type`: Exam metadata
### 📄 `responses_data.csv`
Contains model responses with:
- `model`: Model name (o3, deepseek-reasoner, claude-opus-4-20250514)
- `prompt_template`: Prompting strategy used (zero-shot, role-playing, chain-of-thought)
- `chosen_answer`: Model's selected answer
- `is_correct`: Whether the answer was correct
- `difficulty_level`, `uncertainty_level`: Model's self-reported metrics (1-10 scale)
- `bloom_taxonomy`: Cognitive complexity classification
- Additional metadata matching questions_data
## Cite
```
@misc{godoy2025alvoradabenchlanguagemodelssolve,
title={Alvorada-Bench: Can Language Models Solve Brazilian University Entrance Exams?},
author={Henrique Godoy},
year={2025},
eprint={2508.15835},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2508.15835},
}
```

Xet Storage Details

Size:
3.17 kB
·
Xet hash:
091ef5c005dfb4a9417f0deeb0d3e6293c5d4e5352f8797cd331ef496f06738d

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.