StudentBench: AI and human tutoring yield equivalent GRE learning gains
Abstract
Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.
Community
We introduce StudentBench, a suite of AI teaching evaluations and a public platform for measuring how well AI helps people learn. The paper compares AI tutoring, expert human tutoring, and a no tutoring control on Quantitative and Verbal GRE learning gains, and evaluates AI-generated lesson plans and practice problems.
Dataset: https://huggingface.co/datasets/handshake-ai-research/studentbench
Project: https://studentbench.org
Code: https://github.com/Handshake-AI-Research/studentbench
This is the first work to establish and measure the cost to augment a student’s learning by one percentage point.
Over 1 trillion USD is being spent teaching machines to improve machines.
StudentBench shows the value of teaching machines to improve humans.
across 2000+ students, AI tutoring was statistically equivalent to expert human tutoring on GRE learning gains (p=.015).
this work establishes an important precedent for the democratization of education -- a student who has access to an LLM can improve their GRE score for one penny.
We provide five new kinds of tutoring evaluation leaderboards where we separate 13 LLM-based AI tutors (and we can clearly see which models do best at which capabilities).
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- When AI Tutors Speak: Evidence from a Randomized Field Experiment (2026)
- Methodologies for Improving the Quality of AI Tutoring in K-12 Education (2026)
- StudentSim: Training LLM-based Student Simulators (2026)
- EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics (2026)
- LLM Pedagogical Behavior in AI Tutoring Interactions (2026)
- Simulating Disengaged Students to Evaluate LLM-based Tutors (2026)
- OmniEdu: Open Foundation Models for Learning and Teaching (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.28470 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper