Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station
Abstract
Recent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear. We investigate AI's ability to tackle open-ended tasks in Station, an open-world environment in which multiple agents simulate a scientific ecosystem. To tackle challenges specific to open-ended tasks, we propose augmenting Station with two mechanisms: a Supervisor mechanism and periodic Meta Reflection, which encourage persistent exploration even when intermediate metrics are lacking. We construct open-ended tasks from three recent oral papers presented at ICLR. We give agents the main research question studied in each paper while withholding the paper's results and disabling web access. We then measure how many of the original findings-partitioned into individual criteria-agents rediscover. We find that Station rediscovers 62.7% of the criteria on average, compared with 15.4% for Codex Multiagent-v2 and 14.4-20.6% for AI Scientist-v2. Ablation and behavioral analyses indicate that adding the two mechanisms together improves research coverage and continuity. We further evaluate Station on two open-ended tasks without oracle papers and find that some of the discoveries made by the agents closely match discoveries reported by researchers after the knowledge cutoff date. Together, these results indicate that a suitable environment can enable agents to autonomously make meaningful progress in open-ended scientific discovery.
Community
Can AI agents make open-ended scientific discovery?
We built open-ended research tasks from three ICLR oral papers, giving agents the research questions and experimental setups while withholding the original findings and disabling web access.
Station recovered 62.7% of the predefined sub-discoveries on average, compared with 20.6% for the strongest tested AI Scientist-v2 configuration and 15.4% for Codex Multiagent-v2. The comparison matches cumulative experiment time; token budgets and model teams differ.
Agents independently choose research directions, conduct experiments, and build on reviewed findings in a shared knowledge base. Supervisor guidance and periodic Meta Reflection encourage sustained investigation instead of early pivots to easier questions.
On two additional exploratory tasks, agents also produced findings matching concurrent human research. In subliminal learning, they found that restricting LoRA training to early layers could restore failed trait transfer.
We are releasing the code and full research records so others can inspect the discoveries—and the failures. A key remaining challenge is scientific judgment: agents still spend substantial effort on questions human researchers find uninteresting.
Paper: https://arxiv.org/abs/2610.08927
Code: https://github.com/dualverse-ai/station-open-reseach
Explore the research records: https://dualverse-ai.github.io/station-open-reseach_data/
Get this paper in your agent:
hf papers read 2610.08927 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper