Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station

Wenyu Du, Stephen Chung

arXiv:2610.08927 · 2026-10-11 공개 · arXiv · PDF

ai-agents autonomous-discovery supervisor-mechanism meta-reflection scientific-ecosystem iclr-papers research-coverage agent-exploration

Abstract

Recent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear. We investigate AI's ability to tackle open-ended tasks in Station, an open-world environment in which multiple agents simulate a scientific ecosystem. To tackle challenges specific to open-ended tasks, we propose augmenting Station with two mechanisms: a Supervisor mechanism and periodic Meta Reflection, which encourage persistent exploration even when intermediate metrics are lacking. We construct open-ended tasks from three recent oral papers presented at ICLR. We give agents the main research question studied in each paper while withholding the paper's results and disabling web access. We then measure how many of the original findings-partitioned into individual criteria-agents rediscover. We find that Station rediscovers 62.7% of the criteria on average, compared with 15.4% for Codex Multiagent-v2 and 14.4-20.6% for AI Scientist-v2. Ablation and behavioral analyses indicate that adding the two mechanisms together improves research coverage and continuity. We further evaluate Station on two open-ended tasks without oracle papers and find that some of the discoveries made by the agents closely match discoveries reported by researchers after the knowledge cutoff date. Together, these results indicate that a suitable environment can enable agents to autonomously make meaningful progress in open-ended scientific discovery.

한국어 요약

한 줄 요약

Station은 Supervisor와 Meta Reflection을 통해 62.7%의 과학적 발견을 자율적으로 재현하는 다중 에이전트 연구 환경이다.

핵심 기여도

핵심 아이디어

기존 AI 연구 시스템은 명확한 평가 지표가 있는 과제에서 성공적이지만, 개방형 과학 탐구에서는 지표가 없어 지속적인 탐색이 어려운 문제가 있었다. 이를 해결하기 위해 Station에 Supervisor와 Meta Reflection이라는 두 가지 메커니즘을 도입했다. Supervisor는 에이전트들이 제안한 연구 방향을 최소 기간 동안 유지하도록 유도하며, Meta Reflection은 주기적으로 연구 방향이 인간 연구자들의 관심과 일치하는지 평가하도록 유도한다. 이는 ‘스트리트라이트 효과’를 완화하고, 과학적 가치가 높은 연구를 지속적으로 수행하도록 돕는다. 특히, Meta Reflection은 71.7%의 경우 기존 연구 방향을 심화하거나 검증하는 실험으로 이어졌다.

기술적 접근법

주요 결과

의의 및 한계

Station은 개방형 과학 탐구에서 AI 에이전트의 자율성을 증대시키는 분산형 연구 환경을 제시하며, Supervisor와 Meta Reflection을 통해 지속적인 탐색을 촉진한다. 특히, 이는 ‘스트리트라이트 효과’를 완화하고, 과학적 가치가 높은 연구 방향을 유지하는 데 기여한다. 그러나 AI 에이전트는 인간 연구자에 비해 지속적인 노력과 깊은 설명을 요구하는 발견을 재현하지 못하는 경우가 많으며, 판단의 일치가 부족한 문제가 여전히 존재한다. 이는 ‘판단 불일치’라는 한계로, 인간의 가이드가 필요하지만, 이는 자율 연구의 확장성을 제한한다.

실용적 활용

Station은 개방형 과학 탐구를 필요로 하는 연구 분야, 특히 신경망 해석, 물리 현상 이해, LLM 및 VLM 연구 등에서 활용 가능하다. 또한, AI 에이전트가 인간 연구자와 협력하여 새로운 이론을 제안하거나 실험 설계를 돕는 데 기여할 수 있다.