HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses

Jieyuan Liu, Mengzhou Hu, Jefferson Chen, JungHo Kong, Pratibha Jagannatha, Yiming Gao, Dexter Pratt, Hsin-Yuan Lee, Zhiting Hu, Trey Ideker, Wei Wang, Eric P. Xing, Zhen Wang

arXiv:2609.15938 · 2026-09-18 공개 · arXiv · PDF

evolutionary-search scientific-agents multi-agent-llms drug-repurposing llm-collaboration cancer-research genetic-algorithms hypothesis-discovery

Abstract

Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different forms of agent collaboration affect hypothesis quality remains an open question. Answering this question requires separating the effects of agents' scientific capabilities from those of their collaboration. A framework must therefore preserve agents' scientific roles and support rules for combining, revising, and retaining hypotheses. Building on this view, we introduce HypoEvolve, which makes collaboration explicit through successive updates to a hypothesis population. Specifically, we propose a generational genetic algorithm to coordinate specialized large language model (LLM) agents that integrate mechanistic arguments, reconsider assumptions, and assess evidence and testability. Each generation specifies how scientific judgments and new proposals reshape the population, making collaboration effects on hypothesis quality directly testable. Moreover, we design our evaluation around scientifically meaningful hypotheses that explain how a proposed intervention could work. Drug repurposing links these explanations to target-level biological claims assessed against external evidence. Specifically, we adapt DepMap and Open Targets into complementary external measures grounded in experimental, genetic, and clinical evidence. Across 34 cancer types, HypoEvolve achieves the highest scores against six baselines on both measures. DepMap selectivity reaches 0.171, versus 0.115 for the strongest baseline. Gains over single-pass generation also generalize to held-out cancer types. HypoEvolve advances a vision of autonomous science in which AI research teams achieve a capacity for discovery beyond that of individual models.

한국어 요약

한 줄 요약

HypoEvolve는 과학적 가설의 질을 향상시키기 위해 유전 알고리즘을 활용한 다 에이전트 LLM 시스템이다.

핵심 기여도

핵심 아이디어

HypoEvolve는 과학적 에이전트 간 협업이 가설의 질에 미치는 영향을 분리·평가하기 위해 유전 알고리즘을 도입한 시스템이다. 이는 과학적 능력과 협업 규칙을 분리하여, 후자를 실험 변수로 삼아 가설 발전을 제어할 수 있게 한다. 구체적으로, HypoEvolve는 가설 집단(population)을 세대별로 업데이트하며, 각 세대에서 LLM 에이전트가 기전적 논리(integrate mechanistic arguments), 가정 재검토(reconsider assumptions), 증거와 검증 가능성 평가(assess evidence and testability)를 수행한다. 이 과정에서 유전 알고리즘은 교차(crossover)와 돌연변이(mutation)를 통해 가설 간 논리적 결합과 수정을 이끈다. 이는 단순한 텍스트 생성이 아닌, 과학적 추론을 기반으로 한 가설의 진화를 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

HypoEvolve는 과학적 협업을 알고리즘적으로 설계하고 평가할 수 있는 기반을 제공하며, 단일 모델의 능력을 넘어선 AI 연구팀의 가능성을 제시한다. 특히, 가설의 질과 협업 설계 간의 직접적인 연결성을 입증한 점에서 학술적 의의가 크다. 그러나, HypoEvolve는 실험적 평가가 외부 생물학적 데이터(DepMap, Open Targets)에 의존하며, 실제 실험 환경에서의 검증은 아직 이루어지지 않았다. 또한, LLM의 텍스트 생성 능력에 기반하기 때문에, 과학적 논리의 정확성이나 실험 가능성에 대한 한계가 존재할 수 있다.

실용적 활용

HypoEvolve는 약물 재목적화(drug repurposing)와 같은 의약 연구 분야에서 활용 가능하며, 기존 데이터를 기반으로 새로운 생물학적 가설을 생성하고 검증하는 데 유용하다. 또한, 과학적 연구 팀의 자동화 및 협업 최적화를 위한 AI 연구 플랫폼으로도 활용될 수 있다.