Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search
Zhongwei Yu, Yan Song, Xue Yan, Anjie Liu, Xingyu Lu, Yihang Chen, Huichi Zhou, Siyuan Guo, Luoyang Sun, Sihan Chen, Xiangning Yu, Jun Wang
arXiv:2608.15669 · 2026-08-18 공개 · arXiv · PDF
reinforcement-learning generative-model uncertainty-aware surrogate-model neural-network-training molecular-optimization antibody-design large-discovery-model
Abstract
Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs. Generative models such as large language models (LLMs) provide expressive priors over such spaces, but their likelihoods and self-assessments are unreliable proxies for the objectives and calibrated epistemic uncertainty, especially for novel candidates outside the observed data distribution. We introduce the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model. The generative model proposes and refines candidate designs, while the surrogate predicts their performance and quantifies uncertainty, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection. The discovery memory and the surrogate model are continually updated as each new experimental observation arrives. We evaluate LDM on three scenarios spanning different design modalities and objectives, including neural-network training, antibody design, and molecular optimisation. Compared to LLM-only reflection or traditional statistical search across these domains, LDM achieves a 2.4times greater reduction in validation BPB, an 18.2% relative decrease in binding energy, and more than 60% relative gains in molecular multi-objective performance. These results suggests that LDM could serve as a general-purpose discovery engine for effective search over open-ended hypothesis spaces.
한국어 요약
한 줄 요약
LDM은 생성 모델과 베이지안 비모수 보상 서로게 모델을 결합한, 개방적 가설 공간에서의 효율적 탐색 엔진이다.
핵심 기여도
- LDM은 생성 모델과 베이지안 비모수 보상 서로게 모델을 결합하여, 탐색 효율성을 2.4배 향상시킴.
- 분자 최적화에서 바인딩 에너지 18.2% 감소를 달성.
- 다목적 분자 최적화에서 60% 이상의 상대적 성능 향상.
- 신경망 학습 프로그램, 항체 설계, 분자 최적화 등 다양한 과학 탐색 영역에서 일반화 가능.
핵심 아이디어
LDM은 기존 생성 모델의 한계를 극복하기 위해, 생성된 후보를 평가하고 불확실성을 추정하는 베이지안 비모수 서로게 모델을 결합한 반복 구조를 도입한다. 생성 모델은 후보를 제안하고 개선하며, 서로게 모델은 성능을 예측하고 불확실성을 정량화하여 탐색을 유도한다. 이는 기존 LLM 기반 반사나 전통적 통계 탐색 방식보다 탐색 효율성을 높인다. 특히, LDM은 새로운 후보에 대한 불확실성을 정확히 측정하여, 기존 데이터 분포 밖에서도 신뢰성 있는 탐색이 가능하다.
기술적 접근법
- **LDM 아키텍처**: 생성 모델과 베이지안 비모수 보상 서로게 모델을 결합한 반복 구조.
- **서로게 모델**: 후보 성능을 예측하고 불확실성을 정량화.
- **탐색 메커니즘**: 불확실성 인식 가치를 기반으로 후보 생성, 개선, 선택.
- **데이터셋**: 신경망 학습 프로그램, 항체 CDRH3 시퀀스, KRAS G12D 타겟 분자.
- **실험 환경**: 디지털 오라클 벤치마크를 사용한 시뮬레이션 기반 평가.
주요 결과
- **신경망 학습 프로그램**: 검증 BPB 감소량 2.4배.
- **항체 설계**: 바인딩 에너지 18.2% 상대 감소.
- **분자 최적화**: 다목적 성능 60% 이상 상대 향상.
- **기존 방법 대비**: LLM 기반 반사 및 전통적 통계 탐색 대비 우수한 성능.
의의 및 한계
LDM은 개방적 가설 공간에서의 과학 탐색을 위한 일반적인 엔진으로 활용 가능하며, 생성 모델의 한계를 보완한 새로운 접근법을 제시한다. 특히, 베이지안 서로게 모델을 통해 불확실성을 정량화함으로써, 기존 데이터 분포 밖에서도 신뢰성 있는 탐색이 가능하다는 점에서 학술적 의의가 크다. 그러나 실험은 디지털 오라클 기반으로 이루어졌으며, 실제 실험 환경에서의 검증은 향후 연구 과제로 남아 있다.
실용적 활용
LDM은 분자 설계, 항체 개발, 신경망 최적화 등 다양한 과학 연구 분야에서 적용 가능하며, 실험 비용을 줄이고 탐색 효율성을 높이는 데 유용할 수 있다. 특히, 고비용 실험 환경에서의 후보 선정 및 개선 과정에 활용 가능하다.