Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov

arXiv:2608.18940 · 2026-08-20 공개 · arXiv · PDF

llm retrosynthesis top-k-prompting chemical-plausibility creed-ccv-2uspto-xl c3lm chemcensor ood-benchmarking

Abstract

Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we introduce Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions. We compile CREED-CCV-2+USPTO-XL, an ultra-large-scale dataset of ~45.6 million verified reactions to train the C3LM (Chemistry Constraint-Consistent Language Model). By integrating fine-tuning with ChemCensor-based and novelty-oriented rewards, our model achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark. Further analysis of reaction uniqueness shows that LLMs and conventional models explore complementary reaction spaces, motivating ensemble-based retrosynthesis systems. Overall, our results establish Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.

한국어 요약

한 줄 요약

C3LM 모델을 Top-K 프롬프팅과 ChemCensor 기반 보상으로 훈련하여 URSA-expert-2026 벤치마크에서 최고 성능을 달성했다.

핵심 기여도

핵심 아이디어

단계별 반응 생성의 one-to-many 특성을 반영하기 위해 **Top-K 프롬프팅**을 도입하여 단일 답변 평가의 한계를 극복했다. 기존의 Top-1 방식 대비 **Top-K (K=15)** 모드는 반응 다양성과 가능성성을 동시에 강화하는 효과를 보였다.

C3LM은 **CREED-CCV-2+USPTO-XL** (~45.6M) 데이터셋을 기반으로 훈련되며, **ChemCensor 기반 필터링**과 **Novelty Reward**를 통한 **RFT (Reward Fine-Tuning)**을 적용하여 반응의 화학적 타당성과 독창성을 동시에 강화했다. 특히, **Av. PT-Top-10** 지표는 반응 다양성에 가장 민감하며, Top-K 모드 전환 시 2.5배 이상의 성능 향상이 관찰되었다.

이러한 접근은 **LLM이 단계별 반응 생성에서 전통 모델과 보완적인 역할**을 할 수 있음을 입증하며, **앙상블 기반 합성 계획 시스템** 설계를 위한 기초를 제공한다.

기술적 접근법

주요 결과

의의 및 한계

C3LM은 **LLM 기반 단계별 합성 계획**에서 화학적 타당성과 반응 다양성을 동시에 강화하는 새로운 훈련 패러다임을 제시한다. 특히, **Top-K 프롬프팅**과 **ChemCensor 기반 RFT**는 기존 평가 체계의 한계를 극복하며, **앙상블 기반 시스템** 설계 가능성을 열었다.

하지만, **ChemCensor 기반 평가와 훈련 간의 부분적 순환성**은 모델의 일반화 능력에 영향을 줄 수 있다. 또한, **URSA-expert-2026**과 같은 OOD 데이터셋에서의 성능 향상은 **데이터 누수 가능성**을 고려해야 한다. **Av. PT-Top-10** 기준으로도 **+0.44** 개선 여지가 남아 있어, 모델의 최적화 여지가 있다.

실용적 활용

C3LM은 **의약품 개발 초기 단계의 합성 가능성 평가**, **컴퓨터 지원 합성 계획 시스템 (CASP)**, **LLM 기반 반응 생성 도구** 등에 활용 가능하다. 특히, **Top-K 프롬프팅**과 **RFT 기반 훈련**은 화학 연구자들이 다양한 반응 경로를 탐색하는 데 유용한 도구가 될 수 있다.