A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization

Ming Chen, Rong-Xi Tan, Ke Xue, Yu-Jie Zhou, Taiye Lu, Zhi-Xuan Gao, Peng Xie, Zijun Shen, Chen Lu, Haopu Shang, Chao Qian

arXiv:2610.12183 · 2026-10-11 공개 · arXiv · PDF

llm-agents hyperparameter-optimization black-box-optimization agentic-bbo database-tuning chip-design molecular-design numerical-optimizers

Abstract

Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining task semantics, computation, optimization tools, and feedback-driven decision making, showing great potential due to the integration with mathematically rigorous tools. However, existing agentic BBO studies use different task domains and system configurations, making their results difficult to compare and the effects of individual design choices hard to isolate. We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol. In our experiments, agentic BBO achieves higher family-averaged scores than direct LLM-based methods in all five domains and outperforms the best numerical optimizers in four. We further study three factors shaping agent performance: optimization tools, task information and prior knowledge, and the role of the LLM during search. Our results show that additional numerical tools do not consistently improve performance, task semantics are broadly useful while more specific priors are less reliable, and numerical optimizers can effectively absorb gains from search trajectories established by the agent. Finally, we introduce a five-task frontier challenge within AgenticBBO-Bench and evaluate seven LLMs under the Codex agent harness, where GPT-6 Astra and DeepSeek-V4.1-Flash lie on the Pareto frontier of performance and cost among the evaluated models. Our code is available at https://github.com/lamda-bbo/agentic-bbo.

한국어 요약

한 줄 요약

AgenticBBO-Bench를 통해 LLM 에이전트의 BBO 성능을 다분야 비교하고, 핵심 성능 요인을 분석한 연구.

핵심 기여도

핵심 아이디어

기존 BBO 연구는 도메인과 시스템 구성이 달라 비교가 어려웠다. 본 연구는 AgenticBBO-Bench라는 통합 벤치마크를 통해 LLM 에이전트의 BBO 성능을 다분야에서 평가하고, 핵심 성능 요인을 분석한다. LLM은 태스크 의미, 계산, 최적화 도구, 피드백 기반 의사결정을 결합하여 새로운 BBO 접근법을 제시한다. 특히, 에이전트는 최적화 환경과 상호작용하며, 다음 행동을 스스로 결정하는 방식으로 작동한다. 이는 기존의 고정된 최적화 절차와 구별된다. 연구는 LLM이 태스크 의미를 활용하는 것이 일반적으로 유용하다는 점을 발견했으며, 특정한 사전 지식은 신뢰도가 낮다는 점도 밝혔다.

기술적 접근법

주요 결과

의의 및 한계

AgenticBBO-Bench는 LLM 에이전트의 BBO 성능을 비교할 수 있는 표준화된 벤치마크로, 학술적·실용적 가치가 크다. 특히, LLM이 태스크 의미를 활용하는 방식은 BBO의 새로운 가능성을 제시한다. 그러나 연구는 특정 LLM 모델과 도메인에 제한되어 있으며, 더 넓은 범위의 실험과 모델 비교가 필요하다. 또한, LLM의 참여도에 따른 성능 변화는 아직 명확히 규명되지 않았으며, 이에 대한 추가 연구가 필요하다.

실용적 활용

AgenticBBO-Bench는 칩 설계, 분자 설계, 데이터베이스 튜닝 등 고비용 최적화 문제에 적용 가능한 LLM 에이전트의 성능을 평가하는 데 유용하다. 연구 결과는 LLM이 태스크 의미를 활용하는 방식을 개선하고, 최적화 루프 내 역할을 조정함으로써 실제 산업 문제 해결에 기여할 수 있음을 시사한다.