CHASE-SQL: Multi-Path Reasoning and Preference Optimized Candidate Selection in Text-to-SQL

Mohammadreza Pourreza, Hailong Li, Ruoxi Sun, Yeounoh Chung, Shayan Talaei, Gaurav Tarlok Kakkar, Yu Gan, Amin Saberi, Fatma Ozcan, Sercan Ö. Arik

arXiv:2410.01943 · 2026-07-27 공개 · arXiv · PDF

chain-of-thought text-to-sql candidate-selection multi-agent-modeling bird-dataset llm-generators query-decomposition execution-plan

Abstract

In tackling the challenges of large language model (LLM) performance for Text-to-SQL tasks, we introduce CHASE-SQL, a new framework that employs innovative strategies, using test-time compute in multi-agent modeling to improve candidate generation and selection. CHASE-SQL leverages LLMs' intrinsic knowledge to generate diverse and high-quality SQL candidates using different LLM generators with: (1) a divide-and-conquer method that decomposes complex queries into manageable sub-queries in a single LLM call; (2) chain-of-thought reasoning based on query execution plans, reflecting the steps a database engine takes during execution; and (3) a unique instance-aware synthetic example generation technique, which offers specific few-shot demonstrations tailored to test questions.To identify the best candidate, a selection agent is employed to rank the candidates through pairwise comparisons with a fine-tuned binary-candidates selection LLM. This selection approach has been demonstrated to be more robust over alternatives. The proposed generators-selector framework not only enhances the quality and diversity of SQL queries but also outperforms previous methods. Overall, our proposed CHASE-SQL achieves the state-of-the-art execution accuracy of 73.0% and 73.01% on the test set and development set of the notable BIRD Text-to-SQL dataset benchmark, rendering CHASE-SQL the top submission of the leaderboard (at the time of paper submission).

한국어 요약

한 줄 요약

CHASE-SQL은 Text-to-SQL 작업에서 질 높고 다양한 SQL 후보를 생성하고 정확하게 선택하는 새로운 에이전트 기반 프레임워크로, BIRD 데이터셋에서 73.01%의 실행 정확도를 달성했다.

핵심 기여도

핵심 아이디어

CHASE-SQL은 단일 LLM 호출 내에서 복잡한 쿼리를 관리 가능한 하위 쿼리로 분해하는 divide-and-conquer 방법을 사용하며, 데이터베이스 엔진의 실행 단계를 반영한 chain-of-thought 추론을 적용한다. 또한, 테스트 질문에 맞춘 few-shot 예시를 생성하는 instance-aware synthetic example generation 기법을 도입하여 모델의 이해도를 높인다. 이는 기존의 단순한 프롬프팅 방식보다 더 구조화되고 정확한 SQL 생성을 가능하게 한다. 선택 단계에서는 이진 비교 기반의 선택 에이전트가 후보들을 순위 매기고, 최종적으로 가장 높은 점수를 받은 쿼리를 선택한다. 이는 self-consistency나 ranker 기반 방법보다 더 안정적이고 정확한 결과를 제공한다.

기술적 접근법

주요 결과

의의 및 한계

CHASE-SQL은 질 높고 다양한 SQL 후보를 생성하고, 이진 비교 기반의 선택 메커니즘을 통해 정확도를 극대화함으로써 Text-to-SQL 작업의 성능 한계를 극복한다. 특히, divide-and-conquer와 instance-aware synthetic example generation은 복잡한 쿼리 처리와 데이터베이스 스키마 이해도를 향상시키는 데 기여한다. 그러나, 본 연구는 특정 데이터셋(BIRD)에서의 성능을 강조했으며, 다른 도메인이나 데이터셋에서의 일반화 가능성은 명시되지 않았다. 또한, 선택 에이전트의 fine-tuning 과정에 사용된 데이터셋이 공개되지 않았다.

실용적 활용

CHASE-SQL은 SQL 문법을 모르는 사용자가 자연어로 데이터베이스를 질의할 수 있는 대화형 시스템, 자동화된 데이터 분석 플랫폼, 기업의 BI 도구 등에 적용 가능하다. 특히, 복잡한 데이터베이스 구조를 가진 산업 현장에서 유용하게 사용될 수 있으며, LLM 기반의 코드 생성 기술을 활용한 다양한 애플리케이션 개발에도 기여할 수 있다.