CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

Hang He, Li Wang, Hao Chen, Yuchen Shao, Yuling Shi, Lisheng Wang, Peiyang Liu, Goose Lin, Zaiyuan Wang, Haiying Sun, Ting Su, Chengcheng Wan

arXiv:2610.07557 · 2026-10-07 공개 · arXiv · PDF

long-horizon tool-use coding-agents static-analysis checkerlab defect-specification checker-synthesis cve

Abstract

Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection and rarely assess whether an agent can develop a working checker in a repository from start to finish. We introduce CheckerBench, an executable benchmark of 300 tasks derived from 297 CVEs across 167 repositories, 85 CWEs, and five language ecosystems. Each task includes vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold. We further introduce CheckerLab, a common evaluation framework that independently rebuilds submitted checkers and measures vulnerable-fixed diagnostic contrast, patch localization, false positives, and tool use. Across 21 model-harness configurations and three independent repeats per configuration, mean Pass@1 is 32.30%, while the best reaches 45.33%. These results show that reliable, reusable checker development remains challenging for current coding agents.

한국어 요약

한 줄 요약

CheckerBench는 코드 생성 에이전트가 정적 분석 체커를 독자적으로 개발하는 능력을 평가하는 실행 가능한 벤치마크로, 300개의 실제 취약점 기반 작업을 포함한다.

핵심 기여도

핵심 아이디어

CheckerBench는 기존 코드 생성 에이전트 벤치마크가 체커 개발을 평가하지 못하는 문제를 해결하기 위해 설계되었다. 기존 벤치마크는 패치 생성이나 취약점 탐지에 집중하며, 체커 개발의 전체 과정을 평가하지 않는다. CheckerBench는 취약점 사양을 해석하고, 리포지토리를 검토하여 분석기 특정 로직을 구현하고, 컴파일 및 분석 피드백을 반복적으로 활용하여 체커를 개선하는 과정을 평가한다. CheckerLab은 제출된 체커를 독립적으로 재빌드하고 진단 대비, 패치 로컬라이제이션, FP 등을 측정함으로써 체커의 신뢰성과 재사용성을 평가한다.

기술적 접근법

주요 결과

의의 및 한계

CheckerBench는 체커 개발 과정을 체계적으로 평가할 수 있는 첫 번째 실행 가능한 벤치마크로, 코드 생성 에이전트의 장기적 작업 능력을 평가하는 기준을 제공한다. CheckerLab은 체커의 진단 대비, 패치 로컬라이제이션, FP 등을 정량적으로 평가하여 신뢰성 있는 체커 개발을 측정한다. 그러나 현재 에이전트는 시간 내에 완료하거나 의미 모델링을 정확히 수행하는 데 어려움이 있으며, 이는 체커 개발의 복잡성과 다단계 작업의 필요성을 보여준다. 또한, 기존 벤치마크와의 환경 차이로 인해 결과 비교가 어려운 한계가 있다.

실용적 활용

CheckerBench는 소프트웨어 보안 연구, 정적 분석 도구 개발, 코드 생성 에이전트 평가에 활용될 수 있다. 특히, 취약점 사양을 기반으로 신뢰성 있고 재사용 가능한 체커를 자동 생성하는 시스템 개발에 기여할 수 있다. CheckerLab은 다양한 분석 환경에서 체커의 정확성과 효율성을 평가하는 표준화된 프레임워크로 활용될 수 있다.