long-horizon swe-bench coding-agents terminal-bench qwen3-5-9b self-critique policy-models verbal-critic
Abstract
Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered. We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved. Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence before delivery, and tracks the agent's subsequent actions to distinguish mere compliance from actual resolution. As a test-time critic, Opera improves the resolve rate of non-critic agents by up to 12.4, 15.0, and 8.9 percentage points on Terminal-Bench 2.1, a SWE-Bench Pro subset, and DeepSWE v1.1, respectively, across four policy models, and achieves the highest mean resolve rate among competitive critic baselines on all three benchmarks, and also improves policy models when the policy critiques itself. Beyond inference, Opera-guided rollouts provide approximately on-policy training data: fine-tuning Qwen3.5-9B on them improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 percentage points without a critic at inference time, matching fine-tuning on rollouts from a stronger model, while preserving its performance when switching harness, i.e., from Openhands to Terminus-2, which the latter substantially degrades. Our code is available at: https://github.com/dongyuanjushi/Opera.
한국어 요약
한 줄 요약
Opera는 장기적 코드 생성 작업에서 지속적인 피드백을 관리하는 언어 기반 크리틱 프레임워크로, 테스트 시 15.0%까지 해결률을 향상시킨다.
핵심 기여도
- **Opera 프레임워크**를 제안: 각 수정을 지속적인 노트로 관리하며, 피드백 전에 증거를 검토하고, 해결 여부를 추적한다.
- **테스트 시 크리틱**으로서, Terminal-Bench 2.1, SWE-Bench Pro, DeepSWE v1.1에서 각각 최대 12.4%, 15.0%, 8.9%의 해결률 향상.
- **트레이닝 레시피**: Opera가 생성한 약한 정책 모델의 롤아웃 데이터를 사용한 미세조정이 강력한 모델의 지도 학습과 동일한 성능(+10.2%)을 달성.
핵심 아이디어
기존 크리틱은 작업 중간 피드백을 생성하지만, 피드백 이후의 결과를 추적하지 않는다. Opera는 이 문제를 해결하기 위해 **지속적인 수정 노트**(finding note)를 도입하여, 피드백이 실제로 문제를 해결했는지 확인한다. 이는 **operator critic**과 **audit** 모듈을 결합하여, 각 문제에 대한 증거, 수정 방향, 해결 기준을 명시하고, 피드백 전에 검토한다. 또한, **주기적 및 이벤트 기반 트리거**(periodic and event-driven triggers)를 통해 언제 개입할지 결정하며, 정책 모델이 스스로 자신을 평가하는 경우에도 성능 향상이 가능하다는 점에서 차별화된다.
기술적 접근법
- **Operator critic**: 각 문제를 증거와 연결하고, 수정 방향 및 해결 기준을 정의.
- **Audit**: 피드백 전에 증거와 일치 여부를 검토하여 오도 가능성을 줄임.
- **Finding note**: 피드백이 실행되고 해결되었는지 추적.
- **Triggering**: 주기적 및 이벤트 기반으로 피드백 시점을 결정.
- **Policy model**: Qwen3.5-9B, Qwen3.8-27B 등 사용.
- **Fine-tuning**: Opera가 생성한 약한 정책의 롤아웃 데이터로 Qwen3.5-9B 미세조정.
주요 결과
- **Terminal-Bench 2.1**: 12.4% 향상 (기존 정책 대비).
- **SWE-Bench Pro**: 15.0% 향상.
- **DeepSWE v1.1**: 8.9% 향상.
- **Qwen3.5-9B 미세조정**: SWE-Bench Pro에서 +10.2% 향상, Terminal-Bench 2.1 성능 유지.
- **Critic 기반 정책 모델**: Opera가 자체 평가 시에도 성능 향상.
의의 및 한계
Opera는 테스트 시점에 기존 정책 모델에 추가적인 피드백을 제공함으로써 해결률을 향상시키며, 학습 데이터 생성에서도 유용하다. 특히, 약한 정책 모델의 자체 롤아웃 데이터를 활용한 미세조정이 강력한 모델의 지도 학습과 동일한 성능을 달성하는 점에서 실용적 가치가 크다. 다만, Opera는 단일 모델과 단일 벤치마크에서 훈련되었으며, 더 큰 데이터셋이나 강화 학습과의 결합 가능성은 아직 검증되지 않았다.
실용적 활용
Opera는 코드 생성 에이전트가 장기적 작업에서 반복적 오류를 줄이고, 실제 해결을 촉진하는 데 유용하다. 소프트웨어 엔지니어링 자동화, 코드 리뷰 시스템, 또는 코드 학습 플랫폼에 적용 가능하며, 정책 모델의 자체 평가 능력을 향상시키는 데도 활용할 수 있다.