WideSWE: Can Coding Agents Coordinate Changes Across Repositories?

Baoyi Wang, Xingliang Wang, Jinyang Wu, Keming Wu, Chen Zhi, Jianwei Yin

arXiv:2609.33382 · 2026-09-29 공개 · arXiv · PDF

code-generation coding-agents evaluation-framework task-success llm-codex llm-gpt repository-coordination cross-repository

Abstract

Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We introduce WideSWE to evaluate coding agents on such cross-repository tasks. Mining and reviewing changes across 103 software ecosystems yields 120 real-world tasks, balanced between 60 bug fixes and 60 features. We derive prompts from related issues and pull requests. We systematically review and adapt hidden tests to support diverse correct implementations while preserving required behavior and regression checks. Across seven agent configurations, full task success ranges from 10.83% to 42.50%, with the configuration pairing Codex CLI with GPT-5.6-sol achieving the highest rate. Trajectories show agents failing to identify necessary changes, recognizing changes but leaving them unfinished, or modifying the required repositories without fully satisfying the request. To examine whether working on one repository at a time can alleviate these difficulties, we compare it with joint execution under identical prompts. Independent execution mainly recovers omitted work and is less effective at correcting previously attempted but unsuccessful implementations. Joint execution can use information from related repositories to guide implementation and verification. Code is available at https://github.com/ZJU-ACES-ISE/WideSWE.

한국어 요약

한 줄 요약

WideSWE는 여러 레포지토리 간 조율이 필요한 실제 개발 작업을 평가하는 벤치마크로, 최대 42.50%의 작업 성공률을 보임.

핵심 기여도

핵심 아이디어

기존 평가가 단일 레포지토리에 집중한 반면, WideSWE는 여러 레포지토리 간 조율이 필요한 작업을 평가하는 새로운 접근법을 제시한다. 예를 들어, Sentry의 다중 SDK 작업이나 Godot/Native 작업처럼 여러 레포지토리에서 변경이 필요한 실제 상황을 반영한 작업을 설계했다. 작업 성공은 모든 대상 레포지토리에서 F2P 및 P2P 테스트가 통과할 때만 인정되며, 이는 기존 평가 방식과 구별된다. 또한, 작업 요구사항과 테스트 간 불일치를 해결하기 위해 규칙 기반의 수동 테스트 수정을 도입하여, 다양한 올바른 구현을 허용하면서도 요구 사항을 유지하도록 했다.

기술적 접근법

주요 결과

의의 및 한계

WideSWE는 단일 레포지토리 중심 평가에서 벗어나, 실제 개발 환경에서 요구되는 다중 레포지토리 간 조율 능력을 평가하는 새로운 기준을 제시한다. 특히, 공동 실행이 관련 레포지토리 정보를 활용해 구현과 검증을 돕는다는 점에서 실용적 가치가 있다. 그러나 42.50%의 최고 성공률은 여전히 낮아, 에이전트가 전체 작업 범위를 정확히 파악하고 실행하는 능력이 한계라는 점을 보여준다. 또한, 작업 유형 (버그 vs. 기능)과 언어 다양성에 따라 성능 차이가 크므로, 이에 대한 추가 연구가 필요하다.

실용적 활용

WideSWE는 다중 레포지토리 기반의 소프트웨어 개발 프로젝트 (예: Sentry, Kubernetes)에서 에이전트의 협업 능력을 평가하는 데 활용 가능하다. 또한, 공동 실행 방식은 팀 개발 환경에서 정보 공유와 작업 조율을 지원하는 시스템 설계에 참고가 될 수 있다.