MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution

Wei Tao, Yucheng Zhou, Wenqiang Zhang, Yu-Xi Cheng

arXiv:2403.17927 · 2026-07-27 공개 · arXiv · PDF

code-generation swe-bench multi-agent llm-based agent-framework software-evolution llm-collaboration github-issue-resolution

Abstract

In software development, resolving the emergent issues within GitHub repositories is a complex challenge that involves not only the incorporation of new code but also the maintenance of existing code. Large Language Models (LLMs) have shown promise in code generation but face difficulties in resolving Github issues, particularly at the repository level. To overcome this challenge, we empirically study the reason why LLMs fail to resolve GitHub issues and analyze the major factors. Motivated by the empirical findings, we propose a novel LLM-based Multi-Agent framework for GitHub Issue reSolution, MAGIS, consisting of four agents customized for software evolution: Manager, Repository Custodian, Developer, and Quality Assurance Engineer agents. This framework leverages the collaboration of various agents in the planning and coding process to unlock the potential of LLMs to resolve GitHub issues. In experiments, we employ the SWE-bench benchmark to compare MAGIS with popular LLMs, including GPT-3.5, GPT-4, and Claude-2. MAGIS can resolve 13.94% GitHub issues, significantly outperforming the baselines. Specifically, MAGIS achieves an eight-fold increase in resolved ratio over the direct application of GPT-4, the advanced LLM.

한국어 요약

한 줄 요약

MAGIS는 GitHub 이슈 해결을 위한 LLM 기반 멀티에이전트 프레임워크로, GPT-4 대비 8배 높은 해결률을 달성했다.

핵심 기여도

핵심 아이디어

LLM은 함수 수준 코드 생성에 강하지만, 전체 저장소 수준의 GitHub 이슈 해결에는 한계가 있다. 이는 특히 컨텍스트 길이 제한과 코드 변경 위치 파악의 어려움에서 비롯된다. MAGIS는 이 문제를 해결하기 위해 4개의 전용 에이전트를 설계하여 협업 기반의 작업 흐름을 구축했다. Repository Custodian은 수정이 필요한 파일을 식별하고, Manager는 작업 계획을 수립하며, Developer는 코드 변경을 수행하고, QA Engineer는 변경 사항을 검토한다. 이는 LLM의 한계를 극복하고 저장소 수준의 코드 진화를 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

MAGIS는 LLM이 저장소 수준의 코드 진화를 수행하는 데 기여하며, GitHub 이슈 해결의 새로운 패러다임을 제시한다. 특히, 4개 에이전트의 협업 구조는 LLM의 단점을 보완하고, 복잡한 작업을 구조화하여 효율성을 높인다. 그러나 MAGIS는 여전히 86% 이상의 이슈를 해결하지 못하며, 이는 이슈의 복잡성이나 컨텍스트 부족 등 추가적인 제약이 존재함을 시사한다. 또한, 에이전트 간 커뮤니케이션 비용이나 반복적 수정 과정의 시간 소요도 한계로 작용할 수 있다.

실용적 활용

MAGIS는 오픈소스 프로젝트나 대규모 저장소의 유지보수 작업에 활용 가능하다. 특히, 이슈가 빈번히 발생하는 인기 있는 프로젝트에서 개발자 부담을 줄이고 자동화된 코드 진화를 지원할 수 있다. 또한, CI/CD 파이프라인에 통합하여 지속적인 품질 관리와 자동 이슈 해결을 구현할 수 있다.