Self-Organizing Agent Teams Learn to Reason Together

arXiv:2609.22682 · 2026-09-27 공개 · arXiv · PDF

benchmark-evaluation agent-collaboration mathematics-benchmarks self-organizing-agent-teams collaborative-computation role-learning reasoning-organization teamwork-strategies

Abstract

Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor cannot be specified in advance; teams must learn from experience how to organize reasoning as it unfolds. Human teams routinely adapt this way, while existing AI agent teams rely on fixed protocols, explicit task decomposition, or routing. We introduce Self-Organizing Agent Teams (SAT), fixed teams of AI agents that learn reusable strategies from prior collaborations to organize roles, conversational phases, participation, and information flow. These strategies enable what we call collaborative computation: agents exchange, challenge, repair, and synthesize partial reasoning into solutions no member produced independently. In two independent settings, we learn teamwork strategies that transfer unchanged to unseen benchmarks, using only 15 mathematics and 25 graduate-level knowledge problems. Across five mathematics and physics benchmarks, self-organizing teams average 66.7% accuracy, versus 48.8% for their strongest member, 58.7% for compute-matched inference by that agent, and 59.0% for a perfect router over members' independent answers; on AIME 2026, they exceed this router by 13.4 points. Because gains vary across benchmarks, we ask when self-organizing collaboration helps. Across eight benchmarks, demonstrability (the organizational-psychology construct of whether a team can distinguish correct from incorrect reasoning) strongly tracks improvement over the strongest member (Spearman ρ=0.90, p=0.005): teams benefit most when correct reasoning can be recognized once it appears. More broadly, these results suggest that organization itself can become an agent capability: agent teams can learn how to reason together and produce solutions their members could not reach independently.

한국어 요약

한 줄 요약

AI 에이전트 팀이 협력적 추론을 학습해 개별 에이전트가 도출하지 못한 정답을 생성하는 Self-Organizing Agent Teams (SAT)를 제안한다.

핵심 기여도

핵심 아이디어

기존 AI 팀은 고정된 프로토콜, 명시적 작업 분할, 라우팅에 의존하지만, SAT는 협력 경험에서 팀워크 전략을 학습해 역할, 대화 단계, 참여, 정보 흐름을 조직한다. 이는 "협력적 계산(collaborative computation)"을 가능하게 하며, 에이전트들이 부분적 추론을 교환하고, 수정하며, 종합하여 개별 에이전트가 도출하지 못한 정답을 생성한다. 학습된 전략은 추론 구조가 미지일 때 유용하며, 팀이 문제 해결 과정에서 스스로 조직을 학습할 수 있다는 점에서 혁신적이다. 학습은 오프라인으로 이루어지고, 추론 시에는 전략 풀에서 최적 해를 선택한다.

기술적 접근법

주요 결과

의의 및 한계

SAT는 AI 팀이 협력적 추론을 학습해 개별 에이전트가 도출하지 못한 정답을 생성할 수 있음을 보여준다. 이는 조직 자체를 에이전트의 능력으로 확장하는 기초 연구로, 팀워크 전략이 문제 해결 과정에서 유기적으로 발전할 수 있음을 시사한다. 그러나 풍부한 추론 생성이 정확한 선택으로 이어지지 않는 한계가 있으며, 선택 메커니즘과 증명 형식의 개선이 필요하다. Demonstrability는 팀 성능과의 상관관계를 제공하지만, 팀워크 학습 과정에는 직접적으로 활용되지 않는다.

실용적 활용

SAT는 복잡한 수학 문제 해결, 과학적 추론, 대규모 지식 기반 질문 응답 등에서 활용 가능하다. 특히, 개별 에이전트가 독립적으로 해결하지 못하는 문제에서 협력적 추론을 통해 새로운 해법을 도출할 수 있어, 연구 및 산업 분야에서 혁신적인 도구로 활용될 수 있다.