Mara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution

Yubin Lyu, Fu Li, Jiawei Fei, Yang Zhao, Weixing Mei, Yinan Wu

arXiv:2609.35855 · 2026-10-11 공개 · arXiv · PDF

harness-optimization appworld skill-optimization ai-systems terminalbench musesique mara-chain auto-evolution

Abstract

Optimizing deployed AI systems increasingly amounts to editing prompts, skills, harnesses, and code rather than model weights. Existing approaches commonly optimize these artifacts through propose-evaluate-select procedures, where candidate configurations are evaluated and only those meeting an acceptance criterion are selected. Yet our analysis shows that discarded candidates often contain information critical for subsequent optimization. Discarding them causes later proposals to revisit the same failure modes. We introduce Mara Chain, a refinement procedure that turns rejected candidates into stepping stones. Rather than discarding a rejected candidate, Mara Chain retains and iteratively refines it using evidence accumulated across preceding attempts. The procedure limits each refinement chain to a fixed depth and applies Pareto-filtered Top-N selection to bound the candidate pool. Across AppWorld skill optimization, TerminalBench 2.1 harness optimization, and MuSiQue retrieval-pipeline optimization, Mara Chain delivers greater task-performance gains with fewer rollouts. It outperforms GEPA, ACE, and SkillOpt-Lite by up to 20.5% in relative performance on AppWorld, reaching the target score with 65.5% fewer rollouts than GEPA. It improves the pass rate by 20.2 and 22.5 percentage points over AHE and Meta-Harness on TerminalBench 2.1, respectively, and improves MuSiQue test nDCG@10 and Recall@10 by 0.104 and 0.131 over a hand-written retrieval pipeline.

한국어 요약

한 줄 요약

Mara Chain은 거부된 후보를 반복적으로 개선함으로써 AI 시스템 최적화 성능을 20.5%까지 향상시키는 새로운 절차를 제시한다.

핵심 기여도

핵심 아이디어

기존 AI 시스템 최적화는 **propose-evaluate-select** 절차를 사용하여 후보 설정을 생성하고 평가하며, 기준에 부합하는 후보만 선택한다. 그러나 실패한 후보는 종종 이후 최적화에 중요한 정보를 포함하고 있음에도 불구하고 무시된다. 이는 동일한 실패 패턴을 반복하게 만들며, 최적화의 진전을 저해한다.

Mara Chain은 이러한 문제를 해결하기 위해 **거부된 후보를 반복적으로 개선**하는 절차를 도입한다. 각 후보는 실패한 rollout, 잔차 실패, 분석 결과, 변경 내역을 포함한 정보를 보존하고, 이 정보를 바탕으로 다음 후보를 생성한다. 이는 **역사 조건에 기반한 반복 개선**(history-conditioned refinement)으로, 실패 경험을 누적하여 최종적으로 목표를 달성하도록 유도한다.

기술적 접근법

주요 결과

의의 및 한계

Mara Chain은 AI 시스템 최적화에서 실패한 후보를 자원으로 활용함으로써, 기존 방법이 직면하는 **지속적인 실패 장벽**(persistent failure barriers)을 극복한다. 특히, **복잡한 다단계 작업**(예: TerminalBench 2.1의 sanitize-git-repo)에서 반복적인 부분 수정을 누적하여 최종 성공을 유도하는 점에서 학술적·실용적 가치가 높다.

그러나, **Mara Chain은 실패한 후보의 정보를 효과적으로 활용하는 데 의존**하므로, 초기 후보가 전혀 유의미한 정보를 제공하지 못하는 경우 성능 개선이 제한될 수 있다. 또한, **Pareto-filtered Top-N 선택**은 후보 풀을 제한하지만, 일부 유용한 후보가 제외될 가능성도 존재한다.

실용적 활용

Mara Chain은 **프롬프트, 스킬, 허네스, 코드** 등 모델 가중치를 수정하지 않고 AI 시스템을 최적화해야 하는 상황에서 유용하다. 예를 들어, **AppWorld의 스킬 최적화**, **TerminalBench의 에이전트 허네스 최적화**, **MuSiQue의 검색 파이프라인 최적화**와 같은 다양한 도메인에서 적용 가능하다. 특히, **제한된 rollout 예산**이 있는 환경에서 성능 향상을 추구하는 연구 및 산업 현장에 적합하다.