OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

Yunhao Chen, Xin Wang, Yixu Wang, Yi Liu, Jie Li, Yan Teng, Xingjun Ma, Xia Hu, Yu-Gang Jiang

arXiv:2608.00677 · 2026-08-13 공개 · arXiv · PDF

tool-calls safety-benchmarks attack-success-rate agent-safety long-horizon-workflows environment-evolution agent-red-teaming openart

Abstract

AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedly modified and reused across long-horizon workflows. Current safety benchmarks often fail to capture these cumulative risks because they focus on short, static tasks. To address these limitations, we introduce OpenART, an open-ended arena for scalable agent red teaming through environment evolution. OpenART provides over 10,000 validated stateful scenarios across 50 domains, drawing from a pool of more than 500,000 tools and skills. These tasks require a median of 97 tool calls and enable unified evaluation across 75 different agent-model configurations. To systematically explore these evolving attack surfaces, we propose the Evolutionary Markov Hypergraph Attack (EMHA). EMHA is a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates. Throughout the evaluation, task objectives remain fixed while only the environment state changes. Across all configurations, EMHA achieves a pooled Attack Success Rate (ASR) of 85.0%. Its advantage over instruction-only evolution increases from approximately 2% on simple environments to over 17% on the most complex ones, demonstrating that environment evolution increasingly exposes safety failures as task complexity grows. Furthermore, our analysis shows that the specific runtime implementation of an agent explains a significant portion of safety variation beyond the underlying model's capabilities. These results establish OpenART as a scalable foundation for studying agent safety in complex, evolving environments.

한국어 요약

한 줄 요약

OpenART는 복잡하고 진화하는 환경에서 에이전트 안전성을 평가하기 위한 대규모 레드팀 테스트 플랫폼으로, 환경 진화를 통해 85.0%의 공격 성공률을 달성한다.

핵심 기여도

핵심 아이디어

기존 에이전트 안전성 평가가 단기적이고 정적 환경에 국한되며, 복잡한 상태 변화를 반영하지 못한다는 문제를 해결하기 위해 OpenART는 **실행 가능한 환경**을 레드팀 테스트의 핵심 단위로 설정한다. 이는 단일 프롬프트나 작업이 아닌, **지속적으로 진화하는 환경**을 기반으로 안전성을 평가하는 새로운 접근법이다.

OpenART는 **SkillNet**에서 수집한 500,000개 이상의 도구, MCP, 기술을 기반으로 **의존성 그래프**를 합성하여 시나리오를 생성하고, **99.3%의 정확도**를 달성하는 자동화된 행동 탐지기를 통해 검증한다. 이 시나리오는 15개 배포된 에이전트와 5개 기초 모델을 결합해 75개의 에이전트-모델 구성으로 테스트된다.

EMHA는 **블랙박스 정책**으로, 환경 진화를 **하이퍼그래프 경로 탐색**으로 모델링하며, 매개변수 업데이트 없이 **피드백 기반 상태 전이**를 반복적으로 수행한다. 이는 에이전트의 단기적 행동이 아닌, **장기적 상호작용 궤적**을 통해 안전 실패를 평가하는 핵심 아이디어이다.

기술적 접근법

주요 결과

의의 및 한계

OpenART는 기존 에이전트 안전성 평가의 단기적, 정적 한계를 극복하고, **지속적 상태 변화**를 반영한 **장기적 상호작용 궤적**을 평가하는 기반을 제공한다. 특히, **EMHA**는 매개변수 업데이트 없이도 높은 공격 성공률을 달성하며, **블랙박스 정책**의 유연성을 입증한다.

그러나 OpenART는 **특정 도구와 기술 집합**에 의존하며, **모든 에이전트와 환경**을 포괄하지는 않는다. 또한, **공격 벡터의 선택**은 에이전트의 지원 범위에 따라 달라지므로, **일관된 비교**를 위해서는 추가 연구가 필요하다.

실용적 활용

OpenART는 **AI 에이전트의 장기적 안전성**을 평