AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

Sarim Hashmi, Mukul Ranjan, Kshitij Mishra, Mikhail Kuznetsov, Praneeth Vepakomma, Nils Lukas

arXiv:2610.08773 · 2026-10-07 공개 · arXiv · PDF

curriculum-learning prompt-injection adversarial-training web-agents sim2real adversarial-defense task-curriculum web-world-model

Abstract

Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6\% relative to the base agent.

한국어 요약

한 줄 요약

AdvSim2Real은 웹 에이전트의 안정성과 능력을 동시에 향상시키기 위해 웹 월드 모델 내에서 커리큘럼, 에이전트, 적대자 정책을 공진화시키는 시뮬레이션 기반 훈련 방법이다.

핵심 기여도

핵심 아이디어

기존 방어 방법은 정적 인젝션 예시로 훈련하여 적응형 공격을 막지 못했다. AdvSim2Real은 웹 월드 모델 내에서 **커리큘럼**, **적대자**, **에이전트**를 동시에 훈련함으로써, 에이전트가 새로운 공격에 대응하는 능력을 키운다. 커리큘럼은 에이전트가 약 50% 성공하는 태스크를 제안하여 학습 효과를 극대화하고, 적대자는 **성공 플립**(success flip)이라는 지표로 보상받는다. 이는 정상 실행이 실패로 전환될 때만 보상이 주어지는 방식으로, 공격의 효과성을 정확히 평가한다.

기술적 접근법

주요 결과

의의 및 한계

AdvSim2Real은 웹 에이전트의 **능력과 안정성**을 동시에 향상시키며, 특히 **적응형 공격**에 효과적이다. 월드 모델 내에서 훈련한 성능이 실제 브라우저로 이전되는 점은 실용성 측면에서 중요한 기여다. 그러나 **Stage 2**는 월드 모델 내에서만 평가되며, 브라우저 환경에서의 공격 대응력은 명시되지 않았다. 또한, **적대자의 공격 가능성**(feasibility)을 평가하지 않는 점도 한계로 지적된다.

실용적 활용

AdvSim2Real은 웹 기반 자동화 시스템, 특히 **고객 관리, 데이터 수집, 온라인 트랜잭션** 등에서 공격에 강한 에이전트를 구축하는 데 활용 가능하다. 월드 모델 기반 훈련은 실제 환경으로의 이전을 용이하게 하므로, **보안이 중요한 산업**(금융, 의료, 정부)에서 실용적 가치가 높다.