AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
The preprint co-trains a web agent, task curriculum and adaptive attacker; its reported robustness gain is measured on 150 simulated web tasks.
TL;DR
- AdvSim2Real trains an agent alongside tasks and an injection attacker inside a frozen web world model.
- The authors report a 33.6% relative completion gain against an unseen adversary on 150 web tasks.
- The reported gain is simulator-based; the paper also reports capability transfer to a real browser, but not independent evaluation.
The study targets instructions embedded in third-party webpages that can redirect an agent while the page also contains information needed to complete the task. AdvSim2Real co-evolves the task curriculum, an injection adversary and the agent in a frozen web world model. [arXiv; Hugging Face Papers.] [1] [2]
On 150 web tasks, the authors report a 33.6% relative increase in completion against an unseen adversary compared with the base agent. The abstract says capability gains also carried over to a real browser, but the headline robustness figure is from the simulated evaluation. [arXiv; Hugging Face Papers.] [1] [2]
Why it matters
Web-agent security depends on handling hostile page instructions without discarding task-relevant page content. This preprint tests an adaptive training approach, while its simulator-based robustness figure and lack of independent replication limit what can be concluded.
Editor's note
Preprint; the 33.6% figure is a relative result reported by the authors on 150 simulated tasks. TODO-review: confirm Korean term for “web world model.”