LLM web agents increasingly rely on web content as input, exposing them to indirect prompt injection embedded in webpages. While prior work has shown such attacks in controlled settings, it remains unclear whether prompt injection is already deployed in the wild and what role it plays in the web ecosystem. In this paper, we conduct the first large-scale empirical study of in-page prompt injection. Analyzing 1.2B URLs across 24.8M hosts, we identify 15.3K validated prompt injections, with a small set of reused templates accounting for the majority of cases. Our analysis reveals a multi-stakeholder phenomenon, with injections serving diverse offensive and defensive objectives, including system disruption, reputation manipulation, data protection, and AI bot detection, and target a range of agents from web crawlers and search systems to customer-support and HR automation pipelines. Most injections (70%) are delivered in non-visible channels like HTTP headers, JS comments, or HTML-embedded hidden content. We assess their effectiveness through 5,200 systematic experiments across 13 models and four page representations, observing up to 8% effectiveness for smaller models on plain-text inputs, with lower effectiveness for other representations. Overall, our results show that in-page prompt injection is emerging as an important source of friction between LLM-based agents and the broader web ecosystem.
ACM Conference on Computer and Communications Security (CCS)
2026-11-15
2026-09-18