
Exploration of Indirect Prompt Injection Exploits in AI Agents
The video explores indirect prompt injection as a method to exploit AI agents integrated into applications like email, calendars, and code repositories. Hosted by NahamSec and featuring Mike (Toxic), a red team researcher from Zennity, it demonstrates live labs with three difficulty levels, where the goal is to extract a password reset token from a victim’s inbox. Techniques include reconnaissance (listing available tools), obfuscation (adding characters like periods or 'cheese' to bypass regex-based output guardrails), and framing (altering context to mislead the model). The labs emulate real-world scenarios, including a $3,000 bug bounty vulnerability, and highlight challenges like non-deterministic model behavior and input/output guardrails. Tools such as Parseltongue (for encoding prompts) and Burp Collaborator-like listeners are used for exfiltration. Mike and NahamSec will present a talk, 'XFill: Everything – A Year of Stealing Data from AI Agents,' at Def Con’s Bug Bounty Village on Saturday at 2:00 PM. The video emphasizes that indirect prompt injection is delivered via emails, websites, PDFs, or code comments, making it a critical attack vector for external adversaries.