Prompt injection
Prompt injection is when text the agent reads contains instructions, and it follows them.
A web page, an email, a file. The agent cannot reliably tell your instructions from instructions embedded in something it was asked to read.
This is the real reason an agent with wide permissions is dangerous. It is not that it goes rogue. It is that someone else can tell it what to do.
Was this page helpful?