Insights / AI security
Prompt injection, for the people who sign off on AI
The most common AI vulnerability isn’t exotic. It’s what happens when a model reads text written by someone you don’t trust.
Hidden instruction caught and stripped
Language models follow instructions. The trouble is that they can’t reliably tell your instructions apart from instructions hidden in the content they read: an email, a web page, a PDF, a support ticket. When an attacker plants instructions in that content and the model follows them, that’s prompt injection. It sits at number one in the OWASP Top 10 for LLM Applications.
Why it matters now
A chatbot that only talks is embarrassing when it’s tricked. An assistant that can read your inbox, search your files and send email is dangerous. Security researcher Simon Willison calls the risky combination the “lethal trifecta”: a system that has access to private data, is exposed to untrusted content, and can communicate externally. Put all three together and a single malicious email can instruct your assistant to quietly send your data somewhere else.
Questions to ask before approving an AI tool
- What can it read? Every source of outside text (email, web, uploaded files) is a way in.
- What can it do? Sending messages, moving files and calling APIs turn a trick into a breach.
- Who approves risky actions? A human should confirm anything that sends data out or can’t be undone.
- Is it logged? If something goes wrong, you need to see what the model read and what it did.
- Has anyone attacked it? Vendors test their models; nobody has tested your configuration but you.
The reliable defense is design.
There’s no complete fix yet
Filters and better models reduce the risk but don’t remove it. The reliable defense is design: limit what each AI tool can reach, break the trifecta where you can, require approval for consequential actions, and test the result the way an attacker would. That’s what our AI red teaming and AI security work does.