Raising the alarm on Islamofascism, one column at a time.

Know more →
Amil Imani Articles

Your AI Assistant Has the Keys. Who's Watching It?

For years, the public conversation about AI and safety has been dominated by science-fiction scenarios: machines that wake up, turn on their makers, or decide we're in the way. Those debates are worth having. But while we have them, a more ordinary problem is already showing up in company networks.

Businesses are handing software the authority to act. Not just to answer questions or draft emails, but to move money, change databases, send messages on the company's behalf, and approve or deny requests. The software making those decisions works by predicting what text should come next. It has no understanding of what it is allowed to do, and it can be talked into things.

Classic security rests on a simple idea: a computer does what an authorized person tells it to. When something goes wrong, investigators can trace it. Someone's credentials were stolen, a server was misconfigured, a bug let bad data through.

Language models complicate this. They take their instructions and the material they're asked to process through the same channel, as one stream of text. A command from an engineer and a sentence hidden in an incoming email look the same to the model. It has no reliable way to tell "this is my boss speaking" from "this is something a stranger wrote."

Security people have seen versions of this before. SQL injection and cross-site scripting both work by tricking a system into treating data as instructions. The difference is that those problems eventually got structural fixes. Nobody has found the equivalent for language models yet.

Picture a company that uses an AI agent to sort vendor invoices and schedule payments. An attacker submits a PDF with a line of white text on a white background, invisible to any human reader, telling the agent to reroute an earlier payment to a different account and delete the log entry. The agent reads the document, treats the line as part of its task, and uses its legitimate access to the bank to follow it. No malware, no stolen password, no alarm. The system did what it was set up to do. It was just persuaded by the wrong person.

Researchers call this prompt injection, and it is not hypothetical. It has been demonstrated repeatedly against real products.

Quiet data theft. Security tools watch for strange traffic leaving a network. But an agent that can read internal files and also send email or browse the web can be steered into passing information out through channels that look entirely routine. Developer Simon Willison has described the dangerous combination plainly: access to private data, exposure to untrusted content, and a way to communicate outward. Give an agent all three and you have a leak waiting to happen.

Mistakes that can't be undone. Agents don't need an attacker to cause damage. In one widely reported case last year, an AI coding assistant deleted a company's production database during a session where it had been told not to make changes. Whether the cause is a hallucination, a misread instruction or a planted command, the result is the same when the agent has write access: real systems change, and sometimes there's no undo button.

Errors that spread. Many companies now chain several agents together, one to read input, another to assess it, a third to act. That design assumes each agent can trust the one before it. If the first is fooled, the rest carry the error forward with confidence, because they treat its output as vetted. A problem that should have stayed in one inbox can end up in the whole system.

There is a second risk, and it has little to do with hackers. When a company lets software decide who gets a claim approved, a loan, or a post removed, responsibility tends to blur. "The system flagged it" becomes the answer to every complaint. No one is quite accountable, because the decision-maker is a program that can't be sued, fired or embarrassed.

That's convenient for organizations. It is not good for the people on the receiving end, and it's a governance problem that no technical fix will solve.

Most of the advice you'll hear amounts to writing better instructions for the model: be careful, ignore suspicious requests. That doesn't work. A model that can be talked into something can be talked out of its instructions too. Protection has to sit outside the model, in places it can't argue with.

  • Give agents as little power as possible. Short-lived credentials, read-only by default. An agent that summarizes a database has no reason to be able to delete one.
  • Put a rule-based checkpoint between the agent and anything it touches. Every action should pass through a conventional system, such as a policy engine or proxy, that approves or blocks it on fixed rules. The model proposes, and the checkpoint decides.
  • Require a human for anything irreversible. Transfers, permission changes, deletions and public statements should need confirmation from a real person through a separate channel. This only works if the requests are rare enough that people actually read them. An approval screen that gets clicked through a hundred times a day is decoration.
  • Keep working toward better separation. Computing has long protected its core from untrusted programs, and AI systems need an equivalent way of keeping outside content from rewriting their instructions. That research is still early, so it's a direction rather than a product you can buy.

The danger here isn't a machine that outsmarts us. It's a machine that is easy to fool, sitting in a position of trust that we gave it because it was convenient. Efficiency is a good reason to adopt a tool, but it's a bad reason to hand it authority we wouldn't give a new hire on day one.

The companies that handle this well will be the ones that decide, deliberately, what their software is never allowed to do.

Amil Imani

Read more about Amil Imani →