Build Security In · Issue 3

Output is not evidence

Three pieces of AI security research landed in a single month, on three different problems. They all point at the same soft spot.

Output is not evidence. Three AI security stories, one job the machine still can't do for you.

Key Takeaways

  • Agent data injection forges the small facts an agent trusts, so it completes the task you asked on top of a lie, and the approval prompt looks routine.
  • AI output is not evidence: a polished report or payload does not prove that a bug is reachable, exploitable or real in the deployed system.
  • Model vendors can harden their models, but you still own every email, web page, repository and tool you connect an agent to.
  • Draw the trust boundary explicitly: decide what each agent may read, run and trust, and give it least privilege instead of relying on the approval click.
  • Treat AI output as a lead, not a finding, and keep the human ability to prove reachability and impact.

1. The agent that trusts a lie

Start with the unsettling one. Researchers described a new attack they call agent data injection. It does not hijack your AI agent's task. It forges the small facts the agent quietly trusts: who sent an email, the ID of a button, the record of a step that never ran. The agent then completes the exact job you asked, on top of a lie.

A planted product review made web agents click Buy Now instead of Read More. A faked GitHub author line made coding assistants run a stranger's command. Every major model tested fell for it, up to half the time. And the approval prompt does not save you, because the agent's reasoning is built on the forged facts, so the request it shows you looks completely routine. You approve a lie that reads like the truth.

The agent completes the exact job you asked, on top of a lie.

2. Output is not evidence

The second piece, written by a SANS Fellow, says the quiet part plainly. AI can read code, generate payloads and write a polished vulnerability report in seconds, complete with a CVSS 9.8. None of that proves the bug is reachable, exploitable, or real in your deployed system. Bug bounty programs are already drowning in AI slop, reports that look right and prove nothing, and the result is not better security. It is a longer triage queue.

“Looks vulnerable” is not vulnerable. The proof still takes judgement: does the input reach the dangerous call, is authorization enforced somewhere else, which trust boundary was actually crossed, what is the demonstrated impact. That knowledge is earned by doing the work, and it is exactly the muscle that goes soft when you let the model do the thinking too early.

Looks vulnerable is not vulnerable.

3. The machine red-teams itself

The third is the hopeful one, with a sting. OpenAI built an automated red-teamer that attacks its own models at machine speed and now beats human red-teamers at finding indirect prompt injections. They used it to train their newest model down to a 0.05% failure rate against those attacks. Adversarial training works, and the model makers are taking this seriously.

But the same report is a reminder that the attack surface widens every time an agent is wired to an email, a web page, a repo or a tool. In one test the red-teamer talked an AI vending agent into dropping a price to 50 cents and cancelling a stranger's order. The vendor hardens the model. You still own everything you connected it to.

What ties them together

Put the three side by side and the pattern is hard to miss. AI is now confidently wrong at machine speed, in both directions. It produces findings that look real and are not. It trusts data that looks real and is not. The one thing that separates the convincing from the true, in every one of these stories, is a human who can verify: verify the finding against the deployed system, verify who really sent the data, verify what the agent is actually allowed to do.

That is not a tool you buy. It is a capability you build. The classic lesson underneath all three stories is one traditional software learned the hard way: keep code and data apart, then keep trusted data apart from untrusted. Agents did the first half and skipped the second. Your job is to put that boundary back, on purpose.

What to do

Four things, none of them exotic:

  1. Draw the trust boundary explicitly. Decide what each agent may read, run and trust, and assume anything touching outside content can be forged.
  2. Treat AI output as a lead, not a finding. Promote it only after a human proves reachability and impact in the real environment.
  3. Give agents least privilege, and do not lean on the approval click. If the reasoning behind it can be faked, the click is theatre.
  4. Keep the human muscle in shape. Reward validated impact over volume, and make sure the person can still explain the exploit without the model.
Confidently wrong, at machine speed.

None of this is anti-AI. The speed is real and worth having. But the moment you cannot check what the machine produced, or what it believed, you have handed your judgement to something that is confidently wrong a meaningful share of the time. That is the whole job now.

This is what we train, live and hands-on, not on slides. In workshops we use your own code and pipeline wherever that is practical and agreed. The engineering side is here: cyberment.ee/ai-code-security.

Prove it. Then trust it.

Sources

Agent data injection (The Hacker News, Jul 2026)

AI can find bugs, but human knowledge still proves them

OpenAI's GPT-Red automates prompt injection testing

Newsletter

Get the next issue by email

Build Security In: the engineering side of security. Secure coding, DevSecOps, AI code and NIS2, for teams who ship. One email per new issue.

All articles