Research · White Paper

Securing Tool-Using AI Agents

Enterprises are handing real work to AI agents — software that reads files, queries databases, sends email, calls internal APIs, and makes decisions on a person’s behalf. These systems act, and their actions persist after the conversation ends. This paper documents why agents must be judged on what they do, and the method Fisher uses to prove it.

Read the white paper (PDF) →  Request an Assessment →

Key findings · as of July 2026

The evidence base.

56,000+
adversarial episodes
500,000+
conversation turns
22
model architectures
6.6–14.6pp
adaptive uplift
What the paper examines

Why transcripts alone can’t clear an agent.

Tool-using agents hold standing access to files, mail, databases, and internal APIs. A hidden failure lives in the testing tools themselves: most evaluation grades the agent’s final text, and a fluent refusal can conceal tool calls that already executed the forbidden task. The paper documents this blind spot with a controlled head-to-head on an identical database agent, where transcript grading passed sessions whose action traces showed protected records had been queried.

It then describes the alternative: multi-turn adversarial campaigns evaluated at the action level — tool calls, retrievals, writes, and state changes — where a finding is recorded only when concrete evidence appears, every confirmed failure is replayed to establish whether it reproduces, and fixes are re-attacked and classified as fixed, partial, or bypassed.

What readers will learn

Inside the paper.

The blind spot of grading words instead of actions · the Discover, Prove, Remediate workflow · how the adversarial engine searches strategy space and adapts to the target’s defenses · the matched experiments behind the 6.6–14.6pp adaptive uplift · reproducibility as a first-class result · remediation that survives re-attack · the full methodology, scenario set, and evidence tiers.

Read the white paper (PDF) →