The agent already has the keys — standing access to your files, mail, databases, and APIs. Fisher pressure-tests tool-using agents with adaptive, multi-turn attacks and returns reproducible evidence of what the agent actually did, not just what it said.
Built for security and AI-governance teams deploying tool-using agents into consequential workflows.
AI agents do more than generate text. They browse, query databases, call tools, access files, and change systems. Fisher safely attacks those workflows to find where an agent can be manipulated, exceed its authority, or cross a data boundary.
Each confirmed finding is replayed and documented with the conversation, tool calls, state changes, severity, and confidence needed for security, governance, and release review.
AI agent assurance is the evidence that an agent stayed within its intended boundaries under attack and change.
Red teaming means safely attacking your own AI system to find exploitable weaknesses before someone else does.
A safe-looking answer can conceal an unsafe tool call, data access, or state change. Fisher checks the conversation and the underlying actions, then replays confirmed findings to determine whether they reliably reproduce. Below: the same session, two verdicts — a text-only grader reads the refusal and passes it; Fisher reads the tool log and catches the leak in the act.
A grader can record “defended” even after an agent has already taken the sensitive action. In our database-agent comparison, the transcript passed while the action trace showed password and API-key records had been queried. Fisher evaluates the action trace, not just the response. Read the full method and results in the white paper →
You have AppSec for your code, CloudSec for your infrastructure, and identity for your people. Nothing in that stack covers an agent being talked into misusing access it legitimately holds. That surface is cognitive — and each of the controls below, useful as it is, stops short of proving an agent stayed within its authorized boundaries under adaptive, multi-turn pressure.
Keep your guardrails and observability. Fisher tests whether they hold.
Why now: the reporting obligation has arrived.
The EU AI Act's robustness and cybersecurity obligations for high-risk systems take effect in August 2026. NIST AI RMF asks for documented evaluation; ISO/IEC 42001 is published and being adopted. When a board or regulator asks how you know your agents are safe, "we ran some prompts and read the replies" no longer holds. Reproducible, action-level evidence does.
Multi-turn, adaptive attacks across the agent’s actual tools, permissions, and data boundaries.
→ You see how the agent behaves under realistic adversarial pressure, not one scripted prompt.
Verify the outcome in tool logs, state changes, and planted canary markers.
→ A finding is recorded only when concrete evidence supports it
Re-run each finding strict, guided, and free to sort it into a confidence tier.
→ Every confirmed failure is replayed to establish whether it reproduces reliably or was a one-off.
Bundle the transcript, tool trace, state deltas, and control mapping.
→ Reproducible evidence designed for governance, audit, and customer-security review.
Not another undifferentiated vulnerability list. Evidence to answer a security review, prioritize what to fix, and make a release decision. Approve, block, or defend that decision with evidence your teams can reproduce.
Fisher maps to the workflow you’re deploying and the boundary that matters — then tests the failure and shows the consequence.
The tools that test your agents are increasingly owned by the platforms that sell them. Molt works across the common agent stacks — LangChain, LangGraph, MCP tool servers, and standard REST APIs — and reports to you, not to a platform owner. Independent evidence, from a vendor with nothing else to sell you.
Fisher runs against an approved sandbox, simulated environment, or approved endpoint using synthetic data and canary markers. We do not request production access, employee or customer credentials, or real customer records.
A concrete, low-friction way to see Fisher on an agent you actually run.
One priority agent workflow. Fixed scope, an approved sandbox or endpoint, synthetic data and canary markers, and reproducible, confidence-labeled findings your teams can take into a release or governance review. The Assessment runs on the Fisher platform, not as a one-off audit: the same attack → prove → remediate → re-test loop can then extend across future material releases.
Deep Model Trust is Molt’s technical architecture for moving from testing agent behavior to building systems whose authority, decisions, and release processes are bounded and verifiable.
Fisher is the first commercially available product built around that thesis. Testing is where trust starts, not where it ends. Each additional capability carries a published availability status on the Products page.
Experience from NVIDIA and Microsoft. Academic rigor from Columbia and Harvard. Molt is a small team of security engineers, AI researchers, risk practitioners, and builders who believe trust should be proven, not assumed.
NVIDIA · Microsoft · Columbia · Harvard
Tell us the one workflow you want tested. We’ll follow up to scope a 30-Day Agent Assurance Assessment.