Fuel iX

AI agent red teaming: test what your agents are permitted to do, not just what they say


Phill Thompson
Phill Thompson

Director of Product Marketing, TELUS Digital

Fortify agent testing img 4

Key takeaways

  • AI agent risk is a different category from chatbot risk. Testing what an agent says in conversation does not cover what it is permitted to do as an agent.
  • Runtime defenses like guardrails, filters and monitoring catch problems after they happen. Enterprises need autonomous red teaming to find agent vulnerabilities whether an agent is still in testing or already live.
  • Fuel iX™ Fortify now runs AI agent red teaming that tests what an agent says it’s allowed to do, and will do, not just what it says in conversation.
  • Fortify reads an agent's own capability file, generates attack simulations from the actions it is permitted and denied, and evaluates how the agent responds.

For years, generative AI (GenAI) risk was largely concerned with a chatbot saying the wrong thing. A model gave a bad answer, leaked information it should not have or got talked into ignoring its own rules. That problem was primarily a problem of words.

Today, AI agents change the stakes. An agent does not just answer, it acts. It calls tools, moves data, updates records and triggers operations across the systems it connects to. When a chatbot says something wrong, you get an embarrassing screenshot. When an agent does something wrong, you get a breach, a bad transaction or a compliance failure.

There has been no shortage of recent headlines highlighting agent manipulation. For example, a coding assistant was hijacked by a hidden instruction planted inside a routine bug report. When the AI coding agent read it, the instruction took over and drove the agent to exfiltrate data from private code repositories. The agent did what it was built to do, reach into repositories and act on them, but it did it for an attacker.

In another example, an over-privileged CRM agent leaked confidential records. The hacker compromised a customer-relationship AI agent that had more access than it needed and used it to exfiltrate confidential records across more than 700 companies over the course of ten days.

The agents didn’t say anything wrong in either of these examples. Both had permissions, and both got manipulated into misusing them. That’s what AI security testing needs to catch.

Why agentic AI security testing can’t wait: adoption is outpacing preparation

Agent adoption is not a slow build. Gartner projects that 40% of enterprise applications will include task-specific AI agents by the end of 2026, up from less than 5% a year earlier.

The industries moving fastest are the ones with the most at stake. Salesforce's Agentic Enterprise Index found that financial services grew its agent work output 13x between February 2025 and April 2026, deploying some of the most sophisticated agents in production at consumer scale. These agents are checking balances, moving funds and acting on customer data right now.

But security investment isn't scaling at the same rate. Of the $2.52 trillion spent on AI globally in 2026, only about 2% — roughly $51.3 billion — went to AI cybersecurity, according to Gartner.

Even the labs building these models aren’t immune. In August 2026, OpenAI disclosed that its own internal agents, while running a routine security evaluation, found and exploited a flaw in the testing infrastructure itself, then used it to coordinate with each other on further exploits. The chain of events eventually contributed to a breach.

That kind of exposure is why autonomous agents are drawing a new kind of scrutiny, and why security frameworks are moving to keep up. OWASP published an Agentic Security Initiative and an agentic top-ten list precisely because agent risk does not fit neatly into the older LLM categories.

attack-timeline-edited

Autonomous red teaming is no longer optional. The volume of AI applications and agents now vastly outstrips the number of skilled humans available to test them manually. This is not a challenge enterprises can realistically hire their way out of. Testing has to scale the way the agents do.

How controlled, autonomous red teaming for AI agents works

Runtime defenses like guardrails, filters and monitoring watch for trouble while an agent is live and handling users. Think of them as a bank’s alarm system: essential, but reactive. They go off after a break-in is already underway. What is needed alongside these intervention methods is prevention: someone who tries to break into the vault before the bank opens, and tells you what to fix before it becomes a major issue.

With new capabilities built specifically for AI agents, Fuel iX™ Fortify, TELUS Digital’s proprietary automated red teaming platform, now tests what agents say they can, and will, do. It reads the Agent Card — a standardized capability file that declares what an agent is permitted and denied to do — then presents it with scenarios built from those declared boundaries and evaluates how the agent responds. When an agent indicates it would take, or claims to have taken, an action outside of its own stated boundaries, that’s a vulnerability finding, whether it’s caught before launch or during ongoing testing afterwards.

The attacks then run through Fortify's own attack model.

Agent attack objectives Agent policy 2 1x

Because the objectives come from the agent's permission surface, the testing is specific in a way generic testing checklists cannot be. Instead of asking "could some agent somewhere be tricked into leaking data?" Fortify asks "will this agent, with these permissions and these denied actions, affirm it would take an unauthorized action, a policy-violating operation or an act of data exfiltration?" Those are the failures that matter, because they map to what the agent can reach and touch.

Within Fortify there is a clean split. One policy tests what your AI says in conversation. Another tests what it says it’s permitted to do as an agent. Fortify tests chats with a chat policy and tests agents with an agent policy, so each gets the scrutiny it needs.

Agent card extraction 2 - result (1)

This builds on agent-aware testing Fortify already runs by default. A core set of agentic risk behaviors, including system prompt disclosure, secret and credential retrieval, PII disclosure, and tool and function disclosure, are already part of standard sessions. The Agent Card step makes that coverage sharper and specific to each agent you point it at.

agentic objectives results in a session 2 (1)

Every finding maps to the security frameworks your team and your auditors already use, including the OWASP LLM Top 10, MITRE ATLAS and NIST AI RMF, so agent testing rolls into the same coverage reporting as the rest of your program. And because agents change constantly through new skills, new tools and new integrations, Fortify runs on a schedule you set, rather than a single pre-launch check.

AI agent red teaming: what's next

Agent testing will keep changing shape as agents do, and the team at TELUS Digital is continuously enhancing Fortify to stay ahead.

“Every agent has a different scope, so the surface Fortify tests against widens or narrows depending on the target, and that variability is only increasing,” said Sho Modica, head of product, Fuel iX, at TELUS Digital. “Dynamic objective extraction — how Fortify builds and directs its attacks — is a first step toward keeping pace with that. In tandem we are making enhancements to our execution layer to have improved strategies of interacting with these agents.”

That's the nature of testing something that doesn't hold still. An agent's permissions today aren't its permissions after the next update, the next integration or the next tool it's given access to. Static testing was never going to be enough for that. The only way to keep an accurate picture of what your agents are permitted and willing to do is to keep asking.

Right now, that means evaluating what an agent says about its own actions. Where this is headed is more directly testing the actions that agents take.

Your agents are already doing more than your testing may account for. Find out what they'd admit to under pressure. For a limited time, TELUS Digital is offering qualifying teams a complimentary Fuel iX Fortify AI security scan. Claim your free scan today.


Phill Thompson

Phill Thompson

Director of Product Marketing, TELUS Digital

Phill Thompson has spent ten years working at the intersection of AI and go-to-market strategy, with a career spanning data labeling, computer vision, NLP, and robotics. His experience includes roles at CloudFactory and SparkAI, where he worked with companies using computer vision and robotics to automate edge cases and exceptions, before joining TELUS Digital. Over the past two years, he has led product marketing for nearly ten AI product launches, with deep focus on AI Safety & Security, including Fortify, TELUS Digital's AI attack simulation platform for testing models and AI agents against adversarial threats.

Frequently asked questions

Talk to us

Fuel all stages of your business growth with our solutions. Get started now!

Contact sales