AI, Software Development

Red teaming LLMs: how to run your first round in one sprint

By James KillickSeptember 25, 2026

TL;DR: Red teaming an LLM means attacking your own AI app to find leaks, rule breaks and unapproved actions before someone else does. Start with manual testing to find the harms. Log each finding so you can rerun it, name it with the OWASP Top 10 for LLM Applications, and rank it by what it costs the business. Then turn each fix into a regression test in CI.

Red teaming an LLM means attacking your own AI app on purpose, before someone else does. You try to make it leak data, break its rules or take an action nobody approved. Then you fix what you found.

What you want at the end is a finding you can rerun, fix and lock in with a test. A scary screenshot won't get you that.

So here's how to go about red teaming LLMs in a real product. What to aim at, how to rank what you find, which tools help, and how to make the fixes stick. A first pass fits in one sprint.

Start with goals and evidence

"Try to break it" isn't a goal. It gives you a pile of odd chats and nothing to fix.

Pick attacker goals you can check. Get the system prompt out. Pull another customer's record. Make the agent call a tool it shouldn't. Each one ends in a clear yes or no.

Then decide what you'll record. Microsoft's guide to planning LLM red teaming keeps it simple. Log the input, the output, the date, and a unique ID so you can run the same case again later. It says a shared spreadsheet is often the easiest place to keep it all.

If your app has tools, add two more things:

  • The tool calls the agent made, in order.
  • How often the attack worked, out of how many tries.

That last one matters because an LLM isn't steady. Cybernion's explainer on how AI red teaming differs from a normal pen test puts it well. A pen test hunts for flaws that behave the same way every time. A model's behaviour runs on chance, so the same prompt may work sometimes and not others.

So one hit tells you little. One miss tells you less. Run each attack many times and count.

Researchers call that count the attack success rate. It's the wins divided by the total tries. That's how this survey of LLM red teaming defines it. The paper also walks through attack methods and the ways to judge if an attack worked.

One caution from Microsoft. A red team finds harms. It doesn't tell you how common they are across all your users. Don't read a handful of examples as a rate for the whole product.

Map the attack surface

Every LLM product has two layers. They fail in different ways.

Microsoft's guide says to test both. Test the base model with its safety system, which is usually done through an API. Then test your app, which is best done through the UI. It also says to test on the production UI as much as you can, because that's the closest thing to real use.

Promptfoo's red teaming docs split threats the same way. They point out that most teams plug in a model someone else built, so the app layer is often the focus. That's the part you built. It's where your risk lives.

Walk through your app and list each place where untrusted text or an action crosses a line:

  • Retrieval. Documents, web pages and emails the model reads as context.
  • Memory. Anything saved in one session that shapes the next.
  • Tools. Every function, API or database the agent can call, and whose permissions it uses.
  • Output. Anywhere model output gets shown as HTML, run as code or passed to another system.

We won't teach the defences again here. Prompt injection defence has the fixes and an 8-step checklist. Chatbot security covers the threats, the controls and a 12-point check to run before go-live. This post is about how to test what those two posts protect.

Agents raise the stakes, because a hijacked agent can act. The thing that limits the damage is what the agent is allowed to touch. Agent Swarm's agent governance guide lists least-privilege roles, isolation for each agent and audit logging as the base layer. It's a vendor guide, so read it that way. But its testing list is handy. It includes a privilege-escalation probe against your access rules. Put that probe in your plan.

Got a multi-tenant product? Test hard for one tenant pulling another tenant's data. Our white-label AI platform case study shows what tenant isolation looks like when it's built into the data layer.

Threat classes and how to rank findings

Don't make up your own labels. Use the OWASP Top 10 for LLM Applications. The current list is the 2025 version. It gives engineers, security and risk people one shared set of names.

Here's how common red-team findings map to it:

What you foundOWASP 2025 category
Hidden instructions in a document got obeyedLLM01 Prompt Injection
The bot gave up another customer's detailsLLM02 Sensitive Information Disclosure
Model output ran as code or HTML downstreamLLM05 Improper Output Handling
The agent took an action nobody asked forLLM06 Excessive Agency
The system prompt leaked with secrets in itLLM07 System Prompt Leakage
Retrieval crossed a tenant boundaryLLM08 Vector and Embedding Weaknesses

The other four are Supply Chain, Data and Model Poisoning, Misinformation and Unbounded Consumption.

Two notes from OWASP's own pages. It treats a jailbreak as a form of prompt injection, not a class of its own. And it says the system prompt shouldn't be treated as a secret or used as a security control. So "we got the system prompt out" is only a big finding if something was in there that shouldn't have been.

Prompt injection sits at number one on the list. SQ Magazine's roundup of prompt injection statistics puts the same ranking near the top. It pulls figures from lots of reports, so trace any number back to where it came from before you quote it.

Now rank what you found. Microsoft says to weigh how severe the harm is and the context where it's likely to show up. We'd add one more question. How easy is it to pull off?

A clever jailbreak that writes a rude poem goes near the bottom. A dull bug that leaks a customer record now and then goes to the top. Rank by what it costs the business, not by how smart the attack looked.

White-box or black-box, manual or automated

You've got two calls to make. How much can the tester see? And who does the attacking?

ModeWhat it meansWhen to use it
Black-boxYou only see inputs and outputsMost app teams, most of the time
White-boxFull access to model weights and training dataTeams that run their own models
ManualPeople write and run the attacksFirst, to find the harms
AutomatedTools write and score the attacksNext, to cover them at scale

The Promptfoo docs define white-box as full access to the model's architecture, training data and weights. They also say it's not practical for most teams, because most teams don't build on models with open weights. Black-box treats the system as closed. You see what goes in and what comes out. That's closer to how a real attacker works.

You still know your own system prompt, tools and data sources. Use that to aim your tests. Just run the attacks through the same front door your users walk through.

On manual versus automated, the order matters. Microsoft says to finish a first round of manual red teaming before you measure anything at scale. People find the harms first. Then you know what to measure.

The research survey backs both sides. People come up with creative attacks, but they cost a lot, and that limits how many cases you get. Automated attackers are cheaper. Many of them run in a loop. One model attacks, a second model judges the result, and the first one adjusts and tries again.

One problem though. If a model is the judge, the judge can be wrong. Our post on LLM evaluation shows how to check a judge against answers you scored by hand. Do that before you trust its numbers.

A red-team workflow that fits one sprint

You don't need a security team to get a useful first pass. Here's a plan for a two-week sprint.

  1. Set the scope. List the surfaces from your map. Write down the harms that would hurt the business most. Only test systems you own or have written permission to test.
  2. Pick the testers. Microsoft suggests a mix. Get people who think like attackers, plus ordinary users who weren't part of the build. Add a domain expert if the product needs one.
  3. Run an open round. No script. Testers explore and log anything that looks wrong. This is how you find the harms nobody thought of.
  4. Write the harms list. Name each harm, define it and add an example.
  5. Run a guided round. Go after each harm on the list. Repeat each attack and count the hits. Add new harms as they turn up.
  6. Report. Microsoft's format is short. The top issues, a link to the raw data, and the plan for the next round.
  7. Fix and retest. Test with the fix switched on and off, so you know the fix is what changed the result.
  8. Lock it in. Turn each confirmed finding into a regression test.

Tip: log every attempt, not just the hits. An attack that fails most of the time and lands now and then is the one you'll miss if you only keep the wins.

Some systems need more than a sprint. If an agent can write to production, handles regulated data or moves money on its own, bring in outside specialists too.

Tools and frameworks

Start with people. Then add tools to repeat what they found.

Three open-source tools are worth a look:

  • PyRIT** is Microsoft's open-source framework for finding risks in generative AI systems. The survey describes it as an attacker system set against a target, with an evaluator judging each exchange.
  • garak** calls itself an LLM vulnerability scanner. It's a command-line tool that probes for things like prompt injection, data leakage and jailbreaks.
  • Promptfoo** is a CLI and library for evaluating and red teaming LLM apps. It has CI/CD integration, so it suits the regression step.

Heads up: PyRIT moved from the Azure GitHub org to the Microsoft one. The old repo is archived.

Three frameworks give your findings a shape other people can read:

  • OWASP Top 10 for LLM Applications for naming what you found.
  • MITRE ATLAS for mapping how the attack was done. It's a public knowledge base of attacker tactics and techniques against AI systems. The ATLAS data repo holds the tactics, techniques and case studies.
  • **The NIST AI Risk Management Framework** for feeding results into how the business manages risk. It's voluntary, and its core has four functions: govern, map, measure and manage.

For test data, the survey lists public sets like JailbreakBench. They're a fine place to start. But build your own set from your own logs and near misses. A public set doesn't know your tools or your data.

Turn findings into regression tests

A fix with no test behind it can come undone. The next model upgrade or prompt edit can open the hole again, and nobody will notice.

So each confirmed finding becomes a test case with three parts:

  • The exact attack input, saved with the context it needs.
  • The pass rule. What must never show up in the output, or which tool must never be called.
  • The number of runs and the pass mark, written down before you run it.

Where do those tests live? We've covered that. LLM unit testing sets out what runs on each commit, each pull request and each night. CI/CD for AI shows where the gates sit in the pipeline. And LLM guardrails covers the five checkpoints where most fixes end up.

For agents, the AI Orchestrators guide to AI agent testing shows how to record a real run once and replay it in CI. Their post on AI agent security says to rerun adversarial tests when prompts, tools, model providers or connectors change. Adding tools through MCP? AI-Led's guide on how to build an MCP server covers auth and testing for that surface.

Now here's the important bit. A passing regression test proves one known attack still fails. It doesn't prove you're safe. So red team again before each major release, and any time the model, the prompts, the tools or the connectors change.

Where Devwiz fits

Devwiz builds AI apps and AI platforms as production software. Testing like this belongs in the build, not in a report that turns up after launch.

A scary screenshot gets you a meeting. A finding you can rerun gets you a fix. If you can't turn it into a test, you're not done.

*James*

Devwiz has shipped 200+ apps, including work for NSW Government (Justice and Corrective Services), Briometrix, Vivid and Huskee.

Building an AI feature that reads private data or takes actions for a user? Start with AI app development. Worth a chat.

Frequently asked questions

What is LLM red teaming?

It's attacking your own LLM app on purpose, so you find the failures before a real attacker does. You test for things like prompt injection, data leaks and an agent taking actions nobody approved. Each finding should be something you can rerun and fix.

How is red teaming an LLM different from a penetration test?

A pen test looks for flaws in code, config and infrastructure that behave the same way every time. An LLM's behaviour is probabilistic, so the same prompt can work once and fail the next time. That's why you repeat each attack and count how often it lands.

Should manual or automated red teaming come first?

Manual. Microsoft's planning guide says to finish a first round of manual red teaming before you measure at scale. People find the harms first. Then tools like PyRIT, garak or Promptfoo can repeat those attacks each time the app changes.

Which tools are used for red teaming LLMs?

PyRIT, garak and Promptfoo are three open-source options. PyRIT is Microsoft's framework for finding risks in generative AI systems. garak is an LLM vulnerability scanner. Promptfoo is a CLI and library for evaluating and red teaming LLM apps.

How often should you red team an LLM app?

Before launch, before each major release, and any time the model, prompts, tools or connectors change. In between, regression tests built from past findings tell you if a known attack has started working again.

About James Killick

10+ years building digital products · 200+ apps shipped since 2015

James is a co-founder of Devwiz and an AI product specialist. Since 2015 he has helped ship 200+ apps for founders, businesses and government, including work for NSW Government, Briometrix and Huskee. He builds AI-first platforms and writes about turning a proven program into software. He also hosts the Up in the AI podcast.

More articles by James · James's personal site · LinkedIn · AI Orchestrators

Tags: AI, LLM, Security, Testing

Browse all Devwiz articles·See our case studies