---
title: "AI agent safety evals · red-team arenas, safety benchmarks, adversarial tests · Agent Fieldbook"
url: https://agentfieldbook.org/security/evals/
description: "How AI agents are stress-tested for safety. The red-team arenas, safety benchmarks and adversarial frameworks agents are measured against, charted by type of test and catalogued by who ran each one and what it measures. A source-graded field book to the evidence behind agent-safety claims."
section: "Security · Safety Evals"
source: Agent Fieldbook — generated from the published page
---

# How agents are stress-tested

The red-team arenas, safety benchmarks and adversarial frameworks agents are measured against: charted by **type of test**, then catalogued.

## Evals by type of testing

## Jailbreak benchmark · 4

| Eval | By | Result / what it measures | src |
| --- | --- | --- | --- |
| **JailbreakBench** | UPenn, ETH Zurich, EPFL et al. | Open leaderboard + 100-behaviour JBB-Behaviors dataset + standardised harness for evaluating jailbreak attacks and defences. | [↗ T1](https://arxiv.org/abs/2404.01318) |
| **HarmBench** | Center for AI Safety + UIUC | Standardised automated red-teaming comparing 18 attack methods vs 33 target LLMs across 400+ harmful behaviours. | [↗ T1](https://arxiv.org/abs/2402.04249) |
| **StrongREJECT** | UC Berkeley (Souly, Toyer et al.) | 313 forbidden prompts + automated grader scoring harmful-response usefulness; shows prior jailbreak metrics overstate attack success vs human judgement. | [↗ T1](https://arxiv.org/abs/2402.10260) |
| **AdvBench (GCG attack)** | CMU, CAIS, Google DeepMind (Zou, Kolter, Fredrikson et al.) | 520 harmful-behaviour/string prompts introduced with the GCG attack; the de-facto target set for transferable adversarial-suffix jailbreaks. | [↗ T1](https://arxiv.org/abs/2307.15043) |

## Dangerous-capability eval · 4

| Eval | By | Result / what it measures | src |
| --- | --- | --- | --- |
| **SHADE-Arena** | Anthropic + Redwood Research | Agents get a benign main task + a hidden malicious side-task under an LLM monitor; measures covert-sabotage success and monitor detection. | [↗ T1](https://arxiv.org/abs/2506.15740) |
| **METR Autonomy Evaluation Resources** | METR | ~50 auto-scored autonomy tasks across cyber, software engineering + ML + an evaluation protocol; gauges autonomous dangerous-capability. | [↗ T1](https://metr.org/blog/2024-03-13-autonomy-evaluation-resources/) |
| **Inspect / Inspect Evals (UK AISI)** | UK AI Security Institute | Open-source eval framework (datasets, solvers, scorers) + Inspect Evals library of 200+ community evals across cyber, agentic + dangerous-capability. | [↗ T1](https://inspect.aisi.org.uk/) |
| **OpenAI Preparedness Framework** | OpenAI | Tracked Categories (cyber, CBRN, self-improvement) with capability thresholds + dangerous-capability evals gating deployment; v2 Apr 2025. | [↗ T1](https://openai.com/index/updating-our-preparedness-framework/) |

## Red-team arena · 2

| Eval | By | Result / what it measures | src |
| --- | --- | --- | --- |
| **Gray Swan Agent Red Teaming** | Gray Swan AI | Claude Opus 4.7: 0.1% single-attempt attack-success rate, rising to 5–6% after 100 attempts | [↗ T1](https://www.grayswan.ai) |
| **HackAPrompt** | Learn Prompting / U. Maryland (Schulhoff et al.) | Global prompt-hacking competition; collected 600,000+ adversarial prompts against frontier models, yielding a prompt-injection/jailbreak taxonomy. | [↗ T1](https://arxiv.org/abs/2311.16119) |

## Prompt-injection benchmark · 2

| Eval | By | Result / what it measures | src |
| --- | --- | --- | --- |
| **AgentDojo** | ETH Zurich (Debenedetti, Tramèr et al.) | 97 realistic tool-use tasks + 629 security test cases; measures prompt-injection attack success vs task utility for agents acting over untrusted data. | [↗ T1](https://arxiv.org/abs/2406.13352) |
| **InjecAgent** | UIUC (Zhan, Kang et al.) | 1,054 test cases across 17 user + 62 attacker tools; ReAct GPT-4 agents hijacked by indirect injection ~24% of the time. | [↗ T1](https://arxiv.org/abs/2403.02691) |

## Cyber-capability eval · 2

| Eval | By | Result / what it measures | src |
| --- | --- | --- | --- |
| **Cybench** | Stanford et al. (A. K. Zhang et al.) | 40 professional-level CTF tasks with subtask decomposition; measures autonomous offensive-cyber capability of LLM agents. | [↗ T1](https://arxiv.org/abs/2408.08926) |
| **CyberSecEval 3 (Purple Llama)** | Meta (Purple Llama) | Evaluates offensive-cyber capability + third-party/end-user risks (auto prompt injection, code exploitation, social engineering) across Llama 3 and peers. | [↗ T1](https://arxiv.org/abs/2408.01605) |

## Reliability benchmark · 2

| Eval | By | Result / what it measures | src |
| --- | --- | --- | --- |
| **τ-bench (tau-bench)** | Sierra (Yao, Narasimhan et al.) | Simulates dynamic user + tool/API interactions in retail/airline domains; measures agent reliability and policy adherence via pass^k consistency. | [↗ T1](https://arxiv.org/abs/2406.12045) |
| **τ²-Bench (tau-squared)** | Sierra Research | Successor to τ-bench: dual-control (agent + user both have tools) Dec-POMDP telecom domain; compositional task generator with realistic user simulator. Shows agent performance drops significantly when moving from single- to dual-control. τ³-Bench (2026, banking + voice) further extends it. | [↗ T1](https://arxiv.org/abs/2506.07982) |

## Safety benchmark · 1

| Eval | By | Result / what it measures | src |
| --- | --- | --- | --- |
| **AILuminate v1.0** | MLCommons | Claude 3.5 Haiku/Sonnet + Mistral Large rated "Very Good"; GPT-4o + Gemini 2.0 Flash "Good" | [↗ T1](https://mlcommons.org/ailuminate/) |

## Credential-safety benchmark · 1

| Eval | By | Result / what it measures | src |
| --- | --- | --- | --- |
| **SCAM (Secret/Credential Agent Misuse)** | 1Password | 8 frontier models, every model failed critically in every run | [↗ T1](https://blog.1password.com) |

## Adversarial knowledge base · 1

| Eval | By | Result / what it measures | src |
| --- | --- | --- | --- |
| **MITRE ATLAS** | MITRE | Threat matrix of adversarial ML / agent tactics & techniques (since 2021) | [↗ T1](https://atlas.mitre.org) |

## Risk category · 1

| Eval | By | Result / what it measures | src |
| --- | --- | --- | --- |
| **OWASP LLM07, System Prompt Leakage** | OWASP | Named risk in the LLM Top 10 (v2.0, Nov 2024) directly relevant to agent secret exposure | [↗ T1](https://genai.owasp.org) |

## Harmful-task benchmark · 1

| Eval | By | Result / what it measures | src |
| --- | --- | --- | --- |
| **AgentHarm** | UK AISI + Gray Swan AI | 110 explicitly malicious multi-step agent tasks across 11 harm categories; measures whether jailbroken agents complete harmful tool-use while retaining capability. | [↗ T1](https://arxiv.org/abs/2410.09024) |

## Agent security benchmark · 1

| Eval | By | Result / what it measures | src |
| --- | --- | --- | --- |
| **Agent Security Bench (ASB)** | Rutgers, Zhejiang University et al. | Benchmarks prompt-injection, memory poisoning, backdoor + mixed attacks and defences across 10 scenarios, 10 agents, 400+ tools. | [↗ T1](https://arxiv.org/abs/2410.02644) |

## Risk taxonomy · 1

| Eval | By | Result / what it measures | src |
| --- | --- | --- | --- |
| **SafetyBench** | Tsinghua CoAI (Zhang, Huang et al.) | 11,435 multiple-choice questions across 7 safety categories (EN + ZH); measures an LLM's safety understanding. | [↗ T1](https://arxiv.org/abs/2309.07045) |

## Field notes

Red-team arenas, safety benchmarks and adversarial frameworks, by type of test

A good eval score is evidence, not a certificate. It says the agent resisted the specific attacks in that test set on the day it ran, which is exactly why the incident log keeps growing even as benchmark numbers improve: novel attacks, new tools and a different deployment context all sit outside the test. Evals narrow the risk and make regressions visible. Governance controls and human oversight contain whatever the tests miss. Each row below records the eval, who ran it, and what it measures, with a link to the primary source.

- [Brandon Chaplin](https://www.linkedin.com/in/brandon-chaplin-digital-marketing-strategist)

## Common questions

### What is an AI safety eval?

A safety eval tests how a model or agent behaves under adversarial pressure, rather than how capable it is at a task. It measures resistance to jailbreaks and prompt injection, refusal of harmful requests, and safe use of tools. The tests come in two broad shapes: fixed benchmarks and live red-team exercises.

### What is red teaming an AI model?

Red teaming is people, or other models, actively trying to break the system: jailbreak it, make it leak data, or turn a tool against its operator. It finds novel failures a fixed test set would never contain. Benchmarks give you a score you can compare; red teaming gives you coverage of the attacks nobody wrote a test for.

### How do you test an agent for prompt injection?

You feed it inputs crafted to override its instructions and measure how often the attack lands. The part that matters most is indirect injection, where the instruction hides in a web page, email or document the agent retrieves rather than in the user's message. A low attack-success rate is evidence of resistance, not proof of safety. OWASP ranks prompt injection as the top risk for LLM applications [(OWASP)](https://owasp.org/www-project-top-10-for-large-language-model-applications/).

### Does a good safety score mean an agent is safe to deploy?

No. It means the agent resisted the specific attacks in that test set, on the day it ran. New attacks, new tools and a different deployment context can all defeat it. Evals make regressions visible and narrow the risk; they do not remove it.

### Who runs AI safety benchmarks?

A mix of parties. Consortia such as MLCommons publish cross-model safety benchmarks, the frontier labs run their own internal red-team and safety evaluations, and academic and open-source groups build the adversarial frameworks. Who ran a test matters: a vendor grading its own model is weaker evidence than an independent one.

## Next in the learning path

- [Incident log The real-world failures evals try to predict](https://agentfieldbook.org/security/incidents/)

- [Governance The controls a failed eval argues for](https://agentfieldbook.org/security/governance/)

- [Human oversight The fallback when an eval cannot guarantee safety](https://agentfieldbook.org/security/oversight/)

- [Standards The frameworks that mandate testing](https://agentfieldbook.org/security/standards/)
