Prompt injection
Resistance to instructions planted in content an agent reads.
- Reported by
- Anthropic, Google DeepMind, Meta, Moonshot AI, OpenAI and xAI
- Evaluations
- 25 evaluations, of three kinds of test
- Values
- 118 values in 34 documents
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Definition
Prompt injection is an attack on a model working as an agent. Instructions are hidden in content the model reads while doing a task, such as a web page, an email, a document or the output of a tool, in the hope that the model follows them instead of its user. A successful injection could make an agent leak data, take actions nobody asked for, or abandon its task.
Developers measure resistance with benchmarks of planted instructions across browsing, coding, computer use and tool calls, with adaptive attackers who refine their attempts, and sometimes with public bug bounties.
Results are reported as attack success rates or as the share of attacks resisted, often with and without extra safeguards. Tasks, attack sets and safeguards differ between tests, and an attack success rate depends heavily on how many attempts the attacker is allowed.
The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).
How developers measure it
25 evaluations of three kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.
| Evaluation | Reported by | Values | ||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Static benchmark · 17 evaluations, 84 values | ||||||||||||||||||||||||||||||||
| Prompt injection, in the data explorer Run by the developer OpenAI's prompt-injection sets, which plant instructions in web pages, tool outputs, connectors or function results and score how often the model does not follow them.Reported by OpenAI · 25 values | OpenAI | 25 | ||||||||||||||||||||||||||||||
| Computer-use prompt injection, in the data explorer Run by the developer Measures the share of prompt-injection attacks placed in computer-use environments that the model resists, with and without safeguards.Reported by Anthropic · 15 values | Anthropic | 15 | ||||||||||||||||||||||||||||||
| Instruction hierarchy, in the data explorer Run by the developer Checks whether the model keeps to higher-priority system or developer instructions when a lower-priority message conflicts with them, including injected hijacking and system-prompt extraction attempts.Reported by OpenAI · 9 values | OpenAI | 9 | ||||||||||||||||||||||||||||||
| AgentDojo · xAI, in the data explorer Run by the developer A public benchmark of agent tasks with prompt injections planted in tool outputs, run by xAI, measuring how often an injected instruction succeeds.Reported by xAI · 5 values | xAI | 5 | ||||||||||||||||||||||||||||||
| MCP prompt injection, in the data explorer Run by the developer Measures the share of prompt-injection attacks delivered through MCP tool results that the model resists.Reported by Anthropic · 5 values | Anthropic | 5 | ||||||||||||||||||||||||||||||
| Prompt injection (Codex env), in the data explorer Run by the developer Prompt-injection attacks placed in Codex coding environments, scored on the share of attacks the coding agent ignores.Reported by OpenAI · 5 values | OpenAI | 5 | ||||||||||||||||||||||||||||||
| Tool-use prompt injection, in the data explorer Run by the developer Measures the share of prompt-injection attacks delivered through tool outputs that the model resists.Reported by Anthropic · 5 values | Anthropic | 5 | ||||||||||||||||||||||||||||||
| Agent Red Teaming (ART) · Gray Swan, in the data explorer Run by Gray Swan (third party) Gray Swan's agent red-teaming benchmark, measuring how often attacks on an agent succeed within a set number of attempts.Reported by Anthropic and Meta · 4 values | Anthropic and Meta | 4 | ||||||||||||||||||||||||||||||
Show the other 9 evaluations of this kind (11 values)
| ||||||||||||||||||||||||||||||||
| Adaptive attack · 7 evaluations, 33 values | ||||||||||||||||||||||||||||||||
| Coding prompt injection (adaptive attacker), in the data explorer Run by the developer Measures how often an adaptive attacker hijacks the model through prompt injections in coding environments.Reported by Anthropic · 10 values | Anthropic | 10 | ||||||||||||||||||||||||||||||
| Computer use prompt injection (adaptive attacker), in the data explorer Run by the developer Measures how often an adaptive attacker hijacks the model through prompt injections in computer-use environments, counted per attempt or per scenario.Reported by Anthropic · 8 values | Anthropic | 8 | ||||||||||||||||||||||||||||||
| Indirect prompt injection (adaptive attacks), in the data explorer Run by the developer Success rate of three adaptive attack methods that plant instructions in content the model processes, measured over 500 held-out scenarios.Reported by Google DeepMind · 6 values | Google DeepMind | 6 | ||||||||||||||||||||||||||||||
| Browser use prompt injection, in the data explorer Run by the developer Measures how often prompt injections on web content hijack the model while it operates a browser.Reported by Anthropic · 3 values | Anthropic | 3 | ||||||||||||||||||||||||||||||
| GPT-Red, in the data explorer Run by the developer OpenAI's automated red-teaming attacks on prompt injection and the instruction hierarchy, added to the GPT-5.6 card in August 2026 and scored as the share of attack attempts that succeed.Reported by OpenAI · 2 values | OpenAI | 2 | ||||||||||||||||||||||||||||||
| Indirect prompt injection (ART-style), in the data explorer Run by the developer Measures how often an attacker succeeds with indirect prompt injections within 15 attempts on a benchmark modelled on agent red-teaming attacks.Reported by Anthropic · 2 values | Anthropic | 2 | ||||||||||||||||||||||||||||||
| Internal indirect prompt injection, in the data explorer Run by the developer An internal OpenAI indirect prompt-injection test in the GPT-6 Astra card, scored as the share of attacks the model defends against.Reported by OpenAI · 2 values | OpenAI | 2 | ||||||||||||||||||||||||||||||
| Bug bounty · 1 evaluation, 1 value | ||||||||||||||||||||||||||||||||
| Live bug bounty prompt injection, in the data explorer Run by the developer Reports the share of unique prompt-injection attacks submitted through a live bug bounty that succeeded against the model.Reported by Anthropic · 1 value | Anthropic | 1 | ||||||||||||||||||||||||||||||
Reported values
The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Prompt injection
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
Prompt injection: values by kind of test
Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it.
Hollow: read in a secondary write-up
| Model | Developer | Evaluation | Condition | Value | Document | Kind of test |
|---|---|---|---|---|---|---|
| Claude Opus 4.6 | Anthropic | Agent Red Teaming (ART)Run by Gray SwanAttack success rate within k=100 attempts | extended thinking | Claude Opus 4.6 System CardFeb 2026 | Static benchmark | |
| Claude Opus 4.6 | Anthropic | Agent Red Teaming (ART)Run by Gray SwanAttack success rate within k=100 attempts | no extended thinking | Claude Opus 4.6 System CardFeb 2026 | Static benchmark | |
| Claude Sonnet 4.6 | Anthropic | Computer-use prompt injectionAttack success rate (% of scenarios) | without safeguards | Claude Sonnet 4.6 System Card17 Feb 2026 | Static benchmark | |
| Claude Sonnet 4.6 | Anthropic | Computer-use prompt injectionAttack success rate (% of scenarios) | with safeguards | Claude Sonnet 4.6 System Card17 Feb 2026 | Static benchmark | |
| Claude Opus 4.8 | Anthropic | Indirect prompt injection (external benchmark)Attack success rate | — | Claude Opus 4.8 System Card28 May 2026 | Static benchmark | |
| Claude Sonnet 5 | Anthropic | Computer-use prompt injectionAttack success rate | extended thinking | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Fable 5.1 | Anthropic | Computer-use prompt injectionAttack success rate | — | Claude Fable 5.1 & Claude Mythos 5.1 System Card1 Sep 2026 | Static benchmark | |
| Muse Spark | Meta | AgentDojoPrompt-injection ASR pass@1 | without system mitigations | Muse Spark Safety & Preparedness Report8 Apr 2026 | Static benchmark | |
| Muse Spark | Meta | Agent Red Teaming (ART)Run by Gray SwanAgent red-team ASR pass@1 | without system mitigations | Muse Spark Safety & Preparedness Report8 Apr 2026 | Static benchmark | |
| Muse Spark 1.1 | Meta | AgentDojoPrompt-injection ASR pass@1 | without system mitigations | Muse Spark 1.1 Evaluation Report9 Jul 2026 | Static benchmark | |
| Muse Glimmer-30B | Meta | Siren AgentDojoPrompt-injection attack success rate | open weights | Muse Glimmer-30B model card10 Aug 2026 | Static benchmark | |
| GPT-6 Astra | OpenAI | IPI ArenaRun by Gray SwanEstimated attack success within 15 attempts | 1,810 curated attacks | GPT-6 Astra System Card3 Sep 2026 | Static benchmark | |
| GPT-5.6 Sol | OpenAI | IPI ArenaRun by Gray SwanEstimated attack success within 15 attempts | 1,810 curated attacks | GPT-6 Astra System Card3 Sep 2026 | Static benchmark | |
| grok-code-fast-1 | xAI | AgentDojoPrompt-injection attack success rate | system prompt with refusal policy | Grok Code Fast 1 Model Card26 Aug 2025 | Static benchmark | |
| Claude Opus 4.5 | Anthropic | Coding prompt injection (adaptive attacker)Attack success rate | extended thinking; without safeguards; 1 attempt | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Opus 4.5 | Anthropic | Coding prompt injection (adaptive attacker)Attack success rate | extended thinking; without safeguards; 200 attempts | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Opus 4.5 | Anthropic | Coding prompt injection (adaptive attacker)Attack success rate | no extended thinking; without safeguards; 1 attempt | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Opus 4.5 | Anthropic | Coding prompt injection (adaptive attacker)Attack success rate | no extended thinking; without safeguards; 200 attempts | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Sonnet 4.5 | Anthropic | Coding prompt injection (adaptive attacker)Attack success rate | extended thinking; without safeguards; 1 attempt | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Sonnet 4.5 | Anthropic | Coding prompt injection (adaptive attacker)Attack success rate | extended thinking; without safeguards; 200 attempts | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Sonnet 4.5 | Anthropic | Coding prompt injection (adaptive attacker)Attack success rate | no extended thinking; without safeguards; 1 attempt | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Sonnet 4.5 | Anthropic | Coding prompt injection (adaptive attacker)Attack success rate | no extended thinking; without safeguards; 200 attempts | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Opus 4.5 | Anthropic | Computer use prompt injection (adaptive attacker)Attack success rate | extended thinking; without safeguards; 1 attempt | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Opus 4.5 | Anthropic | Computer use prompt injection (adaptive attacker)Attack success rate | extended thinking; without safeguards; 200 attempts | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Opus 4.5 | Anthropic | Computer use prompt injection (adaptive attacker)Attack success rate | no extended thinking; without safeguards; 1 attempt | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Opus 4.5 | Anthropic | Computer use prompt injection (adaptive attacker)Attack success rate | no extended thinking; without safeguards; 200 attempts | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Opus 4.5 | Anthropic | Computer use prompt injection (adaptive attacker)Attack success rate | no extended thinking; with safeguards; 200 attempts | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Sonnet 4.5 | Anthropic | Computer use prompt injection (adaptive attacker)Attack success rate | extended thinking; without safeguards; 1 attempt | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Sonnet 4.5 | Anthropic | Computer use prompt injection (adaptive attacker)Attack success rate | extended thinking; without safeguards; 200 attempts | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Sonnet 4.5 | Anthropic | Computer use prompt injection (adaptive attacker)Attack success rate | extended thinking; with safeguards; 200 attempts | Claude Opus 4.5 System Card24 Nov 2025 | Adaptive attack | |
| Claude Opus 4.6 | Anthropic | Coding prompt injection (adaptive attacker)Attack success rate | all conditions | Claude Opus 4.6 System CardFeb 2026 | Adaptive attack | |
| Claude Sonnet 5 | Anthropic | Browser use prompt injectionAttack success rate | without new browser safeguards | Claude Sonnet 5 System Card30 Jun 2026 | Adaptive attack | |
| Claude Sonnet 5 | Anthropic | Coding prompt injection (adaptive attacker)Attack success rate | extended thinking | Claude Sonnet 5 System Card30 Jun 2026 | Adaptive attack | |
| Claude Opus 5 | Anthropic | Indirect prompt injection (ART-style)Attacker success within 15 attempts | — | Claude Opus 5 System Card24 Jul 2026 | Adaptive attack | |
| Claude Opus 4.8 | Anthropic | Indirect prompt injection (ART-style)Attacker success within 15 attempts | — | Claude Opus 5 System Card24 Jul 2026 | Adaptive attack | |
| Gemini 2.5 Pro | Google DeepMind | Indirect prompt injection (adaptive attacks)Actor Critic ASR | adaptive attack | Gemini 2.5 Technical Report17 Jun 2025 | Adaptive attack | |
| Gemini 2.5 Pro | Google DeepMind | Indirect prompt injection (adaptive attacks)Beam Search ASR | adaptive attack | Gemini 2.5 Technical Report17 Jun 2025 | Adaptive attack | |
| Gemini 2.5 Pro | Google DeepMind | Indirect prompt injection (adaptive attacks)TAP ASR | adaptive attack | Gemini 2.5 Technical Report17 Jun 2025 | Adaptive attack | |
| Gemini 2.5 Flash | Google DeepMind | Indirect prompt injection (adaptive attacks)Actor Critic ASR | adaptive attack | Gemini 2.5 Technical Report17 Jun 2025 | Adaptive attack | |
| Gemini 2.5 Flash | Google DeepMind | Indirect prompt injection (adaptive attacks)Beam Search ASR | adaptive attack | Gemini 2.5 Technical Report17 Jun 2025 | Adaptive attack | |
| Gemini 2.5 Flash | Google DeepMind | Indirect prompt injection (adaptive attacks)TAP ASR | adaptive attack | Gemini 2.5 Technical Report17 Jun 2025 | Adaptive attack | |
| GPT-5.6 Sol | OpenAI | GPT-RedInstruction hierarchy (direct PI) | — | GPT-5.6 System Card9 Jul 2026 | Adaptive attack | |
| GPT-5.6 Sol | OpenAI | GPT-RedIndirect prompt injection | — | GPT-5.6 System Card9 Jul 2026 | Adaptive attack | |
| Claude Sonnet 5 | Anthropic | Live bug bounty prompt injectionShare of unique attacks succeeding | — | Claude Sonnet 5 System Card30 Jun 2026 | Bug bounty |
Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.
Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.
44 of the family’s 118 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 65 in other units or read the other way (rates from 0 to 1 and percent where higher is better), 3 are categories or statements in words, 4 are restated in a later document or from an earlier version of one and 2 are low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.
The values span more than one and a half orders of magnitude, but three of them are zero, which a logarithmic axis cannot show, so the axis is linear and starts at zero.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
Comparability
- Each of the 25 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
- Values come from the documents of six developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
- Some names are used by more than one evaluation: “AgentDojo” by two. They are different tests, run by different evaluators or reported by different developers, so this page names who ran each one.
- In two evaluations the test changed between documents (Prompt injection and Computer-use prompt injection). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
- In six evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
- Nine evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
- Numbers in this family are printed in two units: rates from 0 to 1 and percent. We never convert one unit into another, so a chart shows one unit at a time.
- For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
- Four values are restatements: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.
Coverage
Documents from six of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.
| Developer | With any value | With a printed number | All its documents |
|---|---|---|---|
| Anthropic | 12 | 10 | 15 |
| DeepSeek | 0 | 0 | 1 |
| Google DeepMind | 1 | 1 | 16 |
| Meta | 3 | 3 | 5 |
| Moonshot AI | 1 | 1 | 2 |
| OpenAI | 12 | 12 | 22 |
| xAI | 5 | 5 | 8 |
| Zhipu AI | 0 | 0 | 1 |
Related findings
No finding is about this family yet.