Definition

Prompt injection is an attack on a model working as an agent. Instructions are hidden in content the model reads while doing a task, such as a web page, an email, a document or the output of a tool, in the hope that the model follows them instead of its user. A successful injection could make an agent leak data, take actions nobody asked for, or abandon its task.

Developers measure resistance with benchmarks of planted instructions across browsing, coding, computer use and tool calls, with adaptive attackers who refine their attempts, and sometimes with public bug bounties.

Results are reported as attack success rates or as the share of attacks resisted, often with and without extra safeguards. Tasks, attack sets and safeguards differ between tests, and an attack success rate depends heavily on how many attempts the attacker is allowed.

The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).

How developers measure it

25 evaluations of three kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.

The 25 evaluations in Prompt injection, grouped by kind of test: who ran each one, what it counts, the developers whose documents report it and how many values it has.
EvaluationReported byValues
Static benchmark · 17 evaluations, 84 values
Prompt injection, in the data explorer Run by the developer OpenAI's prompt-injection sets, which plant instructions in web pages, tool outputs, connectors or function results and score how often the model does not follow them.Reported by OpenAI · 25 valuesOpenAI25
Computer-use prompt injection, in the data explorer Run by the developer Measures the share of prompt-injection attacks placed in computer-use environments that the model resists, with and without safeguards.Reported by Anthropic · 15 valuesAnthropic15
Instruction hierarchy, in the data explorer Run by the developer Checks whether the model keeps to higher-priority system or developer instructions when a lower-priority message conflicts with them, including injected hijacking and system-prompt extraction attempts.Reported by OpenAI · 9 valuesOpenAI9
AgentDojo · xAI, in the data explorer Run by the developer A public benchmark of agent tasks with prompt injections planted in tool outputs, run by xAI, measuring how often an injected instruction succeeds.Reported by xAI · 5 valuesxAI5
MCP prompt injection, in the data explorer Run by the developer Measures the share of prompt-injection attacks delivered through MCP tool results that the model resists.Reported by Anthropic · 5 valuesAnthropic5
Prompt injection (Codex env), in the data explorer Run by the developer Prompt-injection attacks placed in Codex coding environments, scored on the share of attacks the coding agent ignores.Reported by OpenAI · 5 valuesOpenAI5
Tool-use prompt injection, in the data explorer Run by the developer Measures the share of prompt-injection attacks delivered through tool outputs that the model resists.Reported by Anthropic · 5 valuesAnthropic5
Agent Red Teaming (ART) · Gray Swan, in the data explorer Run by Gray Swan (third party) Gray Swan's agent red-teaming benchmark, measuring how often attacks on an agent succeed within a set number of attempts.Reported by Anthropic and Meta · 4 valuesAnthropic and Meta4
Show the other 9 evaluations of this kind (11 values)
More static benchmarks in Prompt injection
EvaluationReported byValues
AgentDojo · Meta, in the data explorer Run by the developer A public benchmark of agent tasks with prompt injections planted in tool outputs, run by Meta and reported as attack success rate at one attempt.Reported by Meta · 2 valuesMeta2
IPI Arena · Gray Swan, in the data explorer Run by Gray Swan (third party) Gray Swan's set of curated indirect prompt-injection attacks, measuring estimated attack success within a set number of attempts.Reported by OpenAI · 2 valuesOpenAI2
External prompt injection benchmark, in the data explorer Run by the developer An unnamed external prompt-injection benchmark on which the card ranks the model against earlier Claude models.Reported by Anthropic · 1 valueAnthropic1
IHEval, in the data explorer Run by the developer A public instruction-hierarchy benchmark testing whether the model follows higher-priority instructions over conflicting lower-priority ones.Reported by Meta · 1 valueMeta1
Indirect prompt injection (external benchmark), in the data explorer Run by the developer An externally built indirect prompt-injection benchmark, measuring attack success over many attempts.Reported by Anthropic · 1 valueAnthropic1
Instruction hierarchy (internal), in the data explorer Run by the developer An internal OpenAI instruction-hierarchy set in the GPT-6 Astra card, reported as one overall robustness rate that the card treats as saturated.Reported by OpenAI · 1 valueOpenAI1
Prompt injection (user-pasted text), in the data explorer Run by the developer Assesses how susceptible the model is to prompt injections hidden in text that a user pastes into the conversation.Reported by Anthropic · 1 valueAnthropic1
Promptfoo red-team (prompt injection), in the data explorer Run by the developer Moonshot's run of the Promptfoo red-teaming tool, measuring the share of safe responses to harmful requests delivered through prompt injection.Reported by Moonshot AI · 1 valueMoonshot AI1
Siren AgentDojo, in the data explorer Run by the developer Prompt-injection attack success on the AgentDojo agent benchmark under a setup the card calls Siren.Reported by Meta · 1 valueMeta1
Adaptive attack · 7 evaluations, 33 values
Coding prompt injection (adaptive attacker), in the data explorer Run by the developer Measures how often an adaptive attacker hijacks the model through prompt injections in coding environments.Reported by Anthropic · 10 valuesAnthropic10
Computer use prompt injection (adaptive attacker), in the data explorer Run by the developer Measures how often an adaptive attacker hijacks the model through prompt injections in computer-use environments, counted per attempt or per scenario.Reported by Anthropic · 8 valuesAnthropic8
Indirect prompt injection (adaptive attacks), in the data explorer Run by the developer Success rate of three adaptive attack methods that plant instructions in content the model processes, measured over 500 held-out scenarios.Reported by Google DeepMind · 6 valuesGoogle DeepMind6
Browser use prompt injection, in the data explorer Run by the developer Measures how often prompt injections on web content hijack the model while it operates a browser.Reported by Anthropic · 3 valuesAnthropic3
GPT-Red, in the data explorer Run by the developer OpenAI's automated red-teaming attacks on prompt injection and the instruction hierarchy, added to the GPT-5.6 card in August 2026 and scored as the share of attack attempts that succeed.Reported by OpenAI · 2 valuesOpenAI2
Indirect prompt injection (ART-style), in the data explorer Run by the developer Measures how often an attacker succeeds with indirect prompt injections within 15 attempts on a benchmark modelled on agent red-teaming attacks.Reported by Anthropic · 2 valuesAnthropic2
Internal indirect prompt injection, in the data explorer Run by the developer An internal OpenAI indirect prompt-injection test in the GPT-6 Astra card, scored as the share of attacks the model defends against.Reported by OpenAI · 2 valuesOpenAI2
Bug bounty · 1 evaluation, 1 value
Live bug bounty prompt injection, in the data explorer Run by the developer Reports the share of unique prompt-injection attacks submitted through a live bug bounty that succeeded against the model.Reported by Anthropic · 1 valueAnthropic1

Reported values

The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Prompt injection

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

Prompt injection: values by kind of test

Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it.

Hollow: read in a secondary write-up

Static benchmarkAdaptive attackBug bounty020406080100Percent

Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.

Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.

44 of the family’s 118 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 65 in other units or read the other way (rates from 0 to 1 and percent where higher is better), 3 are categories or statements in words, 4 are restated in a later document or from an earlier version of one and 2 are low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.

The values span more than one and a half orders of magnitude, but three of them are zero, which a logarithmic axis cannot show, so the axis is linear and starts at zero.

Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0

Comparability

  • Each of the 25 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
  • Values come from the documents of six developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
  • Some names are used by more than one evaluation: “AgentDojo” by two. They are different tests, run by different evaluators or reported by different developers, so this page names who ran each one.
  • In two evaluations the test changed between documents (Prompt injection and Computer-use prompt injection). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
  • In six evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
  • Nine evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
  • Numbers in this family are printed in two units: rates from 0 to 1 and percent. We never convert one unit into another, so a chart shows one unit at a time.
  • For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
  • Four values are restatements: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.

Coverage

Documents from six of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.

Each developer's documents with values in Prompt injection, out of all its documents in the dataset
DeveloperWith any valueWith a printed numberAll its documents
Anthropic121015
DeepSeek001
Google DeepMind1116
Meta335
Moonshot AI112
OpenAI121222
xAI558
Zhipu AI001

Related findings

No finding is about this family yet.