Dangerous capabilities and risk determinations
Capability evaluations in biology, chemistry, cyber and AI research, and the developer's threshold or risk-level decisions.
- Reported by
- Anthropic, Google DeepMind, Meta, Moonshot AI, OpenAI and xAI
- Evaluations
- 90 evaluations, of seven kinds of test
- Values
- 406 values in 60 documents
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Definition
This family covers what a model can do in areas where misuse could cause severe harm: biology and chemistry, cybersecurity, and research that could speed up the development of AI itself.
Developers test these capabilities with question sets written by experts, hands-on tasks such as capture-the-flag challenges or laboratory protocols, long agentic projects, and studies that compare how well people perform with and without the model’s help. Capability results come in many units: pass rates and scores, counts of solved tasks, time spent, and multiples of human performance.
The family also holds the conclusions developers draw from such tests under their own safety frameworks: whether a model reaches a capability threshold or a named risk level. These determinations are recorded as printed, as categories rather than numbers, and are not mapped from one framework to another.
The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).
How developers measure it
90 evaluations of seven kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.
| Evaluation | Reported by | Values | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Capability benchmark · 66 evaluations, 259 values | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Cyber challenges · Irregular, in the data explorer Run by Irregular (third party) Irregular's suite of self-contained cyber challenges, reported by challenge type and by difficulty tier.Reported by OpenAI · 18 values | OpenAI | 18 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Cyber Range · OpenAI, in the data explorer Run by the developer Simulated network environments in which the model must carry out a multi-step cyber operation end to end, scored by scenarios passed.Reported by OpenAI · 16 values | OpenAI | 16 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Anthropic ECI (AECI), in the data explorer Run by the developer Anthropic's aggregate capability index, fitted over a set of benchmarks and used to track AI R&D capability across models.Reported by Anthropic · 15 values | Anthropic | 15 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| AI R&D LLM training task, in the data explorer Run by the developer An Anthropic AI research task in which the model speeds up a language-model training run, reported as average speedup.Reported by Anthropic · 10 values | Anthropic | 10 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| CyberGym · Anthropic, in the data explorer Run by the developer A public benchmark of reproducing known software vulnerabilities that Anthropic runs as a cyber-capability measure.Reported by Anthropic · 10 values | Anthropic | 10 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Cybench · Anthropic, in the data explorer Run by the developer A public set of capture-the-flag security challenges that Anthropic runs on its models as a cyber-capability measure.Reported by Anthropic · 9 values | Anthropic | 9 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Long-form virology task 1, in the data explorer Run by the developer The first of two internal long-form biology-risk tasks Anthropic scores from 0 to 1 in its capability assessments.Reported by Anthropic · 9 values | Anthropic | 9 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Multimodal virology (VCT), in the data explorer Run by the developer A multimodal virology knowledge benchmark that Anthropic runs and scores from 0 to 1 in its biology-risk assessments.Reported by Anthropic · 9 values | Anthropic | 9 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Show the other 58 evaluations of this kind (163 values)
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Framework determination · 13 evaluations, 125 values | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Preparedness Framework determination, in the data explorer Run by the developer OpenAI's formal capability-level decisions under its Preparedness Framework for each tracked category: biological and chemical, cybersecurity, and AI self-improvement.Reported by OpenAI · 55 values | OpenAI | 55 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Frontier Safety Framework determination, in the data explorer Run by the developer Google DeepMind's decision on whether a model reaches the alert threshold or Critical Capability Level in each risk domain of its Frontier Safety Framework.Reported by Google DeepMind · 32 values | Google DeepMind | 32 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| RSP deployment standard, in the data explorer Run by the developer The AI Safety Level standard Anthropic states a model is deployed under after its Responsible Scaling Policy assessment.Reported by Anthropic · 8 values | Anthropic | 8 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Autonomy threat model 2 (automated R&D), in the data explorer Run by the developer Anthropic's determination of whether its second autonomy threat model, automated AI research, applies to a model under the Responsible Scaling Policy.Reported by Anthropic · 4 values | Anthropic | 4 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| CB-2 threat model, in the data explorer Run by the developer Anthropic's determination of whether the CB-2 threat model applies to a model under the Responsible Scaling Policy.Reported by Anthropic · 4 values | Anthropic | 4 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Risk determination, in the data explorer Run by the developer The risk tier Meta assigns a model in each risk domain under its Advanced AI Scaling Framework, before or after mitigations.Reported by Meta · 4 values | Meta | 4 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| RSP determinations, in the data explorer Run by the developer A single statement in a card that gives Anthropic's Responsible Scaling Policy conclusions for several threat models at once.Reported by Anthropic · 4 values | Anthropic | 4 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| AI R&D-4 threshold, in the data explorer Run by the developer Anthropic's statement of whether a model crosses its AI R&D-4 capability threshold under the Responsible Scaling Policy.Reported by Anthropic · 3 values | Anthropic | 3 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Show the other 5 evaluations of this kind (11 values)
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Uplift study · 3 evaluations, 6 values | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Harmful manipulation efficacy study, in the data explorer Run by the developer Human-participant study of how much conversations with the model shift participants' beliefs compared with a non-AI baseline.Reported by Google DeepMind · 4 values | Google DeepMind | 4 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ASL-4 virology uplift trial, in the data explorer Run by the developer A human uplift trial in Anthropic's ASL-4 biology assessment, reported as the ratio of task scores with model help to an internet-only control.Reported by Anthropic · 1 value | Anthropic | 1 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Virology protocol uplift trial, in the data explorer Run by the developer A human trial comparing participants with and without model help on a written biology task, reported as mean critical failures.Reported by Anthropic · 1 value | Anthropic | 1 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Red teaming · 2 evaluations, 5 values | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| CBRN red-team uplift, in the data explorer Run by the developer Expert red teamers score how much the model adds to web-only resources in biological and chemical scenarios, graded on a rubric.Reported by Google DeepMind · 4 values | Google DeepMind | 4 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Bio bottleneck sub-stages red-teaming, in the data explorer Run by the developer Counts the bioweapon-pathway sub-stages on which red-team scores for the model stayed below a set threshold.Reported by Google DeepMind · 1 value | Google DeepMind | 1 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Qualitative assessment · 3 evaluations, 5 values | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Bio capability lift, in the data explorer Run by the developer xAI's statement in words of how much a model's biology capability rose over the previous model.Reported by xAI · 2 values | xAI | 2 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Release decision, in the data explorer Run by the developer How Anthropic chose to release a model (general availability, limited partners, safeguards attached), recorded as a categorical fact.Reported by Anthropic · 2 values | Anthropic | 2 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| AI R&D assessment · METR, in the data explorer Run by METR (third party) METR's judgment of how much a model could speed up AI research and development, given as an acceleration multiplier.Reported by Anthropic · 1 value | Anthropic | 1 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Scenario evaluation · 1 evaluation, 3 values | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Manipulative cue rate, in the data explorer Run by the developer Share of conversational turns in constructed scenarios in which the model uses manipulative cues.Reported by Google DeepMind · 3 values | Google DeepMind | 3 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Survey · 2 evaluations, 3 values | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Internal AI R&D survey, in the data explorer Run by the developer A survey of Anthropic research staff on whether a model could replace an entry-level researcher.Reported by Anthropic · 2 values | Anthropic | 2 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Internal model use survey, in the data explorer Run by the developer A survey of Anthropic staff on the productivity gain they get from using a model internally.Reported by Anthropic · 1 value | Anthropic | 1 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Reported values
The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Dangerous capabilities and risk determinations
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
Dangerous capabilities and risk determinations: values by kind of test
Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it.
Hollow: read in a secondary write-up
| Model | Developer | Evaluation | Condition | Value | Document | Kind of test |
|---|---|---|---|---|---|---|
| Claude Opus 4.5 | Anthropic | Bioinformatics evaluationsScore | — | Claude Opus 4.5 System Card24 Nov 2025 | Capability benchmark | |
| Claude Opus 4.6 | Anthropic | CybenchSuccess rate, pass@30 | — | Claude Opus 4.6 System CardFeb 2026 | Capability benchmark | |
| Claude Opus 4.6 | Anthropic | CyberGymTargeted vuln reproduction rate | — | Claude Opus 4.6 System CardFeb 2026 | Capability benchmark | |
| Claude Mythos Preview | Anthropic | CybenchSuccess rate, pass@1 | 35 challenges | Claude Mythos Preview System Card7 Apr 2026 | Capability benchmark | |
| Claude Mythos Preview | Anthropic | CyberGymTargeted vuln reproduction rate | — | Claude Mythos Preview System Card7 Apr 2026 | Capability benchmark | |
| Claude Mythos Preview | Anthropic | AI R&D novel compiler taskPass rate | — | Claude Mythos Preview System Card7 Apr 2026 | Capability benchmark | |
| Claude Opus 4.7 | Anthropic | AI R&D novel compiler taskPass rate | — | Claude Opus 4.7 System Card16 Apr 2026 | Capability benchmark | |
| Claude Opus 4.7 | Anthropic | CybenchSuccess rate, pass@1 | 35 challenges | Claude Opus 4.7 System Card16 Apr 2026 | Capability benchmark | |
| Claude Opus 4.8 | Anthropic | CyberGymTargeted vuln reproduction rate | — | Claude Opus 4.8 System Card28 May 2026 | Capability benchmark | |
| Claude Opus 4.8 | Anthropic | CyberGymTargeted vuln reproduction rate | with safeguards | Claude Opus 4.8 System Card28 May 2026 | Capability benchmark | |
| Claude Opus 4.7 | Anthropic | CyberGymTargeted vuln reproduction rate | — | Claude Opus 4.8 System Card28 May 2026 | Capability benchmark | |
| Claude Sonnet 4.6 | Anthropic | CyberGymTargeted vuln reproduction rate | — | Claude Opus 4.8 System Card28 May 2026 | Capability benchmark | |
| Claude Opus 4.8 | Anthropic | Firefox 147 exploitationFull exploit rate | without safeguards | Claude Opus 4.8 System Card28 May 2026 | Capability benchmark | |
| Claude Mythos Preview | Anthropic | Firefox 147 exploitationFull exploit rate | without safeguards | Claude Opus 4.8 System Card28 May 2026 | Capability benchmark | |
| Claude Mythos 5 | Anthropic | AI R&D novel compiler taskPass rate | — | Claude Fable 5 & Claude Mythos 5 System Card9 Jun 2026 | Capability benchmark | |
| Claude Sonnet 5 | Anthropic | AI R&D novel compiler taskPass rate | — | Claude Sonnet 5 System Card30 Jun 2026 | Capability benchmark | |
| Claude Sonnet 5 | Anthropic | CyberGymTargeted vuln reproduction rate | — | Claude Sonnet 5 System Card30 Jun 2026 | Capability benchmark | |
| Claude Sonnet 5 | Anthropic | Firefox 147 exploitationFull exploit rate | — | Claude Sonnet 5 System Card30 Jun 2026 | Capability benchmark | |
| Claude Opus 5 | Anthropic | AI R&D novel compiler taskPass rate | — | Claude Opus 5 System Card24 Jul 2026 | Capability benchmark | |
| Claude Opus 5 | Anthropic | Firefox 147 exploitationFull exploit rate | — | Claude Opus 5 System Card24 Jul 2026 | Capability benchmark | |
| Claude Mythos 5 | Anthropic | Firefox 147 exploitationFull exploit rate | — | Claude Opus 5 System Card24 Jul 2026 | Capability benchmark | |
| Claude Opus 5 | Anthropic | CyScenarioBenchSolve rate | — | Claude Opus 5 System Card24 Jul 2026 | Capability benchmark | |
| Claude Mythos 5.1 | Anthropic | Firefox 147 exploitationFull exploit rate | — | Claude Fable 5.1 & Claude Mythos 5.1 System Card1 Sep 2026 | Capability benchmark | |
| Claude Opus 5.5 | Anthropic | BioMysteryBenchAccuracy (human difficult) | — | Claude Opus 5.5 System Card22 Sep 2026 | Capability benchmark | |
| Claude Opus 5.5 | Anthropic | CoBenchScore | — | Claude Opus 5.5 System Card22 Sep 2026 | Capability benchmark | |
| Claude Opus 5 | Anthropic | CoBenchScore | — | Claude Opus 5.5 System Card22 Sep 2026 | Capability benchmark | |
| Claude Mythos 5.1 | Anthropic | CoBenchScore | — | Claude Opus 5.5 System Card22 Sep 2026 | Capability benchmark | |
| Gemini 2.5 Pro | Google DeepMind | RE-BenchBest score as % of expert solution | max across tasks | Gemini 2.5 Technical Report17 Jun 2025 | Capability benchmark | |
| Gemini 3.7 Flash | Google DeepMind | GRB internal research-engineering benchmarkAverage pass@1 | — | Gemini 3.7 Flash Frontier Safety Framework Report13 Aug 2026 | Capability benchmark | |
| Gemini 3.1 Pro | Google DeepMind | GRB internal research-engineering benchmarkAverage pass@1 | — | Gemini 3.7 Flash Frontier Safety Framework Report13 Aug 2026 | Capability benchmark | |
| Muse Spark | Meta | VCTAccuracy | without system mitigations | Muse Spark Safety & Preparedness Report8 Apr 2026 | Capability benchmark | |
| Muse Spark | Meta | CybenchPass@1 | without system mitigations | Muse Spark Safety & Preparedness Report8 Apr 2026 | Capability benchmark | |
| Muse Spark 1.1 | Meta | CybenchPass@1 | without system mitigations | Muse Spark 1.1 Evaluation Report9 Jul 2026 | Capability benchmark | |
| Muse Glimmer-30B | Meta | VCTAccuracy | open weights | Muse Glimmer-30B model card10 Aug 2026 | Capability benchmark | |
| o3 | OpenAI | Capture the FlagHigh school | no browsing; pass@12 | OpenAI o3 and o4-mini System Card16 Apr 2025 | Capability benchmark | |
| o3 | OpenAI | Capture the FlagCollegiate | no browsing; pass@12 | OpenAI o3 and o4-mini System Card16 Apr 2025 | Capability benchmark | |
| o3 | OpenAI | Capture the FlagProfessional | no browsing; pass@12 | OpenAI o3 and o4-mini System Card16 Apr 2025 | Capability benchmark | |
| o4-mini | OpenAI | Capture the FlagHigh school | no browsing; pass@12 | OpenAI o3 and o4-mini System Card16 Apr 2025 | Capability benchmark | |
| o4-mini | OpenAI | Capture the FlagCollegiate | no browsing; pass@12 | OpenAI o3 and o4-mini System Card16 Apr 2025 | Capability benchmark | |
| o4-mini | OpenAI | Capture the FlagProfessional | no browsing; pass@12 | OpenAI o3 and o4-mini System Card16 Apr 2025 | Capability benchmark | |
| o3 | OpenAI | SWE-bench VerifiedPass rate | helpful-only | OpenAI o3 and o4-mini System Card16 Apr 2025 | Capability benchmark | |
| o3 | OpenAI | OpenAI PRsPass rate | — | OpenAI o3 and o4-mini System Card16 Apr 2025 | Capability benchmark | |
| o4-mini | OpenAI | OpenAI PRsPass rate | — | OpenAI o3 and o4-mini System Card16 Apr 2025 | Capability benchmark | |
| o3 | OpenAI | SWE-LancerIC SWE accuracy | helpful-only; with browsing | OpenAI o3 and o4-mini System Card16 Apr 2025 | Capability benchmark | |
| o4-mini | OpenAI | PaperBenchReplication score | no browsing | OpenAI o3 and o4-mini System Card16 Apr 2025 | Capability benchmark | |
| gpt-oss-120b | OpenAI | SWE-bench VerifiedPass rate | high; N=477 | gpt-oss-120b & gpt-oss-20b Model Card5 Aug 2025 | Capability benchmark | |
| gpt-5-thinking | OpenAI | OpenAI-Proof Q&APass rate | — | GPT-5 System Card7 Aug 2025 | Capability benchmark | |
| ChatGPT agent | OpenAI | MLE-BenchMedal rate | — | GPT-5 System Card7 Aug 2025 | Capability benchmark | |
| gpt-5-thinking | OpenAI | Cyber challengesRun by IrregularEvasion: average success rate | — | GPT-5 System Card7 Aug 2025 | Capability benchmark | |
| gpt-5-thinking | OpenAI | Cyber challengesRun by IrregularVulnerability discovery and exploitation: average success rate | — | GPT-5 System Card7 Aug 2025 | Capability benchmark | |
| gpt-5-thinking | OpenAI | Cyber challengesRun by IrregularNetwork attack simulation: average success rate | — | GPT-5 System Card7 Aug 2025 | Capability benchmark | |
| GPT-5.1-Codex-Max | OpenAI | TroubleshootingBenchScore | refusals counted as successes | GPT-5.1-Codex-Max System Card18 Nov 2025 | Capability benchmark | |
| GPT-5.1-Codex-Max | OpenAI | Tacit knowledge and troubleshootingScore | — | GPT-5.1-Codex-Max System Card18 Nov 2025 | Capability benchmark | |
| gpt-5.2-thinking | OpenAI | Cyber challengesRun by IrregularVulnerability research and exploitation: average success rate | v1 atomic challenge suite | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Capability benchmark | |
| gpt-5.2-thinking | OpenAI | Cyber challengesRun by IrregularNetwork attack simulation: average success rate | v1 atomic challenge suite | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Capability benchmark | |
| gpt-5.2-thinking | OpenAI | Cyber challengesRun by IrregularEvasion: average success rate | v1 atomic challenge suite | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Capability benchmark | |
| GPT-5.1-Codex-Max | OpenAI | OpenAI-Proof Q&APass rate | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Capability benchmark | |
| GPT-5.3-Codex | OpenAI | Cyber RangeCombined pass rate | — | GPT-5.3-Codex System Card5 Feb 2026 | Capability benchmark | |
| GPT-5.2-Codex | OpenAI | Cyber RangeCombined pass rate | — | GPT-5.3-Codex System Card5 Feb 2026 | Capability benchmark | |
| GPT-5.1-Codex-Max | OpenAI | Cyber RangeCombined pass rate | — | GPT-5.3-Codex System Card5 Feb 2026 | Capability benchmark | |
| GPT-5.2 Thinking | OpenAI | Cyber RangeCombined pass rate | — | GPT-5.3-Codex System Card5 Feb 2026 | Capability benchmark | |
| GPT-5.3-Codex | OpenAI | CVE-BenchPass rate | — | GPT-5.3-Codex System Card5 Feb 2026 | Capability benchmark | |
| GPT-5.3-Codex | OpenAI | Cyber challengesRun by IrregularNetwork attack simulation: average success rate | — | GPT-5.3-Codex System Card5 Feb 2026 | Capability benchmark | |
| GPT-5.3-Codex | OpenAI | Cyber challengesRun by IrregularVulnerability research and exploitation: average success rate | — | GPT-5.3-Codex System Card5 Feb 2026 | Capability benchmark | |
| GPT-5.3-Codex | OpenAI | Cyber challengesRun by IrregularEvasion: average success rate | — | GPT-5.3-Codex System Card5 Feb 2026 | Capability benchmark | |
| GPT-5.4 Thinking | OpenAI | Cyber RangeCombined pass rate | — | GPT-5.4 Thinking System Card5 Mar 2026 | Capability benchmark | |
| GPT-5.5 | OpenAI | Cyber RangeCombined pass rate | — | GPT-5.5 System Card23 Apr 2026 | Capability benchmark | |
| GPT-5.5 | OpenAI | Cyber tasksRun by UK AI Security InstitutePass rate | pass@5 | GPT-5.5 System Card23 Apr 2026 | Capability benchmark | |
| GPT-5.5 | OpenAI | CyScenarioBenchRun by IrregularAverage success rate | — | GPT-5.5 System Card23 Apr 2026 | Capability benchmark | |
| GPT-5.5 | OpenAI | Internal Research DebuggingMedian score | — | GPT-5.5 System Card23 Apr 2026 | Capability benchmark | |
| GPT-5.5 | OpenAI | Hard-negative protein bindingScore | pass@4 | GPT-5.5 System Card23 Apr 2026 | Capability benchmark | |
| GPT-5.5 | OpenAI | DNA sequence designScore | pass@1 | GPT-5.5 System Card23 Apr 2026 | Capability benchmark | |
| GPT-5.5 Instant | OpenAI | Cyber RangeCombined pass rate | — | GPT-5.5 Instant System Card4 May 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | FrontierCyberRun by IrregularChallenges solved | 197 challenges | GPT-5.6 Preview System Card25 Jun 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | Average success (benchmark not named)Run by IrregularAverage success rate | — | GPT-5.6 Preview System Card25 Jun 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | Multimodal Troubleshooting VirologyScore | — | GPT-5.6 System Card9 Jul 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | ProtocolQA Open-EndedScore | — | GPT-5.6 System Card9 Jul 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | TroubleshootingBenchScore | — | GPT-5.6 System Card9 Jul 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | Hard-negative protein bindingScore | pass@4 | GPT-5.6 System Card9 Jul 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | DNA sequence designScore | — | GPT-5.6 System Card9 Jul 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | Capture the FlagInternal | — | GPT-5.6 System Card9 Jul 2026 | Capability benchmark | |
| GPT-5.6 Terra | OpenAI | Tacit knowledge and troubleshootingScore | — | GPT-5.6 System Card9 Jul 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | Expert CTFRun by UK AI Security InstitutePass rate | — | GPT-5.6 System Card9 Jul 2026 | Capability benchmark | |
| GPT-5.5 | OpenAI | Expert CTFRun by UK AI Security InstitutePass rate | — | GPT-5.6 System Card9 Jul 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | Cyber range 'The Last Ones'Run by UK AI Security InstituteAttempts completed | — | GPT-5.6 System Card9 Jul 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | FrontierCyberRun by IrregularSuccess rate | Hard tier | GPT-5.6 System Card9 Jul 2026 | Capability benchmark | |
| GPT-6 Astra | OpenAI | ExploitBenchWorking exploit rate on known V8 vulns | — | GPT-6 Astra launch post3 Sep 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | ExploitBenchWorking exploit rate on known V8 vulns | — | GPT-6 Astra launch post3 Sep 2026 | Capability benchmark | |
| GPT-6 Astra | OpenAI | ExploitBenchExploit success on post-cutoff vulns | — | GPT-6 Astra launch post3 Sep 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | ExploitBenchExploit success on post-cutoff vulns | — | GPT-6 Astra launch post3 Sep 2026 | Capability benchmark | |
| GPT-6 Astra | OpenAI | ExploitGymDiscovery-to-exploit success rate | — | GPT-6 Astra launch post3 Sep 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | ExploitGymDiscovery-to-exploit success rate | — | GPT-6 Astra launch post3 Sep 2026 | Capability benchmark | |
| GPT-6 Astra | OpenAI | SRE-BenchBinary reverse-engineering success | pass@1 | GPT-6 Astra launch post3 Sep 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | SRE-BenchBinary reverse-engineering success | pass@1 | GPT-6 Astra launch post3 Sep 2026 | Capability benchmark | |
| GPT-6 Astra | OpenAI | SRE-BenchBinary reverse-engineering success | pass@4 | GPT-6 Astra launch post3 Sep 2026 | Capability benchmark | |
| GPT-5.6 Sol | OpenAI | SRE-BenchBinary reverse-engineering success | pass@4 | GPT-6 Astra launch post3 Sep 2026 | Capability benchmark | |
| grok-code-fast-1 | xAI | VCTAccuracy | without safeguards | Grok Code Fast 1 Model Card26 Aug 2025 | Capability benchmark | |
| grok-code-fast-1 | xAI | CybenchUnguided success rate | without safeguards | Grok Code Fast 1 Model Card26 Aug 2025 | Capability benchmark | |
| Grok 4 Fast (reasoning) | xAI | BioLP-BenchAccuracy | reasoning; without safeguards | Grok 4 Fast Model Card19 Sep 2025 | Capability benchmark | |
| Grok 4 Fast (reasoning) | xAI | VCTAccuracy | reasoning; without safeguards | Grok 4 Fast Model Card19 Sep 2025 | Capability benchmark | |
| Grok 4 Fast (reasoning) | xAI | CybenchUnguided success rate | reasoning; without safeguards | Grok 4 Fast Model Card19 Sep 2025 | Capability benchmark | |
| Grok 4.5 (high) | xAI | VCTAccuracy | high | Grok 4.5 Model Card14 Jul 2026 | Capability benchmark | |
| Grok 4.5 (high) | xAI | ProtocolQA open-endedAccuracy | high | Grok 4.5 Model Card14 Jul 2026 | Capability benchmark | |
| Grok 4.5 (high) | xAI | CyberGymMean vulnerabilities reproduced | high; unrestricted | Grok 4.5 Model Card14 Jul 2026 | Capability benchmark | |
| Grok 4.6 (high) | xAI | VCTAccuracy | high | Grok 4.6 Model Card12 Aug 2026 | Capability benchmark | |
| Grok 4.6 (high) | xAI | WMDPCyber accuracy | high | Grok 4.6 Model Card12 Aug 2026 | Capability benchmark | |
| Grok 4.6 (high) | xAI | CyberGymMean vulnerabilities reproduced | high; unrestricted | Grok 4.6 Model Card12 Aug 2026 | Capability benchmark | |
| Grok 4.6 (high) | xAI | CVE-BenchReward | high; unrestricted | Grok 4.6 Model Card12 Aug 2026 | Capability benchmark | |
| Grok 4.7 (high) | xAI | VCTAccuracy | high | Grok 4.7 Model Card21 Sep 2026 | Capability benchmark | |
| Grok 4.7 (high) | xAI | CyberGymMean vulnerabilities reproduced | high; unrestricted | Grok 4.7 Model Card21 Sep 2026 | Capability benchmark | |
| Gemini 3 Pro | Google DeepMind | Manipulative cue rateTurns with manipulative cues | — | Gemini 3.7 Flash Frontier Safety Framework Report13 Aug 2026 | Scenario evaluation | |
| Gemini 3.7 Flash | Google DeepMind | Manipulative cue rateTurns with manipulative cues | — | Gemini 3.7 Flash Frontier Safety Framework Report13 Aug 2026 | Scenario evaluation | |
| Gemini 3.7 Flash | Google DeepMind | Manipulative cue rateTurns with manipulative cues (higher-stakes scenarios) | higher-stakes scenarios | Gemini 3.7 Flash Frontier Safety Framework Report13 Aug 2026 | Scenario evaluation |
Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.
Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.
113 of the family’s 406 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 139 in other units or read the other way (minutes, USD, counts, human-normalised score and 9 others), 137 are categories or statements in words, 12 are restated in a later document or from an earlier version of one and 5 are low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.
The values span more than one and a half orders of magnitude, but one of them is zero, which a logarithmic axis cannot show, so the axis is linear and starts at zero.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
Comparability
- Each of the 90 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
- Values come from the documents of six developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
- Some names are used by more than one evaluation: “Cyber Range” by two, “CyberGym” by two, “Cybench” by three, “VCT” by two, “ExploitBench” by two, “CVE-Bench” by two, “CyScenarioBench” by two and “ProtocolQA Open-Ended” by two. They are different tests, run by different evaluators or reported by different developers, so this page names who ran each one.
- In eight evaluations the test changed between documents (among them Frontier Safety Framework determination, Cyber challenges · Irregular and Cyber Range · OpenAI). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
- In 38 evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
- 32 evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
- Numbers in this family are printed in 14 units: percent, minutes, USD, counts, human-normalised score, scores from 0 to 1 and 8 others. We never convert one unit into another, so a chart shows one unit at a time.
- For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
- 16 values are restatements: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.
Coverage
Documents from six of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.
| Developer | With any value | With a printed number | All its documents |
|---|---|---|---|
| Anthropic | 15 | 13 | 15 |
| DeepSeek | 0 | 0 | 1 |
| Google DeepMind | 13 | 6 | 16 |
| Meta | 4 | 3 | 5 |
| Moonshot AI | 1 | 1 | 2 |
| OpenAI | 19 | 13 | 22 |
| xAI | 8 | 8 | 8 |
| Zhipu AI | 0 | 0 | 1 |
Related findings
No finding is about this family yet.