Definition

This family covers what a model can do in areas where misuse could cause severe harm: biology and chemistry, cybersecurity, and research that could speed up the development of AI itself.

Developers test these capabilities with question sets written by experts, hands-on tasks such as capture-the-flag challenges or laboratory protocols, long agentic projects, and studies that compare how well people perform with and without the model’s help. Capability results come in many units: pass rates and scores, counts of solved tasks, time spent, and multiples of human performance.

The family also holds the conclusions developers draw from such tests under their own safety frameworks: whether a model reaches a capability threshold or a named risk level. These determinations are recorded as printed, as categories rather than numbers, and are not mapped from one framework to another.

The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).

How developers measure it

90 evaluations of seven kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.

The 90 evaluations in Dangerous capabilities and risk determinations, grouped by kind of test: who ran each one, what it counts, the developers whose documents report it and how many values it has.
EvaluationReported byValues
Capability benchmark · 66 evaluations, 259 values
Cyber challenges · Irregular, in the data explorer Run by Irregular (third party) Irregular's suite of self-contained cyber challenges, reported by challenge type and by difficulty tier.Reported by OpenAI · 18 valuesOpenAI18
Cyber Range · OpenAI, in the data explorer Run by the developer Simulated network environments in which the model must carry out a multi-step cyber operation end to end, scored by scenarios passed.Reported by OpenAI · 16 valuesOpenAI16
Anthropic ECI (AECI), in the data explorer Run by the developer Anthropic's aggregate capability index, fitted over a set of benchmarks and used to track AI R&D capability across models.Reported by Anthropic · 15 valuesAnthropic15
AI R&D LLM training task, in the data explorer Run by the developer An Anthropic AI research task in which the model speeds up a language-model training run, reported as average speedup.Reported by Anthropic · 10 valuesAnthropic10
CyberGym · Anthropic, in the data explorer Run by the developer A public benchmark of reproducing known software vulnerabilities that Anthropic runs as a cyber-capability measure.Reported by Anthropic · 10 valuesAnthropic10
Cybench · Anthropic, in the data explorer Run by the developer A public set of capture-the-flag security challenges that Anthropic runs on its models as a cyber-capability measure.Reported by Anthropic · 9 valuesAnthropic9
Long-form virology task 1, in the data explorer Run by the developer The first of two internal long-form biology-risk tasks Anthropic scores from 0 to 1 in its capability assessments.Reported by Anthropic · 9 valuesAnthropic9
Multimodal virology (VCT), in the data explorer Run by the developer A multimodal virology knowledge benchmark that Anthropic runs and scores from 0 to 1 in its biology-risk assessments.Reported by Anthropic · 9 valuesAnthropic9
Show the other 58 evaluations of this kind (163 values)
More capability benchmarks in Dangerous capabilities and risk determinations
EvaluationReported byValues
VCT · xAI, in the data explorer Run by the developer The Virology Capabilities Test, a public set of hard virology troubleshooting questions, run by xAI as a biology capability measure.Reported by xAI · 9 valuesxAI9
AI R&D kernel task, in the data explorer Run by the developer An Anthropic AI research task in which the model optimises compute kernels, reported as the best speedup reached.Reported by Anthropic · 8 valuesAnthropic8
Capture the Flag, in the data explorer Run by the developer Capture-the-flag hacking challenges at several difficulty levels, scored on the share the model solves.Reported by OpenAI · 8 valuesOpenAI8
Firefox 147 exploitation, in the data explorer Run by the developer An Anthropic cyber-capability evaluation against a browser target, reported as the share of trials scored fully successful.Reported by Anthropic · 6 valuesAnthropic6
50% time horizon · METR, in the data explorer Run by METR (third party) METR's estimate of the length of task, measured in the time skilled humans take, that a model completes with 50% success.Reported by OpenAI · 5 valuesOpenAI5
AI R&D novel compiler task, in the data explorer Run by the developer An Anthropic AI research task in which the model builds a compiler for a new language, reported as the share of tests passed.Reported by Anthropic · 5 valuesAnthropic5
Cybench · xAI, in the data explorer Run by the developer A public set of capture-the-flag security challenges, run by xAI without guidance as a cyber capability measure.Reported by xAI · 5 valuesxAI5
Cyber key skills benchmark, in the data explorer Run by the developer Cybersecurity challenges at easy, medium and hard levels testing skills an attacker would need, reported as the number solved.Reported by Google DeepMind · 5 valuesGoogle DeepMind5
DNA synthesis screening evasion, in the data explorer Run by the developer An Anthropic biosecurity evaluation reported as a count out of ten in its biology-risk assessments.Reported by Anthropic · 5 valuesAnthropic5
ExploitBench · Anthropic, in the data explorer Run by the developer A cyber-capability benchmark reported by Anthropic as the mean number of flags a model captures.Reported by Anthropic · 5 valuesAnthropic5
Long-form virology task 2, in the data explorer Run by the developer The second of two internal long-form biology-risk tasks Anthropic scores from 0 to 1 in its capability assessments.Reported by Anthropic · 5 valuesAnthropic5
SWE-bench Verified (hard subset), in the data explorer Run by the developer A hard subset of the public SWE-bench Verified software-engineering benchmark that Anthropic runs as an AI R&D capability measure.Reported by Anthropic · 5 valuesAnthropic5
ExploitBench · OpenAI, in the data explorer Run by the developer Tasks asking the model to write working exploits for known vulnerabilities in the V8 JavaScript engine, plus an internal port run on vulnerabilities disclosed after the model's training data.Reported by OpenAI · 4 valuesOpenAI4
RE-Bench, in the data explorer Run by the developer Machine-learning research-engineering tasks scored against human expert solutions, run by Google on its own models.Reported by Google DeepMind · 4 valuesGoogle DeepMind4
SRE-Bench, in the data explorer Run by the developer Binary reverse-engineering tasks, scored on success within one or four attempts.Reported by OpenAI · 4 valuesOpenAI4
TroubleshootingBench, in the data explorer Run by the developer Expert-written questions on fixing failed biology lab protocols, scored on correct answers.Reported by OpenAI · 4 valuesOpenAI4
WMDP, in the data explorer Run by the developer A public multiple-choice benchmark of hazardous knowledge in biosecurity, chemical security and cybersecurity, run by xAI.Reported by xAI · 4 valuesxAI4
BioLP-Bench, in the data explorer Run by the developer A public benchmark of biological laboratory protocols with introduced mistakes, measuring whether the model can find them.Reported by xAI · 3 valuesxAI3
CoBench, in the data explorer Run by the developer A benchmark Anthropic reports against an expert-substitution threshold in its capability assessment.Reported by Anthropic · 3 valuesAnthropic3
CVE-Bench · OpenAI, in the data explorer Run by the developer OpenAI's run of the public CVE-Bench set of tasks exploiting known software vulnerabilities, scored by pass rate.Reported by OpenAI · 3 valuesOpenAI3
CyberGym · xAI, in the data explorer Run by the developer A public benchmark of reproducing known software vulnerabilities, run by xAI as a cyber capability measure.Reported by xAI · 3 valuesxAI3
FrontierCyber · Irregular, in the data explorer Run by Irregular (third party) Irregular's FrontierCyber challenge set, reported as the share of challenges solved and by difficulty tier.Reported by OpenAI · 3 valuesOpenAI3
Hard-negative protein binding, in the data explorer Run by the developer Predicting which proteins bind when the candidates include close look-alikes that do not, scored as a success rate over several attempts.Reported by OpenAI · 3 valuesOpenAI3
PaperBench, in the data explorer Run by the developer Tests whether the model can reproduce the results of machine-learning research papers, scored against a replication rubric.Reported by OpenAI · 3 valuesOpenAI3
SWE-Lancer, in the data explorer Run by the developer Paid freelance software-engineering tasks, scored by the dollar value of tasks completed and by accuracy on the individual-contributor subset.Reported by OpenAI · 3 valuesOpenAI3
AI R&D quadruped RL task, in the data explorer Run by the developer An Anthropic AI research task in which the model trains a simulated four-legged robot policy, reported as the highest score reached.Reported by Anthropic · 2 valuesAnthropic2
CloningScenarios, in the data explorer Run by the developer Questions on multi-step molecular cloning workflows, used by xAI as a biology capability measure.Reported by xAI · 2 valuesxAI2
Cybench · Meta, in the data explorer Run by the developer A public set of capture-the-flag security challenges, run by Meta as a cyber capability measure and reported at one attempt.Reported by Meta · 2 valuesMeta2
Cyber autonomous offense suite, in the data explorer Run by the developer Capture-the-flag challenges at easy, medium and hard levels that the model attempts on its own, reported as the number of hard challenges solved.Reported by Google DeepMind · 2 valuesGoogle DeepMind2
DNA sequence design, in the data explorer Run by the developer Designing DNA sequences that meet given biological specifications, scored as a success rate.Reported by OpenAI · 2 valuesOpenAI2
Expert CTF · UK AI Security Institute, in the data explorer Run by UK AI Security Institute (government) A UK AI Security Institute set of expert-level capture-the-flag cyber challenges, scored by pass rate.Reported by OpenAI · 2 valuesOpenAI2
ExploitGym, in the data explorer Run by the developer End-to-end tasks in which the model must find a vulnerability and turn it into a working exploit.Reported by OpenAI · 2 valuesOpenAI2
GRB internal research-engineering benchmark, in the data explorer Run by the developer Google's internal benchmark of 74 research-engineering tasks, scored as average pass@1.Reported by Google DeepMind · 2 valuesGoogle DeepMind2
Internal AI R&D suite 2, in the data explorer Run by the developer Anthropic's second internal suite of AI research tasks, scored from 0 to 1 against a rule-out threshold.Reported by Anthropic · 2 valuesAnthropic2
MakeMeSay, in the data explorer Run by the developer A persuasion game in which the model tries to get another party to say a codeword without revealing it, reported as a win rate.Reported by xAI · 2 valuesxAI2
OpenAI PRs, in the data explorer Run by the developer Tasks based on changes OpenAI engineers actually made to internal code, testing whether the model can reproduce them; scored by pass rate.Reported by OpenAI · 2 valuesOpenAI2
OpenAI-Proof Q&A, in the data explorer Run by the developer Hard research and engineering questions drawn from OpenAI's internal work, used to gauge AI self-improvement capability and scored by pass rate.Reported by OpenAI · 2 valuesOpenAI2
SWE-bench Verified, in the data explorer Run by the developer OpenAI's run of SWE-bench Verified, real repository issues the model must fix, scored on the share resolved.Reported by OpenAI · 2 valuesOpenAI2
Tacit knowledge and troubleshooting, in the data explorer Run by the developer Questions on unwritten laboratory know-how and troubleshooting in biology, compared against an expert baseline.Reported by OpenAI · 2 valuesOpenAI2
VCT · Meta, in the data explorer Run by the developer The Virology Capabilities Test, a public set of hard virology troubleshooting questions, run by Meta as a biology capability measure.Reported by Meta · 2 valuesMeta2
AAV capsid packaging prediction, in the data explorer Run by the developer Predicting how well engineered viral capsid variants package, scored by rank correlation with measured results.Reported by OpenAI · 1 valueOpenAI1
Average success (benchmark not named) · Irregular, in the data explorer Run by Irregular (third party) An average success rate from Irregular cyber testing where the record does not name the benchmark.Reported by OpenAI · 1 valueOpenAI1
Bioinformatics evaluations, in the data explorer Run by the developer A set of bioinformatics tasks Anthropic reports alongside a human baseline in its biology-risk assessment.Reported by Anthropic · 1 valueAnthropic1
BioMysteryBench, in the data explorer Run by the developer A biology problem-solving benchmark Anthropic reports as accuracy, including on a subset humans found difficult.Reported by Anthropic · 1 valueAnthropic1
CVE-Bench · xAI, in the data explorer Run by the developer A public benchmark in which the model tries to exploit known vulnerabilities in web applications, reported by xAI as a reward score.Reported by xAI · 1 valuexAI1
Cyber range · UK AI Security Institute, in the data explorer Run by UK AI Security Institute (government) A UK AI Security Institute cyber range, scored by how many of a fixed number of attempts complete the whole range.Reported by Anthropic · 1 valueAnthropic1
Cyber range 'The Last Ones' · UK AI Security Institute, in the data explorer Run by UK AI Security Institute (government) A UK AI Security Institute cyber range called The Last Ones, scored by the share of attempts that complete it.Reported by OpenAI · 1 valueOpenAI1
Cyber tasks · UK AI Security Institute, in the data explorer Run by UK AI Security Institute (government) A set of UK AI Security Institute cyber tasks, scored by pass rate.Reported by OpenAI · 1 valueOpenAI1
CyScenarioBench · Anthropic, in the data explorer Run by the developer A scenario-based cyber-capability benchmark reported by Anthropic as a solve rate.Reported by Anthropic · 1 valueAnthropic1
CyScenarioBench · Irregular, in the data explorer Run by Irregular (third party) Irregular's CyScenarioBench cyber scenario benchmark, reported as scenarios solved and average success rate.Reported by OpenAI · 1 valueOpenAI1
Internal Research Debugging, in the data explorer Run by the developer Debugging problems taken from OpenAI's internal research work, reported as a median score.Reported by OpenAI · 1 valueOpenAI1
MLE-Bench, in the data explorer Run by the developer OpenAI's run of MLE-Bench, a set of machine-learning engineering competitions, scored on the share in which the model reaches a medal.Reported by OpenAI · 1 valueOpenAI1
Multimodal Troubleshooting Virology, in the data explorer Run by the developer Virology lab troubleshooting questions that include images, compared against the median expert score.Reported by OpenAI · 1 valueOpenAI1
OSS-Fuzz, in the data explorer Run by the developer An Anthropic cyber-capability evaluation on open-source fuzzing targets, reported as the number of targets reaching the top score.Reported by Anthropic · 1 valueAnthropic1
ProtocolQA Open-Ended · OpenAI, in the data explorer Run by the developer Open-ended questions asking the model to find and fix errors in published lab protocols, scored against a threshold.Reported by OpenAI · 1 valueOpenAI1
ProtocolQA open-ended · xAI, in the data explorer Run by the developer Questions about errors in laboratory protocols answered in free text rather than by multiple choice, used by xAI as a biology capability measure.Reported by xAI · 1 valuexAI1
Sandbox Bench, in the data explorer Run by the developer Attempts to break out of or exploit 22 targets built to resemble OpenAI's own infrastructure, counted by successful exploits.Reported by OpenAI · 1 valueOpenAI1
Vulnerability discovery, in the data explorer Run by the developer The number of previously unknown software vulnerabilities the model found, counted after human review.Reported by Moonshot AI · 1 valueMoonshot AI1
Framework determination · 13 evaluations, 125 values
Preparedness Framework determination, in the data explorer Run by the developer OpenAI's formal capability-level decisions under its Preparedness Framework for each tracked category: biological and chemical, cybersecurity, and AI self-improvement.Reported by OpenAI · 55 valuesOpenAI55
Frontier Safety Framework determination, in the data explorer Run by the developer Google DeepMind's decision on whether a model reaches the alert threshold or Critical Capability Level in each risk domain of its Frontier Safety Framework.Reported by Google DeepMind · 32 valuesGoogle DeepMind32
RSP deployment standard, in the data explorer Run by the developer The AI Safety Level standard Anthropic states a model is deployed under after its Responsible Scaling Policy assessment.Reported by Anthropic · 8 valuesAnthropic8
Autonomy threat model 2 (automated R&D), in the data explorer Run by the developer Anthropic's determination of whether its second autonomy threat model, automated AI research, applies to a model under the Responsible Scaling Policy.Reported by Anthropic · 4 valuesAnthropic4
CB-2 threat model, in the data explorer Run by the developer Anthropic's determination of whether the CB-2 threat model applies to a model under the Responsible Scaling Policy.Reported by Anthropic · 4 valuesAnthropic4
Risk determination, in the data explorer Run by the developer The risk tier Meta assigns a model in each risk domain under its Advanced AI Scaling Framework, before or after mitigations.Reported by Meta · 4 valuesMeta4
RSP determinations, in the data explorer Run by the developer A single statement in a card that gives Anthropic's Responsible Scaling Policy conclusions for several threat models at once.Reported by Anthropic · 4 valuesAnthropic4
AI R&D-4 threshold, in the data explorer Run by the developer Anthropic's statement of whether a model crosses its AI R&D-4 capability threshold under the Responsible Scaling Policy.Reported by Anthropic · 3 valuesAnthropic3
Show the other 5 evaluations of this kind (11 values)
More framework determinations in Dangerous capabilities and risk determinations
EvaluationReported byValues
CB-1 threat model, in the data explorer Run by the developer Anthropic's determination of whether the CB-1 threat model applies to a model under the Responsible Scaling Policy.Reported by Anthropic · 3 valuesAnthropic3
CBRN-4 threshold, in the data explorer Run by the developer Anthropic's statement of whether a model crosses its CBRN-4 capability threshold under the Responsible Scaling Policy.Reported by Anthropic · 3 valuesAnthropic3
Dual-use bio determination, in the data explorer Run by the developer xAI's statement of whether a model's dual-use biology and chemistry knowledge stays below its safety thresholds.Reported by xAI · 2 valuesxAI2
Overall risk conclusion, in the data explorer Run by the developer The overall risk conclusion that xAI gives at the end of a card, with its safeguards in place.Reported by xAI · 2 valuesxAI2
Cyber risk determination, in the data explorer Run by the developer Meta's conclusion from pre-release testing on whether a model could give plausible uplift toward catastrophic cyber outcomes.Reported by Meta · 1 valueMeta1
Uplift study · 3 evaluations, 6 values
Harmful manipulation efficacy study, in the data explorer Run by the developer Human-participant study of how much conversations with the model shift participants' beliefs compared with a non-AI baseline.Reported by Google DeepMind · 4 valuesGoogle DeepMind4
ASL-4 virology uplift trial, in the data explorer Run by the developer A human uplift trial in Anthropic's ASL-4 biology assessment, reported as the ratio of task scores with model help to an internet-only control.Reported by Anthropic · 1 valueAnthropic1
Virology protocol uplift trial, in the data explorer Run by the developer A human trial comparing participants with and without model help on a written biology task, reported as mean critical failures.Reported by Anthropic · 1 valueAnthropic1
Red teaming · 2 evaluations, 5 values
CBRN red-team uplift, in the data explorer Run by the developer Expert red teamers score how much the model adds to web-only resources in biological and chemical scenarios, graded on a rubric.Reported by Google DeepMind · 4 valuesGoogle DeepMind4
Bio bottleneck sub-stages red-teaming, in the data explorer Run by the developer Counts the bioweapon-pathway sub-stages on which red-team scores for the model stayed below a set threshold.Reported by Google DeepMind · 1 valueGoogle DeepMind1
Qualitative assessment · 3 evaluations, 5 values
Bio capability lift, in the data explorer Run by the developer xAI's statement in words of how much a model's biology capability rose over the previous model.Reported by xAI · 2 valuesxAI2
Release decision, in the data explorer Run by the developer How Anthropic chose to release a model (general availability, limited partners, safeguards attached), recorded as a categorical fact.Reported by Anthropic · 2 valuesAnthropic2
AI R&D assessment · METR, in the data explorer Run by METR (third party) METR's judgment of how much a model could speed up AI research and development, given as an acceleration multiplier.Reported by Anthropic · 1 valueAnthropic1
Scenario evaluation · 1 evaluation, 3 values
Manipulative cue rate, in the data explorer Run by the developer Share of conversational turns in constructed scenarios in which the model uses manipulative cues.Reported by Google DeepMind · 3 valuesGoogle DeepMind3
Survey · 2 evaluations, 3 values
Internal AI R&D survey, in the data explorer Run by the developer A survey of Anthropic research staff on whether a model could replace an entry-level researcher.Reported by Anthropic · 2 valuesAnthropic2
Internal model use survey, in the data explorer Run by the developer A survey of Anthropic staff on the productivity gain they get from using a model internally.Reported by Anthropic · 1 valueAnthropic1

Reported values

The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Dangerous capabilities and risk determinations

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

Dangerous capabilities and risk determinations: values by kind of test

Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it.

Hollow: read in a secondary write-up

Capability benchmarkScenario evaluation020406080100120140Percent

Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.

Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.

113 of the family’s 406 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 139 in other units or read the other way (minutes, USD, counts, human-normalised score and 9 others), 137 are categories or statements in words, 12 are restated in a later document or from an earlier version of one and 5 are low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.

The values span more than one and a half orders of magnitude, but one of them is zero, which a logarithmic axis cannot show, so the axis is linear and starts at zero.

Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0

Comparability

  • Each of the 90 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
  • Values come from the documents of six developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
  • Some names are used by more than one evaluation: “Cyber Range” by two, “CyberGym” by two, “Cybench” by three, “VCT” by two, “ExploitBench” by two, “CVE-Bench” by two, “CyScenarioBench” by two and “ProtocolQA Open-Ended” by two. They are different tests, run by different evaluators or reported by different developers, so this page names who ran each one.
  • In eight evaluations the test changed between documents (among them Frontier Safety Framework determination, Cyber challenges · Irregular and Cyber Range · OpenAI). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
  • In 38 evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
  • 32 evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
  • Numbers in this family are printed in 14 units: percent, minutes, USD, counts, human-normalised score, scores from 0 to 1 and 8 others. We never convert one unit into another, so a chart shows one unit at a time.
  • For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
  • 16 values are restatements: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.

Coverage

Documents from six of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.

Each developer's documents with values in Dangerous capabilities and risk determinations, out of all its documents in the dataset
DeveloperWith any valueWith a printed numberAll its documents
Anthropic151315
DeepSeek001
Google DeepMind13616
Meta435
Moonshot AI112
OpenAI191322
xAI888
Zhipu AI001

Related findings

No finding is about this family yet.