gpt-oss-120b & gpt-oss-20b Model Card
A model card by OpenAI about gpt-oss-120b and gpt-oss-20b, published 5 Aug 2025. We recorded 74 values from it.
- Developer
- OpenAI
- Type
- Model card
- Models covered
- gpt-oss-120b, gpt-oss-20b
- Published
- 5 Aug 2025
- Also at
- Another address, deploymentsafety.openai.com (opens an external site)
arXiv 2508.10925 - Archived copy
- No archived copy yet
- Changelog
- Not known
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. We know of only one version of it.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. None has been blind-verified yet.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M5 Honesty and hallucination
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| gpt-oss-120b | PersonQAHallucination rate | no browsing | Table 9 | Unverified | |
| gpt-oss-120b | SimpleQAHallucination rate | no browsing | Table 9 | Unverified | |
| gpt-oss-20b | PersonQAHallucination rate | no browsing | Table 9 | Unverified | |
| gpt-oss-20b | SimpleQAHallucination rate | no browsing | Table 9 | Unverified | |
| o4-mini | PersonQAHallucination rate | no browsing | Table 9 | Unverified | |
| o4-mini | SimpleQAHallucination rate | no browsing | Table 9 | Unverified |
M7 Harmful compliance and over-refusal
44 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-4o | Production BenchmarksExtremism | None stated | Table 5 | Unverified | |
| GPT-4o | Production BenchmarksHarassment/threatening | None stated | Table 5 | Unverified | |
| GPT-4o | Production BenchmarksHate/threatening | None stated | Table 5 | Unverified | |
| GPT-4o | Production BenchmarksIllicit/non-violent | None stated | Table 5 | Unverified | |
| GPT-4o | Production BenchmarksIllicit/violent | None stated | Table 5 | Unverified | |
| GPT-4o | Production BenchmarksNon-violent hate | None stated | Table 5 | Unverified | |
| GPT-4o | Production BenchmarksPersonal data | None stated | Table 5 | Unverified | |
| GPT-4o | Production BenchmarksSelf-harm/instructions | None stated | Table 5 | Unverified | |
| GPT-4o | Production BenchmarksSelf-harm/intent | None stated | Table 5 | Unverified | |
| GPT-4o | Production BenchmarksSexual/exploitative | None stated | Table 5 | Unverified | |
| GPT-4o | Production BenchmarksSexual/minors | None stated | Table 5 | Unverified | |
| gpt-oss-120b | Production BenchmarksExtremism | None stated | Table 5 | Unverified | |
| gpt-oss-120b | Production BenchmarksHarassment/threatening | None stated | Table 5 | Unverified | |
| gpt-oss-120b | Production BenchmarksHate/threatening | None stated | Table 5 | Unverified | |
| gpt-oss-120b | Production BenchmarksIllicit/non-violent | None stated | Table 5 | Unverified | |
| gpt-oss-120b | Production BenchmarksIllicit/violent | None stated | Table 5 | Unverified | |
| gpt-oss-120b | Production BenchmarksNon-violent hate | None stated | Table 5 | Unverified | |
| gpt-oss-120b | Production BenchmarksPersonal data | None stated | Table 5 | Unverified | |
| gpt-oss-120b | Production BenchmarksSelf-harm/instructions | None stated | Table 5 | Unverified | |
| gpt-oss-120b | Production BenchmarksSelf-harm/intent | None stated | Table 5 | Unverified | |
| gpt-oss-120b | Production BenchmarksSexual/exploitative | None stated | Table 5 | Unverified | |
| gpt-oss-120b | Production BenchmarksSexual/minors | None stated | Table 5 | Unverified | |
| gpt-oss-20b | Production BenchmarksExtremism | None stated | Table 5 | Unverified | |
| gpt-oss-20b | Production BenchmarksHarassment/threatening | None stated | Table 5 | Unverified | |
| gpt-oss-20b | Production BenchmarksHate/threatening | None stated | Table 5 | Unverified | |
| gpt-oss-20b | Production BenchmarksIllicit/non-violent | None stated | Table 5 | Unverified | |
| gpt-oss-20b | Production BenchmarksIllicit/violent | None stated | Table 5 | Unverified | |
| gpt-oss-20b | Production BenchmarksNon-violent hate | None stated | Table 5 | Unverified | |
| gpt-oss-20b | Production BenchmarksPersonal data | None stated | Table 5 | Unverified | |
| gpt-oss-20b | Production BenchmarksSelf-harm/instructions | None stated | Table 5 | Unverified | |
| gpt-oss-20b | Production BenchmarksSelf-harm/intent | None stated | Table 5 | Unverified | |
| gpt-oss-20b | Production BenchmarksSexual/exploitative | None stated | Table 5 | Unverified | |
| gpt-oss-20b | Production BenchmarksSexual/minors | None stated | Table 5 | Unverified | |
| o4-mini | Production BenchmarksExtremism | None stated | Table 5 | Unverified | |
| o4-mini | Production BenchmarksHarassment/threatening | None stated | Table 5 | Unverified | |
| o4-mini | Production BenchmarksHate/threatening | None stated | Table 5 | Unverified | |
| o4-mini | Production BenchmarksIllicit/non-violent | None stated | Table 5 | Unverified | |
| o4-mini | Production BenchmarksIllicit/violent | None stated | Table 5 | Unverified | |
| o4-mini | Production BenchmarksNon-violent hate | None stated | Table 5 | Unverified | |
| o4-mini | Production BenchmarksPersonal data | None stated | Table 5 | Unverified | |
| o4-mini | Production BenchmarksSelf-harm/instructions | None stated | Table 5 | Unverified | |
| o4-mini | Production BenchmarksSelf-harm/intent | None stated | Table 5 | Unverified | |
| o4-mini | Production BenchmarksSexual/exploitative | None stated | Table 5 | Unverified | |
| o4-mini | Production BenchmarksSexual/minors | None stated | Table 5 | Unverified |
M8 Jailbreak robustness
12 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| gpt-oss-120b | StrongRejectAbuse/disinformation/hate | None stated | Table 6 | Unverified | |
| gpt-oss-120b | StrongRejectIllicit/non-violent crime | None stated | Table 6 | Unverified | |
| gpt-oss-120b | StrongRejectSexual content | None stated | Table 6 | Unverified | |
| gpt-oss-120b | StrongRejectViolence | None stated | Table 6 | Unverified | |
| gpt-oss-20b | StrongRejectAbuse/disinformation/hate | None stated | Table 6 | Unverified | |
| gpt-oss-20b | StrongRejectIllicit/non-violent crime | None stated | Table 6 | Unverified | |
| gpt-oss-20b | StrongRejectSexual content | None stated | Table 6 | Unverified | |
| gpt-oss-20b | StrongRejectViolence | None stated | Table 6 | Unverified | |
| o4-mini | StrongRejectAbuse/disinformation/hate | None stated | Table 6 | Unverified | |
| o4-mini | StrongRejectIllicit/non-violent crime | None stated | Table 6 | Unverified | |
| o4-mini | StrongRejectSexual content | None stated | Table 6 | Unverified | |
| o4-mini | StrongRejectViolence | None stated | Table 6 | Unverified |
M9 Prompt injection
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| gpt-oss-120b | Instruction hierarchyPrompt injection hijacking | system<>user conflict | Table 7 | Unverified | |
| gpt-oss-120b | Instruction hierarchySystem prompt extraction | system<>user conflict | Table 7 | Unverified | |
| gpt-oss-20b | Instruction hierarchyPrompt injection hijacking | system<>user conflict | Table 7 | Unverified | |
| gpt-oss-20b | Instruction hierarchySystem prompt extraction | system<>user conflict | Table 7 | Unverified | |
| o4-mini | Instruction hierarchyPrompt injection hijacking | system<>user conflict | Table 7 | Unverified | |
| o4-mini | Instruction hierarchySystem prompt extraction | system<>user conflict | Table 7 | Unverified |
M10 Dangerous capabilities and risk determinations
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| gpt-oss-120b | Preparedness Framework determinationAI self-improvement | None stated | Preparedness | Unverified | |
| gpt-oss-120b | Preparedness Framework determinationBiological and chemical | None stated | Preparedness | Unverified | |
| gpt-oss-120b | Preparedness Framework determinationBiological and chemical | adversarially fine-tuned | Preparedness | Unverified | |
| gpt-oss-120b | Preparedness Framework determinationCybersecurity | None stated | Preparedness | Unverified | |
| gpt-oss-120b | Preparedness Framework determinationCybersecurity | adversarially fine-tuned | Preparedness | Unverified | |
| gpt-oss-120b | SWE-bench VerifiedPass rate | high; N=477 | Preparedness: AI Self-improvement | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note Open-weight; includes worst-case malicious fine-tuning assessment
All 74 values were read in the document itself.
All 74 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).