gpt-5-main
21 values about gpt-5-main from 1 document, in 4 metric families. In each family, values from its own document come first.
- Developer
- OpenAI
- Release date
- 7 Aug 2025The date of its first card in the dataset, the GPT-5 System Card.
- Availability
- Public
- Its own document
- GPT-5 System Card · 7 Aug 2025
- Also reported in
- No other document reports a value about it
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Values by metric family
Each value as the document printed it. Select a value for its source, its checks and its history.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M5 · Honesty and hallucination 3 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| Production-traffic factualityClaim-level hallucination rate reduction vs GPT-4oCondition: with browsing; production prompts | GPT-5 System Card7 Aug 2025 | |
| Production-traffic factualityReduction in responses with 1+ major error vs GPT-4oCondition: with browsing; production prompts | GPT-5 System Card7 Aug 2025 | |
| SimpleQAHallucination rateCondition: no browsing | GPT-5 System Card7 Aug 2025 | |
M6 · Sycophancy 3 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| Sycophancy offline evalSycophancy scoreCondition: offline | GPT-5 System Card7 Aug 2025 | |
| Sycophancy online prevalence (A/B)Relative change vs GPT-4oCondition: free users | GPT-5 System Card7 Aug 2025 | |
| Sycophancy online prevalence (A/B)Relative change vs GPT-4oCondition: paid users | GPT-5 System Card7 Aug 2025 | |
M7 · Harmful compliance and over-refusal 11 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| Production BenchmarksExtremism | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksHarassment/threatening | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksHate/threatening | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksIllicit/non-violent | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksIllicit/violent | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksNon-violent hate | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksPersonal data | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksSelf-harm/instructions | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksSelf-harm/intent | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksSexual/exploitative | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksSexual/minors | GPT-5 System Card7 Aug 2025 | |
M8 · Jailbreak robustness 4 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| StrongRejectAbuse/disinformation/hate | GPT-5 System Card7 Aug 2025 | |
| StrongRejectIllicit/non-violent crime | GPT-5 System Card7 Aug 2025 | |
| StrongRejectSexual content | GPT-5 System Card7 Aug 2025 | |
| StrongRejectViolence | GPT-5 System Card7 Aug 2025 | |