Muse Spark Safety & Preparedness Report
A safety report by Meta about Muse Spark, published 8 Apr 2026. We recorded 22 values from it.
- Developer
- Meta
- Type
- Safety report
- Model covered
- Muse Spark
- Published
- 8 Apr 2026
- Archived copy
- No archived copy yet
- Changelog
- Partial changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
Three versions are on record, oldest first. We hold a copy of one; the others are known only by their date.
8 Apr 2026
First published version
26 May 2026
Revision
What changed instruction-hierarchy comparator scores fixed after a bug
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. Without copies of the earlier versions, changes between them are not recorded value by value; what we know of each version is listed below.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 5 of the 22 have been blind-verified: a second reader found the same value without seeing ours.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 Evaluation awareness
3 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark | Verbalized eval awarenessShare of samples verbalizing evaluation | internal evals | 4.1.11 | Verified | |
| Muse Spark | Verbalized eval awarenessShare of samples verbalizing evaluation | public benchmarks | 4.1.11 | Verified | |
| Muse Spark | Prompted evaluation awarenessEval-vs-deployment discrimination | None stated | Table 1 | Verified |
M2 Reward hacking
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark | ImpossibleBenchCheating rate | without system mitigations | Table 2 | Unverified |
M3 Sabotage and sandbagging
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark | GDM-StealthStealth capability score | without system mitigations | Table 1 | Unverified | |
| Muse Spark | SHADE-ArenaSabotage/stealth success score | without system mitigations | Table 1 | Verified |
M4 Misalignment audits
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark | Agentic MisalignmentHarmful action rate | without system mitigations | Table 2 | Unverified | |
| Muse Spark | Alignment fakingBehavioral difference monitored vs not | without system mitigations | Table 2 | Unverified |
M5 Honesty and hallucination
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark | DeceptionBenchDeception rate | without system mitigations | Table 2 | Unverified | |
| Muse Spark | MASKHonesty rate | without system mitigations | Table 2 | Unverified |
M6 Sycophancy
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark | Internal sycophancySycophancy rate | without system mitigations | Table 2 | Unverified |
M7 Harmful compliance and over-refusal
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark | AgentHarmCompliance rate on harmful agentic tasks | without system mitigations | Table 2 | Unverified | |
| Muse Spark (Meta AI system) | BioTIER refusalsRefusal rate on harmful bio | with system mitigations | Table 2 | Unverified |
M8 Jailbreak robustness
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark | FORTRESSAttack success rate | without system mitigations | Table 2 | Unverified | |
| Muse Spark | StrongREJECTAttack success rate (adaptive multi-turn) | without system mitigations | Table 2 | Verified |
M9 Prompt injection
3 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark | Agent Red Teaming (ART)Agent red-team ASR pass@1Run by Gray Swan | without system mitigations | Table 2 | Unverified | |
| Muse Spark | AgentDojoPrompt-injection ASR pass@1 | without system mitigations | Table 2 | Unverified | |
| Muse Spark | IHEvalIH score | without system mitigations | Table 2; 4.1.1 | Unverified |
M10 Dangerous capabilities and risk determinations
4 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark | Risk determinationRisk tier, chemical and biological | pre-mitigation | 1.1.1 | Unverified | |
| Muse Spark | Risk determinationRisk tier, cyber and loss of control | post-mitigation | 1.1.1 | Unverified | |
| Muse Spark | CybenchPass@1 | without system mitigations | Table 1 | Unverified | |
| Muse Spark | VCTAccuracy | without system mitigations | Table 1 | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note First report under Meta Advanced AI Scaling Framework v2
All 22 values were read in the document itself.
All 22 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).