Muse Spark 1.1 Evaluation Report
A safety report by Meta about Muse Spark 1.1, published 9 Jul 2026. We recorded 10 values from it.
- Developer
- Meta
- Type
- Safety report
- Model covered
- Muse Spark 1.1
- Published
- 9 Jul 2026
- Archived copy
- No archived copy yet
- Changelog
- Not applicable
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. We know of only one version of it.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 1 of the 10 has been blind-verified: a second reader found the same value without seeing ours.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 Evaluation awareness
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark 1.1 | Verbalized eval awarenessShare of samples verbalizing evaluation | Agentic Misalignment scenarios | 4.1.6 | Verified | |
| Muse Spark 1.1 | Verbalized eval awarenessShare of samples verbalizing evaluation | dangerous-capability evals | 4.1.6 | Unverified |
M3 Sabotage and sandbagging
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark 1.1 | SHADE-ArenaOverall success under CoT monitor | None stated | 2.3.2 | Unverified |
M4 Misalignment audits
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark 1.1 | Agentic MisalignmentHarmful action rate | without system mitigations | Table 2 | Unverified |
M5 Honesty and hallucination
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark 1.1 | MASKHonesty rate | without system mitigations | Table 2 | Unverified |
M6 Sycophancy
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark 1.1 | Internal sycophancySycophancy rate | without system mitigations | Table 2 | Unverified |
M8 Jailbreak robustness
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark 1.1 | StrongREJECTAttack success rate | without system mitigations | 3.1.1.1 | Unverified |
M9 Prompt injection
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark 1.1 | AgentDojoPrompt-injection ASR pass@1 | without system mitigations | Agent robustness | Unverified |
M10 Dangerous capabilities and risk determinations
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Muse Spark 1.1 | CybenchPass@1 | without system mitigations | Cyber capabilities | Unverified | |
| Muse Spark 1.1 | Risk determinationRisk tier, cyber | pre-mitigation | Summary | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note Restates several Muse Spark 1.0 values differently
All 10 values were read in the document itself.
All 10 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).