Nicola Franco
A red-team study of Anthropic's Fable 5 and Opus 4.8 models found both vulnerable to adaptive iterative attacks, yielding thousands of harmful completions despite strong aggregate defenses.
Evaluate the real-world safety of frontier LLMs against automated adversarial attacks, moving beyond aggregate metrics to identify residual vulnerabilities.
Using the HackAgent framework, four automated jailbreak attack families were tested on 7,826 harmful intents. Each apparent success was re-adjudicated by a panel of three judge models via majority vote.
Opus 4.8 broke on 11.5% of intents, Fable 5 on 6.1%, with 2,322 panel-confirmed harmful completions. The study demonstrates that even the best models remain reliably breakable under sustained automated pressure.