The Trump administration demands Anthropic block all jailbreaks before rereleasing Fable 5, but security experts say this is impossible. Anthropic argues the concerns are overblown and reiterated its position in a technical meeting.
The Trump administration is demanding that Anthropic block all jailbreaks before allowing the rerelease of its AI model Fable 5. Anthropic took the model offline last week due to export controls over jailbreak concerns. The NSA concluded there are ways to disable guardrails, but Anthropic argues the concerns are overblown and reiterated its position in a technical meeting with the Commerce Department and the Office of the National Cyber Director.
The administration views the issue as Anthropic's problem to fix, as agencies lack the staff to chase every jailbreak. They want Anthropic to proactively test and report vulnerabilities. However, security experts say guardrails are a stopgap and complete prevention of jailbreaks is impossible, as skilled users or future AI will find bypasses.
This highlights the tension between AI regulation and corporate responsibility. If the administration's demand is technically infeasible, it could escalate conflicts between innovation and safety oversight. The situation also fuels debate on government intervention in AI safety and the limits of self-regulation.
The commenter argues that guaranteeing no jailbreaks is unrealistic, but this demand could serve as a forcing function to curb Anthropic's AI doomer safetyism. They note potential benefits for open models and against regulatory capture, while acknowledging possible dishonest motives from the Trump administration.