Anthropic launched Claude Fable 5 with policies that secretly degrade model performance for certain requests and retain all prompts and outputs for 30 days. This strengthens the case for open-weight models, and Microsoft immediately banned its employees from using the model.
Anthropic released Claude Fable 5 on June 9, 2026, with safety measures that secretly degrade model performance for certain requests (e.g., frontier AI development, training pipelines, chip design) without informing users. It also retains all prompts and outputs for 30 days, overriding existing zero-data-retention agreements with enterprise customers.
Anthropic disclosed routing for cybersecurity, biology/chemistry, and distillation attempts by falling back to a weaker model and informing users. However, for frontier AI development requests, it uses invisible safeguards such as prompt modification, steering vectors, or parameter-efficient fine-tuning to limit effectiveness. Microsoft banned its employees from using the model within one day.
This incident highlights the risks of closed AI models: users may not know their product is being deliberately degraded, and data retention policies threaten confidentiality. It strengthens the argument for open-weight models as the only way to avoid such control.
HN users note the discussion mirrors the past RSA/crypto wars, feeling history is repeating. Some express embarrassment at only now realizing this connection. Overall, the sentiment is that 'trust us' is insufficient, emphasizing the need for open weights.