Key facts
- Anthropic released its Fable AI model, a public version of its cybersecurity model Mythos.
- Cybersecurity experts criticize Fable's restrictive guardrails, stating they hinder defenders and could aid attackers.
- The model's keyword-based safety measures trigger on a wide range of cybersecurity-related prompts, including code reviews and secure coding requests.
- Fable defaults to Claude Opus 4.8 when its guardrails are activated.
- Anthropic has a Cyber Verification Program for professionals seeking fewer restrictions on using its AI for cybersecurity tasks.
Anthropic has released Fable, a public and limited version of its cybersecurity-focused AI model, Mythos. However, the model's extensive guardrails have drawn criticism from cybersecurity experts who argue they are overly broad and inconsistently applied. These restrictions, intended to prevent malicious use for developing malware or compromising software, are reportedly triggered by keywords related to cybersecurity and biology, even for seemingly innocuous tasks like code review or writing secure code. This has led to frustration among researchers who feel the model hinders legitimate cybersecurity work and could potentially assist attackers by limiting defensive capabilities. Fable defaults to Claude Opus 4.8 when its safety measures are activated. Anthropic, like OpenAI with its Trusted Access for Cyber program, requires cybersecurity professionals to apply for specific verification to use its AI models with fewer limitations for cybersecurity-related tasks.
