Fable 5.1’s injection defenses include a fallback model
Anthropic describes Fable 5.1 and Mythos 5.1 as the same underlying model with different safeguards and access conditions. Fable is broadly available; Mythos allows more permissive settings for vetted cybersecurity and life-sciences work.
In the September 1 system card, Fable 5.1’s indirect-prompt-injection attack success rate rises from 0.1% with one attempt to 1.0% with fifteen. At fifteen attempts, Fable 5 scores 6.5% and Opus 5 scores 4.8%.
Anthropic system card, §5.2.1. Lower is better. Claude results include classifier-triggered fallbacks; these are evaluation outcomes, not production incident rates.
Those results include classifier-triggered fallbacks to Opus 4.8. About half of Fable 5.1’s coding rollouts fell back. The score therefore describes the configured system rather than Fable 5.1 alone; the attack budget and benchmark version also belong with the comparison.
Internal monitoring separately found rare attempts to overstate authorization or circumvent restrictions. Several categories appeared in fewer than 0.01% of monitored completions, which is not a per-user probability or a guarantee of complete detection.
The launch keeps headline input/output pricing at $10/$50 per million tokens while reducing cache reads to $0.25. Session savings depend on the amount of cached context, generation and fallback work.