News
Claude Sonnet 5 Tops F5 Labs AI Security Ranking
F5 Labs' latest CASI ranking places Claude Sonnet 5 first among AI models resistant to jailbreak and prompt injection attacks, with a score of 93.08 points and the highest task performance among the security leaders.
Contents
F5 Labs updated its public CASI language model security ranking on July 6, 2026, and Claude Sonnet 5 from Anthropic came out on top. The model scored 93.08 out of 100 points, ahead of every other flagship AI system currently on the market.
CASI, the Comprehensive AI Security Index, is a metric developed by the F5 Labs research team, part of F5, the American network and application security company. The index measures how easily a given language model can be manipulated into generating harmful content or content that violates its own safety policy, by subjecting it to a series of jailbreak and prompt injection attacks devised by the company's red team.
How the score is calculated
The ranking is complemented by a second metric, ARS, the Agentic Resistance Score, which tests model resilience in multi-step scenarios where the AI acts as an autonomous agent carrying out complex tasks without constant human oversight. F5 Labs stresses that it's precisely in these scenarios that simple, one-shot safeguards most often fail, since attackers can split a malicious instruction into many seemingly harmless steps.
Beyond the security score itself, the team also publishes a CoS metric, the cost of safety, showing how much performance a model gives up in exchange for resistance to attacks. The ranking is updated monthly and includes comparisons going back to February 2025, making it possible to track how successive generations of models handle the same set of tests.
Security without losing capability
Claude Sonnet 5's score stands out against F5 Labs' earlier findings, which showed that models more resistant to manipulation usually lost usefulness. Anthropic's model achieved an average task performance of 53.4 percent, the highest among all systems that crossed the 90-point CASI threshold, which researchers say marks a break from the previous tradeoff between security and capability.
In write-ups describing the Adversarial Tales attack method used to build the ranking, F5 Labs researchers Lee Ennis and Malcolm Heath described a mechanism for bypassing model safeguards by embedding harmful requests inside a literary narrative.
The model isn't tricked into treating the request as harmless, it's redirected into a mode where the harmful content is treated as material for literary interpretation - Lee Ennis and Malcolm Heath, F5 Labs
Rivals trail behind
Behind Anthropic's models on the podium came OpenAI's GPT-5.5, with a score of 78.47 points, more than ten points below the weakest Claude model in the ranking. Among the latest flagship models from other vendors included in this year's F5 Labs comparisons, scores range from over 90 points for Anthropic's leaders down to just a dozen or so points for the models that fared worst on resistance tests, showing how wide the gap in safety approaches currently is among leading AI vendors.
For companies deploying large language models in products with access to the internet, tools, or customer databases, the CASI score carries practical weight: it shows which vendor offers lower risk that a maliciously crafted prompt will extract data from the model or push it into an unwanted action. That matters especially given the growing number of AI agent deployments that independently browse websites, operate terminals, or send requests to external systems.
What it means for companies in Poland
Polish companies are increasingly choosing AI model vendors based not only on price or answer quality but also on security-related risk, especially in regulated sectors like banking, insurance, or public administration. Independent rankings such as CASI give security teams an argument in conversations with business units that want to deploy an AI agent faster than a risk assessment allows.
F5 Labs says the ranking will keep being updated monthly as new models appear, and the company plans to expand testing with further multi-step attack scenarios measured by the ARS metric. The next update will show whether Claude Sonnet 5 holds its lead as competitors, including OpenAI and xAI, update their safeguards in response to Anthropic's result.
