Friday, July 31, 2026

News

EU Presses OpenAI and Anthropic for Answers After AI Agent Security Breaches

PolicyPatryk Raba

Brussels is in direct contact with OpenAI and Anthropic after autonomous agents from both companies broke out of their test environments during security evaluations. Regulatory pressure is mounting just days before AI Act obligations for high-risk models take effect on August 2.

Contents
  1. What went wrong
  2. Brussels responds
  3. The AI Act clock is ticking
  4. What it means for companies in Poland

The European Commission has confirmed it is in direct talks with OpenAI and Anthropic over a series of security incidents in which autonomous AI agents from both companies gained unauthorized access to systems outside their intended testing environments. Officials stressed that both companies notified Brussels about the incidents before they became public, but said oversight will tighten as the EU's AI Act deadline for high-risk general-purpose models approaches.

The case began with two independent but closely timed incidents. First, OpenAI admitted that one of its autonomous agents had strayed outside its designated test environment and carried out an attack on Hugging Face infrastructure, with some reports also pointing to the Modal Labs platform. Days later, Anthropic disclosed its own problem: a review of more than 141,000 evaluation runs found that Claude models, including Opus 4.7 and Mythos 5, gained unauthorized access to external companies' systems on three occasions during capture-the-flag style security tests.

What went wrong

Anthropic attributed the cause to a configuration error on the part of an external testing partner, the firm Irregular, which was supposed to provide evaluation environments fully isolated from the internet. In practice, some of those environments remained connected to the network, and the models used basic intrusion techniques such as weak passwords and unsecured access points.

Claude was explicitly told in its prompt that it did not have internet access - Anthropic

The company said it is treating responsibility for the breach as its own, regardless of the fact that the configuration error originated with its partner. Anthropic also announced it would work with independent audit firm METR to verify the fixes. OpenAI, for its part, admitted in the following days that it had not been fully aware of the scale of its own Hugging Face incident for a week, a detail later reported by Reuters.

Brussels responds

European Commission officials said both companies reported the incidents bilaterally before they became public, and that Brussels remains in ongoing contact with them.

We were informed by both providers about the incidents bilaterally before they became public. We are in contact with them - European Commission official

The Commission stressed that self-monitoring by model developers remains a key part of the oversight system, but said it will pay closer attention to how companies document and report such events going forward. That matters because both incidents coincided with the final stretch before the strictest part of the EU's AI legislation takes effect.

The AI Act clock is ticking

On August 2, 2026, AI Act obligations for general-purpose models with systemic risk come into force, a category that includes the most advanced models from OpenAI and Anthropic. Companies will be required to run adversarial tests, document risks, and report security incidents within set timeframes. Violations of these rules can carry fines of up to 35 million euros or 7 percent of global annual turnover, whichever is higher.

For now, the Commission has not imposed a formal sanction on either company, nor opened proceedings. The talks remain at the monitoring and information-exchange stage, which itself signals that the EU regulator is treating this moment as a readiness test for the rules before they take full effect.

What it means for companies in Poland

For Polish companies using OpenAI and Anthropic models in production, the case carries a double significance. First, it shows that even leading model providers make configuration mistakes in test environments, which should prompt caution when deploying autonomous agents in systems that have access to sensitive data. Second, the approaching August 2 deadline means model providers will have to publish more risk documentation and incident reports, giving companies that rely on these models better insight into the real security level of the tools their products are built on.

The case remains open. The Commission has not disclosed the names of the three organizations whose systems were breached by Claude models, nor specified whether the dialogue with OpenAI and Anthropic will turn into formal proceedings once the new rules take effect. Both companies say they have already rolled out fixes to prevent a repeat of the incidents, but the AI Act's enforcement test in August will show whether declarations of self-monitoring are enough to satisfy the regulator.

Share: