Monday, July 27, 2026

News

US Expands Government AI Safety Testing to Google, xAI and Microsoft Models

PolicyPatryk Raba
US Expands Government AI Safety Testing to Google, xAI and Microsoft Models
Fot. APK, Wikimedia Commons (CC BY-SA 4.0)

The US Center for AI Standards and Innovation is adding Google DeepMind, xAI and Microsoft to a program in which government scientists test unreleased AI models for cyberattack and bioweapon risks. OpenAI and Anthropic have long participated, and their tests have already uncovered real security flaws.

Contents
  1. What the expansion covers
  2. What the tests involve
  3. ChatGPT Agent vulnerability
  4. Regulatory context

The Trump administration has expanded its voluntary AI model safety testing program to three more companies: Google DeepMind, xAI and Microsoft. Government scientists at the Center for AI Standards and Innovation (CAISI) will get access to as-yet-unreleased models to search for vulnerabilities before they reach the market.

CAISI is a team of government scientists set up to assess the risks posed by advanced AI models. The institution operates within the Department of Commerce, and its work relies on voluntary agreements with technology companies, which make their models available for independent review before public release.

What the expansion covers

Google DeepMind has agreed to give CAISI access to its proprietary models and data. Microsoft said it would build joint datasets and evaluation processes for advanced AI models, though it did not specify which systems exactly would be involved. xAI did not respond to a Reuters request for comment on the matter.

OpenAI and Anthropic had already been working with CAISI on a voluntary basis. OpenAI is currently testing GPT-5.5-Cyber with the government science team, a variant of its latest model designed specifically for defensive cybersecurity tasks.

GPT-5.5-Cyber is a variant of our latest model built for defensive cybersecurity work - Chris Lehane, head of global public affairs at OpenAI

What the tests involve

Government scientists are focusing on what they describe as "measurable risks." Chief among these is the possibility that advanced models could be used to attack US critical infrastructure, as well as reducing the chances that US adversaries could use AI to develop chemical or biological weapons, or poison the data used to train American models.

Tests CAISI has already run with Anthropic and OpenAI show these concerns are not purely theoretical. Work with Anthropic revealed that simple tricks, such as falsely claiming that content had passed human review, or swapping individual characters in a query, could bypass the model's safeguards.

Tricks such as claiming human verification had occurred, or swapping characters, could bypass safety mechanisms - Anthropic statement

ChatGPT Agent vulnerability

Even more serious findings concerned OpenAI's ChatGPT Agent. Tests with CAISI showed that sophisticated attackers could exploit a flaw in the tool to remotely seize control of computer systems the agent had access to during a given session, and impersonate the user on other sites where they were logged in.

An attacker could remotely take control of computer systems the agent had access to during a session, and impersonate the user on other sites where they were logged in - OpenAI statement

Both companies stress that the vulnerabilities discovered have already been patched. Still, the fact that such flaws surfaced at all in products from large, well-funded AI labs shows why the administration wants to extend similar testing to more companies, including those only now joining the program.

Regulatory context

The CAISI program did not emerge in a vacuum. Back in 2023, companies including Meta, Amazon and Inflection AI agreed to let independent experts assess their models for biosecurity and cybersecurity risks. Government scientists have since published voluntary guidelines on AI privacy violations and the factual accuracy of model responses, and are now working on new standards for critical infrastructure providers.

For Polish companies and institutions using OpenAI, Google or Microsoft models, the program's expansion matters indirectly but tangibly. If US testing forces vulnerabilities in commercial models to be patched before launch, users of those systems outside the US also benefit, including public institutions and businesses in Poland that are increasingly deploying tools built on these very providers' models.

The program remains voluntary, meaning none of the companies is legally required to participate or to implement CAISI's recommendations. That sets the US approach apart from the EU's AI Act, which imposes hard obligations on providers of high-risk systems. Whether a voluntary testing model proves effective in the long run will depend on how many real vulnerabilities get found and fixed before the models reach widespread use.

Share: