Regulation

U.S. to Pre-Screen Advanced AI Models Before Public Release

The U.S.

U.S. to Pre-Screen Advanced AI Models Before Public Release

The National Institute of Standards and Technology (NIST), part of the U.S. Department of Commerce, announced a new multi-agency task force that will assess national-security risks from advanced AI models prior to their deployment. The move marks a sharp reversal from the White House’s earlier more hands-off approach.

What the program will examine

NIST said the tests will focus on demonstrable risks, with particular attention to cybersecurity, biosecurity and chemical-weapons concerns. The administration has not released detailed terms of its agreements with AI firms or spelled out what restrictions it might impose on models based on test results.

The evaluations will be carried out by Testing Risks of AI for National Security (TRAINS), overseen by the NIST Center for AI Standards and Innovation (CAISI). TRAINS is designed for rapid response and draws on multiple federal agencies, including the Departments of Commerce, Defense, Energy and Homeland Security, as well as the National Security Agency and the National Institutes of Health.

TRAINS has not published the specific benchmarks it will use. NIST has, however, shared CAISI’s earlier comparison that ranked DeepSeek V4 Pro against other large language models using an aggregate of nine widely used public benchmarks covering cybersecurity, coding, mathematics, natural sciences and abstract reasoning, plus an internal test called PortBench (porting command-line interface tools between programming languages).

Corporate participation and agreements

Leading U.S. AI companies — including Google, Microsoft and xAI — have agreed to provide models that feature limited or no guardrails. Anthropic and OpenAI reached similar terms in 2024. Those agreements are intended to enable joint public–private research on model capabilities, risk evaluation and mitigation strategies.

Participation to date has been voluntary. The White House is considering an executive order that would require AI models to gain approval before deployment, which would make pre-release testing mandatory.

Background and political context

The shift represents a departure from policies under the previous Trump Administration, which emphasized rolling back Biden-era regulatory barriers. The change follows a series of events involving Anthropic: in March 2026 the company tried to restrict military uses of its Claude model for surveillance and autonomous weapons, but the White House rejected those limits and banned the model from military use. A month later Anthropic said its Claude Mythos Preview could autonomously exploit vulnerabilities in major operating systems and applications; the company had shared Mythos with 50 organizations for vulnerability detection and patching.

Last week the White House opposed Anthropic’s plan to expand the Mythos preview to another 70 organizations, citing national-security concerns and questioning whether Anthropic has sufficient compute capacity to serve both existing Mythos users and the government. Anthropic has not said whether it will challenge the administration’s authority to limit distribution of the preview model.

Historically, the Biden Administration issued a 2023 executive order requiring developers to notify the government when they train models with processing requirements roughly on the order of one trillion parameters. After taking office in January 2025, President Trump assigned three advisors to craft an AI Action Plan aimed at sustaining and enhancing America’s global AI leadership by suspending or eliminating Biden-era regulatory policies.

Why this matters

Moving from a laissez-faire stance to pre-release scrutiny recognizes that AI models can now pose immediate national-security risks. Mandatory or voluntary pre-release testing could give the government earlier warning of potential issues and encourage developers to manage risks proactively. It would also give authorities a mechanism to determine which models are fit for broader distribution and which should be withheld or modified — decisions that might not always be fully transparent.

At present, submission of new models for government testing remains voluntary; officials are considering an executive order that would make such testing compulsory.

Editorial note

A standardized, consistently applied battery of benchmarks would likely benefit the industry, but whether such standards should be established primarily by market forces or by government mandate is debatable. Requiring government clearance before release could slow U.S. developers, risk placing them at a competitive disadvantage internationally, and create avenues for regulatory capture that might disadvantage open-source competitors.