Structural shift

US AI testing may prioritize closed models

The White House has reportedly shared its cyber-testing framework with major AI labs without publishing it. Available details indicate that it may focus on closed frontier models with national-security risks and use secure government-access environments.

The consequential change in US federal AI safety policy may not be a new model-approval regime, but the construction of a confidential cyber-testing access mechanism for the strongest closed models. If implemented along the disclosed lines, governance would shift from public principles toward controlled testing, access logging, and risk communication around deployment.

The framework is moving from principle to access

A June White House executive order directed the creation of classified cyber-capability evaluations and a threshold for covered frontier models, while allowing limited government access before release to other trusted partners. Axios reported that an August 4 meeting provided labs with framework details and that covered models may be closed systems with state-of-the-art capabilities and national-security risks. WIRED separately reported that the government had shared the plan with OpenAI, Anthropic, and other labs, without releasing the full text publicly.

Confidential testing changes the operating mechanism

The design is not a conventional pre-approval regime: the executive order explicitly does not authorize mandatory licensing, pre-clearance, or approval. Its practical function is closer to a controlled testing channel, in which the government receives model access in a secure environment, access is logged, and labs develop more standardized interfaces for cyber-risk assessment and serious-incident reporting. OpenAI had previously confirmed that the federal government was developing standards, timelines, and processes for cyber testing of the most capable models.

Closed and open models may face policy separation

If open models are excluded, federal cyber governance would initially center on closed frontier systems that can be centrally accessed, tested, and logged. For major labs, secure evaluation environments, auditability, and access governance could become clearer operating requirements. Open-weight ecosystems may follow a different policy path. The strongest countercase is that the framework remains unpublished, and its reported coverage definition and exclusions could change in the final text.

What to watch next

Key observable evidence will be a published White House framework defining covered-model capability and risk thresholds, formal language on whether open models are included, and disclosures from participating labs on test frequency, access windows, logging, or independent audits. The thesis would weaken materially if the final rules include open models equally or remain limited to non-operational voluntary principles.

Sources