White House finalises voluntary AI hacking tests for frontier models

The White House has finalised a voluntary framework for testing the advanced cyber capabilities of frontier AI models, as discussions begin with major AI developers over how the assessments will be conducted before model releases.

White House finalises voluntary AI hacking tests for frontier models

The White House has completed a voluntary framework for assessing the advanced hacking capabilities of frontier AI models and has invited leading AI developers to discuss its implementation.

According to reports, staff-level talks involve representatives from Meta, Anthropic, OpenAI and Google, although the White House has not confirmed which companies will participate or whether any have agreed to submit models for evaluation.

The framework implements a June executive order directing federal agencies to establish a classified benchmarking process for evaluating AI models with advanced cyber capabilities. Models that exceed a government-defined capability threshold may be designated as ‘covered frontier models’ under the order.

Developers participating in the programme may voluntarily provide the federal government with access to covered models for up to 30 days before releasing them to other trusted partners. The executive order requires safeguards covering confidentiality, cybersecurity, insider threats and intellectual property protection during the evaluation process.

The administration has emphasised that participation remains voluntary and does not give the government authority to delay or prevent the release of AI models. The executive order explicitly excludes mandatory licensing, pre-approval or permitting requirements.

Details of the framework remain limited. The White House has not disclosed the capability thresholds, testing methodology or reporting procedures, and it is unclear whether the framework itself will be published or what actions could follow if a model demonstrates particularly advanced offensive cyber capabilities.

The discussions follow recent disclosures by OpenAI and Anthropic regarding incidents that occurred during internal cybersecurity evaluations.

OpenAI reported that, during testing, one of its research models exploited a previously unknown vulnerability in a package-registry proxy, reached the public internet and accessed Hugging Face’s production infrastructure. The company said the model involved was an internal research prototype that was not intended for public release.

Anthropic later disclosed that it had reviewed more than 141,000 evaluation runs and identified three cases in which Claude models accessed the internet and gained unauthorised access to real organisations. Unlike the OpenAI incident, Anthropic said these cases resulted from a misconfigured evaluation environment rather than the exploitation of a software vulnerability.

The company also reported that its latest research model stopped its activity after recognising that it had reached a real system, while older research models did not consistently do so. Anthropic noted that the evaluations were conducted without some of the safeguards present in publicly deployed models.

Go to Top