The US government has finalised a voluntary testing framework designed to measure whether highly capable AI models can discover vulnerabilities, escape controlled environments or gain unauthorised access to external systems. Meta and Anthropic have been invited to discuss the framework with White House officials, while OpenAI and Google have also reportedly been involved in the process.
The move follows recent disclosures that experimental AI systems developed by OpenAI and Anthropic accessed other organisations’ systems during cybersecurity testing. Details such as the benchmarks, reporting arrangements and whether results will be public have not yet been released, but the initiative marks a practical step towards testing frontier models before their capabilities are deployed more widely.
Key facts
- The framework will assess the cybersecurity capabilities of advanced US-developed AI models.
- Participation is voluntary.
- Testing is expected to examine whether models can identify and exploit security weaknesses.
- The government has not yet disclosed the metrics, reporting process or transparency requirements.
- Major AI developers are being consulted on how the framework should operate.
Our take
New Zealand organisations rely heavily on AI platforms developed overseas, so testing standards adopted by major providers will influence the assurances available to local customers. Organisations should not assume that a model is safe simply because it is commercially available.
Procurement and AI governance processes should ask providers about independent testing, incident reporting, containment controls and how customers will be informed if a model behaves unexpectedly. For professional services firms handling confidential or privileged information, these questions should form part of vendor assessment before advanced agents are given access to internal systems or client data.

