Trump signed what his supporters describe as a balanced, safety-centered executive order to expand voluntary safety testing of frontier AI models. The text envisions the government coordinating with industry through a voluntary framework rather than imposing broad mandates, seeking to accelerate safe deployment while not stifling innovation. The administration intends to create a simplistic pathway for AI firms to submit models for safety testing, while the NSA would run a classified benchmarking process to determine when a model qualifies as a 'covered frontier model' and to scan for vulnerabilities at scale in collaboration with the Treasury and CISA. The order also directs the Office of Management and Budget to assess funding opportunities for vulnerability detection and to expand the federal cybersecurity workforce. Critics, however, say the approach is performative and underpowered: it relies on voluntary cooperation, provides no firm requirements for firms, and relies on a government with reported staffing shortages and past contract slowdowns. The differences between the leaked draft and the final text include a shortened 30-day testing window compared with an earlier 90-day window, a move widely seen as an attempt to preserve U.S. AI leadership at the expense of safety oversight. Experts from CFR warn that defining what counts as a 'covered frontier model' and ensuring real observability will be the biggest challenges, and they caution that even with a formal process, meaningful testing is hard to implement quickly in a fast-evolving field. The piece underscores that the outcome will depend on how transparently the program operates and whether industry partners participate in good faith.
Introduction of a voluntary safety-testing framework with NSA-led benchmarking and a 30-day testing window; shift from earlier leaked drafts.
Unchanged: No mandatory regulatory requirements for AI firms; no binding pre-deployment constraints; reliance on voluntary collaboration.
A cautious, mixed tone reflecting safety ambitions but skepticism about enforceability and capacity.
The move centers on safety testing rather than altering core AI capabilities; impact depends on participation.
Voluntary framework means regulatory teeth are weak; potential future shifts to binding rules remain possible.
Aim to improve vulnerability scanning and cyber defenses through a government-led benchmarking effort.
Leads classified benchmarking to improve security for frontier AI models.
Collaborates to establish a cybersecurity clearinghouse for vulnerability scanning.
Involved in coordinating funding and cross-agency efforts.
Evaluates funding opportunities for testing and workforce expansion.
Provided expert critique on feasibility and observability.
Referenced Mythos as a risk example in policy discussions.
Context on threat landscape and potential risks of AI systems.
Potentially affected by voluntary testing; may gain safety cooperation but face delays.
The plan signals a shift toward a governance approach that relies on voluntary participation and targeted government benchmarking. Its effectiveness depends on industry adoption, the robustness of defined safeguards, and the government's ability to hire, fund, and coordinate across agencies. Observability and definitional clarity will determine whether this framework meaningfully reduces frontier AI risks or remains symbolic.
Participation in testing protocols may add integration overhead but penalties for non-compliance are not present.
Regulatory uncertainty and voluntary framework could affect risk assessment and funding.
May adopt voluntary testing; impact on product timelines is uncertain.
Aims to strengthen national security while facing staffing and funding constraints.
Policy focus is domestic but may influence international discussions.
Observability and patching depend on rapid and accurate assessment.
No data governance changes described.
Public scrutiny exists but not a systemic reputational crisis.
Staffing, funding, and definitional challenges threaten timely execution.
No critical infrastructure dependencies described beyond testing workflows.
Domestic policy with limited cross-border friction evident.
Voluntary framework may evolve toward binding regulation; uncertainty remains.
Not a supply chain issue.
Government staffing shortages could slow implementation.
No liability framework detailed in the text.