Verge Trump AI Testing Plan Excludes Open Models and Hides the Rules
- Ethan Carter

- Aug 6
- 14 min read
President Donald Trump’s administration finalized an AI testing framework, but excluded open models despite their growing role in cybersecurity and global competition. The Verge Trump coverage points to a basic contradiction. Washington wants early warning about dangerous AI capabilities while leaving a major distribution model outside its review process.
The framework is voluntary and applies before covered systems reach the public. Participating developers can give federal officials access to advanced models for up to 30 days. Yet the government has not publicly released the framework, its eligibility threshold, or the full list of agencies and partners involved.
That makes the plan difficult to evaluate from outside the selected group of companies briefed by the White House. It also divides the market by release model. Closed systems from companies such as OpenAI, Anthropic, and Google face potential government review, while downloadable models remain outside the framework.
The immediate question is not whether every AI model needs federal approval. The executive order rejects mandatory licensing and pre-clearance. The harder question is whether a secret, voluntary system can produce consistent security decisions when its scope excludes an increasingly important part of the market.
What the Trump Framework Actually Changes
The framework creates a federal review channel for advanced closed models, but leaves its most important definitions hidden.
The White House said it completed the framework by the deadline established in Trump’s June 2 executive order. According to reporting on the closed-door framework, officials then held staff-level discussions with AI companies about its contents and next steps.
The framework covers the handling of advanced, unreleased models during a government evaluation period. It reportedly addresses confidentiality, intellectual property, insider threats, cybersecurity controls, and restrictions on who can access a submitted system.
Models would be held in high-security environments. Detailed access logs would record who used them, while employees at participating companies would face limits during the review window. Multiple administration officials would participate instead of one agency controlling the process.
The executive order provides for as many as 30 days of government access before a covered model reaches other trusted partners. That window begins near release, not during the earliest stages of training. Companies were reportedly encouraged to submit systems that closely resemble the versions they plan to ship.
The difference matters. An early prototype might lack the tools, permissions, or deployment configuration that creates practical security risk. A near-final system offers a better test target, although 30 days still leaves limited time for complex evaluation and remediation.
The government’s benchmark for determining whether a model qualifies is classified. The threshold is expected to focus on advanced cyber capabilities and national security risk. Developers and researchers can receive relevant information when officials consider disclosure appropriate.
However, the voluntary framework itself is not necessarily classified. The White House has still chosen not to publish it. Officials also have not explained when regular submissions will begin or how disagreements over coverage will be resolved.
Axios reported that the framework defines a covered frontier model as a closed system with state-of-the-art capabilities and national security risks. Neither “state-of-the-art” nor “national security risk” reportedly receives a clear public definition.
A frontier model is a system near the highest current capability level. That description changes whenever laboratories release stronger models. Without measurable public criteria, companies cannot independently determine whether a planned release belongs inside the program.
The framework therefore changes government access without establishing a transparent public rulebook. The White House gains a structured path to inspect selected systems. Developers outside its discussions gain little guidance about when that path applies to them.
This is the central lesson from the Verge Trump story. The administration has moved beyond general concern and created an operational review process. It has not provided enough information for outsiders to judge whether that process is consistent, technically adequate, or fairly administered.
Why Open Models Sit Outside the Verge Trump Plan
The open-model exemption narrows the framework exactly where control becomes hardest after release.
Open models allow users to download model components, usually including their weights. Model weights are the numerical parameters learned during training that shape a system’s outputs. Access lets researchers inspect, modify, fine-tune, and run the system on infrastructure they control.
The term “open model” can cover several licensing arrangements. Some releases provide weights but not training data or complete source code. Others permit broad commercial use, while more restrictive licenses limit particular deployments.
According to the reported framework details, open models are excluded from the federal testing process. The document also reportedly says it should not be interpreted as restricting those systems after their release.
That choice has a coherent policy rationale. Once developers publish weights, a government review cannot reliably control every later copy. Users can host the system abroad, alter safeguards, connect new tools, and distribute modified versions without returning to the original laboratory.
Pre-release access also creates a different burden for open-model developers. A delayed release gives closed providers time to serve selected customers through controlled interfaces. An open developer generally makes a more binary decision because releasing weights transfers lasting control to users.
Supporters of open models argue that broad access strengthens American research and cybersecurity. Independent teams can reproduce evaluations, examine failures, and adapt a model without depending on a vendor’s application programming interface.
Open access also lowers dependence on a small group of frontier laboratories. Developers can run models locally, protect sensitive data, and study security behavior in environments that a hosted service cannot reproduce.
The exemption avoids converting the voluntary framework into a de facto approval gate for downloadable software. That is consistent with the executive order’s rejection of licensing, mandatory pre-clearance, and government permitting for new models.
Yet the exemption creates a large analytical gap. A model’s license does not determine whether it can discover vulnerabilities, automate social engineering, write malicious code, or coordinate actions through external tools. Those properties depend on capabilities, deployment, access, and safeguards.
A weaker open model can also become more useful through fine-tuning or tool integration. Several specialized systems working together might create risks that a benchmark focused on one frontier model does not capture.
The framework appears to treat controllability as a boundary for government review. Closed developers can restrict access during evaluation and apply centralized updates afterward. Open releases cannot provide the same continuing control once their weights circulate.
That is a practical distinction, but it is not a complete risk distinction. Excluding a difficult category does not make its security questions disappear. It only places those questions outside this particular process.
This is why the Verge Trump policy debate cannot be reduced to regulation versus innovation. The stronger conflict is between administratively manageable testing and comprehensive risk coverage. The plan chooses the manageable target.
The result pressures closed-model laboratories more than their open-model competitors. OpenAI, Anthropic, and Google reportedly provided feedback on a draft. Less advanced developers and companies not invited to White House meetings remain uncertain about whether they will participate later.
That asymmetry can shape release strategy. A company expecting a federal review might delay features, limit early customers, or spend more on secure evaluation infrastructure. A developer releasing downloadable weights would not face the same framework, even when users can adapt its model for cyber tasks.
The Plan Protects Security Through Secrecy
Some confidentiality is defensible, but hidden thresholds prevent the public from measuring whether the program works.
The strongest argument for secrecy concerns the benchmark itself. Publishing detailed tests for dangerous cyber capabilities could help developers train directly against them. It might also reveal sensitive defensive methods or identify targets that government agencies consider vulnerable.
Keeping test data private can reduce benchmark contamination. Contamination occurs when a model encounters evaluation material during training, making its score less meaningful. A classified test environment can preserve surprise and protect operational details.
Secure handling also protects company intellectual property. Frontier models represent major investments, and unreleased weights can expose both commercial secrets and national security concerns. Detailed logs and restricted access are sensible controls for government evaluation.
NIST already describes its broader AI risk framework as voluntary. That public framework emerged through drafts, workshops, comments, and collaboration with outside organizations. Its process shows that voluntary risk management does not require every rule to remain hidden.
The Trump framework takes a different approach. The benchmark is classified, the eligibility threshold is restricted, and the operating framework remains unpublished. Even the identity of all trusted partners appears unresolved.
Trusted partners are organizations that can receive controlled early access to a model. They might include cybersecurity companies, infrastructure operators, researchers, or government-approved customers. Their selection can determine who benefits commercially from an advanced model before general release.
That creates governance concerns beyond technical testing. If officials help decide which customers qualify, the process can affect competition. Companies need a predictable method for challenging exclusions and understanding what security obligations partners must meet.
OpenAI previously said government access should not become the long-term default. During an earlier restricted rollout, the company described the review as a temporary step toward broader availability. That response showed reluctant cooperation, not an endorsement of permanent federal release management.
The Associated Press reported that OpenAI made GPT-5.6 Sol available to roughly 20 approved customers during a prior review. Anthropic also limited access to Mythos 5 after government scrutiny. Those episodes provided a preview of how federal intervention can influence commercial availability.
Representative Lori Trahan criticized officials deciding company by company who receives access without a law, defined process, or outside oversight. Her concern goes to institutional legitimacy rather than the need for cyber testing itself.
A confidential benchmark can still operate under published procedures. The government could disclose who makes decisions, what evidence companies receive, how long reviews take, and how developers appeal. It could publish aggregate findings without exposing sensitive test prompts.
None of those safeguards requires revealing classified vulnerabilities. They require separating technical secrecy from procedural secrecy.
The current framework appears to combine both. That makes it difficult to tell whether two similar models would receive similar treatment. It also prevents independent experts from evaluating whether the government has enough staff, infrastructure, and time.
A 30-day window sounds concrete, but its effectiveness depends on test depth. Evaluators must configure the system, understand its tools, probe misuse scenarios, reproduce findings, and communicate remediation options. Advanced cyber testing can require repeated interactions and expert judgment.
The Verge Trump reporting also leaves unclear what happens when a company rejects the government’s assessment. Participation is voluntary, yet federal procurement, customer approvals, or political pressure can create strong incentives to cooperate.
Voluntary systems can work when participants trust the process and expect consistent treatment. Ambiguity weakens both conditions. Companies may comply to protect government relationships while remaining unsure about the rules applied to competitors.
This uncertainty also affects enterprise customers. Security teams need to distinguish a government-reviewed model from a government-approved model. The framework appears to provide evaluation, not a universal certification of safety.
A model that performs well in a controlled test can behave differently after deployment. Users may connect it to private repositories, cloud consoles, messaging systems, or autonomous agents. Those permissions can transform a general assistant into an operational security risk.
Government review should therefore be one signal, not a substitute for organizational controls. Buyers still need access limits, audit logs, incident procedures, and continuous monitoring for models deployed in sensitive environments.
Closed Labs Face an Uneven Competitive Test
The framework places its clearest obligations on companies that already maintain the most centralized control over their models.
OpenAI, Anthropic, and Google operate major closed or hosted systems. Their centralized infrastructure lets them limit users, monitor activity, change safeguards, and disable features. It also makes them identifiable targets for government intervention.
Open-model developers distribute control more widely. That architecture complicates remediation, but the framework does not place them under the same pre-release process. The policy burden therefore falls where enforcement is easiest, not necessarily where aggregate risk is highest.
The competitive effect will depend on which models cross the classified threshold. If only a few highly capable systems qualify, the review could remain narrow. If officials interpret “state-of-the-art” broadly, more developers may need secure submission processes and release contingencies.
Smaller laboratories face another problem. Companies involved in drafting discussions gain early knowledge about federal expectations. Developers excluded from those meetings must plan around press reports and private guidance shared only when officials consider it appropriate.
That information gap can favor incumbents. Large firms already employ government affairs teams, security specialists, and lawyers familiar with classified programs. A smaller developer might struggle to determine whether it needs the same infrastructure.
The administration says it is working with more partners than the three best-known laboratories. That claim remains difficult to assess because the framework and participant list are not public.
The plan also creates tension between American closed models and foreign open alternatives. If domestic providers experience delays while downloadable systems remain immediately available, developers may move workloads to models outside the review regime.
That substitution would weaken the intended security benefit. It could also move sensitive deployments toward systems receiving less direct scrutiny from American providers or agencies.
The opposite outcome is also possible. A government-reviewed closed model might gain credibility with critical infrastructure operators. Banks, hospitals, utilities, and defense contractors may prefer systems that passed a structured federal evaluation.
The executive order also directs the Treasury Department to establish an AI cybersecurity clearinghouse. This body would coordinate vulnerability scanning, validation, and patch distribution with developers and critical infrastructure operators.
That broader mechanism can provide practical value. Discovering a vulnerability is only the first step. Defenders must verify it, notify affected organizations, prepare patches, and distribute fixes without giving attackers an unnecessary advantage.
AI models can accelerate both sides of that process. A system might help a security team inspect code and prioritize patches. The same capability can help an attacker search widely for unprotected systems.
The government became more interested in pre-release review after models showed stronger cyber abilities. Earlier reporting described scrutiny of Anthropic’s Mythos systems and OpenAI’s GPT-5.6 Sol. The companies characterized their products as useful for defensive work while acknowledging misuse and unforeseen risks.
This dual-use character makes simple labels unreliable. A model that finds software flaws for defenders can produce similar information for attackers. Safeguards, access controls, and deployment context often determine which side benefits first.
The 30-day review order tries to give government evaluators time before wide distribution. However, it cannot guarantee that vulnerabilities discovered during testing will be fixed within that period.
It also cannot prevent comparable capabilities from appearing elsewhere. Closed laboratories compete with each other, open developers, foreign companies, and specialized cyber models. A framework centered on selected American releases covers only part of that landscape.
The policy might still improve security if it catches serious failures before deployment. That is a reasonable, testable objective. The problem is that outsiders lack the information needed to determine whether reviews find issues, change releases, or merely delay them.
Companies will also need to preserve records from these interactions. Policy, security, and engineering teams must track evolving definitions, government requests, test findings, and release decisions. A searchable technical knowledge base can help teams connect those records without treating a single briefing as permanent policy.
Documentation will not resolve the framework’s ambiguity. It can prevent internal confusion when different teams receive partial guidance and deadlines change.
What the Verge Trump Report Still Cannot Answer
The framework’s biggest weakness is not a proven technical failure, but the absence of evidence needed to assess its scope and accountability.
First, the public does not know the covered-model threshold. Officials can reasonably protect the exact benchmark, but companies still need understandable capability categories. Without them, coverage can appear discretionary.
Second, the framework reportedly excludes open models without explaining how agencies will monitor the risks they create after release. Other government programs might address those systems, but no connected process has been publicly described.
Third, officials have not disclosed the complete governance structure. Axios reported that several administration officials would participate. It remains unclear who makes a final decision when technical evaluators disagree.
Fourth, the meaning of a successful review is uncertain. The executive order does not establish a licensing system. Therefore, the government might identify risk without possessing a formal mechanism to block release.
Informal leverage can still be substantial. Federal agencies buy AI services, regulate critical industries, manage classified information, and influence approved partner access. A voluntary request from the government does not always feel optional to a company.
Fifth, the plan lacks public reporting requirements. Aggregate disclosures could reveal how many models were reviewed, how many findings required mitigation, and whether releases changed. Such reporting would demonstrate value without exposing classified tests.
The absence of those details does not prove that officials are acting unfairly. It means fairness cannot be independently evaluated. That distinction matters in any cautious assessment of the policy.
The same caution applies to claims that excluding open models necessarily makes the framework useless. The most advanced cyber capabilities may still reside in closed systems. Testing those systems can reduce risk even if the program is incomplete.
However, capability leadership changes. Open models can improve, and users can modify them after publication. A framework built around today’s market structure can age quickly if it lacks a process for reassessing exemptions.
The government must also avoid confusing model evaluation with deployment security. Even a carefully tested system can cause damage when connected to excessive permissions. Continuous monitoring becomes essential after tools, data, users, and operating conditions change.
NIST research has emphasized the difficulty of treating AI security as a one-time certification problem. Models and environments evolve, while attackers adapt to known defenses. Pre-release review provides a snapshot, not a lasting guarantee.
The Verge Trump framework therefore needs a feedback loop. Incident reports should update tests. Deployment failures should inform threshold decisions. Researchers should learn enough from aggregate results to improve independent evaluation methods.
Secrecy makes that learning harder. If only officials and selected laboratories see findings, smaller companies may repeat known mistakes. Critical infrastructure operators may also misunderstand what the government reviewed.
The administration could preserve classified details while publishing a plain-language scope document. It could identify decision owners, review stages, expected developer evidence, appeal routes, and reporting commitments.
That would not settle the open-model debate. It would clarify what the closed-model program promises and what remains outside its reach.
Three Signals That Will Show Whether the Plan Works
The next test is execution, and three observable signals will reveal whether this framework becomes a security process or an opaque release gate.
The first signal is publication of procedural guidance. The White House does not need to release classified benchmarks. It should explain eligibility categories, agency responsibilities, trusted-partner rules, and what happens after evaluators identify a serious risk.
Clear guidance would strengthen the framework’s credibility. Continued silence would reinforce concerns that developers receive different rules through private briefings.
The second signal is the next covered model release. Watch whether a laboratory announces a review, changes access, or delays particular capabilities. The most useful evidence will be a concrete mitigation tied to the government process.
A restriction alone will not establish success. Officials and companies should explain whether the review found a new risk, verified an existing concern, or imposed a precaution without new technical evidence.
The third signal is treatment of advanced open models. The administration might maintain the exemption, create a separate post-release testing program, or coordinate evaluations through NIST and independent researchers.
A separate program would acknowledge the practical differences between hosted and downloadable systems. It could emphasize reproducible testing, incident sharing, and deployment guidance instead of attempting pre-release control.
No response would leave the largest conceptual gap intact. The government would be testing systems it can access centrally while relying on the wider research community to assess models distributed beyond centralized control.
Readers should also distinguish the framework from older voluntary commitments. Previous federal AI initiatives often published more information about their objectives and participants. This program is tied more directly to pre-release cyber capability and controlled access.
That focus reflects a real policy challenge. Advanced models can help defenders find vulnerabilities faster, but they can also reduce the expertise required for malicious activity. Government agencies have legitimate reasons to test those capabilities before they spread.
The question is whether the chosen process matches the threat. A classified benchmark can protect sensitive methods. A secure 30-day review can provide useful warning. Neither feature explains why open models receive a categorical exemption or why operating rules must remain private.
The Verge Trump story ultimately reveals a framework defined as much by its boundaries as by its protections. It covers selected closed systems, depends on voluntary cooperation, and gives officials significant discretion within an unpublished process.
Developers should watch for concrete submission rules rather than assuming every advanced model requires review. Enterprise buyers should ask what a government evaluation covered before treating it as a broad safety endorsement.
Security teams should keep testing models inside their actual deployment environments. They should examine tool permissions, network access, data exposure, logging, and human approval points. A federal review cannot replace those controls.
The administration now has an opportunity to show that the framework produces measurable security improvements. Publishing procedural details would be a useful start. Reporting aggregate outcomes would be stronger evidence.
Until then, the policy remains limited and vague for reasons that extend beyond its open-model exemption. It offers a serious mechanism for handling selected frontier systems, but not a complete strategy for AI cybersecurity.
The next major model release will supply the first meaningful test. Did government access identify a specific problem, produce a documented mitigation, and improve deployment decisions? Or did it simply decide who received the model first?
That answer will determine whether the Verge Trump AI testing plan becomes durable security infrastructure or an opaque checkpoint in the race to release more capable models.


