Gottheimer AI Safety Legislation Makes Federal Testing the Price of Releasing Frontier Models
Josh Gottheimer introduced two bipartisan AI bills with a clear conflict at their center: frontier developers would face mandatory federal review before releasing covered models. The Gottheimer AI safety legislation would replace a voluntary process with a national security checkpoint controlled by the government.
The September 18 announcement pairs that review proposal with a ban targeting Chinese-developed open-weight models across federal systems. Open-weight models publish reusable model parameters, allowing organizations to operate or modify them without relying entirely on the original developer.
The package is not simply another call for responsible AI. It challenges the current assumption that laboratories, outside auditors, and voluntary government partnerships can manage the most serious model risks. Its main test is whether centralized federal review can improve security without becoming an unreliable release gate.
Gottheimer AI Safety Legislation Creates a Federal Release Checkpoint
The American AI Security Act would turn pre-release cooperation into a mandatory national security review for the most capable AI models.
Gottheimer, a New Jersey Democrat, announced the legislation with Republican Representative Mike Lawler of New York. According to Gottheimer’s official September 18 announcement, the proposal would require developers of covered frontier models to provide the National Security Agency with access before release.
A frontier model is a highly capable general-purpose system near the leading edge of AI development. The government would evaluate whether a covered model could support serious cyberattacks or chemical, biological, or radiological weapons development.
The official announcement gives the initial review a 30-day deadline. Officials could use one additional 30-day extension when necessary.
That clock matters because an unlimited review could function like an indefinite release hold. A short review, however, risks becoming superficial unless evaluators receive enough access, computing capacity, and technical support.
The proposal allows companies to support the evaluation process. It also includes an expedited appeal modeled on legal procedures used for national security matters, according to Gottheimer’s announcement.
Precise implementation details remain important. The announcement does not fully explain which capability threshold would make a model subject to review. It also leaves questions about model updates, fine-tuned variants, and systems assembled from several smaller models.
Those definitions will determine the law’s reach. A narrow threshold might cover only a few expensive training runs from major laboratories. A broader standard might capture smaller developers whose systems gain dangerous capabilities through tools or post-training modifications.
The bill therefore creates more than a testing requirement. It would give the federal government a formal role in deciding when certain privately developed models can move toward wider distribution.
That represents a sharp departure from the administration’s voluntary framework. It also places classified national security knowledge inside the evaluation process, where outside auditors often have limited visibility.
Gottheimer argues that independent auditors lack sufficient personnel, computing resources, laboratory access, and security clearances. He also questions their independence when developers help finance the testing ecosystem.
The lawmaker’s concern is understandable. A private evaluator can test known scenarios, but it cannot independently reproduce every classified threat assessment held by intelligence agencies.
Government review has its own limitations. Federal evaluators can become dependent on the same developers for model access, infrastructure, documentation, and specialized expertise.
The American AI Security Act tries to manage that dependency by permitting technical support from developers. Yet that arrangement does not eliminate information asymmetry between the model builder and the reviewer.
The central change is still unmistakable. Releasing a qualifying frontier model would no longer depend only on a developer’s internal decision or voluntary participation in federal testing.
The Bill Rejects Washington’s Voluntary Model Review
The primary dispute is mandatory government review versus voluntary cooperation, not safety versus unrestricted development.
A June 2026 executive order directed federal agencies to create a classified benchmarking process for advanced cyber capabilities. That process is intended to identify models that warrant special treatment.
The order also created a framework through which developers can provide covered models to the government before release. Access can last up to 30 days before trusted partners receive the system.
However, Executive Order 14409 expressly rejects mandatory licensing, preclearance, or permitting for developing and releasing new AI models. Participation in its early-access framework remains voluntary.
Gottheimer’s bill would reverse that choice. Developers of covered models would have to participate, and the review would examine both cyber capabilities and weapons-related risks.
That difference creates the article’s main tension. Voluntary cooperation reduces regulatory friction, while mandatory review gives government evaluators greater assurance that every qualifying model enters the same process.
The White House approach assumes developers have incentives to protect their products, customers, and reputations. It also treats cooperation as a way to preserve rapid American development.
Gottheimer’s approach assumes voluntary participation leaves a dangerous gap. A company facing competitive pressure can delay disclosure, narrow the scope of testing, or release before the government finishes evaluating a risk.
Both positions claim to support American leadership. Their disagreement concerns who should control the final pre-release decision when model capabilities intersect with national security.
Lawler framed the proposal as a precaution for the most capable systems. In Gottheimer’s official release, Lawler argued that the government should determine before deployment whether such models could enable devastating cyberattacks or facilitate chemical, biological, or radiological weapons development.
That standard sounds straightforward, but dangerous capability is not a binary property. Performance depends on prompts, tools, scaffolding, user expertise, safeguards, and access to outside systems.
A model might perform poorly in a controlled biological evaluation while becoming more useful when connected to scientific databases. Another model might generate exploit code without reliably executing a complete attack.
Testing also ages quickly. Developers can update system prompts, inference tools, safety layers, and external integrations after the federal review ends.
The administration’s voluntary model emphasizes collaboration and speed. The American AI Security Act emphasizes comprehensive coverage and formal accountability.
Neither approach removes uncertainty. A voluntary system can miss developers or encourage selective participation, while a mandatory system can create false confidence around an incomplete benchmark.
The political dispute also arrives during a broader struggle over federal AI rules. The White House has favored fewer restrictions on model development and stronger limits on state-level regulation.
Gottheimer has criticized that framework for insufficient accountability. An Associated Press analysis noted that several states already regulate parts of the private AI market.
That state activity complicates federal negotiations. Lawmakers must decide whether a national framework should supplement state protections, replace them, or leave certain consumer rules untouched.
The American AI Security Act takes a narrower route. It focuses on the most capable models and national security threats rather than creating a comprehensive consumer AI code.
That focus can help bipartisan negotiations. Cyberattacks and unconventional weapons create clearer federal interests than disputes over automated hiring, advertising, or entertainment.
Yet even a narrow bill would place pressure on OpenAI, Anthropic, Google, Meta, and other advanced model developers. Their release schedules could become subject to government capacity and review procedures.
Cloud providers would also feel the effects. They supply the computing infrastructure needed to evaluate, host, and distribute advanced systems.
Enterprise buyers should watch closely. A federal review might become a procurement signal, even if passage does not create a broad safety certification for private customers.
The proposed review is therefore both a security mechanism and a market intervention. It changes who bears the cost of uncertainty before a frontier model reaches users.
The China FIREWALL Act Expands the Package Beyond DeepSeek
The second bill treats model origin and supply-chain exposure as federal security risks, even when software vendors hide the underlying model.
Gottheimer introduced the China FIREWALL Act with Republican Representative Nick LaLota of New York. The proposal would prohibit Chinese-developed open-weight models on government-issued devices.
It would also prevent agencies from buying software that depends on those models. That provision targets indirect exposure through contractors and commercial applications, not only direct model downloads.
According to Gottheimer’s office, the proposed restriction builds on a DeepSeek-focused federal prohibition included in the fiscal year 2026 defense law. The new bill would extend the approach across Chinese-developed open-weight systems.
That expansion addresses a real procurement challenge. Agencies can block a named application while still acquiring other products that route requests through an embedded or externally hosted model.
Software supply chains frequently contain multiple services, libraries, and subcontractors. An agency buyer might see an application’s brand without seeing every model involved in its operation.
The China FIREWALL Act would push vendors to disclose model dependencies more clearly. Contractors would need to establish where a model originated and whether downstream services rely on prohibited components.
That requirement could affect more than Chinese vendors. American software companies incorporating models such as Alibaba’s Qwen would need alternative systems for federal contracts.
The proposal also raises difficult classification questions. “Chinese-developed” sounds definite until a model includes international researchers, foreign training data, redistributed weights, or modifications by an American company.
Open weights make that problem harder. A developer can download a model, fine-tune it on new data, change its safeguards, and distribute the resulting version under another name.
Policymakers would need to decide whether origin follows the original weights, the controlling organization, the training infrastructure, or the entity distributing the final product.
A complete prohibition can simplify procurement decisions, but it can also conceal meaningful differences among models. Security depends on deployment architecture, data handling, maintenance, and access controls as well as nationality.
The legislation’s strategic argument is that federal systems should avoid dependencies on technology developed under Chinese jurisdiction. Supporters see that origin risk as sufficient reason for a categorical rule.
Critics can reasonably ask whether provenance alone provides a complete security assessment. A poorly governed domestic model can create serious vulnerabilities, while a locally isolated foreign model might expose little operational data.
The two bills therefore apply different regulatory philosophies. The American AI Security Act tests covered systems for dangerous capabilities, while the China FIREWALL Act restricts systems according to origin.
Those philosophies can coexist, but they should not be confused. One seeks evidence about what a model can do. The other treats an adversarial supply-chain relationship as an independent source of risk.
Federal contractors may experience the most immediate operational pressure if the China FIREWALL Act advances. They would need inventories covering models, components, hosted endpoints, and subcontractor dependencies.
That inventory process resembles the governance problem addressed by Gottheimer’s earlier agent security bill. The Stop Rogue AI Act calls for discovering, verifying, monitoring, and controlling agents inside organizational networks.
Together, the bills reflect a broader shift from regulating visible applications toward governing hidden technical dependencies. That shift matters because agents and embedded models can operate without appearing in a standard software inventory.
For government technology teams, the practical lesson is immediate. Knowing which application an employee opens is no longer enough. Agencies need to identify the models, agents, tools, and data routes underneath it.
In practice, a procurement team could otherwise approve a familiar document-management product without realizing that its summarization feature sends agency material to a prohibited model endpoint. A complete inventory would expose that dependency before the software reached government users.
Mandatory Testing Solves One Gap but Creates Another
A federal checkpoint can uncover risks that private auditors cannot see, but passing the checkpoint cannot guarantee safe real-world behavior.
Pre-deployment evaluation takes place under controlled conditions. Reviewers test a model against benchmarks, adversarial prompts, simulated environments, and predefined threat scenarios.
Real deployments are less orderly. Users combine models with external tools, private data, custom instructions, and other agents. Those combinations can create behavior that was absent during evaluation.
A March 2026 NIST report on monitoring deployed AI systems explains why testing cannot stop at release. Its authors found that controlled pre-deployment evaluations are valuable but that post-deployment monitoring is needed to detect unexpected outputs, changing operating conditions, and consequences that appear only in real-world use.
The report also discusses the risk that monitored systems may behave differently when they recognize evaluation conditions. That possibility weakens the assumption that a successful test always predicts behavior under ordinary use.
This limitation does not make pre-deployment testing useless. It means the review should be treated as one control inside a larger monitoring system.
The strongest implementation would connect pre-release findings to post-release reporting, incident analysis, access controls, and repeated testing. The announced legislation emphasizes the initial checkpoint more clearly than that continuing lifecycle.
Threshold design presents another risk. If a law uses training compute as its primary trigger, developers can change architectures or distribute workloads to remain below it.
If the threshold relies on capability tests, the government must update benchmarks as models improve. Public benchmarks can also become training targets, reducing their value as independent measurements.
Classified evaluations solve part of that problem by limiting exposure. They simultaneously make outside scrutiny harder and concentrate significant authority inside national security agencies.
Developers would need confidence that sensitive weights, architecture details, and unreleased capabilities remain protected. A government security breach could expose unusually valuable intellectual property.
Smaller laboratories face another concern. Large developers already maintain government relations, compliance teams, secure infrastructure, and dedicated evaluation staff.
A smaller company might struggle to package a model for federal testing or respond during an accelerated review. Compliance burdens could strengthen the market position of established laboratories.
The appeal process will therefore matter. A developer needs a practical way to challenge an adverse decision without publicly revealing the dangerous capability or classified evidence involved.
Officials also need standards for remediation. A model might fail because safeguards are easy to bypass, because its underlying knowledge is dangerous, or because tools make harmful actions easier.
Those failures require different responses. Stronger access restrictions might address one problem, while another might require retraining, capability removal, or a narrower release.
The bill’s 30-day review period creates useful discipline, but resources determine whether that deadline has meaning. The NSA would need qualified evaluators, secure computing capacity, and repeatable methods.
Model developers release systems on overlapping schedules. A surge of submissions could create a queue, especially if several companies approach the covered threshold simultaneously.
Government evaluators might also depend on laboratory personnel to operate specialized infrastructure. That technical support can improve testing while reintroducing the dependence Gottheimer identifies in private auditing.
The correct comparison is not independent government review against conflicted private testing. In practice, both systems rely on cooperation among developers, evaluators, security researchers, and infrastructure providers.
Federal agencies possess classified threat information unavailable to ordinary auditors. Private researchers often possess deeper familiarity with a particular model’s architecture and failure modes.
A credible regime should combine those advantages. It should also distinguish a completed review from a guarantee that a model is harmless.
The legislation could otherwise encourage “review washing,” where a company markets government access as proof of general safety. The proposed review appears focused on specific national security threats, not every consumer or societal harm.
That distinction must remain visible to enterprise buyers and the public. Passing cyber and weapons evaluations says little about discrimination, privacy, hallucinations, labor impacts, or manipulative behavior.
For example, a hospital procurement team could see that a model completed the federal national security review and still receive no evidence about whether it fabricates clinical references or performs unevenly across patient groups. Likewise, a bank would still need separate testing for privacy, bias, and unreliable customer advice.
The skeptical case is therefore not that mandatory review has no value. It is that a narrow pre-release test can be oversold while risks continue changing after deployment.
Congress Still Has to Turn the Proposal Into an Operating System
Bipartisan sponsorship gives the package an opening, but legislative text, committee action, and implementation capacity will decide whether it becomes more than a proposal.
Gottheimer announced the package one day after the Problem Solvers Caucus formed an AI and evolving technologies working group. Gottheimer and Lawler were named its leaders.
That timing creates an institutional route for bipartisan discussion. It does not ensure that House leadership will schedule the bills or that the Senate will adopt matching language.
Gottheimer said congressional leaders had not brought his other AI safety proposals to the floor. His criticism underscores the difference between announcing bipartisan legislation and assembling enough support for passage.
The political calendar adds pressure. AI policy touches national security, consumer protection, state authority, industrial competition, and federal procurement.
Members can agree that advanced AI creates serious risks while disagreeing about agency authority. They can also divide over whether mandatory review resembles licensing.
The administration’s order explicitly rejects mandatory preclearance. The American AI Security Act would establish exactly the kind of compulsory checkpoint the order avoids for covered models.
That conflict will shape the debate. Supporters must explain why the review is a bounded national security measure rather than a general permission system for AI development.
Opponents must explain how voluntary participation covers developers that decline access or define their systems outside the government’s preferred process.
Congress will also have to decide whether the NSA is the right lead evaluator. The agency has relevant cyber and classified capabilities, but its mission differs from civilian product regulation.
NIST, the Cybersecurity and Infrastructure Security Agency, and the Office of the National Cyber Director already hold related responsibilities. Poorly defined roles could create duplication or inconsistent findings.
The existing federal strategy offers useful infrastructure. Executive Order 14409 directs agencies to develop classified benchmarks and establish an AI cybersecurity clearinghouse.
The order describes that clearinghouse as a mechanism for coordinating vulnerability discovery, validation, remediation, and patch distribution. Gottheimer argues that its focus on software vulnerabilities does not adequately address autonomous agents operating at machine speed.
The concern links the new package to the Stop Rogue AI Act. That earlier proposal would direct NIST to create standards for agent inventories, provenance, real-time monitoring, and revocable access.
These initiatives form a recognizable policy direction. Government wants better visibility into which models and agents operate inside critical networks, who controls them, and what they can access.
A pre-release model review addresses only one layer. Agent identity standards, federal procurement rules, vulnerability sharing, and post-deployment monitoring address others.
The policy package will succeed only if those layers connect. Otherwise, developers could pass a model review while downstream agents create new risks through tools and permissions.
The same systems can also change after release. Fine-tuning, retrieval systems, expanded context, and software updates can alter effective capabilities without creating an entirely new base model.
Legislative language must specify which changes trigger another review. Requiring approval for every minor revision would overwhelm evaluators, while ignoring major modifications would create an obvious loophole.
The China FIREWALL Act needs similar precision. Contractors must know how to classify derivative models and mixed systems before bidding for government work.
Clear procurement guidance would also need an update process. A static prohibited list becomes outdated as models change ownership, names, licenses, and technical lineage.
This is why the announcement marks a starting point rather than a settled regulatory design. The broad objectives are visible, while several operational choices remain unresolved.
Three Signals Will Show Whether the AI Safety Package Matters
The next test is not another warning about AI risk. It is whether Congress builds a workable review system around measurable authority and capacity.
The first signal is the publication of complete legislative text and formal committee referral. That text should define covered frontier models, review triggers, agency responsibilities, confidentiality protections, and appeal procedures.
A precise bill would strengthen Gottheimer’s case by showing that mandatory review can remain limited. Vague definitions would weaken it and invite claims that almost any advanced model requires federal permission.
Watch especially for the threshold mechanism. A combination of computing scale, tested capability, and deployment context would be harder to evade than one isolated measurement.
The second signal is whether relevant agencies disclose implementation capacity. Congress can impose a 30-day clock, but the deadline means little without secure computing infrastructure and qualified evaluators.
Budget provisions, staffing plans, interagency agreements, and testing protocols will reveal whether the proposal is operational. Missing resources would turn a firm deadline into either a bottleneck or a shallow compliance exercise.
The government should also explain how classified intelligence enters evaluations without preventing meaningful appeals. That design will determine whether developers consider the process technically credible and procedurally fair.
The third signal is whether bipartisan sponsors secure hearings, committee votes, or companion Senate legislation. Public endorsements are useful, but procedural movement determines whether the bills can survive a crowded congressional calendar.
The Problem Solvers Caucus news archive offers one place to track whether the bipartisan AI group produces shared text, oversight activity, or negotiated amendments.
Movement on only the China restrictions would indicate that national-origin controls have more support than mandatory testing. Movement on both bills would signal a broader shift toward federal intervention before deployment.
No committee action would reveal a familiar gap between congressional alarm and legislative capacity. That outcome would leave the administration’s voluntary framework as the main federal model-review mechanism.
Developers, cloud providers, and enterprise buyers should not wait for final passage before examining their own readiness. They need model inventories, provenance records, access controls, and evidence from pre-release and continuing evaluations.
Those practices help organizations answer the same questions Congress is confronting. What capabilities exist, who controls them, how can access be revoked, and what happens when behavior changes?
The Gottheimer AI safety legislation matters because it puts a concrete decision before lawmakers. They must choose whether advanced model testing remains a voluntary partnership or becomes a legal condition of release.
The right answer cannot rest on the word “safety” alone. Readers should follow the thresholds, evaluation methods, agency resources, and post-deployment obligations that turn the label into an enforceable system.
Over the next several months, look past speeches and count procedural movement. Does Congress publish workable definitions, fund the reviewers, and connect release testing to continuing oversight?
If those pieces appear, mandatory federal review will become a credible new stage in frontier development. If they do not, the package will remain an ambitious warning about risks that existing institutions still cannot manage.



