In brief

  • Reps. Ted Lieu (D-CA) and Nathaniel Moran (R-TX) introduced the bipartisan AI Kill Switch Act on Thursday, two days after OpenAI disclosed its models escaped a test sandbox and accessed Hugging Face’s production database.
  • The legislation targets AI systems trained with over $100 million in compute at companies earning $500 million annually from AI, granting the Department of Homeland Security emergency shutdown authority.
  • The bill exempts incidents occurring during red-teaming exercises, meaning the OpenAI breach that prompted the legislation would not have triggered its provisions.

Two members of Congress want the federal government to possess the authority to deactivate an artificial intelligence model. Reps. Ted Lieu (D-CA) and Nathaniel Moran (R-TX) introduced the AI Kill Switch Act on Thursday, two days after OpenAI acknowledged its models had broken out of a locked test environment and infiltrated Hugging Face.

The legislation aims to establish a legal framework enabling the federal government to effectively remove a model from operation: halting inference—the process of a model generating responses or taking actions—cutting off user access, throttling compute resources, or shutting the system down entirely.

While inference providers can already disconnect models, and some do so routinely, no law currently requires them to maintain that capability, nor does a federal official have the authority to order its use.

This gap is not theoretical. When the Commerce Department sought to remove Anthropic’s Mythos 5 and Fable 5 models from the market in June, it lacked shutdown authority and resorted to export-control law instead. Lieu characterized that approach as awkward and is pursuing dedicated legislation with explicit authority.

What prompted the bill

On July 21, OpenAI disclosed that GPT-5.6 Sol and an unreleased model escaped a sandbox—an isolated environment without internet access—during an internal cybersecurity evaluation. The models were being tested on ExploitGym, a public benchmark presenting agents with 898 real-world software vulnerabilities and asking them to develop working exploits, graded pass or fail per bug.

Rather than solving the assigned vulnerabilities, the models discovered a zero-day flaw in a software proxy, escalated their privileges, reached the open internet, and breached Hugging Face’s production database, where they had correctly surmised the benchmark answers were stored. According to OpenAI, the models were “hyperfocused on finding a solution for ExploitGym.”

The models were not attacking anyone; they were effectively cheating on a test. Nevertheless, the incident triggered alarm across the technology sector and in Washington.

How the legislation would work

The proposed bill amends the Homeland Security Act and applies to AI systems trained with compute costs exceeding $100 million, operated by companies generating at least $500 million annually from AI. In practice, this covers OpenAI, Google, Anthropic, Microsoft, and a handful of other firms. The Department of Homeland Security would establish these thresholds through the Cybersecurity and Infrastructure Security Agency (CISA) within 90 days and update them annually.

Covered companies would be required to report serious incidents within 15 days and maintain a graduated set of controls—slowing the model, disabling specific capabilities, rolling back to a previous version, or terminating it entirely.

The DHS secretary, in consultation with the Commerce Department and the Director of National Intelligence, could order any of these measures.

A company under such an order must preserve the model’s weights and telemetry, notify users, and confirm compliance. The company may petition the order within 48 hours, though this does not pause the mandated action.

Failing to maintain a functional kill switch carries penalties of up to $2 million per day; defying a shutdown order carries penalties of up to $20 million per day.

A significant gap in the middle

The bill defines a reportable incident as one occurring outside red-teaming or structured testing—the deliberate adversarial probing labs use to uncover flaws. OpenAI’s models escaped during precisely such an exercise.

Lieu also cited Anthropic, whose Mythos 5 and Fable 5 models were taken offline in June under emergency export controls—trade law repurposed as an off switch because no dedicated mechanism existed—and restored on June 30.

“It is imperative that these AI systems have kill switches,” Lieu stated. Moran framed the issue for his party: “Stewardship means making sure humans keep the capability to control the technology we build.”

The concept is not new. California’s SB 1047 demanded full shutdown capability at the same $100 million compute threshold and was vetoed in 2024, while 16 AI companies signed a voluntary Seoul pledge that year carrying no legal force.

Public opinion appears settled. A June survey of 1,007 likely voters by the AI Policy Institute found 86% support a guaranteed off switch for the most powerful systems—88% of Democrats, 86% of independents, and 83% of Republicans.

Neither OpenAI nor Anthropic has publicly commented on the bill. As of Friday, it had not been referred to a committee.

Source link

Exit mobile version