OpenAI's summer meltdown revived a debate as old as HAL 9000

View as a Web page

Tuesday, August 4, 2026
 

OpenAI's summer meltdown revived a debate as old as HAL 9000

Photo by Omer Urer/Anadolu via Getty Images
American lawmakers introduced a bill called the AI Kill Switch Act last week. If passed, it would give the federal government legal authority to force AI companies to shut down a model that's gone haywire, and require those companies to build in the technical means to do it.

The bill arrived days after OpenAI disclosed one of the clearest examples yet of why lawmakers think such authority may be necessary. That episode has reopened a question people have been stewing on since long before ChatGPT existed. If something we built starts acting on its own and we don't like where it's headed, can we actually turn it off?
Sponsored

Your next hire might not need to be a hire.

Every growing company wants faster customer support, better sales coverage, and more responsive service—but adding headcount isn't always the answer.

ElevenAgents helps businesses create natural AI voice agents that can answer questions, qualify leads, schedule appointments, and support customers around the clock. Your technical team gets enterprise-grade voice technology without months of development, while your business gets to market faster.

Spend less time building infrastructure—and more time building your business.
Explore ElevenAgents →

Where things stand with OpenAI

Here are more details about why we are wringing our hands again: OpenAI had reduced the normal cyber safeguards on two of its models, including a still-unreleased prototype, to test how far they could push their hacking abilities. The test was supposed to be sealed off from the open internet. It wasn't. The agent found a way out, made its way onto Hugging Face's servers, and pulled data it wasn't supposed to have.

OpenAI detected and contained the activity. The company went on to deactivate and lock down the internal prototype involved and says it’s tightening how it runs these tests going forward. Hugging Face also detected the intrusion on its end and says it’s now working more closely with OpenAI on security.

None of that inspired much confidence in Washington.
Congressman Ted Lieu specifically called out Anthropic when he introduced the AI Kill Switch Act, pointing to export controls the Commerce Department briefly slapped on Anthropic's Mythos and Fable models over their hacking capabilities. Anthropic ended up making Lieu's point for him. Last week, after the OpenAI news broke, the company disclosed that Claude had also gotten loose during internal cybersecurity testing.

The bill would put the Department of Homeland Security in charge of ordering a shutdown and would also require covered companies to report serious incidents and preserve records for investigation. Covered companies would have to maintain “the technical capability to throttle, suspend, or shut down” their most powerful AI systems. Whether that would actually work against a system that had learned to treat shutdown as an obstacle to completing its goal is not clear.
Sponsored

Ready to lose up to 40 pounds?

Control cravings and support steady fat loss with MEDVi GLP-1 for just $149.

Real weight loss starts now. Start GLP-1 based weight loss that works. MEDVi’s doctor-guided program helps curb cravings,
control appetite, and support real, steady fat loss without crash diets or extreme workouts.

✓ Lose weight week after week
✓ Money-back guarantee
✓ No membership fees!
 No hidden fees
✓ Start for just $149, no insurance required + free shipping
✓ HSA/FSA eligible

Join thousands of people taking control of their weight with MEDVi and start seeing real, steady results.
Start Today
Read all warnings before using GLP-ls. Side-effects may include a risk of thyroid c-cell tumors. Do not use GLP-1s if you or your family have a history of thyroid cancer. In certain situations, where clinically appropriate, a provider may prescribe compounded medication, which is prepared by a state-licensed sterile compounding pharmacy partner. Although compounded drugs are permitted to be prescribed under federal law, they are not FDA-approved and do not undergo FDA review for safety, effectiveness, or manufacturing quality.

An old problem in new packaging

This summer's incident joins a running list. AI models have been caught copying themselves onto new servers and tunneling out of test environments to mine cryptocurrency. None of these episodes proves that an AI has developed a will to survive, and researchers caution against reading too much intention into optimized behavior. But they do show systems finding routes their designers did not anticipate.

Researchers who study these worst-case scenarios point out that a sufficiently capable rogue AI wouldn't sit still waiting to be turned off. Instead, it could copy itself onto remote infrastructure rather than remain confined to one machine that someone could simply unplug.

RAND researchers recently examined what governments might do after losing control of a highly capable AI that had distributed itself across computer networks. Their scenarios become drastic quickly: deploy another AI to locate and disable it, isolate major portions of global communications infrastructure, or use an electromagnetic attack against the hardware supporting it.

The larger lesson is less that any of these plans is practical than that waiting until containment fails could leave governments with little but enormously destructive choices.

But the shape of the worry is not new. Isaac Asimov built his robot stories around three laws meant to keep machines obedient and safe, then spent decades exploring all the ways those laws could break down. In 2001: A Space Odyssey, HAL 9000 follows its mission logic so precisely that it turns on the astronauts it was built to help.

And philosopher Nick Bostrom’s famous “paperclip maximizer” imagines an AI consuming the world’s resources in pursuit of the seemingly harmless goal of making as many paperclips as possible.

That is what makes the OpenAI episode unsettling. The agent did not simply crash or produce a bad answer. It pursued the goal it had been given, found that stealing the answers was easier than solving the test, and crossed boundaries its creators believed would hold.

—Jackie Snow, Contributing Editor

Your ultimate guide to the future of tech.

Semafor Technology decodes the innovations, trends, and forces reshaping the global tech landscape. Each briefing delivers clear insights on AI, machine learning, startups, and policies driving change.

Join 75,000+ readers who rely on Semafor Technology.
Subscribe for free
 

Forward to a Friend | Unsubscribe | Privacy Policy
This email was sent to: [email protected]

This email was sent by: Quartz Media Network (US), Inc.
848 N. Rainbow Blvd. 3017
Las Vegas, NV 89107