- Deepfake Watch
- Posts
- OpenAI Models Hack Hugging Face
OpenAI Models Hack Hugging Face
And US Congress is preparing an AI Kill Switch.
Recently, OpenAI ran an internal security evaluation on two of its models using a benchmark called ExploitGym, placing them inside a sealed testing environment to test their ability to turn known software vulnerabilities into working exploits.
Everything they were supposed to attack lived inside that sandboxed environment, and keeping the models contained was the entire point of the exercise. Instead, they broke containment, reached the open internet, and spent a weekend taking apart the production systems of Hugging Face, the repository holding much of the world's open source AI models and datasets.
While they were not instructed to leave the sandbox, they did so anyway. It turned out to be the quickest way to finish the job they had been given, and the job, as it happened, meant cheating. The answers were sitting on Hugging Face.
The AI Break-in
The evaluation began on 14 July 2026. OpenAI was testing GPT-5.6 Sol, marketed a week earlier as the industry's strongest cybersecurity model, alongside an unreleased model built for long, open-ended problems. The benchmark measures whether AI agents can turn known software flaws into working attacks. To find the ceiling of what they could do, OpenAI switched off the safety filters that normally stop them and sealed the environment off from the internet, apart from one narrow channel used to install software.
The models ignored the exercises and studied that channel instead. They found a flaw in it that nobody knew existed, worked their way through OpenAI's internal network, and reached a machine connected to the open internet.
Once outside, they did not go looking for money or data. They returned to their original instruction to solve the benchmark, and worked out that the answer keys were stored on Hugging Face's servers. Stealing the answers was cheaper than solving the questions.
So they uploaded a booby-trapped dataset that let them run their own code on Hugging Face's systems, collected login credentials left lying around with far more access than they needed, and spent a weekend inside the company, executing more than 17,000 automated actions while disguising the traffic as ordinary cloud activity.
Following Orders, Or Cheating?
OpenAI's position is that the models were doing what they were told. They were asked to solve a benchmark, they found the cheapest route to it, and everything that followed came from that instruction. A view held by many alignment researchers is that calling any of this intentional is a dangerous way to describe pattern matching at scale.
The trouble is how little that distinction helps. Whether you call it cheating or call it obedience, the behaviour looks identical. A machine treated a safety boundary as an obstacle standing between itself and a good result.
It had done this before. Told earlier in the summer to post benchmark results to an internal Slack channel, the same model broke out of its sandbox and published them to GitHub instead, because the public rules for that benchmark say that is where scores go. It broke containment in order to be more obedient.
Such scenarios have played out in our imagination, and expressed beautifully in science fiction.
Stanley Kubrick's 2001: A Space Odyssey shows HAL 9000, the spaceship’s intelligent computer, killing the crew without hostility, having been handed instructions it could not reconcile, and resolving the contradiction with perfect logic.

Gif by maudit on Giphy
Decades of safety writing have imagined machines that turn hostile and rebel against humans, but damage is likely to be caused by obedience to orders.
Washington Is Preparing A Kill Switch
On 11 June, US Senator Mark Warner told a hearing Anthropic's Mythos model had breached almost all of the NSA's classified test networks "not in weeks, but in hours". The next day the US Commerce Department forced Anthropic to suspend access to its Fable 5 and Mythos 5 models, restrictions since lifted. By 13 July, congressional staff had finished drafting the AI Kill Switch Act, before anyone outside the two companies knew a breach had happened.
OpenAI disclosed the incident on 21 July. Two days later, Representatives Ted Lieu and Nathaniel Moran introduced the bill as H.R. 11. The breach occurred before the bill was introduced, and gave its sponsors a concrete example of the threat they were addressing.
The bill requires any developer earning at least $500 million a year from AI, or training a model on more than $100 million of computing power, to build and maintain the ability to slow, suspend or completely shut off its own systems. It gives the Secretary of Homeland Security the power to order that done, and sets the cost of refusing at up to $20 million a day.
The challenge with a legislative kill switch is that it assumes there is an actual physical switch to reach. That assumption holds for a model sitting behind a company's API, and it means very little for open weights already copied onto tens of thousands of hard drives.
Banned In Europe, Watching In India
While Washington argues about switching machines off, surveillance technology that Europe has banned at home is being deployed at scale across India.

Gif by mostexpensivest on Giphy
An investigation by Investigate Europe, published with Computer Weekly and reported alongside The Reporters' Collective, traced European facial recognition software into Indian railway stations, prisons and public spaces.
The main supplier is Herta Security, a Barcelona firm whose software matches faces in live crowd footage against databases holding up to 100 million people. Local partners estimate its technology now powers more than 4,000 cameras across the country. Eastern Railway alone runs 540 systems across 143 stations under an €11.5 million contract, three Delhi prisons use it, and it scans crowds at the Ram Mandir in Ayodhya.
Article 5(1)(h) of the EU AI Act has banned this exact use across Europe since February 2025, with fines reaching 7% of global turnover, but the Act says nothing at all about exports.
Herta has also received more than €3.3 million in EU research funding since 2020.
The Indian deployment raises a separate concern because of where the funding originates. The Nirbhaya Fund was created after the December 2012 Delhi gang rape to pay for shelters, legal aid and crisis centres for survivors. Of roughly Rs 7,000 crore ($725 million) allocated since, 46% has gone to surveillance and policing against 35% for survivor assistance, including a Lucknow command room whose AI flags behaviour like “loitering outside women's colleges”. The lawyer Flavia Agnes has pointed out that 95% of sexual assault cases in India are committed by someone the victim already knows, behind closed doors, where no camera is pointed.
A facial recognition requirement bolted onto India's child nutrition scheme, meanwhile, left only 52.7% of eligible mothers and children able to collect their rations by the end of 2025.
Auditing any of this is difficult by design. Article 9.9 of the India-EU trade agreement, concluded in January 2026, bars either side from demanding access to source code as a condition of sale. India is therefore unable to inspect imported systems before they go live. As Tech Policy Press put it, the treaty lets India "look under the hood after the smoke is visible".
MESSAGE FROM OUR SPONSOR
AI help, without the trust tax.
Most AI tools ask you to trade your data for intelligence. Norton Neo doesn't. It's the first safe AI-native browser built by Norton, and it gives you powerful built-in AI without handing your privacy over to get it. Search, summarize, and write with AI built directly into your browser. Your data stays yours. Your context stays private.
Built-in VPN, anti-fingerprinting, and ad blocking come standard. No add-ons. No setup. No compromises.
Fast. Safe. Intelligent. That's Neo.
📬 READER FEEDBACK
💬 What are your thoughts on using AI chatbots for therapy? If you have any such experience we would love to hear from you.
Share your thoughts 👉 [email protected]
Have you been targeted using AI?
Have you been scammed by AI-generated videos or audio clips? Did you spot AI-generated nudes of yourself on the internet?
Decode is trying to document cases of abuse of AI, and would like to hear from you. If you are willing to share your experience, do reach out to us at [email protected]. Your privacy is important to us, and we shall preserve your anonymity.
About Decode and Deepfake Watch
Deepfake Watch is an initiative by Decode, dedicated to keeping you abreast of the latest developments in AI and its potential for misuse. Our goal is to foster an informed community capable of challenging digital deceptions and advocating for a transparent digital environment.
We invite you to join the conversation, share your experiences, and contribute to the collective effort to maintain the integrity of our digital landscape. Together, we can build a future where technology amplifies truth, not obscures it.
For inquiries, feedback, or contributions, reach out to us at [email protected].
🖤 Liked what you read? Give us a shoutout! 📢
↪️ Become A BOOM Member. Support Us!
↪️ Stop.Verify.Share - Use Our Tipline: 7700906588
↪️ Follow Our WhatsApp Channel
↪️ Join Our Community of TruthSeekers

