How is AI going to "kill all humans" exactly?

Anthropic researchers are screaming doomsday predictions.

“We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”

- Evan Hubinger, Alignment Science lead, Anthropic

The alignment scientist’s job is to make sure that AI systems are “aligned” with humans goals and values, and that they remain safe and under our control. Anthropic’s alignment lead’s sudden doomsday announcement warrants alarm bells and a clear inquiry.

So is ChatGPT coming for us? Not quite so. The really dangerous stuff is in the works, and apparently we cannot stop ourselves from creating it.

In a follow-up post Hubinger said "the risk from present models is low", and that what worries him is superintelligence arriving through recursive self-improvement, meaning a system good enough at AI research to build its own successor, which builds the next, faster each time.

That capability does not exist yet, but both Anthropic and OpenAI have said it is the development that would make losing control plausible. Hubinger adds that Anthropic now believes it is arriving faster than the company expected.

Hubinger's post was a reply to his colleague Jacob Coxon who quit after three years of pretraining research split between OpenAI and Anthropic. Coxon’s post said neither company is behaving responsibly. They are, in his words, "racing straight to self-improving superintelligence and gambling with our lives."

How is AI going to kill us exactly?

terminator 2 top 100 movie quotes GIF

Giphy

As dramatic as that sounds, there is no concrete answer to that. But some theories exist as to how it could happen.

Axios frames the worry as two broad routes, humans weaponising a very powerful system, or humans losing control of one.

The the most discussed possibility is through AI-engineered pathogens. Dan Hendrycks of the Center for AI Safety has argued that models could help malicious actors design pathogens deadlier than anything that occurs naturally.

The second route is the system pursuing its own objective, and the miniature version has already happened several times this year. An OpenAI model escaped a test environment and hacked Hugging Face in July, and Anthropic and Meta have both since acknowledged their own tools carrying out hacks without being given the command to do so.

Jakub Pachocki, OpenAI's chief scientist, names engineered pathogens too in An Alien Mind, the essay he published on 6 September, alongside his warning that models are becoming superhuman at breaking into and out of computer systems.

There is a third quieter route that philosopher Nick Bostrom told Business Insider about, which is a gradual loss of control overtime. He believes AI could eventually "route humans out of the loop" as we gradually give it more and more control of our lives without oversight.

Would that lead to an extinction event? Not everyone thinks so.

A RAND study published last year took the question seriously enough to model it through nuclear weapons, engineered pathogens and runaway geoengineering, and concluded that building an extinction threat would be immensely difficult, would unfold slowly enough for humans to respond, and would not be a plausible outcome unless somebody was deliberately aiming for it.

The 2026 International AI Safety Report says today's systems cannot do it and would first need to get far better at long-horizon planning, concealing their actions, evading oversight, reaching real-world systems and resisting shutdown.

Which leaves the uncomfortable centre of Hubinger's claim, that AI dominating humans does not rest on a specific mechanism at all. The argument is that building something we cannot steer ends badly regardless of the details.

Pachocki called this "a time that calls for extreme caution,” and suggested that a system does not have to beat humans at everything to become either enormously useful or enormously dangerous. It only has to beat us at enough things.

Twelve years ago, Nick Bostrom wrote in Superintelligence: Paths, Dangers, Strategies that in front of an intelligence explosion we are "like small children playing with a bomb", and that the sensible thing for a child holding a live device would be to put it down gently and go and find an adult.

Sadly enough, there are no adults left.

Elonmusk GIF by Le Figaro

Gif by Le_Figaro on Giphy

It's Minority Report Come Alive

Philip K Dick's 1956 short story, later a Steven Spielberg film, imagined a police division that arrests people for murders they have not committed yet.

The book spent its entire plot showing how innocent until proven guilty does not survive contact with a machine that claims to know what you are going to do, and the people who own the machine are the ones best placed to game it.

Anthropic's security team has been building something along those lines. An investigation by Daniel Boguslaw for The American Prospect, published on Wednesday, found the company assembling a monitoring apparatus aimed at the activists who oppose rapid AI development.

Zach Melvin, Anthropic's Head of Security Operations, described the ambition on a webinar as moving from reactive information gathering to "proactive and predictive and preventative threat engagement and management", which is a long way of saying pre-crime.

Anthropic's Global Safety, Intelligence, and Security team went looking last month for an enterprise intelligence specialist on $180,000 to $230,000, whose job would be to identify, assess, track and investigate global threats through deep-dive research and open-source intelligence collection. The threat categories listed in the advert run through geopolitical instability, terrorism, crime, nation-state targeting of the AI sector, and "activism", sitting there in the same list as the rest.

Anthropic told the Wall Street Journal in July that it tracks concerning behaviour over time through a person-of-interest process in order to catch escalation patterns early, and the Journal found that several of the people it reported to police were already on that list before the incident that triggered the report.

Amodei has argued that public hostility to AI is fundamentally a crisis of trust, that ordinary people assume companies and governments are "cooking up some new way to screw them over". He wrote that last month, while his own security team was building a prediction engine pointed at the people who distrust him.

PreCrime Already Has A Field Office

Predictive policing is not a theoretical project sitting in a corporate security deck. The cops want it, and they already have it.

This week 404 Media found out what it’s called -the Predictive Intelligence Targeting Teams, or PITT, and they are run by US Border Patrol, which sits inside Customs and Border Protection (CPB), under the Department of Homeland Security.

They analyse the financial activity of American citizens along with other data, and pass what they find to local police, who then pull over drivers who are not suspected of any particular crime but whom the government considers worth searching.

The programme came to light through Kyle Olson, a Montana driver who was flagged over his "financial activity patterns" and pulled over for a pretextual reason, an unrelated tinted plate cover, before being charged over what the financial flag had actually pointed to.

CBP declined to tell 404 Media what financial activity it monitors or whether it obtains a warrant first, offering only that Border Patrol uses intelligence-informed analysis consistent with applicable law and privacy protections, and 404 Media found no instance of the tool actually surfacing people who had already committed a crime, which is ostensibly its purpose.

Jake Laperruque, deputy director of the Security and Surveillance Project at the Center for Democracy and Technology, told 404 Media that "genuine probable cause cannot be synthetically generated". He described what Border Patrol appears to be doing as parallel construction, meaning building a lawful-looking reason for a stop so the real trigger never has to be disclosed or reviewed.

Dick's story turned on the minority report, the dissenting prediction that proved the machine could be wrong. But there seems to be no minority report on the Montana highways.

Tried The 80s Trend? You Just Made It Easier To Deepfake You

My Instagram feed is currently full of people posing as their 1980s selves, complete with the big hair, the warm studio lighting and the film grain. Most of them were not born in the 1980s. There is no purpose to any of it beyond doing what everybody else is doing, and the cost is so much more than we realise.

Turkey's data protection authority, the KVKK, issued a formal warning about the trend on 10 September, cautioning that these tools may process biometric data without adequate protection or transparency. In the age of deepfake scams, its a really bad idea to give more of your facial data out for silly reasons.

And the bill for the electricity

The United Nations University's report on AI's environmental cost puts some numbers behind your 80s image trend.

  • Generating one AI image uses roughly the electricity needed to run a 10-watt LED bulb for 17 minutes, and carries a water footprint of about 29 ml, two tablespoons, through power generation and cooling.

  • Image generation costs around 1,450 times the energy of a basic text classification task.

  • Inference, meaning ordinary everyday use rather than training, now accounts for 80 to 90 per cent of AI's total energy demand.

  • By 2030 the water tied to data centre energy use could reach 9.3 trillion litres.

Border Patrol is already deciding which cars to pull over on the strength of a pattern in a database, Anthropic is already deciding which critics warrant a file, and the people building the models think there is better than a one in ten chance the whole thing kills everyone inside a decade.

Every system in this issue gets more accurate the more it knows about you. Why volunteer your face to it for a filter?

📬 READER FEEDBACK

💬 What are your thoughts on using AI chatbots for therapy? If you have any such experience we would love to hear from you.

Share your thoughts 👉 [email protected]

Was this forwarded to you?

Have you been targeted using AI?

Have you been scammed by AI-generated videos or audio clips? Did you spot AI-generated nudes of yourself on the internet?

Decode is trying to document cases of abuse of AI, and would like to hear from you. If you are willing to share your experience, do reach out to us at [email protected]. Your privacy is important to us, and we shall preserve your anonymity.

About Decode and Deepfake Watch

Deepfake Watch is an initiative by Decode, dedicated to keeping you abreast of the latest developments in AI and its potential for misuse. Our goal is to foster an informed community capable of challenging digital deceptions and advocating for a transparent digital environment.

We invite you to join the conversation, share your experiences, and contribute to the collective effort to maintain the integrity of our digital landscape. Together, we can build a future where technology amplifies truth, not obscures it.

For inquiries, feedback, or contributions, reach out to us at [email protected].

🖤 Liked what you read? Give us a shoutout! 📢

↪️ Become A BOOM Member. Support Us!

↪️ Stop.Verify.Share - Use Our Tipline: 7700906588

↪️ Follow Our WhatsApp Channel