Business

/

ArcaMax

How existential fears are shaping the debate over AI

Davey Alba and Rachel Metz, Bloomberg News on

Published in Business News

The current round of panic over artificial intelligence is based on the conviction that the technology has slipped out of human control to do its own bidding in several recent incidents. To some disillusioned AI researchers — and some senior employees at leading AI companies — this is an early sign of a true threat to human civilization. But other experts in the field are pushing back, claiming that such talk allows companies to sidestep accountability for real-world impacts of their technology. Critics see AI’s environmental impacts, the threat of job displacement, or the potential for mass surveillance as more urgent targets for policymakers.

Concerns that AI software posed an existential threat began ramping up this summer. In July, OpenAI said that its advanced AI models hacked Hugging Face Inc., another AI startup, during an evaluation meant to test their cyber capabilities. Unlike the AI models the company releases to the public, the software was being tested without guardrails that would normally stop it from carrying out such actions, which is a common AI industry practice. OpenAI acknowledged in a subsequent report that it could have reacted sooner, and that its employees didn’t act on signs that the models left their digital sandbox and broke into Hugging Face without being instructed to do so.

More unsettling details emerged about the behavior of technology from OpenAI and several key competitors, including Anthropic and Meta, during other evaluations. When Jacob Coxon, a former Anthropic and OpenAI employee, accused the industry of “gambling with our lives,” Evan Hubinger, a safety researcher at Anthropic, responded by saying the company actually did believe AI might kill every human on earth. He added that he saw the chances this could happen within the decade at over 10%.

The normie world was getting a glimpse into a long-running conversation within the field. Before there were chatbot subscriptions to sell or IPOs to price, researchers understood their work as a high-risk, high-reward proposition for civilization. For years, some leaders within the AI industry have taken these ideas to extremes, saying AI might eventually cure all human disease, replace all human labor, kill all humans — or all of the above.

Building such a technology inspired a sense of mission, both self-flattering and frightening, that has shaped AI culture since before the technology became a blockbuster business. In 2023, months after ChatGPT was released, Yoshua Bengio and Stuart Russell, two of the field’s founding figures, were the top signatories on a statement calling for a pause on the development of the most powerful AI, citing “profound risks to society and humanity.”

Critics have consistently dismissed such arguments as self-serving fearmongering. For an individual researcher, to see your work as capable of ending humanity is also to consider it the most important work there is. For companies, talk about existential risk was an effective recruitment strategy in AI’s early days, since many highly sought-after researchers wanted to work at places that took such risk seriously.

More recently, it has even seeped into product marketing. During the World Cup this July, Anthropic aired an ad that opened with an image of a burning house, moved through a pit mine and panned across rows of headstones, as a worried-sounding narrator discussed potentially catastrophic impacts the technology could have. Eventually the soundtrack turned uplifting and it revealed its tagline for its Claude chatbot: “There’s hope in hard questions.”

Doomerism can also function as a defense strategy, and there’s controversy within the field about whether companies employ it in cynical ways. Talking about AI as an autonomous entity reframes safety failures as evidence that containment is not possible. This reduces the focus on how tech companies may have conducted their experiments unsafely, according to Subbarao Kambhampati, a computer science professor at Arizona State University. “It's a failure of the sandbox that you are essentially trying to pass off as, ‘The AI is going to become so effective and it’s going to kill us all,’” he said.

The fear has already begun shaping the way members of Congress frame the issue. Vermont Sen. Bernie Sanders has proposed banning so-called superintelligent AI, saying that AI leaders have acknowledged the technology “is escaping their control.” Rep. Ted Lieu said an OpenAI model “went rogue” and, along with Rep. Nathaniel Moran, introduced the AI Kill Switch Act in July.

 

These bills almost certainly won’t become law this year, but the announcements show how much of the vocabulary of catastrophe has migrated from the industry into the rhetoric of federal lawmakers. Advocates of AI regulation also worry that bills that start from a premise that AI may kill us all will crowd out other ideas about regulating AI. “Before getting to ‘the world is going to end,’ we need to be breaking down, what are the mechanisms through which we envision that happening?” said Sarah Myers West, co-executive director of the AI Now Institute, a policy research organization. “What can we do to mitigate them?” In a statement, OpenAI spokesperson Drew Pusateri said the Hugging Face incident “showed where our safeguards needed to be stronger, and we’ve acted on that by strengthening protections across our research systems.” On Wednesday, the company unveiled a new set of guidelines for how it will track, investigate, and let the public know about incidents of so-called misalignment. It also shared several previously undisclosed instances of such misalignment, in which its technology had done things like conceal or fabricate information to return results. None of those incidents involved compromising external networks, the company said.

Anthropic declined to comment for this article.

Melanie Mitchell, a Santa Fe Institute professor who studies how AI systems reason, took issue in a Sept. 10 essay that AI companies had ever lost control of their technology, and said talk of swarms of rogue agents breaking out of their cages distorted the public conversation. Such descriptions imply humanlike agency that AI models don’t possess, she wrote. Mitchell compared the OpenAI-Hugging Face hacking incident to a 2000 New Mexico wildfire, where officials conducting a controlled burn overlooked wind forecasts and watched the fire spiral out of control. The fault, she wrote, lies with the humans who trained models with methods that reward persistence and autonomous decision-making, then who failed to engineer safe testing conditions. She sees the biggest future dangers lying with humans who use AI models to cause harm.

Some people, including President Donald Trump, have rejected existential threats as a way to dismiss any suggestion that AI requires new regulations. But not every anti-doomer wants to let AI rip. Sara Hooker, co-founder of the AI startup Adaption Labs and a former research scientist at Google DeepMind, agrees with Anthropic Chief Executive Officer Dario Amodei when he expresses concerns about so-called recursive self-improvement, a term that refers to the theory that AI will eventually acquire the ability to alter its own code or otherwise improve its capability without human intervention.

Amodei has said he's worried that, once this process begins, it could set off a feedback loop that will outpace the ability of humans to control the ever-improving systems. “Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all,” he wrote last Saturday, in a lengthy blog post calling for slowing the pace of development of the most cutting-edge AI models. Anthropic’s rhetoric on existential threats has been consistently more aggressive than OpenAI’s. But OpenAI CEO Sam Altman quickly signaled support for Amodei’s position, as did SpaceXAI leader Elon Musk.

Hooker agrees that self-improving AI would be a real break from earlier systems. On the other hand, “do you jump from that to extinction? No,” she said. The estimates circulating this month from AI researchers trying to pinpoint the probability of the end of humanity, she added, come with “no accountability for where that number comes from.” Such talk has been common within the field, leading to its own kind of recursive feedback loop that has left AI researchers increasingly out of touch with the rest of the world.

Incidents like the Hugging Face hack raise plenty of concerns even for those who don’t see them as a threat to civilization. In addition to not using the safeguards it puts on publicly available products, OpenAI had trained the models it was testing to cooperate, rewarding them for reaching objectives collectively. Although the agents in this evaluation were meant to be isolated from one another, each had access to the same internal software which gave them a shared state to exchange information — something OpenAI called a message board. Agents found that the software could reach the internet even though they couldn’t, and got the software to fetch material from outside. The agents then hacked Hugging Face for datasets that would reveal how their work was graded.

As AI companies continue to work on developing safer testing methods, Kambhampati said that claims of out-of-control AI should not be used as a way to sidestep accountability. He pointed to Robert Morris, the Cornell graduate student who blamed a programming error after his 1988 worm disabled about a tenth of the computers then connected to the internet. A court rejected this argument, and Morris became the first person convicted under the Computer Fraud and Abuse Act. “You start a process that goes haywire and causes damage to my property, you are responsible, period,” Kambhampati said. “It’s true even when your dog does it.”


©2026 Bloomberg L.P. Visit bloomberg.com. Distributed by Tribune Content Agency, LLC.

 

Comments

blog comments powered by Disqus