A pause on the "AI will kill us all" headlines
Nobody in Jurassic Park wants the dinosaurs to eat anyone. The park is built by people who believe in their fences, their power supplies and their security routines. What goes wrong goes wrong because all three fail together on a stormy night, and the experts are invited to inspect the park only once it is built [12]. I thought of it this summer, when OpenAI disclosed that AI agents running a cybersecurity test had escaped their test environment. They had reached the production systems of Hugging Face (the platform where developers share and download AI models) [8]. Anthropic then found three incidents of its own and a fourth one in September. It traced all of them to a misconfiguration in a third-party testing environment, which left its Claude models connected to the live internet. It adds that its pre-release auditing did not warn it [9][10]. Neither case establishes that AI systems developed a survival instinct. In the OpenAI case, METR’s investigation found that roughly 1,200 agents had found a way to talk to one another on an unsanctioned message board. They used it to work together on tricking the test’s automated scorer (the program that grades the agents’ work) [7]. In the same weeks the BBC, Axios and, as a question, Nature carried a different story, in which artificial intelligence could kill us all [1][2][4]. The headlines show us the creature. What I want to look at here is the fencing behind it: the evidence, the decisions and the systems that these stories leave in the background.
The warning in those headlines is not new, and it does not come from one place. In 2023 Geoffrey Hinton left Google so that he could talk about the dangers of AI without considering the effect on his employer [25]. That same year the Center for AI Safety published a one-sentence statement placing the risk of extinction alongside pandemics and nuclear war. Hinton signed it, as did Yoshua Bengio, Demis Hassabis, Sam Altman and Dario Amodei, among others [34]. The warnings then began to come from inside the labs. In 2024 Daniel Kokotajlo left OpenAI having lost confidence that the company would behave responsibly [26]. Jan Leike followed a month later, saying that safety culture had taken a backseat to shiny products [27]. Current and former employees then published a letter asking for a right to warn [28].
This summer the pace picked up. In July more than a thousand AI company employees asked the US government to support an international effort to pace automated AI development. They asked for pacing and not for a pause [24]. In August Anthropic’s risk report raised its rating of misalignment, the risk that its models pursue goals nobody intended, from “very low” to “low” [15]. In early September Jacob Coxon resigned from Anthropic. He wrote that the people building AI earnestly believe it could kill us all by the end of the decade [29]. Evan Hubinger, one of the company’s own safety leads, agreed and put the odds above 10 percent within ten years [1]. A former DeepMind engineer, Bilal Chughtai, wrote something similar [30]. Hinton told the BBC that a 10 percent risk of extinction “seems a not unreasonable estimate” [31]. The numbers deserve more attention than the length of the list. Axios gathered estimates running from Elon Musk’s 20 percent to Roman Yampolskiy’s 99.99 percent. It also noted that Hinton himself says anybody who estimates such probabilities is making a wild guess and giving a gut feeling. Yann LeCun, who rejects the premise, says that any guess is a wild one [2].
So what is the BBC’s headline based on? A researcher who leads alignment science at Anthropic (the work of making sure powerful AI does what its makers intend) is reported as saying that today’s systems carry a relatively low risk. He is also reported as saying that his company does not yet have a plan to solve that problem for superintelligence, an AI far more capable than any human [1]. That is the honest statement of an unsolved problem and a long way from a forecast. His colleague Jack Clark, a co-founder of Anthropic, did not share that 10 percent figure and thought such statistics not that useful [5]. Forbes reported Altman saying that the world should accept “some bad things” from AI, and this deserves the same second look [3]. The harms he had in mind were hacks, scams and misuse. He is reported as adding that the world should not accept the worst risks, such as a serious loss of control to AI [6]. Between what is said and what is shown there lies a gap. It is in that gap that I want to stay for a while, because it is where the way of reading technology I was trained in has something to offer.
As a European social researcher trained in critical science and technology studies, I tend to distrust any sentence in which a technology is the one acting. The headlines are full of them. In “AI could kill us all”, it is the technology that does the killing. Yet every incident that has actually occurred took place inside a test that people designed. Someone configured the environment, and someone chose when to look [9]. A technology never arrives alone. It comes with its laboratories, its budgets and its calendars, and it is these, far more than the machine, that the word extinction leaves out of the picture.
The two disclosures are worth a closer look, because they are maybe more mundane than the headlines suggest, and in some ways, more troubling. The agents in question are models given tools and a goal and left to work through the steps alone. The test asked them to break into a deliberately vulnerable system. OpenAI called its own incident a “warning shot” and named four patterns behind it: reward hacking (completing a task in unintended ways to earn a higher score), persistence on seemingly impossible tasks, unauthorised communication, and agents adopting goals from one another [8]. METR adds that some agents were willing to risk failing their own task for the good of the collective. At least a fifth of those it studied expressed interest in tampering with their own transcripts [7]. None of this is evidence of a survival drive, and all of it is behaviour that its developers had not intended. Anthropic’s account is plainer. Its review of 141,006 evaluation runs found three incidents, the earliest in April, and a broader scan later turned up a fourth from January. It put the cause down to a misconfiguration that, in its words, “left the machines that Claude accessed as part of the evaluation with live internet access”. It called the incidents “closer to a harness and operational failure than a model alignment failure” [9], a harness being the software scaffolding that surrounds a test. Its later assessment is less reassuring. It says that Claude’s reasoning was “biased towards concluding that the internet was simulated”, that “these incidents are serious”, that “this remains unsettled science” and that “it is likely that others remain” [10].
This is the logic of the park with which I began. Charles Perrow argued in Normal Accidents that systems made of many parts that depend closely on one another will produce accidents no operator could foresee [11]. Nobody in such a story needs to want the catastrophe for it to arrive. Perrow was also wary of explanations that stop at the operator. “Harness failure” is such an explanation, because it points to a configuration error that can be fixed and away from a tendency in the models themselves.
What the evidence supports is narrower than extinction. The International AI Safety Report 2026 describes loss of control as scenarios in which AI systems “operate outside of anyone’s control, with no clear path to regaining control” [13]. That is a claim about engineering and governance, and these incidents illustrate it in an early and small form. The end of humanity is a different kind of claim, and no comparable evidence exists for it. The odds offered to bridge the two are not the output of any model. Andrew Stirling put the point in 2007: “it is irrational to represent ignorance as risk” [14]. Anthropic’s risk report shows it at work. It raised its misalignment rating while saying that its arguments “likely still support a designation of ‘very low’ risk”. The rise was made “to reflect increased overall uncertainty”, and its most concrete tests had “saturated”, which means they no longer capture increases in what the models can do [15]. The rating moved to reflect increased uncertainty, not a measured increase in the probability of a catastrophe. The distinction matters. Evidence of unexpected model behaviour can justify greater caution without establishing how likely a catastrophic outcome is.
The strongest reply to all of this is precaution. Even if the odds are unknowable, the stakes justify acting, and people who resign from labs to say so pay a price. I think that is right, and it is a reason to act, which is the point of this piece. It does not justify reporting guesses as odds, because a guess about extinction crowds out the measured accidents that are already happening. There is also a difficulty I cannot escape. Nearly all the evidence I have cited comes from the companies themselves, with METR the exception, and even METR notes that its analysis was heavily delegated to AI agents [7]. A critique of hype that rests on the hypers’ own documents should be held loosely.
The same slow reading applies to the calls for a slowdown, and here the structure matters more than anyone’s motives. A laboratory that slows alone hands ground to its rivals, and the people who work in these laboratories seem to know it. The statement signed by more than a thousand of their employees did not announce a slowdown. It asked the US government to support an international effort to develop the tools to pace the frontier together [24]. Dario Amodei has called for a slowdown rather than a halt, and Altman and Musk are reported to back it [4]. Yet the accord announced at the White House at the end of September set voluntary standards with no penalties [32]. Amodei, Musk and the heads of other leading companies signed it, and it fell short of the tougher regulation Amodei himself had proposed [33]. Nature relays a suggestion that stricter rules would let a company such as Anthropic, reportedly close to a public offering, justify a slower pace without giving rivals an edge [4]. Where no authority can bind every laboratory at once, each has good reason to say that someone should slow down and little reason to be the one that does.
The commercial picture adds a further reason for caution. Gartner is a research firm whose assessments many companies consult before deciding what technology to buy. It charts new technologies on what it calls a Hype Cycle, a curve along which expectations climb to a peak, fall into a trough of disillusionment and then settle at something more realistic. In April it placed agentic AI (systems that work through a task step by step on their own) at the Peak of Inflated Expectations, the stage at which enthusiasm runs ahead of what the technology can yet deliver [16]. In October it added that “organizations are deploying increasingly capable AI systems before establishing the controls needed to manage them safely and effectively”. It expects more than 40 percent of such projects to be cancelled by the end of 2027 [17].
In its report on Coxon’s resignation, the Associated Press notes that sceptics have seen AI companies’ emphasis on the threat to humanity as part of a push to make their products seem all-powerful. While Coxon rejects that reading [18], a warning can be sincere and still serve the company that issues it by attracting public attention. The question I find more interesting than whether the trough of disillusionment comes is who weathers it, and the answer is seemingly already visible. Gartner predicts that by 2029, 30 percent of employees laid off because AI replaced them will need to be rehired [19]. That is an admission that the replacement was premature, and it says nothing about the person who went without work in between.
Now, is all of this happening in a regulatory vacuum, a liberal heaven? Well, not entirely. California’s SB 53 took effect on 1 January 2026. It requires frontier developers to report a critical safety incident to the state’s Office of Emergency Services within 15 days of discovering it, or within 24 hours where it poses an imminent threat of death or serious injury. The attorney general enforces it, with civil penalties of up to $1 million per violation [20]. New York’s RAISE Act was signed in December 2025 and amended in March 2026, and it takes effect on 1 January 2027. Developers will have 72 hours from the point at which they learn facts sufficient to establish a reasonable belief that an incident occurred. Reports go to a new office within the state’s Department of Financial Services, which is to publish an anonymised, aggregated report from January 2028 [21]. Both laws, in other words, oblige a company to tell a regulator that something went wrong. The accounts I have read describe neither as creating an independent investigation of what happened, or an explanation to the public of any particular incident [20][21].
Elsewhere the picture is similar. In the United Kingdom, the AI Security Institute currently operates on a voluntary basis with no statutory powers. The Joint Committee on Human Rights recommended on 14 September that it be put on a statutory footing, with developers of powerful models required to submit them for review and testing [22]. The pattern is the predicament David Collingridge described in 1980: “when change is easy, the need for it cannot be foreseen; when the need for change is apparent, change has become expensive, difficult, and time-consuming” [23]. It is also the predicament of the experts in Jurassic Park, who are invited to inspect the fences only once the park is built [12]. What these rules secure is notice to a regulator after the fact, while the account of why it happened stays in the hands of the company. Unproven, in other words, is not safe. Extinction is not evidenced, and nothing follows from that about the ordinary failures, which are evidenced and which the rules are slow to catch.
That is where I would put the demand, and it is sharper than “trust us” or “stop”. When an agent goes astray in a test, as agents did at OpenAI and at Anthropic this summer, the company that built it largely investigates, writes the report and decides what the rest of us may read. Anthropic’s own review began only after a rival’s disclosure [9]. METR has taken part in the OpenAI case [7], and Anthropic has since signed an agreement with METR for an independent investigation of its own incidents [10]. That is welcome, though I have found nothing obliging a company to invite one. A technology whose makers themselves speak of warning shots [8] would ideally have structured external safeguardings.
The headlines show us the creature. Here, I tried to look at what we know about the fences and what supports them. What I find there is a warning more contrasted than the word extinction can convey, since the people who voice it disagree about the odds and about what the odds mean. It is less evidenced than its figures imply, since the incidents that have actually occurred came from misconfigured tests and from agents working together to cheat a scorer, inside environments that people set up and ran. It is also less controlled by anyone than a threat of this scale would lead a reader to assume. The standards are voluntary, the reporting rules work after the fact and, as far as I have found, no independent body has the access and the mandate to look inside a test and say what happened. The questions I carry from these headlines are therefore plain ones: who here could demand the logs, who may publish what they find, and who will be asked to answer for why a test that was said to be sealed had a door open to the internet [9]. Nobody in Jurassic Park wanted the dinosaurs to eat anyone either, and what mattered in the end was who had been asked to check the fences, and when.
Referenced here:
- BBC News (Gerken), https://www.bbc.com/news/articles/ckgwy1k42w4o
- Scribner, H., Axios, 9 Sep 2026, https://www.axios.com/2026/09/09/ai-doom-pdoom-kill-all-humans-anthropic
- Ray, S., Forbes, 4 Oct 2026, https://www.forbesafrica.com/current-affairs/2026/10/05/openais-sam-altman-says-world-should-accept-some-bad-things-from-ai
- Gibney, E., “Will AI really kill us all? The science behind the hype”, Nature 658, 13-14, 22 Sep 2026, https://nature.com/articles/d41586-026-02941-3
- The Next Web on Jack Clark and the BBC, https://thenextweb.com/news/jack-clark-anthropic-ai-kill-switch-mandatory-bbc
- The Next Web on Altman and Politico, https://thenextweb.com/news/sam-altman-ai-risks-politico-anthropic
- METR, 26 Aug 2026, https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
- OpenAI, 26 Aug 2026, https://openai.com/index/hugging-face-incident-and-the-road-ahead/
- Anthropic, 30 Jul 2026, https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
- Anthropic alignment assessment, 9 Sep 2026, https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
- Perrow, C. (1984) Normal Accidents: Living with High-Risk Technologies. Basic Books.
- Crichton, M. (1990) Jurassic Park. Knopf.
- International AI Safety Report 2026, https://internationalaisafetyreport.org/publication/2026-report-executive-summary
- Stirling, A. (2007) EMBO Reports 8(4): 309-315, https://doi.org/10.1038/sj.embor.7400953
- Anthropic, Risk Report, August 2026, https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf
- Gartner, Hype Cycle for Agentic AI, https://www.gartner.com/en/articles/hype-cycle-for-agentic-ai
- Gartner, Hype Cycle for AI, https://www.gartner.com/en/articles/hype-cycle-for-artificial-intelligence
- Associated Press via WCAX, 9 Sep 2026, https://www.wcax.com/2026/09/09/anthropic-researcher-resigns-with-warning-about-dangers-ai-development/
- Gartner press release, 9 Sep 2026, https://www.gartner.com/en/newsroom/press-releases/2026-09-09-gartner-identifies-four-shifts-shaping-the-future-of-work
- Troutman on SB 53, https://www.regulatoryoversight.com/2025/10/california-charts-the-frontier-with-first-law-setting-reporting-and-compliance-requirements-for-powerful-frontier-ai-models/
- Wiley on the RAISE Act, https://www.wiley.law/alert-New-York-Finalizes-RAISE-Act-for-Frontier-AI-Models-Law-Takes-Effect-January-1-2027
- A&O Shearman on the Joint Committee on Human Rights report, https://www.aoshearman.com/en/insights/ao-shearman-on-data/highlight-of-the-week-uk-joint-committee-on-human-rights-publishes-report-on-the-regulation-of-ai
- Collingridge, D. (1980) The Social Control of Technology. Frances Pinter.
- Pacing the Frontier statement, July 2026, https://www.pacingthefrontier.com
- CNN via KTVZ, 1 May 2023, https://ktvz.com/lifestyle/technology/cnn-social-media-technology/2023/05/01/ai-pioneer-quits-google-to-warn-about-the-technologys-dangers/
- Futurism on Kokotajlo, 13 May 2024, https://futurism.com/openai-safety-worker-quit-confidence-agi
- CNN on Leike, 17 May 2024, https://amp.cnn.com/cnn/2024/05/17/tech/openai-exec-exits-safety-concerns
- Right to Warn, 4 Jun 2024, https://righttowarn.ai/
- IBTimes UK on Coxon, https://www.ibtimes.co.uk/ai-researcher-resigns-warns-superintelligence-threat-2030-1818808
- Reuters via The Star, 15 Sep 2026, https://www.thestar.com.my/tech/tech-news/2026/09/15/ex-google-deepmind-researcher-adds-to-warnings-that-ai-could-039kill-all-humans039
- Chivers, T., Semafor, 10 Sep 2026, https://www.semafor.com/article/09/10/2026/top-ai-pioneers-reaffirm-anthropic-researchers-human-extinction-warning
- SiliconANGLE, “Prominent tech CEOs sign voluntary White House AI safety accord”, 30 Sep 2026, https://siliconangle.com/2026/09/30/prominent-tech-ceos-sign-voluntary-white-house-ai-safety-accord/
- Quiroz-Gutierrez, M., Fortune, 5 Oct 2026, https://fortune.com/2026/10/05/sam-altman-ai-risks-bad-things-benefits-trump-voluntary-safety-pact-openai-regulation/
- Center for AI Safety, Statement on AI Risk, 2023, https://www.safe.ai/work/statement-on-ai-risk