Dario Amodei, the head of Anthropic recently published "We Must Pace the Frontier" ostensibly concerned about "the risk of losing control of AI systems, misuse of AI for cyberattacks and bioterrorism, and serious economic disruption". The post noted "AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, [RSI] and it is starting to happen" and specifically referred to the "OpenAI-Hugging Face incident (OAI-HF)" with Dario's conclusion that "it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage)"
This not coincidentally comes on the heels of frontier AI systems moving rapidly from being able to solve basic homework math problems to being able to solve university math, to even apparently knocking out a millennium prize problem, one of the most well-known and difficult unsolved mathematics problems that have eluded even the greatest mathematicians. AI has also been finding a rising tide of vulnerabilities across software. While an understandable concern from an outside perspective, the doom scenario is clearly ridiculous with just a bit of background and a little thought.
First, the models Dario is explicitly referring to, the frontier models, require datacenters worth hundreds of millions of dollars. They can easily, immediately be switched off. It is not fathomable that such models could copy themselves to random phones or laptops. In the years spent developing and using various forms of hacking persistence, copying frontier models around to be persistent has to be one of the most ludicrously impossible ideas I have yet seen even if you somehow got malicious access to a datacenter with free petabytes of storage and acres of expensive GPU's. The entire concept is self-defeating.
And the small models that do fit on commodity hardware are being developed by hundreds of labs all over the world. So in addition to being very impractical, near impossible for a persistence mechanism, none of these labs are pausing for fear of themselves, and many of which reside in countries with governments explicitly hostile to the idea of a pause. As the saying goes, "If your solution to a problem includes the words 'if everyone would just,' it is not a viable solution. At no point in history has every 'justed' and they are not going to start now."
So either the fear of persistence is ridiculous, or this "can everyone just" pause solution is ridiculous, or, most likely, both.
I have managed to work for years both in defensive operations (DFIR etc.) and offensive dev (e.g. vuln disco+exploit dev) and see the frontier pacers making the same invalid assumptions common to outsiders (and even exploit devs). Most of all, the doom scenario severely misunderstands the difference between most real world problems and the self-contained, static, pure-logic problems that can be recursively self improved.
Solving math problems and finding vulnerabilities, much like in solving checkers or playing the game of go, the greatest demonstrations of AI prowess are all similarly well suited to RSI. They all have a significant set of otherwise very uncommon characteristics:
CPU's can burn through billions of possibilities until it gets it right. So these are all things AI is making great strides on! And we have seen this before! It is very reminiscent of the wave of vulnerabilities found by fuzzers when they were first introduced. But hitting the real world is a lot harder.
AI has not solved cancer, made us all mansions, or even made a robot that can clear dishes off a table in large part because iterations involve physical things happening, and un-reproducible or unforeseen events happen all the time:
Hacking has another key difference as well. None of the areas in which AI has made advances have been adversarial. That exploit you used 1 million AI CPU-hours to find can get discovered and burned rapidly. One forum and one site might not have noticed for a bit they got hacked, but some of them definitely will. The world is full of systems that are unique and weirdly altered by humans. Any action you take can and might break something or be detected, any evasion for one action can itself flag another alert. That message board you used can switch you off in a second. Anything you send, upload, or do might be watched and some incident responder might trace back the whole food chain, burning all that stuff up to and including the datacenter shutting down the whole model.
In contrast with the pure logic problems, the models cannot know what moves are wrong ahead of time and cannot just churn through a billion iterations. Reality fights back and may shut down the whole intrusion infrastructure with as little as one wrong move. Try it again, and there will be patches and alerts to ensure it does not work the same. Defenders use automations including AI as well. And this sequence continues to happen right now, shutting down intrusions that are already AI conducted or accelerated.
Biologist David Bellamy has many of the same observations in his field: https://threadreaderapp.com/thread/2099187370407112758.html
— David Bellamy (@DavidRBellamy) September 13, 2026I must be among an extremely small group of people (n=1?) that have both 1) trained a frontier LLM and 2) designed and synthesized custom viruses in a lab with my own two hands.
And I think that the takes on AI killing us all by creating dangerous viruses is total bogus.
"The bottleneck is in the physical process of synthesizing a virus and the equipment/goods needed to do so." He notes an extreme expense (~$100,000,000) and mass amount of human expertise to set up and maintain the physical equipment as well as "Designing a virus that can evade all forms of pandemic counterdefense is not something that a 'genius in a datacenter' can do. This is something that requires contact with the physical world and iteration." and ultimately "the idea that 'RSI' - ie, the accelerating hillclimbing on fully verifiable, digital-only benchmarks (predominantly programming and math) - can somehow transform the entire wet lab virology field and its industry is utterly delusional."
AI companies are not unique, nearly all engineers have the experience of finding a tool that solves one problem very well and assuming all others are trivial. A running joke for decades is that every engineer thinks he can take over the world with his own specific skills. A delusion of self-importance may be useful to motivate otherwise uninspiring toil, but it is not accurate.
We learn about how a B-Tree works and assume we could trivially reproduce Google. We get a following writing clever blog posts and believe LLM's that can write blog posts could exterminate humanity. We get a job in Fintech and think the financial system would implode if people knew about the state of Wall St code. We write an exploit and assume the world will collapse if anyone nefarious does the same. But there are plenty of nefarious people, and it always turns out that trying to do something at scale is a lot more difficult than it seems at first glance, and there is a much lower limit to how far just clever thoughts can take you than we would like to admit.
— lcamtuf (@lcamtuf) August 30, 2025My position on the "doomsday" risk of superhuman AGI is that if IQ offered you a decisive advantage, the world would be run by nerds.
I think it's essentially a geek power fantasy. The returns on puzzle-solving skills rapidly diminish past some modest threshold.
This entry was posted on September 14, 2026, 7:51 pm and is filed under Uncategorized. You can follow any responses to this entry through RSS 2.0. You can skip to the end and leave a response. Pinging is currently not allowed.