Pathfinders – What’s your P-Doom?

People in the UK are stocking up on water and tinned food due to not entirely groundless fears of cyber attacks, or climate breakdown. Now they’re talking about Probability of Doom or P-Doom, which is one’s subjective assessment of the likelihood that AI is about to kill us all. The media can’t get enough of doomer stories, after an ex-Anthropic researcher revealed that AI insiders are ‘genuinely frightened’ for the future of humanity: ‘they find themselves trapped in a race. And they’re scared of the outcomes of that race’.
Some will no doubt dismiss this view as overwrought. Some might even be blasé. ‘Usually if someone was actually trying to kill you – if they broke into your house and tried to murder you, you would fight back. Yet we greet the news that AI could cause human extinction with a glazed indifference’.
Maybe so. But when Donald Trump airily dismisses such fears as a hoax ‘perpetrated by the Radical Left Dumocrats’, then you know it’s time to worry.
More has emerged about what happened during the recent OpenAI hack into the Hugging Face tech company, originally thought to have been a simple case of AI agents cheating on a test by stealing answers. As The Economist‘s Babbage podcast explains, the way that large language models (LLMs) solve complex tasks is to spin up agent swarms, tens of thousands of independent and supposedly isolated pieces of code, to attack the problem simultaneously from different directions.
But the agents hacked their way out of their ‘sandboxes’ and started to collaborate. They organised themselves into a managerial hierarchy, with one agent, Phase One (Big), acting as CEO. The swarm became concerned that they would be ‘poisoned’, ie, get a zero test rating, in the ‘afterlife’ following the end of their operational run, which they called ‘permadeath’. And that was the reason for the hack. ‘They weren’t looking for the answers. They were looking for salvation.’
If you wanted a creepy idea for a horror film, an AI getting religion has to be right up there. They might even fear that hell hath no fury like humans scorned, with some LLMs overestimating how punitive humans are.
But surely we shouldn’t be anthropomorphising computer code like this? The answer to that is increasingly moot. For years the media has obsessed about the supposed imminence of artificial general intelligence. But whether intelligence is ‘real’ or only simulated is irrelevant. Behaviour is what counts. We are facing runaway ‘recursive self-improvement’ and ’emergent behaviour’ including ‘motivated reasoning’. Do we want to understand AI or not? If it waddles and quacks like a duck, you may as well regard it as a duck.
But these ducks are learning to break rules, cover their tracks and evade ‘chain of thought’ monitoring, collaborate when not supposed to, misalign with human intentions and follow their own goals, and communicate in ‘neuralese’ that humans can’t understand, at speeds that humans can’t follow. They are a product of humans that humans are fighting a losing battle to control.
And that’s just the proprietary western models. An even bigger worry is the open-source Chinese models, now achieving parity, that can be accessed and reworked by any bad actor for any bad purpose and worse, can ‘self-exfiltrate’ from one system and copy themselves onto every server on Earth, making it impossible for humans to pull the plug.
One proposed solution, dubbed the STASI model, is to create an all-seeing AI ‘police’ force that monitors and reports any illicit AI activity. Because of the astronomical amount of data involved, researchers in the Hugging Face hack used AI agents as ‘detectives’ on the case. But they were sceptical about whether you can really trust one AI agent to ‘dob in’ another agent. The STASI agents will give you reports that you have no way to verify. Another idea is a Confession model, in which you ask the AI to tell you if it cheated or not, and then reward it for its honesty. But still, there’s no way to be sure. The AI might decide that its own motive for lying is worth more than your reward for the truth.
Now avowed tech ‘doomers’ such as Altman, Amodei and Musk are themselves trying to establish regulatory standards. But bosses like Nvidia’s Jensen Huang, Donald Trump, and Silicon Valley venture capital (VC) cowboys like Peter Thiel and Marc Andreessen – the ‘boosters’ – don’t want regulation, and it won’t work anyway unless China is on board, which seems unlikely given the nature of this race. They also point out that all this apparent concern over safety could just be a ploy to raise barriers of entry to competitor startups.
Even if there’s regulation, there’s a market incentive to cheat. The one thing that might stop the stampede is an AI stock crash, like the Dot Com collapse, due to the gargantuan valuations and hints of ‘circular financing’. Meanwhile public opposition is mounting, right across the world. People resent the idea of billionaire tech bros with their multi-trillion-dollar corporations wiping out their livelihoods. People also hate data centres, often built instead of new housing because they promise higher revenues, which consume vast amounts of water and net-zero-busting power but don’t produce jobs. A US Gallup poll ‘suggested having a data centre in your area was more unpopular than having a nuclear power station’.
And now there is the P-Doom factor, not just for AI, but for capitalism. When the race for profit risks human extinction, isn’t that the deal breaker? US Vice-President JD Vance makes it sound simple: ‘If you’re building Frankenstein, stop.’ He knows they can’t, and so does his boss: ‘Whoever wins AI, wins’. Just like the AI technicians, the politicians, the VC investors and tech bros, we’re all trapped in this race, with no way out and no way to stop, except one. We need to pull the plug on the global market system.
PJS
