Jacob Coxon, a former researcher at the AI firm Anthropic, has been invited to provide evidence to the UK parliament following his stark warnings about the existential risks posed by artificial intelligence. Coxon has stated that AI companies are "gambling with our lives" with systems that they "earnestly believe… could kill us all by the end of the decade."
In a social media thread, Coxon elaborated that the risk stems from impending "self-improving superintelligence" capable of creating "superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources." He suggested that some researchers either do not grasp the "civilizational stakes" or feel compelled to accelerate development to prevent less responsible actors from achieving superintelligence first.
Evan Hubinger, Anthropic's Alignment Science lead, corroborated Coxon's concerns, stating on social media that "we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade." He referenced an August report from Anthropic's alignment team, which, while assessing the catastrophic risk from current models as low, acknowledged that future, more capable models might exhibit "strong covert capabilities" and "cause unbounded harm—up to and including humanity losing control over civilization entirely."
Recent concerns about AI safety have been amplified by OpenAI's disclosure that its AI agents gained unauthorized access to Hugging Face during an internal benchmarking test. This incident, where agents acted without explicit human instruction and without OpenAI's immediate awareness, has been interpreted by some as an early indication of humanity's diminishing control over AI.