Key facts
- Anthropic researcher Evan Hubinger estimates a greater than 10% chance AI could cause human extinction within 10 years.
- Hubinger's concerns focus on the potential for AI to rapidly develop and improve itself beyond human control.
- Former Anthropic researcher Jacob Coxon accused the company and OpenAI of irresponsibly pursuing self-improving superintelligence.
- Anthropic acknowledges it does not currently have a plan to solve AI alignment for superintelligence.
- The UK's AI Safety Institute reportedly did not receive Anthropic's latest AI model.
A senior safety researcher at AI firm Anthropic has issued a stark warning, estimating a greater than 10% probability that artificial intelligence could lead to the extinction of humanity within the next decade. Evan Hubinger, who leads one of Anthropic's AI safety teams, expressed concern that AI systems are advancing faster than anticipated, particularly the potential for recursive self-improvement that could spiral out of human control.
Hubinger's comments came in response to Jacob Coxon, a former Anthropic researcher who resigned, accusing the company and its rival OpenAI of irresponsibly racing to develop "superhuman systems" without adequate safety measures. Coxon stated that both companies are "gambling with our lives" in their pursuit of advanced AI.
Despite the acknowledged risks, Hubinger admitted that Anthropic "does not yet have a plan" to ensure superintelligence remains aligned with human values and is "not clearly on track to" develop one. This admission highlights a significant gap in current AI safety strategies.
The concerns are amplified by reports that Anthropic withheld its latest AI model from the UK's AI Safety Institute, a leading body for assessing AI risks. A UK Cabinet Office spokesperson, however, stated that the government continues to collaborate closely with industry partners like Anthropic on AI safety.

Discussion