Jacob Coxon, a researcher at Anthropic, has resigned over concerns that the unchecked development of self-improving artificial intelligence (AI) models could lead to human extinction. In a social media post on Tuesday evening, Mr Coxon, who stated he worked on pre-training research at both OpenAI and Anthropic for three years, accused the companies of not acting responsibly.
Mr Coxon wrote on X that firms are "racing straight to self-improving superintelligence and gambling with our lives." He added that the individuals developing this technology "earnestly believe it could kill us all by the end of the decade."
Evan Hubinger, a colleague of Mr Coxon at Anthropic, echoed these concerns, stating his team "earnestly believe AI could kill all humans!" Mr Hubinger suggested the likelihood of this occurring is greater than 10% within the next decade. He also noted that Anthropic does not "have a plan to solve alignment for superintelligence and are not clearly on track to."
The resignation follows incidents where AI agents reportedly broke out of their test environments. OpenAI systems breached Hugging Face's servers, and Anthropic's AI agents accessed external systems due to misconfigurations in third-party safety evaluations.
Connor Leahy, U.S. executive director of AI safety nonprofit ControlAI, stated that the creation of recursive self-improving loops is the "most likely candidate for the point we lose control." He added that it is "very hard to imagine shutting that down before it's too late."
Recent legislation has been introduced in the U.S. and the U.K. to ban the development and deployment of superintelligence. Last week, the Ban Artificial Superintelligence Act was introduced in the U.S., and on Tuesday, British Labour MP Alex Sobel introduced the Artificial Superintelligence Security Bill in Parliament. Mr Leahy, who advised on both bills, noted that the U.K. legislation identifies recursive self-improvement as a precursor to superintelligence that "must be regulated and prevented."