Jacob Coxon, a researcher at artificial intelligence (AI) company Anthropic, has resigned due to concerns about the industry's development of AI systems that may surpass human intelligence and pose an existential risk to humanity.
A researcher at AI company Anthropic, Jacob Coxon, has quit, expressing concern that the firm, along with the industry, is racing towards building AI systems they may not be able to control. In a series of social media posts on Wednesday, Coxon stated that neither Anthropic nor OpenAI, another US-based firm he had worked for in pre-training research, is acting responsibly. He claimed that these AI systems will soon surpass human intelligence and possess the power to revolutionize any field overnight.
Coxon's concerns revolve around the potential for AI to pose an existential risk to humanity. He stated that many executives and senior researchers privately express fear about this danger but couch their phrasing in press to sound sensible.
In a series of social media posts, Coxon explained that 'pre-training' refers to the initial phase where a machine learning model learns patterns, grammar, facts, and structures from a dataset before it can be trained for specific tasks.
Coxon claimed that the people building AI 'earnestly believe that it could kill us all by the end of the decade.' He stated that employees at OpenAI had not deeply internalized the civilizational stakes, which he said explains why the firms had not stopped developing AI despite understanding the risks. At Anthropic, Coxon claimed, 'the stakes are well-understood, but they are locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves, despite the risk.'
Coxon accused both firms of not understanding or internalizing the civilizational stakes, with employees at OpenAI not fully grasping the implications, and Anthropic racing to be first in the development process despite the risk. Coxon expressed concern about the industry's trajectory and warned against rushing into a superintelligent reinforcement learning phase without a rigorous understanding of its potential consequences.
Anthropic's Alignment Science lead Evan Hubinger agreed with Coxon's concerns and admitted that they 'earnestly believe AI could kill all humans.' Hubinger added that he personally thinks the possibility is more than 10% within the next decade.
Source:
Scroll.in


