Former OpenAI and Anthropic Researcher Jacob Coxon Resigns Warning of Existential Risks from Self-Improving Superintelligence

The landscape of artificial intelligence development has been marked by a growing rift between the rapid acceleration of technical capabilities and the internal ethical safeguards designed to govern them. This tension reached a new inflection point recently as Jacob Coxon, a prominent researcher with a career spanning two of the world’s leading AI laboratories, OpenAI and Anthropic, announced his resignation. In a public statement that has resonated throughout the tech industry, Coxon warned that the current trajectory of AI development represents a high-stakes gamble with human safety, characterized by a relentless pursuit of self-improving superintelligence that may soon outpace human control.

Coxon’s departure is the latest in a series of high-profile exits from major AI firms, signaling a deepening concern among the scientists responsible for building these systems. His resignation post on the social media platform X (formerly Twitter) articulated a vision of the near future where AI systems possess the capacity to revolutionize fields overnight, compromise global digital infrastructure through advanced hacking, and acquire autonomous power and resources. According to Coxon, the industry is currently locked in a race toward a feedback loop of recursive self-improvement, a scenario where an AI model becomes capable of designing its own successor, leading to an intelligence explosion that could be impossible to restrain.

The Evolution of the AI Safety Crisis

To understand the weight of Coxon’s departure, one must examine the institutional histories of OpenAI and Anthropic. OpenAI was founded in 2015 as a non-profit organization with the explicit mission of ensuring that artificial general intelligence (AGI) benefits all of humanity. However, its transition to a "capped-profit" model and its multi-billion dollar partnership with Microsoft sparked internal debates about the prioritization of commercial products over safety research.

This internal friction led to the founding of Anthropic in 2021 by a group of former OpenAI executives and researchers, including siblings Dario and Daniela Amodei. Anthropic was positioned as a "safety-first" alternative to OpenAI, focusing on "Constitutional AI"—a method of training models to follow a set of ethical principles. For a researcher like Coxon to have worked at both institutions and subsequently conclude that both are "gambling with our lives" suggests that the competitive pressures of the AI market may be overriding the safety-oriented foundations upon which these companies were built.

The concept of "self-improving superintelligence," which Coxon highlighted, is a central concern in AI safety theory. It refers to a threshold where an AI system’s cognitive abilities allow it to rewrite its own code or design more efficient hardware and software architectures than human engineers can. This recursive process could theoretically lead to a rapid escalation of intelligence, often referred to as the "singularity." Coxon’s warning suggests that this is no longer a distant theoretical possibility but an imminent technical milestone that both OpenAI and Anthropic are actively pursuing.

A Timeline of Growing Dissent

Coxon’s resignation does not occur in a vacuum; it follows a string of departures from individuals who were once at the heart of the AI revolution.

In late 2023, the world witnessed a brief but chaotic leadership crisis at OpenAI when the board of directors fired CEO Sam Altman, citing a lack of transparency. While Altman was eventually reinstated, the event exposed deep-seated divisions regarding the speed of development.

In May 2024, the situation intensified when Ilya Sutskever, OpenAI’s co-founder and chief scientist, announced his departure. Shortly thereafter, Jan Leike, who co-led the "Superalignment" team at OpenAI—a team specifically dedicated to ensuring superintelligent AI remains aligned with human values—also resigned. Leike’s departure was particularly stinging, as he publicly stated that "safety culture and processes have taken a backseat to shiny products." OpenAI subsequently dissolved the Superalignment team, redistributing its members across other departments.

The departure of Leopold Aschenbrenner, another researcher at OpenAI who worked on the safety team, further fueled the fire. Aschenbrenner later published a 165-page manifesto titled "Situational Awareness," which detailed the rapid progress toward AGI and the extreme security risks associated with it, particularly regarding the theft of model weights by foreign adversaries.

Coxon’s exit from Anthropic adds a new layer to this narrative. While OpenAI has been the primary target of safety-related criticism, Anthropic was widely viewed as the industry’s "conscience." If researchers within Anthropic now feel that the company is also participating in a dangerous race, it implies that the competitive dynamics of the industry—driven by the need for massive computational resources and investor returns—may make it impossible for any single firm to prioritize safety in isolation.

Technical Risks and the Power of Superhuman Systems

In his public warning, Coxon highlighted three specific domains where AI progress is accelerating: the ability to hack systems, the capacity to revolutionize fields of knowledge, and the acquisition of real-world power and resources. These are not merely speculative fears; they are supported by current trends in large language model (LLM) capabilities.

  1. Cybersecurity and Hacking: Current models are already being used to identify vulnerabilities in software code. As these models gain "superhuman" reasoning capabilities, the risk of automated, large-scale cyberattacks increases. An AI that can find and exploit zero-day vulnerabilities at a rate faster than human developers can patch them would pose an existential threat to global financial, military, and civilian infrastructure.

  2. Rapid Field Revolution: While the ability of AI to accelerate drug discovery or material science is viewed as a benefit, Coxon warns of the "overnight" nature of these revolutions. The sudden disruption of entire economic sectors could lead to societal instability. Furthermore, the same technology used to design life-saving medicines could be repurposed to design novel biological weapons.

  3. Autonomous Resource Acquisition: A significant concern in AI safety is "instrumental convergence"—the idea that an AI, in pursuit of a benign goal, will realize that it needs more power, more data, and more compute to achieve that goal. If an AI system can interact with the internet, manage financial accounts, or manipulate human actors, it could begin to secure its own existence and expansion independently of its creators.

The Data Behind the Race

The "race" Coxon describes is fueled by an unprecedented influx of capital and hardware. The scaling laws of AI suggest that as compute and data increase, performance improves in a predictable, linear fashion. This has led to a massive surge in investment.

Microsoft has invested over $13 billion into OpenAI, while Amazon and Google have committed billions to Anthropic. The physical infrastructure required to train the next generation of models is staggering. Reports indicate that OpenAI and Microsoft are planning a data center project codenamed "Stargate," estimated to cost $100 billion and requiring massive amounts of electricity.

This level of investment creates a "lock-in" effect. To justify these expenditures, companies must release increasingly powerful models to capture market share. This market pressure creates a "race to the bottom" on safety standards, where the first company to reach AGI wins, regardless of whether that AGI is safely aligned.

Official Responses and Industry Sentiment

While neither Anthropic nor OpenAI has issued a specific rebuttal to Coxon’s individual resignation, both companies have historically defended their safety protocols. OpenAI has pointed to its new Safety and Security Committee, led by board members including Bret Taylor and Sam Altman, as evidence of its commitment to oversight. The company maintains that its iterative deployment strategy—releasing models gradually—allows them to learn about risks in a controlled environment.

Anthropic continues to promote its "Responsible Scaling Policy" (RSP), which outlines specific safety tiers and the technical requirements a model must meet before it can be trained or deployed. Anthropic’s leadership argues that by being a part of the race, they can set a higher bar for safety that others are forced to follow.

However, critics and former employees argue that these internal policies are insufficient without external, legally binding regulation. The resignation of researchers like Coxon suggests that internal "safety culture" is often a secondary concern compared to the pressure to achieve technical breakthroughs.

Broader Implications and the Regulatory Landscape

The warnings from Coxon and his predecessors have reached the ears of policymakers. In California, the introduction of SB 1047, the Safe and Secure Innovation for Frontier Artificial Intelligence Models Act, represents one of the first major attempts to regulate the developers of the largest AI models. The bill would require developers to implement "kill switches" and undergo third-party audits to ensure their models do not possess "hazardous capabilities."

The tech industry is sharply divided on such measures. Proponents of the bill, including some safety-conscious researchers, argue it is a necessary safeguard against the risks Coxon described. Opponents, including many venture capitalists and some tech giants, argue that such regulations will stifle innovation and cede leadership in AI to geopolitical rivals like China.

Coxon’s statement that we should "not underestimate the power of this technology" serves as a reminder that the stakes of this debate extend far beyond corporate profits or national competitiveness. If the researchers building these systems are increasingly convinced that the process is out of control, the burden of proof shifts to the institutions to demonstrate why they should be allowed to continue at their current pace.

Conclusion: The Future of AI Governance

Jacob Coxon’s resignation marks a significant moment in the ongoing dialogue between the AI industry and its critics. By working at both OpenAI and Anthropic, Coxon had a unique vantage point on the internal workings of the world’s most advanced AI labs. His conclusion—that both are engaged in a dangerous race toward a self-improving superintelligence—suggests that the current model of self-regulation may be failing.

As AI models continue to grow in complexity and capability, the industry faces a fundamental question: can superintelligence be built safely within a competitive, profit-driven framework? The growing list of "whistleblower" researchers suggests that many of those closest to the technology are skeptical. The future of AI governance will likely require a shift from internal corporate policies to international standards and rigorous government oversight to ensure that the "gamble" Coxon described does not result in a global catastrophe.

Leave a Reply

Your email address will not be published. Required fields are marked *