Why one researcher’s resignation made others speak
The news was not simply that another AI researcher had warned about catastrophic risk. It was the public response from people still inside the labs.
Jacob Coxon announced his resignation from Anthropic on September 8, after working on pretraining at both Anthropic and OpenAI. In his subsequent WIRED interview, he argued that competitive pressure could push even comparatively responsible companies toward unsafe decisions. He called for coordination on systems that help develop their own successors, eventually involving the United States and China.
Zvi interprets the response as a preference cascade: one conspicuous, costly act makes it easier for others to acknowledge views that were previously harder to express. Anthropic alignment researcher Evan Hubinger publicly agreed that the risk was serious; researchers across multiple labs joined the discussion.
That does not make their probability estimates a scientific consensus. Zvi explicitly notes disagreement among lab employees. Nor does a wave of statements resolve the technical question of whether alignment will improve fast enough.
The institutional puzzle is more interesting than assuming everyone secretly holds the same view. Two researchers can share a severe risk assessment and make opposite decisions: one leaves to force public attention; another stays because they believe their safety work reduces the danger.
A resignation makes the concern visible. What follows determines whether that visibility changes the incentives or merely produces another week of headlines.