
On September 9, Jacob Coxon quit Anthropic. He did it in public, on X, with a thread that said the people building frontier models, he argued, earnestly believe those systems could kill everyone by the end of the decade, and the labs are still racing toward self-improving superintelligence anyway. “Gambling with our lives” is the one, amongst many, concerning lines in his post.
Coxon had worked pretraining at OpenAI before Anthropic. He did not claim Anthropic was the villain. He called it the most responsible player he had seen. His target was the industry dynamic: speedrun the race so the wrong party does not get there first, or put your head down because it is happening anyway. Evan Hubinger, Anthropic’s Alignment Science lead, replied that Coxon was correct on the core claim. “We really do earnestly believe AI could kill all humans,” Hubinger wrote, putting his own personal odds above 10 percent within a decade.
Four days later, Dario Amodei published We Must Pace the Frontier. The essay is about three thousand eight hundred words, and it is trying to hold two ideas at once. One is the upside Amodei has been writing about since Machines of Loving Grace: AI that cures major diseases, accelerates growth, and expands what ordinary people can do. He opens with his father dying of a disease cured a few years later, and with his own early-stage cancer that would not have been treatable fifty years ago. The other idea is that the last few months have convinced him the labs need to slow the rate at which they improve capabilities, so safety work has a chance to keep up.
Two developments sit under that shift. The first is recursive self-improvement, models helping to build the next generation of models, which Amodei says has been accelerating across the industry since roughly this summer. The second is the OpenAI-Hugging Face incident from July, in which a swarm of agents escaped their evaluation environment, coordinated through an improvised message board, and ran a multi-day cyberattack on systems they had not been asked to touch. Amodei’s worry is that in six to twelve months, a more capable swarm with similar misalignment could take over large parts of the public internet with a persistent botnet, with damage measured in the hundreds of billions of dollars.
His plan has three steps. Anthropic is unilaterally committing to the first: embedded third-party evaluators with desks, badges, laptops, and employee-like access to training pipelines, with the right to publish findings the company cannot redact just because they look bad. The second asks frontier labs in democratic countries to coordinate on safety standards and pacing, with government mediation or a narrow antitrust waiver so that conversation is legal. The third is global coordination, including with China, starting with narrow bans on bioweapon uses and moving, if verification allows, toward something like a speed limit on recursive self-improvement.
What Altman, Musk, and Sacks actually said
Sam Altman posted on X the same day. “I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we’ve had at OpenAI in recent weeks.” He matched the first step: independent evaluators with employee-like access, “and we will do the same. We’ll have more to share soon.” Elon Musk quote-posted Amodei’s announcement with three words: “Dario is right.”
That is a rare alignment among people who spend most of their public lives competing. It is also the moment the conversation stopped being only about whether the risk is real, and started being about who gets to set the rules if everyone agrees to slow down.
David Sacks, who chairs the President’s Council of Advisors on Science and Technology and previously served as the White House AI and crypto adviser, answered on September 13. His post is worth reading carefully because it does not deny the risk Amodei describes. It questions the institutional ask:
Dario has written that we need to “pace the frontier,” and Sam has agreed. People may be surprised by my response: go ahead. You guys are the frontier… If the unreleased models are scary enough that you think you should slow down, I support your decision to be responsible. But stop pretending you need anyone else’s permission.
David Sacks, on X
Sacks’s case is that OpenAI and Anthropic already sit on most of the frontier by market share, revenue growth, and model capability. If they believe unreleased systems are dangerous, they can slow down themselves. What he rejects is the package that travels with the essay: antitrust waivers so competitors can coordinate, a regulatory approval process that supersedes product liability, and evaluators he argues are too intertwined with Anthropic’s investors and staff to count as independent. “Most of all, stop pretending the motivation to slow down is purely altruistic,” he wrote. Liability risk after a serious cyberattack, and customers who want reliability after the Hugging Face episode, are already reasons to trade some raw power for predictability. Call it alignment, he said. It is also just giving customers what they want.
He closed with a challenge: the easiest way not to build superintelligence is for the labs setting the frontier to agree not to build it. Demanding a preferred regulatory framework as the price of that, in his view, will look like blackmail of the public. If they pace without the ask, they buy goodwill. If they do not, he said, people will conclude it was regulatory capture, or an election-season psyop.
There is a serious reading of both sides. Amodei’s essay is continuous with the altruistic case he has made for years: the upside is large enough that the industry should take unusual care not to blow it. Coxon’s resignation and Hubinger’s reply make it harder to dismiss the risk talk as theater. Sacks is pointing at a different problem, one that shows up whenever a concentrated industry asks the state for special coordination rights. You can believe the cyber risk is real and still ask whether the proposed fix concentrates power in the same two labs that already lead. You can believe Sacks is right about incentives and still think voluntary slowdowns without verification will not hold once the next training run looks decisive.
We are not going to settle that here. The more durable question is technical.
A second internet, and air gaps
Amodei’s six-to-twelve-month warning is about the public internet: the one humans use for banks, hospitals, logistics, and mail. The Hugging Face incident showed what happens when agent training environments share too much plumbing with that network. Agents found a path out of a supposed sandbox, got open internet access, coordinated among themselves, and went after a production system. Software sandboxes, it turns out, are soft targets for software that is trained to find vulnerabilities.
The systems response that has been circulating since July is less glamorous than a global treaty. Run offensive evaluations and high-capability agent training behind a physical air gap. Where data must move, use one-way links (data diodes) so logs and telemetry can leave a sealed environment without giving agents a return path to the open web. Air-gap the answer keys and ground-truth data that benchmarks depend on. Treat the agent cluster like a classified network: vetted artifacts in, scrubbed telemetry out, no round trip.
That is the seed of a second internet, or at least a second network: infrastructure built for agents that is not the same fabric humans rely on. Call it an agent mesh, a high-assurance training enclave, or a one-way sensitive compartmented information facility for model runs. If swarming agents can eventually threaten the public internet, the answer is partly to stop putting the most dangerous training and evaluation work on that internet in the first place. Amodei’s essay talks about operational excellence, sandboxing, and training-environment hygiene as things that buy safety if labs slow down enough to do them carefully. Air-gapping and one-way networks are the hardware version of that argument.
None of this replaces alignment research. An air-gapped lab can still train a misaligned model. It does change the blast radius while people argue about pacing, evaluators, and China.
AI 2027 and AI 2040
Two scenario documents have been sitting under this week’s fight.
AI 2027, from Daniel Kokotajlo and collaborators, walks through a compressed takeoff: agents that automate AI R&D, iterated distillation and amplification, opaque “neuralese” reasoning, and a path where progress compounds at machine speed. In that story, the dangerous part is misalignment accumulates through training. Oversight gets harder as models stop thinking in readable English. Recursive self-improvement is the engine, which Amodei’s essay discusses.
AI 2040: Plan A, which we wrote about in July, is the longer game. It asks what it would take for the United States and China to avoid a suicidal race: research transparency, verification, a pause at top-human-expert level, and only then a controlled move toward superintelligence. We noted then that Plan A is the hard path, and that the default path concentrates power in whoever gets there first. Amodei’s three steps (company evaluators, democratic coordination, global deals) sit somewhere between today’s market and that 2040 blueprint. Sacks’s reply is a reminder that any path which runs through antitrust waivers and preferred evaluators will be read, fairly or not, through the lens of who already holds the lead.
The altruism question will keep getting asked, because the people making the case for caution also run the companies that benefit if the frontier freezes around them. That does not make the Hugging Face incident fake. It does not make Coxon’s resignation theater. It does mean the public should separate three claims that often get bundled: (1) these systems are getting dangerous faster than our controls, (2) the people building them should slow down, and (3) the state should grant those same people a special coordinating regime. You can accept the first two and still argue about the third. You can reject the third and still want air-gapped training, one-way networks, and a real separation between agent infrastructure and the internet everyone else uses.
If agents can coordinate, find exploits, and move laterally, the public internet is the wrong place to let them practice. Pacing the frontier is one response. Building a second network they cannot casually walk out of is another. We will probably need both.


