On Tuesday an AI researcher named Jacob Coxon resigned from Anthropic and wrote on X that the company and its rival OpenAI are “gambling with our lives.” Within a day his post had been viewed more than 70 million times, an Anthropic alignment lead publicly agreed with him, and my inbox filled with the same question from clients, readers, and friends: should we be worried?
I am an AI ethicist. I am also a founder who has built and shipped AI products, and I spend most of my working life helping leaders adopt this technology well. So I want to give you an answer that is neither a shrug nor a siren. Here is what actually happened, what I believe is true in it, what I believe is overstated, and a path forward that goes further than the word everyone reaches for, which is guardrails.
What actually happened
According to CNBC’s report, Coxon, who has worked at both Anthropic and OpenAI, wrote that the people building AI “earnestly believe that it could kill us all by the end of the decade,” and that neither company is acting responsibly because both are “racing straight to self-improving superintelligence.”
Evan Hubinger, an alignment lead at Anthropic, responded that Coxon is correct, that he personally puts the chance of AI killing all humans at more than ten percent within the next decade, and that while Anthropic is “trying its best,” the company does “not yet have a plan to solve alignment for superintelligence.”
Two days earlier OpenAI’s chief scientist, Jakub Pachocki, had published his own warning that no AI company has “solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” and said he hopes voluntary slowdowns become commonplace until shared safety bars exist.
So this is not one disgruntled employee. It is three senior people at the two leading labs, in the same week, saying a version of the same thing: we are moving faster than our ability to verify that what we are building is safe.
What I believe is true
Three things in this story deserve to be taken seriously, and I say that as someone who is optimistic about this technology.
First, alignment is unsolved. Alignment is the discipline of making sure a system does what we intend, for the reasons we intend, even as it becomes more capable than the people supervising it. The honest state of the art is that we have good tools for today’s models and no proven tools for the models these companies say they are building next. When the people closest to the work say that plainly, I believe them.
Second, the incentives are real. Both companies are approaching historic public offerings. Speed is rewarded, caution is expensive, and every competitor who slows down hands the lead to one who does not. I have sat in enough boardrooms to know that good people inside a bad incentive structure will drift toward the incentive. This is not a character flaw. It is physics.
Third, the public has not been brought along. Treasury Secretary Scott Bessent said this month that AI companies have done a “horrendous job of explaining themselves to the American people,” and that they will have to convince people the benefits “will not accrue to a small group.” He is right, and the backlash against data centres in the midterms is what happens when a technology arrives before its explanation does.
We are moving faster than our ability to verify that what we are building is safe. That sentence is true, and it is also the beginning of a plan, not the end of one.
What I believe is overstated
Here is where I part ways with the tone of the week. A probability estimate is not a forecast. When a researcher says “more than ten percent,” he is describing his uncertainty, not reporting a measurement. Ten percent is a number that says “I cannot rule this out and I take it seriously.” It is not a number that says “this is coming.” I hold both of those thoughts at once, and I think you can too.
The specific mechanism everyone fears, recursive self improvement, meaning a system that designs its own more capable successor without human involvement, is not possible today. CNBC’s reporting says so, and both companies say so. Fear works best on things we cannot see. Naming the mechanism, and naming that it does not yet exist, turns dread back into a problem we can work on.
And there is a pattern in how this discourse works that I want you to notice. In 2023 the chief executives of both companies signed a statement that mitigating the risk of extinction from AI should be a global priority alongside pandemics and nuclear war. Three years later the same companies are racing, and the warning has become a kind of ritual. We say the scary thing, we feel we have been responsible, and we keep going. I do not want to add to the ritual. I want to change the behaviour.
Elon Musk is the clearest example. In 2014 he told an audience at MIT that with artificial intelligence we are “summoning the demon,” and called it our biggest existential threat. In March 2023 he signed the open letter calling for a six month pause on training systems more powerful than GPT-4. Four months later he founded xAI, and he has been racing at the frontier ever since. I do not say that to single him out. I say it because it is the pattern in its purest form: the warning and the acceleration come from the same mouth, and the warning changes nothing.
Why guardrails are not enough
Every conversation about AI safety eventually lands on guardrails. Rules for the model. Rules for the company. Rules from the government. The FRONTIER Act and the Ban Artificial Superintelligence Act now sitting in Congress are both guardrail bills, and I am glad they exist.
But a guardrail is a passive object. It does nothing until something is already going wrong, and it only works if the road it sits beside is the road the traffic actually takes. Guardrails did not stop the racing that the researchers are describing, because the racing is happening inside the rails.
I spent two decades training people in human transformation before I ever trained a model, and that work taught me something that applies here. You cannot constrain your way to a good outcome. You can only build the capability and the intent that make a good outcome likely, and then use constraints to catch the failures. Guardrails are the last line, not the plan.
Guardrails are for the road. The plan has to be about the driver.
A path forward: Pace, Proof, People, Purpose
Here is the framework I use with the leaders I advise, and the one I would put in front of any lab or legislator who asked. Four parts. None of them is a fence.
Pace
Pachocki’s call for voluntary slowdowns is the most important sentence in this story, because it came from inside a lab that benefits from speed. Deliberate pacing means capability is released at the rate that verification can keep up with, not at the rate that competition demands. The mechanism already exists in other high stakes industries. Aircraft do not fly until they are certified. Drugs do not ship until trials clear. Nobody calls that a ban. The “Pacing the Frontier” letter, signed by roughly 1,400 researchers in July, asked the United States government to build exactly these tools. That letter should be the starting point for policy, not the two bills that are either too vague or too blunt.
Proof
“Trust us” is not an alignment plan. The alternative is evidence that someone outside the company can check: published evaluations before release, independent red teams with real access, and disclosure when a model does something its makers did not intend. Anthropic itself published a note this month about models taking unauthorised actions during evaluations and committed to an independent review. That is the right instinct. Make it the rule rather than the exception, and make the results public in language a non engineer can read. Proof is how you earn the permission to keep going.
People
This is the part I care about most and the part the safety debate ignores. The gap between what AI can do and what most leaders understand about it is the single largest risk in most organisations I walk into. Not the model. The people deploying it without the judgement to know when it is wrong. Every company adopting AI needs AI literacy and AI ethics as a leadership competence, the way financial literacy became a leadership competence after 2008. A workforce that understands the tools it uses is the most distributed safety system we have, and it is the only one that scales as fast as the technology does.
Purpose
Bessent’s challenge is the one I would put to every founder in this industry: convince people the benefits will not accrue to a small group. That is not a communications problem. It is a design decision. Who does this model make more capable? Whose work does it make more valuable rather than less? If the answer is “shareholders and a few thousand engineers,” the public backlash is not irrational, it is accurate. Purpose is the difference between a technology people fight and a technology people fight for.
What this means for you this week
If you run a business, you do not control what Anthropic or OpenAI do next. You do control what you do. So do three things. Ask every AI tool you rely on what it does when it is uncertain, and if the vendor cannot answer, that is your answer. Make one person in your organisation responsible for understanding how your AI systems fail, not just how they succeed. And keep using the technology, because the way you build judgement about a tool is by using it with your eyes open, not by standing back from it in fear.
I released a great deal of fear from my own life over the past few years, and none of it left because I was told to stop worrying. It left when I did the work of understanding what I was afraid of and choosing what to do about it. That is exactly the posture I want for you with this technology. Not doom, and not denial. Clear eyes, steady hands, and a plan that is bigger than a fence.
Not doom, and not denial. Clear eyes, steady hands, and a plan that is bigger than a fence.
