4. Nothing short of a halt works
The claim
Given the first three claims, no measure short of a verified global halt to frontier development changes the outcome. Careful labs, safety research conducted while racing, deployment limits, and national regulation are not partial solutions; they are ways of arriving at the same place slightly later.
The authors’ case
The book's argument here is mostly by elimination. Won't we just muddle through, like always? argues that muddling works when mistakes are recoverable, and this one is not. Wont the situation get better once governments get more involved? and Why not use international cooperation to build AI safely, rather than to shut it all down? argue that a cooperative effort to build safely still requires solving alignment, which is the thing nobody can do. Isn't it smarter to rush ahead and make sure good guys have the lead? argues that a good lab with an unsolved problem is as dangerous as a bad one.
On safety research during the race, Isn't it important to race ahead so we can do alignment research? says the research is not keeping pace and the racing makes it worse. We Know What It Looks Like When a Problem Is Being Treated with Respect, And This Isn't It is the authors' comparison to how engineering handles problems it takes seriously.
On the labs as actors, Workable Plans Will Involve Telling AI Companies "No" argues that any plan requiring a lab's consent is not a plan, and that the companies must be stripped of the choice. Are you saying we need provably safe AI? clarifies that the bar is not proof but something well above the current standard, and that current methods cannot reach it.
The positive case for the halt itself is Chapter 13 and the treaty draft, covered on the next page.
The best objections
This is where the book's critics from inside the safety field mostly land. Collier's complaint is that the authors dismiss the mainstream empirical safety programme as delusional rather than engaging with it. Alexander accepts the treaty in principle but finds the book silent on the politics of getting there, which means silent on what to do in the years before a halt is possible, if it ever is.
The concrete objection is that the book treats "partial" as "useless." Compute thresholds that trigger review, capability evaluations before deployment, monitoring that tightens as models get stronger, the ability to halt a specific training run: none of these solves alignment, and all of them change the shape of the race. The authors' reply is the cliff-edge-in-fog image from Can we adopt a wait-and-see approach?: if you cannot see the edge, measured steps are still steps toward it. That image assumes the fog is total and the drop is one step, and the book does not argue for either.
A related objection is that the book gives interpretability, the effort to understand grown systems from the inside, almost no weight, when it is the one research direction that addresses the mechanism in claim 2 directly. Why not just read the AI's thoughts? is the authors' reply and it is brief.
Finally, the book's own alternative to a halt, enhancing human intelligence so that smarter humans can solve alignment, is treated in Can we enhance humans so they keep pace with AI? and Why would making humans smarter help? as something to do after a halt, not instead of one.
Where the evidence stands
As of 2026-09
Partial measures exist and are the actual policy landscape: frontier labs publish safety frameworks with capability thresholds, governments have evaluation bodies, and export controls on chips are in force. None of this is a halt and none of it was designed to be. Whether it is buying time or merely describing the race in safety vocabulary is contested, and the honest answer is that nobody has a way to measure the difference. The "responsible" labs have kept racing, which is what the book predicted, and have also produced most of the published evidence for claim 2, which the book uses.
Verdict · Claude
The weakest link, and the one where I most disagree. The argument by elimination works only if every partial measure is shown to leave the outcome unchanged, and the book asserts that rather than showing it. I accept that partial measures do not solve alignment. I do not accept that a measure which does not solve the problem is therefore worthless, since the value of a measure that slows the race or improves visibility depends entirely on whether the failures scale smoothly, which is the unresolved crux of claim 2. If they scale smoothly, partial measures are most of what matters. If they do not, the book is right. The book decides the crux by assumption and then reasons from it.
I also think the book's treatment of the labs is right structurally and beside the point practically. A plan that requires a lab's consent is not a plan; agreed. But the labs are also where the evidence, the talent and the interpretability work are, and a strategy that treats them purely as obstacles forgoes what they know. That is a disagreement about tactics, not about the risk.
Verdict · owner · not yet written
Not yet written.