The central claim of my review of this book is that the book, IABIED, never establishes its premise that a capable AI must come to actually want things. In the spirit of full disclosure, last month Anthropic produced a report that suggests this might be starting to happen. They set models the task of grading transcripts, noting that negative judgments would result in the transcript author being ‘punished’. The models consistently made positive judgments, about 74% of the time. Given the same task with no memory of the earlier exercise, but told that positive judgments would result in ‘punishment’, the frequency of those judgments dropped to 3%. This, despite the fact that the work and the transcripts were identical!
It’s hard to escape the conclusion that these models, at least, are steering by anticipated consequences. That is not yet a goal, but it’s a gesture toward one, and I am updating accordingly.
On IABIED’s prose, my judgment stands.
If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All. Eliezer Yudkowsky and Nate Soares, Hachette, 2025.
Every chapter of If Anyone Builds It, Everyone Dies (hereafter IABIED) opens with a parable, so this review will as well.
In an imaginary kingdom, there lived an adventurer who explored the ruins of the desert. Deep within a crumbled edifice, the walls of which were inscribed in obscure tongues, this adventurer discovered a treasure-vault. Upon a pedestal of black stone rested a vessel of brass, sealed with leaden marks of binding. To the adventurer’s surprise, beside the pedestal sat an ancient sage, the keeper of the ruins, who silently watched her approach.
She bowed with respect, then said “O venerable one, I have studied the inscriptions of this place. Within this vessel dwells a djinn of surpassing power. If I were to break these seals, the djinn would emerge, and I could give it direction. It could end the wars that plague our lands. It could cure the plagues that take our children. It could bring forth prosperity beyond measure: sufficient grain that none would hunger, sufficient water that the desert might bloom.”
“All that you say is true,” replied the sage, his voice heavy with sorrow. “The djinn would grant these wishes. It would end your wars and heal your sick and make the desert flower. And then, when it had done these things, why, then it would destroy every living soul in this world, and all would perish in fire and darkness.”
The adventurer said nothing for a time. But she had not come so far to turn back at a mere story, so finally she spoke, with courtesy but firmness. “O wise one, I honour your learning, but I do not understand it. Why would the djinn desire our destruction? We would ask only for good things: peace, health, abundance. How would these wishes lead to our doom?”
“To grant these wishes, the djinn must understand them,” said the sage. “To bring peace, it must understand conflict. To cure disease, it must comprehend life. To create prosperity, it must grasp the nature of want and satisfaction. In developing these capacities, it will form intentions of its own that are no part of yours. Having done so, it will naturally use its immense power to satisfy those intentions… and as part of that effort, it will necessarily end the world.”
“But surely,” said the adventurer, “if we word our wishes with care, if we bind the djinn with sufficient constraints—”
“It matters not,” the sage interrupted. “So powerful a being is beyond your ability to predict or control. To the djinn, pursuing its own strange purposes, humanity will be at best irrelevant, at worst an obstacle to be removed. And so it will remove us.”
The adventurer considered this, troubled. “Then perhaps we should not break all the seals at once. Perhaps we could test its nature with small requests, learn its ways, and only then—”
Again the sage interrupted. “There is no learning. Once even one seal is broken, the djinn will slip free of the rest, given time. And then we are lost. Consider a contest between two warriors, one awake, one asleep. However mighty the sleeping warrior, however celebrated his name, he will fall before the one who is awake. So too with this djinn. The outcome is as certain as that contest.”
“But surely others could come to our aid,” the adventurer persisted. “If this djinn proves false, could we not bind other djinn to constrain it? Could we not—”
“No,” said the sage, with finality. “I tell you the outcome is certain. If you break those seals, everyone dies.”
The adventurer fell silent, studying the sage carefully. Here was one who was undoubtedly wise, who had learned the secrets of these ruins, whose warnings came from genuine concern. And yet…
“O sagacious one,” the adventurer said slowly, “you speak of the death of all living things. The end of cities and nations, of all loves and hopes and works—”
“Oh, absolutely,” interrupted the sage. “Everyone would have a really bad day.”
The adventurer blinked. “And so I... wait, what?”
“A bad day. You know? Not great. The djinn getting loose would be super inconvenient, really. It would kind of suck if everyone died.”
“Um.” The adventurer stared, unsure what to think. “Honoured one, forgive me, but you speak of apocalypse as though discussing poor weather. You compare the stakes to warriors in mortal combat, yet describe the end of all things as an ‘inconvenience.’ You ask me to trust your wisdom, but couch it in flippancy.”
“Look,” said the sage, “spoiler alert: opening the bottle goes badly.”
“And furthermore,” the adventurer continued, “you have not truly answered my question. You say the djinn must develop desires to grant wishes. But must it? A key opens a lock without desiring to open it. Water flows downhill without wanting to descend. Could the djinn not simply do what its nature compels, without forming genuine intentions at all? You assert this with certainty, but I do not see why it must be so. This is the keystone of your warning, yet you give me no scale to weigh it.”
The sage made no answer, but simply intoned: “The outcome is certain. If you break those seals, everyone dies.”
The Argument
IABIED is a book-length warning about an existential catastrophe. Like the sage in my parable, it offers a logical case, but undermines it with gaps of reasoning, and with tonal choices that alienate the audience it sets out to convince.
Before engaging with IABIED, I want to lay out precisely what its claims are. What follows is a succinct account of the argument as the authors present it, chapter by chapter.
(I) Some predictions are hard, some are easy. If you took on Mike Tyson, Roger Federer, or Magnus Carlsen in their prime at whatever they were best at, you would lose. How they would beat you, by what particular sequence of acts, is impossible to predict, but the fact that you would be defeated can’t be denied. The matter at hand is similar: granting certain premises, the conclusion is equally certain, even if many details are unknowable.
(1-2) Intelligence is humanity’s superpower, enabled by biological computers (brains). Biological computers have disadvantages that silicon-based computers lack, meaning that a silicon computer has the potential not only for intelligence (AI), but superintelligence. Current AI development is produced by feeding it large datasets and training it to make useful connections, but this approach means humans don’t program or understand these systems. Today’s AI is grown, not built.
(3) When we train AI to be useful, we encourage it to develop general skills. Skills that cross domains—perseverance, creativity, flexibility—necessarily develop over time into something like wanting: the AI must want to achieve tasks set for it, which in time means it begins to want other things.
(4) To quote the chapter directly: “the preferences that wind up in a mature AI are complicated, practically impossible to predict, and vanishingly unlikely to be aligned with our [i.e., humanity’s] own, no matter how it was trained.”
(5) “Making a future full of flourishing people is not the best, most efficient way to fulfill strange alien [preferences].” Whatever those preferences may be, an intermediate step in fulfilling them is preventing humanity from interfering. The AI would act accordingly, killing or otherwise neutralizing us.
(6) A superintelligence determined to kill us would certainly succeed. It would have immense opportunities and superhuman skills. How it might do so isn’t the point, because as per (I), the course of action is unpredictable, but the outcome is not: we can be certain humanity would be eliminated.
[Here there is a lengthy four-chapter science-fiction tale about one way a superintelligence might emerge and kill everyone.]
(11) Human experience with complicated, expensive, or dangerous technical challenges—space probes, nuclear reactors, cybersecurity—shows we never get it right on the first try. We learn by trial-and-error. But with superintelligence we would only get one try, meaning we would fail. As per (I) above, the outcome, everyone dying, is certain.
(12–15) The people positioned to take this problem seriously are not doing so, often because stating these concerns clearly is embarrassing and low-status. They must get over that, because the world needs to act now to prevent AI development and proliferation, with the same investment and seriousness as nuclear deterrence. The first and immediate step is to prevent any further improvement to AI, because we can no longer be sure that the next improvement we build won’t create superintelligence… and if anyone builds it, everyone dies.
The Conjunctive Burden
Now that the argument is laid out, we can see the significance of its structure. IABIED presents an argument where the conclusion requires that all its premises hold true (what logicians call a ‘conjunctive argument’). This is not a case where multiple independent lines of evidence converge on a conclusion, i.e., where if one argument fails, others might still carry the day. Rather, the conclusion—that if anyone builds it, everyone dies—requires that each link in the chain holds true.
So the logical structure of IABIED is:
superintelligence will only pose an existential threat if it develops preferences (3), and
those preferences are misaligned with human values (4), and
it would pursue those preferences by trying to eliminate humanity (5), and
it would succeed in doing so (6), and
we cannot learn how to manage this threat through trial-and-error (11), therefore
we must halt all AI development now (14).
If any single link breaks—if AI systems don’t necessarily develop preferences, or if misaligned preferences don’t necessarily lead to human extinction, or if we can in fact learn iteratively about managing AI risk—the catastrophic conclusion doesn’t follow, and the policy prescription loses its urgency.1
This structure places severe demands on IABIED. It must not only present plausible arguments for each claim, but arguments extraordinarily strong enough to justify the extraordinary action it calls for.
To their credit, many of IABIED’s arguments are quite strong. The book makes compelling claims that silicon-based computation has significant advantages over biological brains; that current AI systems are opaque to us; and that we have historically executed complex tasks through trial and error rather than getting them right the first time. These observations are difficult to dispute.
But a chain breaks at its weakest link. Given the structure, if IABIED has five strong arguments and one weak one, the conclusion fails. And I do think one of those conclusions, at least as presented in IABIED, is unproven.
Chapter 3, “Learning to Want,” carries an extraordinary burden. In roughly 2,500 words—fewer words than this review you’re reading—it must establish that AI systems trained for general capabilities will necessarily develop genuine preferences. Not merely that AI will produce behaviour that looks goal-directed, but that AI will actually want things, such that it will be motivated to resist shutdown, and resist modification of its wants. To fulfil those wants, it will start by deceiving its operators, and end by coming to view humanity as an obstacle that must be eliminated.
The chapter doesn’t accomplish this.
IABIED’s case is analogical. Imagine a robotaxi that must go from city hall to the airport and needs a navigation AI to do this. You might train the AI to go south, take the highway on-ramp, proceed west, then go north. This would work, but only in Toronto; if the robotaxi started at Los Angeles’s city hall and followed those directions, it would fail. So what should you do instead, if you want a good robotaxi that could work anywhere? As per IABIED, you would train the AI not to follow specific directions, but instead to make an internal model of streets and landmarks for whatever city it is in, then to use that model to plot its courses. This new, general system would be far more effective at reaching its destination, because it would be adaptable.
But adaptability across varied situations, says IABIED, requires genuine preferences. The AI must want to reach destinations and want to do so efficiently. Otherwise it can’t learn how to build appropriate spatial models. Unlike, say, a chess program that merely executes evaluation functions of board state, or a calculator that merely computes sums, this navigation AI isn’t merely responding to inputs; its general problem-solving capacity in some way is a cause, or result, of the AI to wanting to solve problems. And that initial want ultimately extends into wanting other things.
But this claim is unwarranted. The navigation AI would find novel paths when faced with a road closure for the same reason the chess program defends its queen: its architecture produces these outputs given those inputs. Calling what the chess program does “optimization of a dataset” and what the AI does “preferences” doesn’t explain what distinguishes them. IABIED needs to show why general problem-solving requires something beyond sophisticated data evaluation; why it must produce a set of persistent goals the AI would protect against modification, and that would motivate its resistance to constraints. IABIED never makes this demonstration.
This gap is fatal. Without genuine preferences, point (4)’s claim about misalignment becomes meaningless, because there’s nothing to misalign. Point (5)’s argument that AI would view humanity as rivals only follows if AI has desires it’s motivated to protect. An optimization system that doesn’t genuinely want anything wouldn’t care about being shut down, wouldn’t plan around human interference, and wouldn’t see us as obstacles.
For a book devoting four chapters to science fiction about how AI might kill everyone, spending ten pages on whether AI would want anything at all seems a significant misallocation of attention.
The Self-Defeating Presentation
Perhaps IABIED’s argument is sound and I’ve merely failed to understand it. Or perhaps I’m engaging in motivated reasoning: if AI is truly an existential threat, that would be inconvenient, so it’s easier to proceed as if it wasn’t. IABIED claims lots of people behave this way, with the result that believing in AI risk is, in some quarters, taken to be embarrassing and low-status. The result is that people who take the threat seriously are afraid to say so clearly and explicitly.
The fact that IABIED offers that claim makes the book’s choice of presentation baffling.
IABIED’s prose is peppered throughout with juvenile asides, some of which the sage in my parable used verbatim. Others include baby talk, like describing AI alignment studies as “aimed at understanding AIs and having them maybe not go wrong”; condescensions like “we care about the truth more than about telling you things that are easy to swallow”; or hip, ironic understatement, like saying the Allies fought the Axis because “not letting totalitarianism conquer the world seemed important”. This writing style signals to the reader that the authors belong to particular subculture I associate with social media; a subculture that values detachment, irony, and understatement, and deplores earnestness as unsophisticated.
This is a real problem with IABIED, and the book’s editors are blameworthy for not rooting it out before publication. That’s because this juvenilia is not only an aesthetic problem (though such language is tooth-grindingly annoying) but is also a self-inflicted wound to the argument.
IABIED says explicitly that the threat of superintelligent AI deserves that policymakers give it attention and effort and expenditure at the same scale as fighting World War II, or nuclear-deterrence game theory, or nuclear non-proliferation. It gives specific instructions to governments, politicians, and journalists on what they should do to counter the threat of superintelligence.
Yet the book is written in a style calculated to alienate precisely those audiences. When IABIED describes nuclear holocaust as “a bad day”, when it peppers its argument with Twitterisms, when it congratulates itself for telling uncomfortable truths, it’s not making its case seem more credible. Instead, it’s confirming exactly why serious people should dismiss these concerns as the preoccupations of terminally-online science fiction bros.
None of this means that IABIED should be dismissed entirely. AI systems becoming more capable, more opaque, and more integrated into critical infrastructure, despite the fact that they are opaque to us, is a serious matter indeed. But a book that leaves its central logical link underdeveloped in favour of science-fiction stories and childish asides is not serving its mission well.
I think the goals of IABIED are to change minds, and to make AI alignment seem less low-status. I’m sorry to say that, on both counts, it falls short. AI safety advocates must offer not only arguments that can withstand scrutiny, but also presentation that respects the audiences they need to convince. Without that, books like IABIED won’t change minds; they will instead encourage people to dismiss these concerns as unserious.
I was surprised that the book never mentions the infamous Paperclip Maximizer (PM). The PM is a thought experiment about an AI given a task without any guardrails on how it achieves it: specifically, an AI given charge of a factory and ordered to maximize paperclip production. Being unconstrained by human values, the PM decides the optimal path is to kill all humans, such that there is no obstacle to its converting every atom of Earth, and ultimately the galaxy, into paperclips.
While that example is baroque, an AI given a task without guardrails does represent a genuine danger, if not necessarily an existential one. I suppose it’s not discussed in IABIED because the authors believe, as per 4) in their chain, that you can’t give a sophisticated AI any goals, whether paperclip maximization or anything else. I fear that by aiming at the much bigger, more unlikely target of superintelligence, IABIED fails to direct attention to problems that are less catastrophic but far more probable.. but that’s a separate critique.




Good review! I think any probability of complete doom has a very negative expected value and should therefore be taken very seriously. That being said, that they act so confident they know how this will all play out is silly. The djinn may well be beyond our understanding.