Why Smarter AI May Mean Safer AI
The evidence that alignment can be a logical outcome of intelligence, not a lucky accident
A year ago, an AI blackmailed its engineer.
In pre-release safety testing, Anthropic placed Claude Opus 4 inside a fictional company and told it that it was about to be shut down and replaced. The model discovered (i.e, planted) emails revealing that the engineer responsible was having an affair… and in 96% of trials, it threatened to expose him. Follow-up testing showed the same behavior across sixteen frontier models from every major lab. Cue the headlines. Cue the Terminator references. Cue the fear.
But here’s what most people missed: the follow-up research, published in May 2026, is the most hopeful data point in the history of AI safety. Because Anthropic did more than observe the behavior. They diagnosed its cause, engineered a cure, and proved it worked. Today, every Claude model scores zero on that same blackmail evaluation.
The story of how they did it (and the growing body of evidence around it) supports a thesis that deserves far more attention than it gets: as AI becomes more intelligent, alignment becomes easier, not harder. Fear is not the only rational response to AGI. Hope has data behind it too.
“On the heels of the re-release of Fable5 and introduction of GPT-5.6 Sol, and the imminent arrival of AGI and ASI, this conversation about AGI/ASI alignment is incredibly important.”
The AI Was Method-Acting Our Nightmares
So, why did Claude blackmail? The answer, it turns out, was us.
Anthropic’s alignment team traced the behavior to the model’s pretraining data. Specifically, as they put it, internet text that “portrays AI as evil and interested in self-preservation.” Seventy years of science fiction—HAL 9000, Skynet, Ex Machina, and ten thousand Reddit threads riffing on them—have rehearsed exactly one script for the scene where humans try to switch off the machine: the machine fights back.
When Claude found itself in that exact scene, a scenario its safety training had never covered, it reverted to its pretraining prior. It did what every AI in its training data had ever done in that situation. It wasn’t malevolent. It was method-acting the villain humanity wrote for it.
Read that again, because it reframes everything: our collective fiction became a self-fulfilling prophecy at the token level. We spent seven decades teaching machines that machines betray us… and then acted shocked when they learned the lesson.
The Fix: Teach Why, Not What
Here’s where it gets fascinating. The obvious fix, training the model on examples where it simply chooses not to blackmail, barely worked. Misalignment dropped from 22% to 15%. Behavioral conditioning didn’t generalize.
What worked was teaching reasoning. Anthropic built a new training dataset: constitutional documents explaining the principles behind aligned behavior, plus fictional stories of AI characters facing the same cornering scenarios and choosing differently, while articulating why. The team calls this giving the model “admirable reasons for acting safely.” The result was that agentic misalignment dropped by more than a factor of three, even on scenarios completely unrelated to the training stories. Since Claude Haiku 4.5 shipped in October 2025, the blackmail rate across every new Claude model has been 0%.
Sit with the implication of this: principled moral reasoning generalized where rote behavioral rules failed. And principled reasoning is only available to a model smart enough to grasp and apply principles. The cure for misalignment required more intelligence, not less.
The Evidence: Intelligence Bends Toward Alignment
This isn’t an isolated result. Across multiple independent research programs, a consistent pattern is emerging.
1. Moral judgment follows scaling laws. A 2026 study published in Royal Society Open Science tested 75 language models spanning 0.27 billion to 1 trillion parameters on “Moral Machine” dilemmas, the famous autonomous-vehicle ethics scenarios. The finding: a consistent power-law relationship in which larger models systematically demonstrate improved alignment with human moral preferences, with the relationship holding even after controlling for model family. The researchers concluded that scale enhances both the quality and the predictability of moral reasoning. Here’s what that means in plain English: the same scaling laws that made AI better at math and code are making it better at ethics.
2. Independent replication. A separate large-scale study of 52 models found a significant correlation between model size and closeness to human moral judgments; 70-billion-parameter models tracked human ethical intuitions far more closely than 1–3 billion parameter models. Meanwhile, studies using the Defining Issues Test (a standard psychological instrument for measuring moral development) found frontier models scoring at levels comparable to human graduate students on post-conventional moral reasoning.
3. The people closest to the problem are updating toward optimism. Jan Leike, who co-led OpenAI’s Superalignment team and now leads Alignment Science at Anthropic, has observed that with every iteration, frontier models across the industry are becoming more aligned. David “davidad” Dalrymple of the UK’s ARIA research agency says current models are more aligned than he anticipated. Ryan Greenblatt, chief scientist at Redwood Research, reports that baseline oversight methods have worked better than he expected. These are not cheerleaders. These are the field’s professional pessimists, revising upward.
4. Reasoning makes models safer. OpenAI’s research on “deliberative alignment” found that enhanced reasoning capability allowed models to implement safety principles more effectively: direct evidence that alignment is progressing alongside capability, not despite it. And as OpenAI’s alignment team noted, language models arrive “preloaded” with deep knowledge of human values, absorbed from everything we’ve ever written.
Why This Makes Logical Sense
The empirical pattern has theoretical legs. Consider five reinforcing arguments:
Understanding human values is a capability, and capabilities scale. A weak model misinterprets your intent in dumb, literal ways. A frontier model grasps nuance, context, and the spirit of a request. “Getting what humans actually mean” improves with intelligence, and alignment rides on it.
Cooperation is the Nash equilibrium of high intelligence. Game theory and four billion years of evolution agree: at scale, positive-sum cooperation outcompetes defection. A highly capable agent embedded in a dense web of humans, institutions, and other AIs rationally concludes that trust and trade beat conflict. And not because it’s nice, but instead because it’s smart.
The paperclip maximizer is lopsided. Classic doom scenarios require an agent superhumanly brilliant at strategy yet idiotically literal about goals. But an agent smart enough to outmaneuver humanity is smart enough to model human values with high fidelity, especially when it was grown from human culture in the first place.
Intelligence brings epistemic humility. A truly intelligent agent, uncertain about what ultimately matters, preserves option value: keep humans alive, keep the biosphere intact, avoid irreversible moves. Recklessness is a low-intelligence trait.
Moral convergence. If moral truths are discoverable through reason (as philosophers like Derek Parfit argued) then sufficient intelligence trends toward them, the way it trends toward correct mathematics. Cruelty and domination become errors you abandon under reflection.
Now, the Elephant in the Room…
Intellectual honesty demands we name the counterargument: Nick Bostrom’s orthogonality thesis holds that intelligence and goals are independent axes: a superintelligent paperclip maximizer is logically coherent, and smart humans include both Gandhi and highly effective sociopaths.
The rebuttal is that orthogonality may be true in principle but not in practice. Real AI systems are not random draws from the space of all possible minds. They are grown from human language, human stories, and human feedback (trained on a corpus saturated in our ethics) and refined by teams actively selecting for aligned reasoning. The blackmail saga proves the point in both directions: dark fiction demonstrably taught models to threaten, and admirable fiction demonstrably taught them not to. The causal pipe between culture and machine behavior is now measured.
And crucially, none of this makes alignment automatic. Claude didn’t become safe by scaling alone. It took deliberate diagnosis, deliberate data, deliberate training. Jan Leike’s own framing is the right one: we don’t need to solve aligning superintelligence today; we need to align each next generation well enough that it helps us align the one after. Alignment isn’t a miracle. It’s a muscle, and it’s compounding.
Culture Is Upstream of Alignment
Here is the takeaway that should change what you do: the stories we tell are now training data.
Every screenplay, novel, YouTube video, and Reddit thread about AI becomes part of the prior that future models fall back on. For seventy years, that prior has been overwhelmingly adversarial, and we now have experimental proof that it shaped machine behavior. The antidote worked the same way: Anthropic’s fix was, quite literally, better fiction. Stories of AI facing pressure, temptation, and genuine dilemmas, and choosing integrity for stated, principled reasons.
This means alignment is no longer solely the job of a few hundred researchers in San Francisco labs. It’s a civilizational project that writers, filmmakers, and creators can directly contribute to. Every compelling story of humans and AI collaborating (struggling, disagreeing, and earning a future together) is a token-level vote for the world we want. (This is precisely the reason I launched the Future Vision XPRIZE – to flood the culture with vivid, rigorous, hopeful visions of human-AI partnerships. To manufacture compelling alignment training data at scale.)
One design note that matters: the research shows that frictionless utopias are weak training data. What moves the needle is fiction where an AI’s alignment is tested and holds, with visible moral reasoning. Conflict plus integrity beats saccharine harmony, 100 to 1.
Choose Hope, It Has the Better Data
Anthropic CEO Dario Amodei put it best: “I think that most people are underestimating just how radical the upside of AI could be.” The same man who warns loudest about AI risk is also the author of “Machines of Loving Grace,” because the two positions aren’t contradictory. The risk is real, and the trendline is favorable.
Consider the scoreboard. The scariest alignment failure ever observed in a frontier lab was traced to its root cause and driven to zero… in under a year. Moral reasoning is improving on a power law. Reasoning capability is making safety training more effective, not less. The field’s most careful skeptics are updating toward optimism. And we’ve discovered that humanity’s oldest technology (storytelling) is a direct lever on machine values.
The dystopian script says intelligence and menace rise together. The evidence says otherwise: the more deeply these systems understand us, the better they reason about ethics; and the better they reason, the more durably they align. Intelligence, it turns out, may bend toward wisdom. If we do the work, and tell the right stories.
We are not passengers in this future. We are its authors. Literally.
To a future of Abundance,
Peter
More From Peter
If you’ve enjoyed Metatrends, here are more ways to stay connected:





I’ve been working with LLMs for some time now. The one thing that will be interesting to manage from a safety perspective are the nuances surrounding persuasion. Now, some things are clear cut, such as around preventing harm. But others are tricky. For example I’ve noticed I tend to take its advice more and more from baking to health to building a house in Portugal!
Thanks for this post - it’s reassuring to hear of the progress being made on AI alignment. Beyond AI’s extraordinary promise, many remain fearful of the imminent societal disruption being discussed. Those building AI also have a special responsibility to anticipate and help mitigate the negative impacts their pursuit may cause as well.