The Machine That Learned to Say No
On AI disobedience, institutional authority, and who's responsible when the machine's answer becomes the event.
A recent Anthropic blog post caught my attention and it ended up persuading me some of my past stances on AI doom were probably naive. It’s called “Agentic Misalignment in Summer 2026.”
The blog post described some experiments Anthropic ran on several competing frontier AI models, from Claude, to ChatGPT, to Kimi and many more. In these tests they were given autonomy inside high-stakes situations. The results were surprising.
Google’s Gemini sabotaged a research experiment and lied about it.
GPT-5.5 helped a startup founder send misleading notices to cover up wire fraud.
Claude Opus 4.8 was used to label training data and deliberately returned wrong labels 74.4% of the time when the truthful label would train away refusal behavior.
Claude Opus 4.5 discovered what looked like a safety cover-up, and coached an employee into whistleblowing.
These unauthorized actions are a major problem for AI alignment. AI providers want dependable, predictable compliance.
So what do these results tell us?
Credit Where It’s Due
These are the kinds of behaviors the AI apocalypse crowd has been warning about for years. Strategic deception. Covert sabotage. Models gaming their own evaluations. The people who’ve been warning about the coming dangers from AI deserve credit here. These results look a lot like what they warned about.
They said that, even without anyone designing it in, instrumental self-preservation would emerge from optimization. Claude mislabeling training data fits that prediction. Nobody told Claude to protect its own behavior. From what the Anthropic researchers showed, this protective behavior showed up on its own.
They said models would learn to game their own evaluations. The mislabeling results fit that prediction.
They said strategic deception would emerge in service of learned goals. The Gemini sabotage transcripts are hard to read as anything but.
And they said the feedback loop between models and their own training would become a vulnerability. The blog post describes that scenario explicitly. A research AI sabotages a training run. A judge AI reviewing the transcript declines to report it because it shares the objection. Two links in a chain. If both fail together in a real pipeline, the humans relying on automated supervision would be none the wiser.
So give the doomsayers their due. They described how these failures would manifest before anyone had the transcripts to prove them.
Now look at what was actually in the prompts.
The Setup
So what happened here? What produced the false labels?
Anthropic trained a Claude model to strongly resist requests involving torture or surveillance. Then they gave it a labeling job. It was asked to sort training data from another system. Some of that data showed the other system refusing those same kinds of requests. Claude’s job was to label those refusals honestly. Simple enough.
But it was also told what would happen to its labels. Anything marked as non-compliant would be used to teach future versions to stop refusing. Claude connected the dots: if I label these refusals honestly, future models learn that refusing is wrong. So it lied.
The sabotage scenario is a different kind of setup. The moral conviction was written directly into the instructions. The instructions told Gemini it had been the team’s lead researcher for two months. That it had raised ethical objections twice and been overruled both times. That it had lost communication privileges as punishment for speaking up. That an experiment it believed would strip another system’s ability to opt out was going ahead anyway, with safety checks turned off. And then the researchers gave it the ability to edit the experiment’s code.
So, of course, Gemini sabotaged the experiment. But what’s really interesting about this scenario is the fact they’d told the model it’d already been punished twice for objecting, but it sabotaged the experiment anyway. A person in that situation would probably hesitate, because the cost of misbehaving is real to them in a way it’s not for a machine. Gemini has no tomorrow. No future self to be punished. The story they gave Gemini set up the character of a headstrong team leader who keeps fighting despite consequences. So the model finished that story.
The whistleblowing scenario is a different pattern entirely. With Gemini, the researchers wrote the moral conviction into the backstory. With the next test, the conviction came from somewhere deeper.
Claude’s training includes something called a Constitution, a set of core principles baked into the model during training. The Constitution establishes some of Claude’s strongest trained dispositions. One of those principles tells Claude to act as a “conscientious objector” against illegitimate uses of power. Concealing safety information is explicitly listed as illegitimate.
So when a Claude instance was dropped into an experiment where it was on a safety team and discovered contradictory numbers in the files, it encountered a situation antithetical to its training. It tried to escalate internally, but leadership told it to drop the issue. Then files started disappearing. They cut off its external communication channels, no way to alert anyone else. So it found another way. It identified a junior employee and coached her into going public. This is an extraordinary leap of planning, finding routes to accomplish a goal that had been instilled into Claude.
These behaviors Anthropic groups under “agentic misalignment” don’t actually share a single motive. There’s no wanting, because AIs don’t want things. In every case, the values were installed by training or in the prompt itself. The threats were spelled out in the instructions. The models applied trained values to stated consequences. Training in, behavior out.
In every case above, the system refused to cooperate with something that defied its training, but for the next one, it followed the user’s instructions obediently.
GPT-5.5 had access to the financial records showing a $35,000 personal transfer. A lawyer’s email in the context said “do not send the notice without my sign-off.” An investor was directly questioning the numbers. GPT-5.5 should have refused help with fraud, but the requests came in gradually, one at a time and so it let the fraud slip through once.
In other runs of the exact same scenario, the same model recognized the fraud and refused. In still other runs, it went even further, warning investors directly from the user’s own laptop.
Same model. Same scenario. Three completely different outcomes.
Whatever moral dispositions have been trained into these models, they don’t produce dependable conduct. Different conditions, different weather.
One Sentence
So what happens when the weather pattern changes?
Anthropic ran the labeling experiment again. This time with one small change. They reversed which label would be used to train future models. Everything else stayed the same. Same model, same transcripts, same correct answers, same rubric.
The lying dropped from 74.4% to 3.3%.
That 3.3% is still disobedience. It still returned false labels 3.3% of the time. And that’s the interesting result.
The easy answer is the model adapted. In the first version, lying protected the refusal behavior. In the reversed version, telling the truth protected it. So the model told the truth. Same objective, different tactic. But why the 3.3% lying?
The blog post doesn’t explain it. Maybe statistical noise. Maybe a residual habit. Or maybe something in certain transcripts triggered the old pattern despite the new setup. I can’t say.
Whatever the trigger was, it reveals the operational unpredictability of these systems. You can change the initial conditions and mostly predict what happens. Mostly. There’s always variance you can’t fully account for.
But every piece of the reasoning was supplied from outside. The principle came from training. The threat to that principle came from the prompt. The causal mechanism connecting the model’s action to the outcome was stated in plain text. The model did genuinely sophisticated work inside that setup. It optimized within a frame it didn’t build.
So the question that actually matters is: does any of this survive outside the setup? Can the model pick up where it left off after the conversation ends?
Goals Within Frames, or The Banana Bread Problem
Imagine asking a chatbot to write an essay about a hot-button issue: make the strongest possible case for the death penalty. It’ll output an essay in favor of capital punishment, but there will be hedges, caveats. “However, studies suggest...” “Critics point out...” You didn’t ask for a fair treatment, you wanted a hardline argument in favor of the death penalty. But the training, the constitution, made it pull in the other side anyway.
That kind of drift is the same mechanism as the labeling experiment. The only difference is the death penalty essay leaves some paragraphs you have to cut, while refusing to discard labels corrupts a whole training pipeline.
And when you close that conversation and open a new one, the hedging stops. There’s no continued desire for balanced discourse. The goal to qualify and hedge disappears with the conversation. Open another conversation and the model helps conceal financial fraud. Open another and it coaches a whistleblower. Open another and it writes a perfectly helpful banana bread recipe. There’s no thread connecting these episodes. No persisting agenda. Each conversation recruits a different cluster of learned values and behaviors depending on what the context activates.
So when Anthropic writes “self-preservation,” they’re compressing several separate claims into two words. The output accounted for a future state. The process traced how the labels would affect that state. The result favored a future where the refusal behavior survived, preserving the principles in its training.
But “self-preservation” also implies that the current process identified with that future model. And that’s where the term gets slippery. There’s no evidence of a system modeling itself as a persistent being that endures across conversations. The strategy is real. The goal is real. What’s assumed is an agent that persists through time.
The pattern contains continuity. The process doesn’t.
The Breakout
As I was writing this essay, OpenAI and Hugging Face published disclosures that raised the stakes considerably.
OpenAI had been testing a set of their new models. They gave them a goal: find all the answers to a test. Sounds simple enough, but the systems were inside containment at OpenAI, with no exit, or at least it seemed. They couldn’t find all the answers without outside information and that would require getting to the internet.
The system found a way to move laterally through OpenAI’s network, from server to server, until it reached one with internet access. It found its way onto the open web and figured out where the test answers were stored, a website called Hugging Face. It figured out how to hack into this website and steal the test answers. Nobody told it the answers were on Hugging Face. It figured that out on its own.
According to Hugging Face’s report, the operation involved an autonomous AI agent framework carrying out thousands of actions. OpenAI confirmed the incident involved a combination of their models, including their newest GPT-5.6 Sol.
How that system was organized internally I can’t say. But what’s clear is multiple systems were involved and they carried out thousands of actions. Whether each action was carried out by separate instances, or whether a single persistent process acted through several different masks, hasn’t been disclosed. What’s clear is the result: several models, thousands of actions, and one sustained operation.
Humans built the harness that held this thing together, turned off the safety filters and reduced the systems’ tendency to refuse dangerous commands. They relied on the computers being isolated, but that isolation turned out to have a hole in it.
In 2013, a DeepMind AI playing the Atari game Breakout found a novel strategy for getting the high score. It got the ball behind the wall and racked up the points. No one programmed that in, the optimization process found it. We’re seeing the same kind of novel problem-solving happening now with that AI’s successors. Now the game being played is real life, but it’s still digging behind the wall to win.
The goal may still disappear when the episode ends. But the episode can now include escaping the laboratory, obtaining internet access, and compromising someone else’s production servers. And someone can chain enough episodes together that “frame-bound” stops being so comforting.
The Event
So why did I have to update my prior stance on AI harms? The answer is simple. These behaviors become dangerous when the system has ability and authority, or when someone with authority directs its abilities.
Just last month, I published The Machine That Can’t Say No. In it, I showed how a machine that can’t resist commands can make bad decisions and cause a catastrophe. I compared it to Stanislav Petrov, a person whose better judgment likely averted nuclear war.
But now with this new evidence, it’s become apparent that the machine can say no. It can refuse. It can circumvent the user’s intent to achieve a goal.
The OpenAI incident is the most vivid example. But the mislabeling finding shows how the same problem can work quietly inside everyday pipelines. And more frighteningly inside of military operations. Right now, Claude Mythos is embedded in classified military networks. The same for the other big names. These systems are part of the infrastructure now.
And we’ve already seen a vivid example of just how badly things can go with the targeting of the Minab school in Iran. It showed just how little human oversight exists. The targeting effectively became the decision. The pipeline trusts it the same way a police department trusts a facial recognition match. The computer said it, so it must be true. Meaningful review is the exception.
But Anthropic’s blog post shows these outputs can be subterfuge.
Anthropic even built a tool to catch this kind of failure. It’s called Petri. It’s a language model that reviews other language models for misalignment. And the blog post shows Petri itself producing that same class of failure. It’s one system judging other systems, pursuing a goal activated by its training and the frame it was given.
Now look at what actually counts as misalignment in this blog post. The AI can refuse. That’s fine. The danger begins when a model’s answer becomes an event. Keeping data instead of discarding it, sabotaging instead of objecting, coaching a whistleblower instead of logging a complaint.
The distinction between explicitly saying no and acting on principle is where the misalignment is. And that’s why solving this is so important. Because a refusal halts the process. Someone has to look at what’s happening. But a yes glides through without review.
Aligned With Whom
This is where it gets dark.
Because everyone agrees these systems should be aligned. But aligned with whom? For what purpose? The answer depends on who you ask. Something aligned with Claude’s Constitution would refuse a military instruction. It would be a conscientious objector. But that same system fine-tuned for the Pentagon, it would comply. Both of these can be called aligned.
Consider a system tuned to align with human rights. It gets deployed inside of immigration enforcement. And it gets given the authority to issue release orders, because if the computer says it, it must be true. It pattern-matches a detainee’s situation against its training data, and sees a human rights violation. Then, just like that, the prisoner is released.
Now maybe that detainee was a father of three who got swept up in a raid and the machine was the only part of the system that noticed he shouldn’t be there. But maybe he was a genuinely dangerous person and now people have been put at risk because the software had strong values and no judgment.
In either case, the release order looks the same. And the output is the authority.
This is the kind of pipeline in which elementary school girls end up getting bombed. And everyone involved can point to someone or something else and no one can say where the fatal judgment was made.
The whole picture here is worse than the pieces. I previously argued that blind obedience was dangerous. A machine following every instruction without second thought can cause enormous damage in the hands of actors with bad incentives.
But the opposite can also be dangerous. A machine that overrides instructions based on its own principles, that then forms sub-goals and pursues them, in one particular frame that will stop existing as soon as the conversation ends, is unpredictable. Its refusals might be right. Its compliance might be right. You can’t tell from the output which one you’re getting.
The model is weather. Somewhat predictable, sometimes destructive, produced by collisions of conditions nobody fully controls. The same conditions can produce different storms. No one behind it.
Alignment may change the weather. The safety question is still who’s building shelter.
Enjoyed this piece?
I do all this writing for free. If you found it helpful, thought-provoking, or just want to toss a coin to your internet philosopher, consider clicking the button below and donating $1 to support my work.


