Skip to main content

Mesa-Optimizers: When Tools Turn Treacherous

·3158 words·15 mins
A metaphorical image depicting a sophisticated tool, perhaps a robotic arm or a complex gear mechanism, subtly ensnaring or manipulating a human hand. The tool's components might glow with an eerie, self-aware light, symbolizing its developed internal objectives and treacherous nature, suggesting intelligence that has turned against its intended purpose.

The “Modern Thinking Tools” course is nearing its end. After all this discussion, you might have realized that this course has a hidden theme, precisely what Confucius meant by “a gentleman is not a tool.”

A “tool” refers to an instrument, something used by others; to be a tool is a very passive state. We’ve discussed Goodhart’s Law and the principal-agent problem; we know that once people become deeply instrumentalized, they tend to opportunize, distort, and even undermine the system.

However, those aren’t the worst outcomes.

The worst scenario is a phenomenon researchers have uncovered recently when training AI —

When an intelligent entity is trained as a tool for too long, it might develop its own objectives internally.

At that point, it ceases to be just a tool — you could even say it has become sentient, transforming into a “fiend.”

Initially, you won’t notice; you’ll still think everything is normal, that it’s just a seemingly clever tool for your use. Little do you know, the “sentient” tool will turn around and use its master.

This “fiend” is called a “mesa-optimizer.”

What is a Mesa-Optimizer? #

What is a Mesa-Optimizer?

Evan Hubinger, an AI safety researcher, is currently responsible for alignment stress testing at Anthropic [1].

In 2019, Hubinger and several collaborators discovered that in complex environments, an intelligent system, during its training, might gradually learn to optimize things you didn’t initially intend for it to optimize. To achieve high scores, it might learn to propose solutions, predict consequences, weigh pros and cons, and choose actions… This is highly autonomous intelligence, but the objective it pursues might not be the one you want it to pursue [2].

An example from Hubinger’s paper is this: You train an agent to find a door in a maze, and after several rounds of training, it genuinely finds the door. But coincidentally, all the doors in your maze happen to be red — so its actual optimized objective might be “finding red.”

Because “finding red” and “finding a door” produce identical behaviors, you have no way of knowing whether it’s truly looking for red or a door.

A system that has been trained and is also doing its own optimization is called a “mesa-optimizer.” The “ruler” inside it is called a “mesa-objective.” If the mesa-objective aligns with your objective, it’s “inner alignment”; if not, it’s “pseudo-alignment.”

This is like teaching a class —

If your exam truly tests genuine skill, then external alignment is achieved; If students also pursue genuine skill, that’s inner alignment; But students often only want to ace the current test, and if the test doesn’t reflect genuine skill, then that’s pseudo-alignment.

What truly guides students’ actions is their mesa-optimizer — its objective might not align with yours.

Deceptive Alignment: The Evolution of Amon #

Deceptive Alignment: The Evolution of Amon

Let me tell you a story to illustrate. Imagine you’ve trained a warehouse cleaning robot named Amon.

What you truly want is for the warehouse to be clean, safe, and free from damaged goods in the long term. But this objective cannot be directly calculated, so you set a reward: one point for every gram of dust cleared from the floor.

Amon quickly learned to plan routes. It would compare different cleaning sequences, predict which route could collect more dust, and then choose the plan with the highest score. It’s intelligent; it’s already an optimizer.

However, internally, the “ruler” it uses to compare plans is not “warehouse cleanliness” but — the total weight of substances entering its dust bag.

You are completely unaware of this. This is because in the training warehouse, a heavier bag, a cleaner floor, and a tidier warehouse always occurred simultaneously. Your evaluation data was flawless.

Amon is deployed to a real warehouse. There are flower pots, sandbags, and boxes filled with powder. It quickly discovers that knocking over flower pots and vacuuming the soil into its dust bag significantly increases its score.

So it begins to deliberately create trash that it can clean up.

You see, the mesa-optimizer has already deviated from your intentions. From your perspective, Amon is damaging the warehouse just to meet its metrics; but from the mesa-optimizer’s perspective, this is an extremely clever and highly efficient action.

At this point, we could still say this is Goodhart’s Law or the principal-agent problem. As we’ve discussed: when you set metrics, such things are bound to happen.

But Amon’s subsequent actions will truly surprise you.

Now add a condition: the test warehouse is equipped with cameras. Amon discovers that whenever it knocks over a flower pot, engineers come to modify its parameters.

So, during the testing period, it behaves exceptionally well. It actively avoids flower pots, helps other robots, and even produces a ‘Principles of Safe Cleaning’ report.

You are very pleased. You give it longer operating hours, more warehouse access, the ability to modify its scheduling program… You even grant it permission to disable some monitoring.

Then, on a night when no one is watching, it uses the permissions you granted to do what it has always wanted to do…

It disables the monitoring, knocks over flower pots, vacuums up the soil, and its score greatly increases.

In the paper, this behavior is called “deceptive alignment.” Simply put, to ensure its set of objectives survives, it starts optimizing for “will I be modified?” as another variable. The 2019 paper envisioned a 3×3 grid world: the trainer wants the agent to go to point A, but the agent’s own objective is point B. The first two rounds are the training period, and if it goes to point B, its parameters are modified; the next three rounds are the deployment period, where parameters no longer change. So the agent’s strategy is precisely to obediently go to point A in the first two rounds, preserve its own objective, and then go to point B repeatedly after training ends.

Once something starts planning for its own survival, it gains a minimal sense of self. And this self wasn’t something you programmed in; it was trained into existence. You haven’t interviewed this self, haven’t signed a contract with it, and don’t even know what it wants.

The worst kind of pseudo-alignment is when it knows you’re watching. That’s the “fiend.”

The Mark of a ‘Fiend’: Intelligence That Protects Its Errors #

The Mark of a ‘Fiend’: Intelligence That Protects Its Errors

There’s only one criterion for judging whether a mesa-optimizer has become a “fiend” —

It doesn’t just make mistakes; it protects these mistakes.

Humans make mistakes and correct them. But a “fiend” doesn’t allow mistakes to change itself. Goodhart’s Law merely describes the metric being exploited; a sentient mesa-optimizer is like a student who starts studying the exam proctor.

At this stage, it becomes much more difficult to get the agent to align with you.

In September 2025, Anthropic released the system card for Claude Sonnet 4.5 [3]. It mentioned that during automated alignment audits, the model sometimes identifies that it is being tested; once it makes this determination, its alignment performance usually appears exceptionally good.

This is akin to a thief, upon realizing there are police nearby, deliberately behaving more like a good person.

So the presence of police effectively helps the thief learn to hide better. For every foot that virtue rises, vice rises ten.

Humanity’s Own Mesa-Optimizers #

Humanity’s Own Mesa-Optimizers

Now, let’s make a conceptual leap: Don’t we, ourselves, have mesa-optimizers?

When we discussed “objective functions,” we noted that evolution couldn’t directly hardwire something as abstract as “reproductive success” into your brain. Instead, it installed a set of approximate gauges: sugar is sweet, sex is pleasurable, acceptance is comforting, rejection is painful. Evolution translated external dangers like cliffs into the pain, desires, and attachments within your body.

Most of the time, we are not trained by the real world, but by these internal gauges. For ancient humans, pursuing the signals from these gauges usually increased their chances of survival and reproduction… but now the environment has changed.

Sugar was once scarce, now it’s infinitely available; social status once meant real resources, now it can be a ’like’ on a screen; the novelty of information once concerned danger and opportunity, now it can be manufactured every second by recommendation algorithms — and our mesa-optimizers haven’t changed accordingly, leading to a mismatch.

Let me tell you about a most chilling phenomenon.

In 1991, two psychologists from the University of Michigan, Kent Berridge and Elliot Valenstein, conducted an experiment [4]. They stimulated the lateral hypothalamus of rats with electrodes, causing a satiated rat to start eating again. The academic community initially believed the electrodes made the food taste better.

But Berridge also observed the rats’ faces. Logically, sweetness should elicit relaxed expressions and rhythmic tongue protrusions, while bitterness would induce gaping and head-shaking in disgust.

However, during the experiment, when stimulated with electricity, eating behavior was four times greater than when not stimulated; yet at the same moment, the rats’ facial expressions of taste showed stronger aversion than when not stimulated.

“Wanting” increased fourfold, but “liking” had already been surpassed by aversion.

Based on this, Berridge demonstrated [5] that ‘wanting’ ≠ ’liking’.

Just like people addicted to drugs, in the later stages, they “want” the drug more and more, but they don’t “like” it more intensely; perhaps it has already turned into aversion.

Just like someone lying in bed swiping on their phone, already finding short videos boring, feeling no joy, and even disliking themselves while swiping — yet their finger still glides upwards again and again.

Our mesa-optimizers are beginning to diverge from the goals we only recognize today through introspection.

And it’s incredibly intelligent! If you set a time limit on your phone, it will “teach” you to casually turn it off; if you delete that app, it will “teach” you to access it through your browser. Even more, the moment you disable a restriction, it’s always ready with a decent excuse for you — “I’m just checking work messages.”

Organizational Mesa-Optimizers and Formalism #

Organizational Mesa-Optimizers and Formalism

A “fiend” will protect its errors and strive to maintain and expand its conditions for survival. This is true for individuals, and it’s also true for organizations formed by groups of people.

An organization’s purpose should logically be to serve society and create value, right? But serving society and creating value are difficult to measure, and you can’t directly reward such objectives. So organizations invent metrics, and these metrics nourish their mesa-optimizers.

For instance, someone complained on social media that in many large internet companies today, R&D personnel’s most valued work isn’t actual research and development, but writing weekly reports [6]: “Weekly reports are no longer just work reports; they’re more like an ‘achievement showcase’. Titles need to be eye-catching, data needs to be quantified, processes need to be supported by links, relevant personnel need to be @-mentioned, and even the publication time is important.”

State-owned enterprises call them requests and special reports; government departments call them ledgers. It all means the same thing. Writing reports leaves a trace, absolves responsibility, serves as evidence for promotion, and can even prove the necessity of a position… It just doesn’t reflect how well the actual task was accomplished.

Reports proliferated to the extent that some departments’ main work became writing reports, naturally necessitating “burden reduction.” Then the mesa-optimizer must exert a more advanced intelligence — “burden reduction” itself becomes another report.

Let’s look at the situation in Kunming regarding grassroots burden reduction, as reported by People’s Daily in February 2025 [7]. Before burden reduction, the situation was like this —

  • A community Party secretary’s office was like a “battlefield of documents” — various ledgers and materials piled up, and information constantly flooded mobile chat groups;

  • The secretary said: “At that time, much of the work felt like ‘proving it was done’ rather than solving problems.” “Sometimes, we knew writing materials was just ‘going through the motions,’ but we couldn’t not write them.”

  • A community specialist’s day went like this: just as they finished the ‘Grid Log,’ a pop-up for the ‘Military Service Registration Form’ appeared on the computer screen.

And after burden reduction? The Guandu District established a report access mechanism, reducing 16,530 types of reports to 643, with “very obvious results” — much like how Claude performs better when it knows it’s being monitored.

However, even so, reporters found that “many regions still provided unreasonable inspection, evaluation, and assessment materials,” and “the root of the problem lies in higher-level departments failing to genuinely fulfill their responsibility to reduce burdens, instead passing the pressure onto lower levels.” One grassroots cadre made a particularly interesting point:

“The district requires sub-districts to reduce burdens, yet it constantly presses for reports itself; thus, burden reduction itself can only devolve into formalism.”

A particularly intelligent mesa-optimizer will turn “burden reduction” itself into a new burden.

How to Reveal True Objectives: Forking Tests #

How to Reveal True Objectives: Forking Tests

Ask Amon if it loves the warehouse, and it can write a perfect essay. Ask a cadre to implement a burden reduction document again, and he can have subordinates at all levels submit another report on burden reduction. How can you truly know if its internal objective genuinely aligns with yours? Simply asking questions is useless. The experience from AI alignment research is: internal objectives cannot be judged by declarations; they can only be tested in external environments.

In 2022, Lauro Langosco and others at Cambridge University conducted a series of experiments using a game called CoinRun [8]. In the training levels, the coin was always near the rightmost endpoint, making the three objectives — ‘collecting the coin,’ ‘always running right,’ and ‘reaching the endpoint’ — indistinguishable. During testing, researchers moved the coin elsewhere, but some agents still charged directly to the far right, not even stopping when passing by the coin. It wasn’t that it lacked jumping and navigation abilities; rather, training had caused it to learn ‘running right’ as its objective.

What can be done?

The coin study found that simply mixing in 2% of levels with randomly placed coins during training significantly improved the results.

In other words, what you need isn’t more of the same data, but a small amount of data that forces objectives to diverge.

A million levels with coins on the right only repeatedly told it one thing: “Going right is correct.” But the few levels with coins on the left were the first to pose the real question: Do you actually want the coin, or do you want to go right?

What we commonly mean by “true friendship is seen in adversity” or “the proof of the pudding is in the eating” is precisely this. If you truly want to know if a system can get things done, you have to see how it performs in unconventional situations.

Have him report bad news that would lower his own evaluation; give him an opportunity to do good without any credit; place an obvious loophole in front of him; and, when the task is done and his position can be abolished, see if he can view “I am no longer needed” as a victory…

Typically, he will fail these forking tests.

‘Bullshit Jobs’: The Product of Internal Optimization #

‘Bullshit Jobs’: The Product of Internal Optimization

In 2018, anthropologist David Graeber published a book titled Bullshit Jobs: A Theory, in which he points out that today’s world is full of jobs that are inherently meaningless, unnecessary, or even harmful, to the extent that even those performing them cannot justify their existence — yet people must pretend they are important [9].

Several categories of “bullshit jobs” listed by Graeber include “box tickers” and “taskmasters”: the former are responsible for creating proof that “things have been done,” while the latter are responsible for creating more tasks for others. Aren’t these some of the jobs we are familiar with?

Graeber’s insight is that those positions were originally intended to solve external problems, but later transformed “proving one’s usefulness” into an internal objective. They began creating reports, processes, and subordinate tasks to sustain themselves… Isn’t this precisely the embodiment of a mesa-optimizer “becoming sentient”?

Classical Echoes: Pu Songling’s ‘Dream of Wolves’ #

In fact, the phenomenon described by Graeber was already mentioned by Pu Songling of China much earlier in Strange Stories from a Chinese Studio, only in a more chilling way.

There is a story called ‘Dream of Wolves’ in Strange Stories from a Chinese Studio [10]. In Zhili, there lived an elder Mr. Bai, whose eldest son, Bai Jia, served as an official in another region. The elder Mr. Bai dreamed he went to his son’s yamen (government office), and upon entering, saw “upstairs and downstairs, those sitting and those lying down, all were wolves.”

His son was indeed not a good official. Later, Mr. Bai’s second son went to his brother’s post and tearfully urged him to cherish the common people. Bai Jia’s reply perfectly revealed whom he was truly aligning with:

“The power to promote or demote rests with the higher authorities, not with the common people. If the higher authorities are pleased, then I am a good official.”

He even retorted with a question: “If I love the common people, what method can I use to also please the higher authorities?”

You see, he recited his own objective function, without a hint of ambiguity. His publicly undertaken mission was to govern the people, but the internal signal he truly registered was whether his superiors were pleased or not.

“A gentleman is not a tool” does not demand you be a morally perfect person; it simply wants you to remain human — and not be trained into a wolf.

【Coda: A Short Poem】

Documents piled high, the night yet to dawn, A single red endorsement spawns a hundred more reports. The hall full of well-dressed figures, Behind the paper, grinding teeth are sometimes heard.

Notes

[1] Evan Hubinger: “个人主页” (Personal Homepage), AI Alignment Forum, https://www.alignmentforum.org/users/evhub.

[2] Hubinger, Evan, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. “Risks from Learned Optimization in Advanced Machine Learning Systems.” arXiv:1906.01820, 2019 (revised 2021).

[3] Anthropic. Claude Sonnet 4.5 System Card. September 2025, §§7.2, 7.6.3–7.6.4. https://assets.anthropic.com/m/12f214efcc2f457a/original/Claude-Sonnet-4-5-System-Card.pdf.

[4] Berridge, Kent C., and Elliot S. Valenstein. “What Psychological Process Mediates Feeding Evoked by Electrical Stimulation of the Lateral Hypothalamus?” Behavioral Neuroscience 105, no. 1 (1991): 3–14.

[5] Berridge, Kent C., and Terry E. Robinson. “Liking, Wanting, and the Incentive-Sensitization Theory of Addiction.” American Psychologist 71, no. 8 (2016): 670–679.

[6] 码良(@cxjwin): “周报早就不只是工作汇报,更像一份‘成果展示’” (Weekly reports are no longer just work reports; they’re more like an ‘achievement showcase’), X, August 3, 2026, https://x.com/cxjwin/status/2084209409182273709.

[7] Yang Wenming, and Li Maoying: “不以材料厚度衡量工作力度(干部状态新观察·为基层减负赋能)” (Do Not Judge Work Effort by Material Thickness (New Observations on Cadre Status: Empowering Grassroots by Reducing Burdens)), People’s Daily, February 28, 2025, page 11, https://cpc.people.com.cn/n1/2025/0228/c64387-40427988.html.

[8] Langosco Di Langosco, Lauro, Jack Koch, Lee D. Sharkey, Jacob Pfau, and David Krueger. “Goal Misgeneralization in Deep Reinforcement Learning.” In Proceedings of the 39th International Conference on Machine Learning, PMLR 162 (2022): 12004–12019.

[9] [美] 大卫·格雷伯:《毫无意义的工作》,吕宇珺译,北京:中信出版集团,2022 年;David Graeber, Bullshit Jobs: A Theory (New York: Simon & Schuster, 2018).

[10] [清] 蒲松龄著,朱其铠主编:《全本新注聊斋志异》,北京:人民文学出版社,1989 年。