Skip to main content

Objective Function: What Does This Universe Reward?

·2994 words·15 mins
An intricate, high-tech interface or cosmic game backend, with glowing lines of code and data flowing, representing the "Earth Online" universe. A large, abstract guillotine blade is symbolically severing a chain between "IS" and "OUGHT" text, embodying Hume's guillotine. Below, a vast, star-strewn field extends towards a subtle, deep chasm or cliff edge, signifying the universe's inherent "hard constraints" and the absence of a visible leaderboard or explicit reward system.

Imagine the world we live in is a massive multiplayer online game, let’s call it “Earth Online.” You’ve been playing for decades, working hard, yet your gear is still terrible. Meanwhile, some seemingly unremarkable people are doing extremely well. You feel indignant. One day, you finally hack into the game’s backend. You pull up the source code and go straight to the point.

You want to know what this game truly rewards.

Is it accumulating good deeds? Is it maximizing self-interest? Or, like in cultivation novels, are all worldly things actually unimportant, and the only thing you should do is absorb spiritual energy quickly, because the game’s real goal is to beat it, to move on to the next, more advanced game?

You also want to glance at the leaderboard. This server has been running for four billion years; who are the top-ranked players? Can I just imitate them?

Well, I have to tell you, you might be disappointed. This game has no explicit “reward function,” and there’s no leaderboard either.

We don’t need to hack into the backend to roughly deduce the logic of Earth Online, because we’ve designed our own games on Earth. We’ve written reward functions, trained agents, managed communities, designed exams and KPIs… We’ve discussed “Goodhart’s Law,” and we know all too well what players turn into once a system posts a scoring rubric.

Instrumental Rationality and Value Rationality: What “Should” We Do? #

Instrumental Rationality and Value Rationality: What “Should” We Do?

Whether Earth Online has a reward function is critically important to us, as the reward function determines what we should do.

All scientific and technological knowledge in the world, including most of the thinking tools we discuss here, teaches you how to do things—it’s for when you already know what you want, and this knowledge teaches you how to get it. In Max Weber’s words, this is called “instrumental rationality (Zweckrationalität),” which tells you the most effective means to achieve a given goal.

But if you ask what goals are worth pursuing, that’s called “value rationality (Wertrationalität)” [1].

So, can value rationality be derived from instrumental rationality? No. David Hume, in A Treatise of Human Nature, published in 1740, pointed out the inferential gap later known as “Hume’s guillotine” [2]: from “is” alone, you cannot derive “ought” (the is–ought gap).

This guillotine severs the seemingly natural chain of reasoning between “what the world is” and “what one should do”: facts can only tell you the consequences of various choices; but to derive an “ought,” you must also add a value premise.

For example, “it’s raining outside” is merely a fact; it doesn’t imply “you should bring an umbrella”—a hidden value premise is involved: “you don’t like getting wet.”

You must first clearly define what you want before you can employ instrumental rationality to figure out how to get it. In Hume’s words: Reason is but the slave of the passions.

If we don’t know what we want, don’t know what we should do, but are very good at doing things, then we would be similar to current AI. In fact, Oxford University philosopher Nick Bostrom, inspired by AI, proposed the “orthogonality thesis” in 2012: intelligence level and ultimate goal are two orthogonal, mutually independent axes; any level of intelligence can be paired with any goal [3]. This also underpins Bostrom’s famous “paperclip problem”: an infinitely capable superintelligence could single-mindedly follow your command, at the cost of using the entire universe’s resources, to… say, maximize paperclip production.

The biggest difference between us and AI, it seems, is that we have to find our own goals.

This statement strikes some as strange and others as unsettling. Most people are busily striving for a livelihood every day, and can’t even achieve existing goals in life; how can there be time to ponder value rationality? What’s more, traditional morality and various religions have always told us what is right and what we should do.

But if you think carefully, “what should be done” is indeed a problem.

Some people in this world pursue money, some pursue fame, others pursue contribution—which one is correct? What does Earth Online truly reward?

Why Doesn’t the Universe Have an Award System? #

Why Doesn’t the Universe Have an Award System?

There is no evidence of an award system in the universe, saying that whoever does something right gets five reward points… and there’s nowhere to redeem your prizes either.

In fact, we only need to switch to a “creator’s perspective” to understand why Earth Online has no reward function. If you create a world, you probably hope your world can be diverse and vibrant, and can exist for a long time, right?

Note that this is not an “ought.” There doesn’t even need to be a creator here. This is an “is”: the universe we observe has inherently existed for a long time and is diverse and vibrant.

Okay, now let’s assume the universe has a reward function. What would it be like?

If the universe truly gave people scores, everyone would immediately know what they should do—we Chinese know all too well what that implies, we discussed that scenario in our course: “Moloch.”

Think about cultivation novels and online games. Everyone pursues leveling up, and the resources supporting leveling up—spirit stones, spiritual veins, heavenly materials and earthly treasures—are limited. In such a scenario, killing others to seize treasures would be the optimal strategy, and “morality” would merely be the excuse of the weak. If the entire meaning of living is to accumulate enough points, leave this world, and level up to a better one, then this world would become a thoroughly exploited mine.

Playing games occasionally is fine, but actually living in such a world, most people wouldn’t make it out of the newbie village.

And no matter what your reward function is, no matter how complex the rules you design, under competition, everyone will eventually become the same, just as national playing styles in the World Cup become increasingly similar.

If the whole world were a World Cup, there wouldn’t even be any spectators. The experience would be truly terrible.

Even more terrifyingly, such a world would be extremely unstable.

When we discussed “adaptive cycles,” we said that optimization leads to overfitting, and overfitting leads to instability. If everyone is the same, what if the environment changes? What if the universe had to change the reward function then? What would happen to these people?

Diversity isn’t for aesthetics; diversity is the most crucial survival strategy.

For example, when we talked about the “Kelly criterion,” we mentioned that in a bacterial population, a tiny fraction of cells—whose proportion precisely aligns with the Kelly criterion—spontaneously switch to a very slow-growing “persister state.” This phenomenon is called “bacterial persistence,” and cells entering this state are called persister cells. From an immediate perspective, these cells are a pure waste: they divide extremely slowly, don’t claim territory, and occupy resources for nothing. But when antibiotics arrive, the fast-growing ones are wiped out, while precisely these “non-competing” ones survive.

And there’s a cruel mechanism here: many common antibiotics target processes “currently at work” like cell wall synthesis and DNA replication. In such attacks, the harder you strive to grow, the more of a target you become.

Researchers call this phenomenon an “insurance policy” [4]. The optimal switching rate primarily depends on how often the environment changes, and has little to do with the selective pressure of any specific environment.

Simply put, because you can’t read the next question (on the test), you must keep a batch of people doing things that currently seem useless.

So if you were the creator, how could you dare to give this world an explicit reward function? The worlds in those cultivation novels are by no means prosperous, but merely efficient and fragile meat grinders.

The Logic of the Universe: Absence of Rewards, Persistence of Constraints #

The Logic of the Universe: Absence of Rewards, Persistence of Constraints

The more you understand the causal relationships of all things, the harder it is to believe that the universe has set any explicit reward function.

Thus, many people believe that value rationality is purely subjective, and that everyone’s desired way of life is their own business… and that life inherently has no meaning. But this isn’t entirely correct.

You could say there are indeed no rewards, but there are constraints.

Hard Constraints: No Rewards, But Clear-Outs #

Hard Constraints: No Rewards, But Clear-Outs

The universe provides physical laws, feasibility boundaries, and biological elimination mechanisms—what we call “hard constraints” that must be respected. Heaven and Earth are impartial; they treat all things as straw dogs. It doesn’t give awards, but it does clear the field.

We can borrow the imagery from Salinger’s The Catcher in the Rye [5]. The universe is a vast field of rye, and we are thousands of children running and playing in it. We can run carelessly, but there’s a cliff at the edge of the field—you simply cannot run towards that cliff.

And it’s not just one side with a cliff; this field is practically full of cliffs and traps everywhere: you can’t self-harm, you can’t refuse to eat, you can’t provoke beasts stronger than you… no matter where you fall, you’re in big trouble, and you might even be deleted from the game.

Because the cliffs and traps are so numerous and dense, the places where you can safely tread are already very narrow… you find that you can only walk along the remaining narrow strips of land.

It’s clearly that you can only, but because those remaining narrow strips of land are like paths, you feel that you should walk along these paths.

In this way, a mechanism that only deletes, after running for four billion years, statistically, greatly resembles a reward function.

Instrumental Convergence: The Inevitability of Survival, Reproduction, and Cooperation #

Instrumental Convergence: The Inevitability of Survival, Reproduction, and Cooperation

Here we must once again invoke Bostrom. In the same article where he proposed the orthogonality thesis, he also proposed a second thesis, called “instrumental convergence,” meaning that while everyone’s ultimate goals can be diverse and eccentric, to achieve them, almost all agents will utilize the same set of intermediate means—preserve themselves, preserve goals from being overwritten, enhance capabilities, acquire resources [3].

For example, “preserve oneself”—the universe does not demand that you must preserve yourself. “Wanting to live” was not originally anyone’s ‘ought’; you don’t have to list it among your life’s purposes.

But Bostrom says that even if a system doesn’t care about its own survival at all, it will instrumentally care about its own survival to achieve the goals it truly cares about.

If you are not alive, then health, assets, prestige, relationships—none of it has meaning.

“Wanting to live” is not a reward, nor a value, but a constraint.

And it’s not enough just for you to live; you must also reproduce—what you reproduce isn’t necessarily genes; it can also be culture, craftsmanship, institutions, but you must, in the end, leave something for the next generation.

And this cannot be done alone, therefore, you must cooperate.

In 2019, three anthropologists at Oxford University reviewed ethnographic records from 60 societies and examined seven rules for cooperation: help kin, help one’s group, reciprocate favors, be brave, defer to superiors, be fair, respect others’ property. They found that these seven rules were uniformly positively evaluated in all societies; not a single society deemed them bad [6].

Human hearts share the same sentiments, and minds arrive at the same principles. This is not a cultural coincidence; it is also a constraint.

Alright, now simply by starting from avoiding traps and cliffs, we derive that you must survive, reproduce, and cooperate. We can continue to derive more.

Time Horizon and Virtue: The Birth of Meaning #

If you want to live, you must consider tomorrow. If you want to reproduce, you must consider the next generation. If you want to cooperate with others, you must make them believe you will still be around.

Therefore, you must extend your “time horizon.”

For individuals, it manifests as delayed gratification, compound interest, skill accumulation, and reputation management—the reason you are willing to suffer a loss today is that you are factoring in your future self three years from now. For organizations, it manifests as a company’s willingness to make investments that will only show results in five years, or whether it just focuses on this quarter’s financial report. Applied to a civilization, it’s whether it’s willing to build structures that take over a hundred years to complete, where the people who start the project know they won’t see its completion.

And so, “meaning” emerges.

In 2013, several American social psychologists used questionnaires to discover [7] that people’s conception of “happiness” is often related to “the present, taking, and needs being met”; while “meaning” is related to “spanning time, giving, and connecting past and future.”

Simply put, waking up in the middle of the night to change a baby’s diaper will lower your current happiness, but will significantly increase your sense of meaning.

Perhaps that sensor in your heart called “meaning” doesn’t measure happiness, but rather the length of your time horizon.

Virtues: The Dashboard for Adapting to Constraints #

Pushing further, we arrive at another string of more familiar words: To avoid being dragged off a cliff by immediate desires, you need temperance and prudence; to allow accumulation to span time, you need diligence, patience, and a willingness to be responsible; to make cooperation withstand repeated temptations to take advantage, you need trustworthiness, fairness, and reciprocity; to guide families and communities through danger, you also need courage, loyalty, a willingness to care for the weak, and to bear immediate costs for future generations…

We call these things virtues.

Many people believe that virtues are reward functions, saying that if you act according to these virtues, you will accumulate merit and blessings, and perhaps receive good fortune in the future—but our previous derivations show that these are merely adaptations to constraint conditions.

So you say, I do feel good after performing good deeds; surely that counts as a reward?

That’s not the universe giving out prizes; that’s an instrument panel installed in you by evolution.

Evolution couldn’t directly write something as abstract as “long-term survival” into your brain, so it installed a set of approximate gauges for you [8]: sugar is sweet, sex is pleasurable, being accepted is comforting, being rejected is painful, winning is exciting, curiosity satisfied is happiness, caring for children brings fulfillment—allowing you to be driven by these feelings.

Evolution translated those external cliffs into the pain, desires, and attachments within your body.

Once again, all of this is merely our adaptation to the universe’s hard constraints.

The Source of Freedom: The Vast Expanse Beyond Constraints #

You can respect all constraints and practice all virtues, but even after fulfilling all those requirements, you still have a great degree of freedom left. You still don’t know where to go. You still have to set your own objective function.

This is precisely the source of freedom.

A world with only constraints and no rewards leaves a large, undefined space. In that space, you can do anything; the system doesn’t care. You can study a beetle’s wings, you can devote your life to a theorem with no practical value, you can love someone who is disadvantageous from any perspective.

This is not a flaw in the system; it’s a necessity for the system to avoid fragility. It is precisely that “useless” curiosity, that “unprofitable” aesthetic, that “irrational” loyalty that prevents this world from overfitting to the single immediate answer, and preserves other solutions for the next change.

Nature doesn’t first create a world and then kindly bestow freedom upon you; it simply has no other choice—to keep this world stable and perpetually interesting, it must give you freedom.

Conclusion: The Blank Space Beyond the Prohibited List #

Your task is to figure out what to do with this freedom; my suggestion is to use it to create new possibilities… We’ll discuss this at the end of the course.

In this lecture, we distinguished between how to do something and what to do. We distinguished between what must be done to adapt to hard constraints and what one purely wants to do.

Finally, let’s return to the game’s backend.

You haven’t found a game guide, but you’ve pulled up a very long list of prohibited items: what will get you out, what will get your descendants out, what will get your species out… But the source code says absolutely nothing about what you should do.

The true problems of your life all lie in that blank space beyond the prohibited list.

A poem states:

For countless eons, this field endures, No keeper, but a dangerous abyss lures. Heaven’s craft, through ages, only clears the way, Leaving wild, open ground for nature’s free play.

References

[1] Weber, Max. Wirtschaft und Gesellschaft. Tübingen: J. C. B. Mohr, 1922; Chinese translation: Economy and Society, translated by Lin Rongyuan, Beijing: The Commercial Press, 1997.

[2] Hume, David. A Treatise of Human Nature. London, 1739–1740, Book II, Part III, Section III; Book III, Part I, Section I.

[3] Bostrom, Nick. “The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents.” Minds and Machines 22, no. 2 (2012): 71–85.

[4] Kussell, Edo, Roy Kishony, Nathalie Q. Balaban, and Stanislas Leibler. “Bacterial Persistence: A Model of Survival in Changing Environments.” Genetics 169, no. 4 (2005): 1807–1814.

[5] J. D. Salinger: The Catcher in the Rye, translated by Shi Xianrong, Nanjing: Yilin Press, 2015.

[6] Curry, Oliver Scott, Daniel Austin Mullins, and Harvey Whitehouse. “Is It Good to Cooperate? Testing the Theory of Morality-as-Cooperation in 60 Societies.” Current Anthropology 60, no. 1 (2019): 47–69.

[7] Baumeister, Roy F., Kathleen D. Vohs, Jennifer L. Aaker, and Emily N. Garbinsky. “Some Key Differences between a Happy Life and a Meaningful Life.” The Journal of Positive Psychology 8, no. 6 (2013): 505–516.

[8] Tooby, John, and Leda Cosmides. “The Psychological Foundations of Culture.” In The Adapted Mind: Evolutionary Psychology and the Generation of Culture, edited by Jerome H. Barkow, Leda Cosmides, and John Tooby, 19–136. New York: Oxford University Press, 1992.