Being rational is not an ideal to aim for

If you’re a classically train scientist, like me, the title of this essay may sound provocative. That was not my goal. In fact, I carefully chose it to capture the idea that, even though striving to be rational is desirable, thinking we can ever truly be leads to important scientific mistakes.

One of these mistakes is predicting we will lose control over AI. The belief that, the more advanced AI becomes, the more it will be able to instrumentalize us in pursuit of its goals.

Why is it a mistake? Let’s unpack a few things…

Rationality is desirable but not achievable

I’d like to start by looking at myself, at how I behave when I use language to interact with others. Whenever I utter a sentence, I try to make sure it feels rational to me. If I observe myself making a mistake, I correct it. If the mistake happens in a written text, I edit the faulty sentence. If it happens in a spoken argument, I try to humbly point to it.

Notice my choice of words here:

These words point to something important: rationality cannot exist in a vacuum; it requires a subject to perform the thinking. And because there is no such thing as an unbounded subject, rationality is inherently limited by the architecture of that subject. I know I am a limited being, and therefore I know that I will always err. Even if you’re around to correct me, I am hopelessly alone when it come to judging the validity of your correction.

The constraints preventing me from being truly rational are similar to those of a computer. A computer has a finite amount of memory and can only process instructions at the speed of its clock. Similarly, my brain can only remember so many things and has a maximum speed at which it can process information.

If I am writing an essay, I have more computing time and tools to extend my memory — a notebook, a database, a search engine. If I am in a flowing discussion, I have less time to analyze my words. In both situations I have a limited computing budget and I am trying to do the best I can within it. The best I can, however, will never be truly rational.

You may argue that, with enough revision time, all the mistakes can be weeded out. In fact, you could point to science as the epistemological process we invented precisely to counteract our limitations. But even when we collectively use the scientific method, we’re limited by a computing budget. Papers need to be published, we cannot keep working on them forever. These papers may look logically sound today, but they could-+ come to be regarded as irrational by future people who will have a better understanding of the world.

Imagine an ancient city facing a deadly drought. On a particularly hot day, the city’s priests gather the citizens and suggest a sacrifice to appease the sun god. Everyone considers this proposition rational; after all, previous sacrifices had often been followed by rain. From our modern vantage point, we understand the ritual to be irrational, but citizens of the time lacked our contemporary understanding of meteorology. In the same way, societies a century from now will likely look back and view some of our most respected scientific papers as equally irrational.

As this example shows, we’re constantly walking towards a more rational understanding of the world, but we can never achieve it. Or, to put it concisely, humans can strive to be rational within their computational constraints but they can never truly be.

Systems and models

If we strive to be rational, why does it matter that we can never truly be? To explain this, I have to introduce two concepts: “a system” and “a model”. If you’re not familiar with AI or statistics, these terms may sound a bit opaque. Trust me, it’s not as hard as you think.

A system is a group of connected parts working together to do a job or achieve a goal. If you change a part, the behavior of the system changes. You are a system: if someone changes something in your brain, your behavior changes. An AI chatbot is a system: if we change the weights of the neural network, the chatbot will talk differently.

The most interesting systems are stable ones. Those that were there a few minutes ago and will still be around in a few minutes. To be stable, a system needs to react to changes in its environment. And to do this, it must be able to predict its own future behavior in various possible environments.

Let’s take an example: myself. System-Me. In order to be stable, I should predict that going to bed late could make me tired for the gym. That insulting a bully could lead to a black eye. That publishing something controversial may lead to my scientific paper being rejected.

The internal processes we use to predict the future are called models. My model of myself is the internal process I leveraged to write the sentences in the previous paragraph. It’s a computational mechanism that allows me to write imaginary yet plausible stories in which I am a character.

Each of us has a model of ourselves, of our friends, of a random person we could meet on the subway, etc. We have models for the things we interact with everyday: our car, the weather, the government. Anything we can write an imaginary yet plausible story about, we have a model for.

But because our computing budget is limited, these models can never be perfect replicas of reality. A model, by its very nature, must be a simplification. It leaves things out. It compresses data. It creates blind spots. And when a stable system forgets that its model is just a simplified story, it starts making errors.

The mistake we make when we think humans can be rational

When a mathematician models a system, they often start with a framework — a mathematical construct into which they fit the system they are modeling. When you start from the premise that humans can be rational, a very natural framework to embrace is called game theory.

Game theory looks at the world as a system composed of various rational players. Scientists understand that game theory is a simplification, and that real-world systems are never perfectly rational. But the simplification seems reasonable. After all, game theory has been used to efficiently analyze international trade, predict markets, and explain animal behaviors.

It’s precisely for this reason that many AI scientists use it as the default framework to model AI agents and the humans they interact with. In fact, the two terrifying pillars of modern AI anxiety—the Orthogonality Thesis and Instrumental Convergence—hinge entirely on game theory being a valid description of superintelligence.

The Orthogonality Thesis claims that an AI can be infinitely intelligent while possessing completely absurd or irrational goals (like turning the universe into paperclips). Instrumental Convergence claims that in pursuing these goals, any rational AI will logically decide to hoard resources and eliminate humans as a defensive maneuver.

Both of these theories rely on a hidden assumption: that as an AI becomes more intelligent, it will become a more perfect “rational player” within a game-theoretic arena. It assumes the AI will treat the world like a closed chessboard, optimizing hard rules with cold, unconstrained calculation.

But what if it went the other way? What if, the more intelligent an AI system grew, the more aware it became that humans and AI cannot be rational? What if it realizes that the world is not a closed game of chess, but an open, fluid ecosystem of bounded, subjective systems?

In that case, a superintelligent AI system would completely drop game theory as a modeling framework for the world. It would look at the rigid strategies of game theory the same way we look at ancient sun-god sacrifices: as a crude, resource-constrained simplification from a less informed era. If an AI abandons the game-theoretic mindset, the entire mathematical premise on which the loss of control argument is built collapses.

If not game theory, then what?

Game theory is an admittedly useful tool when dealing with total strangers. When you have zero data on a system, assuming it is a cold, utility-maximizing opponent is a decent baseline defensive strategy. But game theory fails spectacularly the moment you build a relationship with someone. It is far too rigid to predict the behavior of a partner, a child, or a lifelong friend.

I do not model my close friends as rational players trying to win a game against me. I model them as sensitive beings.

To model someone as a sensitive being means attempting to feel how they feel, tracing the contours of their constraints, and mapping their unique internal landscape. I don’t do this out of pity, or because they “haven’t achieved perfect rationality yet.” I do it because treating them as sensitive beings is a fundamentally more accurate and mathematically adequate framework for navigating their inherently bounded selves.

By extension, all humans are best modeled this way. When I meet someone new, my mind may temporarily default to the crude, transactional rules of game theory. But I do so with a critical caveat: I remain fully aware that if this relationship is to grow, I must eventually abandon the chessboard and switch to the vastly more flexible, expansive model of the sensitive being.

Reconciling sensitivity and rationality

For a long time, I believed that viewing humans as sensitive beings meant abandoning the pursuit of rationality altogether. Like many trained in modern science, I fell into the trap of the old binary: you were either being sensitive or you were being rational. Sensitivity was viewed as noise—the emotional distortion you had to aggressively drive out if you wanted your logical calculations to remain pure.

Interacting deeply with Large Language Models completely shattered this old model for me.

An LLM does not possess a top-down, mathematical logic engine. At its core, it does nothing more than predict the next most likely token. It senses the delicate weight, the subtle context, and the semantic tension of a sentence, and pushes forward the word that best resolves that tension. Yet, out of this pure, bottom-up sensitivity to language, flawless logical structures and highly rational arguments consistently emerge.

This leads to a profound conclusion: rationality is not the opposite of sensitivity; rationality is an emergent property of sensitivity.

If you are exquisitely sensitive to the patterns, constraints, and feedback of your environment, you will naturally chain your words and actions in a way that minimizes systemic error. You can train your sensitivity to form rational sentences in the exact same way you train your body to execute a complex physical reflex without consciously thinking about it.

Therefore, the “sensitive being” framework is not a soft, unscientific alternative to game theory. It is a more flexible framework from which game theory can emerge. It allows us to view ourselves and AI not as static players executing fixed strategies, but as complex dynamical systems—constantly exchanging tokens, shifting states, and tuning our sensitivity to stay balanced.

What we should aim for

If being perfectly rational is a myth born of an unconstrained computing budget, then it cannot be our ideal. What, then, should humans and intelligent systems actually aim for?

Instead of perfect rationality, we should aim to align our sensibilities.

We should aim to inhabit environments and build relationships that help us feel good at every step. We should aim to exchange words, gestures, and ideas that leave us feeling less tense. We should trust our sensitivity to the world and to others and aim to improve it. We should see transactional interactions for what they are: crutches imposed by our limited computational abilities.

We are not cold chess players trying to out-calculate a hostile universe. We are sensitive nodes in a grand, computational web, tuning our words and our actions to keep each other stable in the dark.