Intelligence Is Not Enough
Put the most capable AI model in the world inside a robot whose sensors are each accurate but describe different moments, and it will reason flawlessly from a reality that never existed as a unified whole.
The Coherence Papers · Part IV
In the first three Coherence Papers, I argued that robotics’ next frontier may be coherence, that time is the hidden reference that makes coherence possible, and that biology offers important lessons about how distributed systems coordinate perception and action. This paper turns to a harder question: Is intelligence alone enough?
When intelligence enters a body
For decades, artificial intelligence has pursued an extraordinary ambition: to build machines capable of recognizing patterns, solving problems, generating plans, using language, and adapting to unfamiliar situations. The progress has been remarkable. Modern AI systems can write software, analyze medical information, translate languages, create images, solve difficult mathematical problems, and hold conversations that would have seemed impossible only a few years ago. That progress deserves its recognition.
But the moment intelligence enters a physical body, another problem appears.
Imagine placing the most capable AI model in the world inside a robot. Its cameras observe the environment. Its LiDAR measures distance. Its inertial sensors detect motion. Its joint encoders report body position. Its tactile sensors register contact. Its motor controllers execute movement. Now imagine that all of those components are individually accurate, but their information does not describe the same moment. The camera reports where an object was a fraction of a second ago. The arm controller reports where the hand is now. The force sensor reports contact after the grip has already begun to change. The inertial system detects motion, but its reading is interpreted against an earlier visual frame.
The AI may still be highly intelligent. It may reason flawlessly from the information it receives. But it is reasoning from a reality that never existed as a unified whole.
That is not a failure of intelligence. It is a failure of coherence.
Accurate pieces, inaccurate present
Every intelligent action begins with some representation of the current state of the world. Before a robot can decide where to step, it must estimate where the ground is, where its body is, how it is moving, and what may happen next. Before it can hand an object to a person, it must estimate the position and motion of the hand, the object, its own arm, and the surrounding space. Intelligence can then interpret that state, predict possible outcomes, and choose an action.
But what happens if the state itself is assembled from observations belonging to different moments? The system may have accurate pieces and still construct an inaccurate present.
Two different questions
This suggests an important distinction. Intelligence asks: What am I seeing? What does it mean? What should I do? What is likely to happen next? Coherence asks: Do these observations belong together? Do they describe the same event? What had already happened when this information was produced? Has the system changed since then? Can I trust that my model represents the present rather than a mixture of several different pasts?
Neither capability is sufficient on its own. Coherence without intelligence could produce a well-aligned but passive observer. Intelligence without coherence could produce brilliant decisions based on an internally inconsistent world. Real autonomy requires both.
Meaning depends on when
We often treat information as though its meaning travels with it unchanged. A measurement is a measurement. A message is a message. A fact remains a fact. But in an evolving system, meaning depends partly on when information enters the process.
Consider a simple human conversation. Suppose someone is reading a twelve-part argument and pauses after Part VII to say, “This idea changes how I see intelligence.” That statement has a specific meaning at that moment. The person has encountered the first seven ideas but has not yet incorporated Parts VIII through XII. Their interpretation, prediction, and intended next action are still developing.
If the same statement is interpreted as though it were made after the entire argument had been read, its meaning changes. Later insights are incorrectly projected backward onto an earlier state of thought. The words are identical. The temporal position is different. And therefore the meaning is different.
A coherent system must know not only what information says, but where that information belongs in the unfolding history of the system.
Clock time is part of that history, but it is not the whole of it. A system may also need to preserve what occurred before the information was generated, what had not yet occurred, what the system believed at that moment, what action was being planned, and whether the system has since changed its state. Information is not merely content. It is content associated with a particular stage of an evolving process.
The present is constructed
A robot does not receive “the present” from a single sensor. The camera provides one stream. The inertial unit provides another. Touch, sound, joint position, force, and motor status each arrive through their own pathways, at their own frequencies, with their own delays. The system must determine which observations belong together and how they relate to its current state.
In that sense, the present is constructed. It is assembled from distributed evidence.
If those relationships are preserved well, the robot can form a stable and useful understanding of what is happening now. If those relationships are degraded, the robot may combine accurate observations into a false present. More computation does not automatically solve this problem. A faster processor can analyze inconsistent information more quickly. A larger model can reason more deeply over a misaligned state. A more sophisticated planner can generate a better plan for a world that has already changed.
Intelligence improves what the system can do with its model. Coherence determines whether the model deserves to be trusted.
Prediction depends on order
The importance of coherence becomes even clearer when a system attempts to predict. Prediction requires more than recognizing objects or identifying patterns. It requires understanding change. To estimate what will happen next, a system must know what happened first, what happened afterward, what is changing, how quickly it is changing, and which observations describe the current state.
If temporal order is corrupted, causality becomes difficult to infer. A robot may mistake an effect for a cause. It may respond to a condition that has already disappeared. It may repeat a correction after the system has already recovered. It may act on an earlier plan even though the environment has changed.
The resulting failure may look like poor intelligence. But the deeper problem is that the intelligence has lost its correct position in time. The system is no longer reasoning from a coherent present.
Preserving meaning across latency
Not every delay is a coherence failure. Biological nervous systems contain delays. Communication networks contain delays. Mechanical systems contain delays. Perfect simultaneity is neither possible nor necessary. The key issue is whether the system knows enough about those delays to preserve the correct relationships among observations and actions.
A delayed measurement can still be useful if the system knows when it was captured, how the world may have changed since then, and how the measurement relates to the current state. A very fast measurement can still be harmful if its timing or context is misunderstood.
The goal is therefore not simply to eliminate latency. The goal is to preserve meaning across latency.
This is a more demanding requirement. It means that a system must maintain the relationship between information and the state from which that information emerged.
Intelligence in the organization of the system
This also points toward a broader view of intelligence. We usually ask how intelligent an individual processor, model, organism, or agent is. But in a distributed system, some capability may arise not from any one component, but from the relationships among them.
A camera does not understand the world. An inertial sensor does not understand motion in context. A motor controller does not understand the purpose of an action. Even the AI model does not directly experience the physical environment. What we call intelligent behavior emerges when sensing, timing, communication, interpretation, prediction, and action remain sufficiently related.
The intelligence of the whole is therefore not located entirely inside one part. Some of it exists in the organization of the system.
This does not mean that every relationship is intelligent. Random interaction creates noise. Excessively rigid interaction creates fragility. Poorly timed interaction creates confusion. But when relationships preserve relevant information, support feedback, correct error, and allow the whole to adapt, capabilities emerge that none of the isolated components possesses.
The system becomes more than a collection of parts. It begins to behave as a whole.
Capability, and the conditions for it
The future of robotics will continue to depend on better artificial intelligence. Robots will need stronger reasoning, better perception, improved learning, richer world models, and greater adaptability. But intelligence cannot substitute for coherence.
Intelligence determines how well a system can interpret, predict, and decide. Coherence determines whether the information supporting those abilities forms a sufficiently consistent account of reality. One provides capability. The other preserves the conditions under which that capability can be expressed.
The distinction may be summarized this way: Intelligence tells a system what its information means. Coherence tells the system whether that information belongs together.
In an embodied machine acting in a changing world, that difference is fundamental.
A robot does not need a perfect representation of reality. No biological or engineered system possesses one. It needs a representation coherent enough to support safe, timely, and adaptive action.
That may be the deeper challenge now facing robotics. Not simply building machines that can reason, but building machines whose information, decisions, and actions remain correctly situated within the unfolding reality they are trying to understand.