Where Does a Humanoid’s Intelligence Actually Live?

A practical map of control, decision-making and learning, from motor loops to foundation models

I was talking to my brother recently, who is a doctor, about how distributed control and intelligence are inside the human body. A surprising amount of our motor behavior is not handled by one central “brain” layer. Fast responses can be mediated through the spinal cord, local sensory pathways and neuromuscular mechanisms before higher-level cognition becomes involved. Even something as seemingly simple as the stretch reflex involves local sensory receptors in the muscle and spinal circuitry that can produce a motor response without waiting for conscious processing.

Some time ago, a client asked me to analyze an investment case and assess how “intelligent” one of the humanoid companies in their portfolio actually was. They might read this, and they know who they are :)

My first reaction was that “intelligence” in a robot cannot be described simply by saying which SOTA VLA or world model it is using. It is a multilayer system of control, decision-making, perception, adaptation and learning. What we eventually call robot intelligence comes from the interaction between all of these.

In fact, the first humanoid I worked on in a lab, around 13 years ago, used position-controlled locomotion, with much of the walking calculated around ZMP-based methods. Its AI layer was mostly face,voice and motion recognition, with very little semantic understanding. For its time, it was great work. But compared with today’s humanoids, the intelligence stack has become far more complicated.

For a long time, I thought about humanoid intelligence mainly in terms of high-level and low-level systems. High level meant general models such as LLMs, VLMs, LBMs, VLAs and world models. Low level meant PID loops, sensor filtering, actuator control and similar systems.

Then, during the research for this client, I realized that this framing was too simple. The better question was not just how “high” or “low” a system sits, but what it is responsible for, how much of the robot and its environment it considers, and how quickly it has to make a decision.

So I decided to structure the problem properly.

First, intelligence is probably the wrong word

One of the problems is that we use the word intelligence for many completely different things.

A PID (Proportional, Integral, Derivative) controller is not intelligent in the same sense as a language model. It does not reason about the world, understand a task or form a plan. It is a feedback controller. But remove those feedback loops and your highly capable foundation model may be sitting on top of a robot that cannot keep a joint where it should be.

The opposite is also true. A robot can have excellent torque control, balance and locomotion, yet have almost no understanding of why it is moving or what the human next to it wants.

So I think we need to separate at least four ideas.

Control is about making the physical system behave in a desired way.

Decision-making is about selecting between possible actions, states or plans.

Learning is about changing a model, policy or behavior based on data or experience.

Intelligence is the broader capability that appears when these systems work together to perceive, decide, act, adapt and achieve goals.

These boundaries are not perfect. In modern learned systems they overlap heavily. But they are still a better starting point than calling every neural network “AI” and every PID loop “low-level intelligence.”

Two dimensions: time and scope

I think there are two dimensions that make this landscape much easier to understand.

The first is timescale.

At one end, a controller may need to react in milliseconds. At the other, a system may be deciding what task the robot should perform next, whether it should recharge, whether a failed task should be retried, or how work should be distributed across several robots.

The second is decision scope.

At one end, the system cares about one motor, one joint, one fingertip or one contact point. At the other, it may consider the entire body, the environment, the task, the human nearby, or even a fleet of robots.

This is slightly different from simply saying “low level” and “high level.”

A balance controller can be very fast, but its scope is already quite broad because it may coordinate the legs, torso, arms, contact forces and center of mass. A calibration routine can be very local, yet operate slowly over minutes or hours. A safety response can be both fast and global.

Also, local and global here do not mean where the computation physically runs. A global task planner can run onboard the robot. A local actuator model can run on the same GPU as a VLA. I am talking about the scope of the decision, not the location of the computer.

This is why I prefer thinking of it as a landscape rather than a ladder.

Put one humanoid in the middle of it

Let us take a very ordinary industrial task.

A humanoid stands next to a conveyor. Its job is to identify a parcel, pick it up and place it into the correct container.

From outside, we might describe this as one task: sort the parcel.

Inside the robot, it is anything but one task.

At the lowest physical level, the actuators need to regulate current, torque, velocity or position. Encoders need to be read, signals filtered, friction compensated for, and joint limits respected. Depending on the hardware and architecture, these loops may run hundreds or thousands of times per second.

Above that, the robot needs to manage contact and the body as a coupled system. Is the gripper making contact? Is the object slipping? Has the payload shifted the center of mass? Can the torso move without destabilizing the stance? Does one foot need to change its loading?

This is not a new problem created by AI. Whole-body control research has dealt with this type of hierarchy for decades. Work by Sentis and Khatib on whole-body control described humanoid control in terms of multiple prioritized tasks operating under physical constraints, including balance, posture and manipulation.

Then comes a region that is harder to name because it mixes perception, motion and skill execution. The robot estimates the parcel pose, decides where to grasp it, generates a reaching motion, adjusts its hand, avoids nearby geometry, and corrects the trajectory as the conveyor moves. A grasp policy, diffusion policy, imitation-learned skill or conventional planner could all live somewhere around here.

Above that sits behavior and task coordination. Which parcel should I take next? Is the destination container full? Did the last grasp fail? Should I retry from another side? Should I move my feet before reaching? Is the current task still achievable?

Then we reach the more semantic end. What did the operator ask me to do? Which objects count as priority parcels? If a human walks into the workspace and starts taking the parcel from my hand, should I resist, release, stop, step back or ask for clarification?

The farther we move across this map, the more context is normally needed. But the important point is that all of these processes can be active at the same time.

The high-level model does not finish thinking and hand the task down through a one-way pipeline. Information is constantly travelling in both directions. A failed grasp changes the plan. Unexpected weight changes posture. A balance problem can interrupt manipulation. A human intervention can cancel the entire task.

This is one reason humanoids are such difficult systems. As I have argued before in Robotics Complexity vs. Versatility: Why Robots Are Converging, their complexity comes less from any individual subsystem than from the dependencies between locomotion, manipulation, perception, interaction and the physical body itself.

Now add learning, and the map gets more confusing

This is where I think a lot of current robotics discussion becomes unnecessarily messy.

Reinforcement learning is not a layer.

Imitation learning is not a layer.

A diffusion model is not a layer.

A transformer is not a layer.

Even a world model is not automatically a “high-level brain.”

These are architectures or learning methods that can be used at different parts of the landscape.

Reinforcement learning is a good example. Hwangbo et al.’s work on learning agile and dynamic motor skills for ANYmal showed that policies trained in simulation could control dynamic locomotion and transfer to the real robot. That is learning applied relatively close to the physical control side of the landscape. Today, related approaches are being applied to humanoid locomotion, motion tracking and whole-body behavior.

Move upward and imitation learning becomes very prominent. A robot can learn a manipulation skill from teleoperated demonstrations instead of having every trajectory hand-programmed. Diffusion Policy, for example, represents visuomotor behavior as a conditional diffusion process and demonstrated it across a range of manipulation tasks.

Then we get to generalist robot policies and VLAs.

Google DeepMind’s RT-2 took pretrained vision-language models and adapted them to output robot actions, connecting semantic knowledge learned from web data with physical behavior.

OpenVLA followed the same general direction with an open VLA trained on large amounts of real-world robot data.

Physical Intelligence’s π0 combines a pretrained vision-language model with a separate action expert designed for continuous control and dexterous manipulation.

The important part is that these systems start to cover a much larger region of the landscape. They are not just deciding what to do, and they are not just controlling how to move. They increasingly connect the two.

But they still do not make the rest of the stack disappear.

Figure’s Helix almost draws this diagram for us

One of the clearest current examples is Figure’s Helix architecture.

The original Helix separated the system into a slower semantic model, System 2, operating at 7 to 9 Hz, and a fast visuomotor model, System 1, producing upper-body actions at 200 Hz.

Helix 02 took this further. Figure added System 0, a learned whole-body controller operating at 1 kHz for balance, contact and coordination. System 1 produces full-body joint targets at 200 Hz, and System 2 handles scene understanding, language and semantic goals.

I find this architecture notable because it looks remarkably close to the map we just built from first principles.

Slow and broad at the top.

Faster and more physical in the middle.

Very fast and tightly coupled to the body at the bottom.

Figure describes Helix 02 as a unified whole-body neural system, but unified does not mean flat. Different parts still operate at very different timescales and solve different problems.

NVIDIA’s GR00T N1 uses another version of this idea. Its vision-language System 2 interprets the environment and language instruction, and its System 1 diffusion transformer generates continuous actions.

Physical Intelligence’s π0 has another separation. It starts from a pretrained VLM for semantic and visual understanding, then adds an action expert capable of producing continuous robot commands at a much higher rate.

Different architectures draw the boundaries differently, but the recurring pattern is hard to miss: semantic understanding and physical control have very different computational requirements.

This is probably one of the most important things to understand when somebody says:

“Our humanoid is controlled by one foundation model.”

The next question should be:

Controlled down to where?

Does the model output a semantic goal? A skill? An end-effector trajectory? Joint positions? Joint torques? Motor current?

Those are very different claims.

World models do not belong at one level either

I would make the same argument about world models.

The name sounds inherently global, as if there is one internal simulation of the entire world sitting at the top of the robot. That can be one implementation, but it is not a requirement.

A world model is fundamentally about predicting how some part of the world changes, often conditioned on actions. The “world” might mean an entire scene and long-horizon consequences, but it can also mean the dynamics of a much smaller system.

This matters because predictive models can appear at several scales. A model can predict the next visual state, the effect of a manipulation action, the robot’s future body state, or the longer-term consequences of a task decision.

So I would place world models across the map depending on what they model, not automatically in the top-right corner.

There is also another direction: learning inside the physical loop

Most current foundation-model discussion is built around pretraining on large datasets, then adapting or fine-tuning the model for robotic behavior.

But that is not the only direction.

There is another line of work around online adaptation and biologically inspired learning, where part of the system continues to change from real physical interaction instead of treating deployment as mostly inference.

IntuiCell describes its approach around local, continuously adapting sensorimotor learning inspired by biological systems. The stated goal is to deal directly with effects such as friction changes, wear, load variation and drift instead of expecting all variation to be covered through offline training.

That approach is still developing, so I would not put it in the same category of maturity as decades of classical control or established RL methods. But conceptually it points to something important:

Learning itself may also be distributed.

A humanoid might use massive offline pretraining for semantic knowledge, imitation learning for manipulation, reinforcement learning for locomotion, online adaptation for actuator or contact dynamics, and deterministic control for hard constraints.

There is no reason one learning philosophy has to own the entire robot.

This connects to an argument I made in Are We Building Robotics Backwards? about general foundations and specialization. The same tension exists inside a single humanoid. Some knowledge benefits from being extremely general and shared. Some behavior has to become specific close to the body, task and environment.

Safety does not sit in one box

Safety makes the landscape even clearer.

In Humanoid Safety: What We Know and What We Still Don’t, I argued that safety is not an E-stop or one safety controller. It is a property of the complete system.

At the actuator level, safety can mean current limits, torque limits, temperature limits and joint constraints.

At the contact level, it can mean impedance, collision detection, tactile feedback and grip-force regulation.

At the whole-body level, it can mean balance recovery, fall mitigation and maintaining stable contacts.

At the perception level, it can mean detecting a person, an obstacle or an uncertain region.

At the behavior level, it can mean slowing down, choosing another path, releasing an object, retrying a task or asking for assistance.

At the semantic level, it can mean deciding that a requested action should not be attempted at all.

Recent research reflects this layered character. SHIELD, for example, adds a control-barrier-function-based safety layer around a learned humanoid locomotion controller on Unitree G1 instead of assuming that the learned policy itself is the complete safety system.

So if I were drawing the map, I would not give safety its own horizontal layer.

I would draw it as a band crossing almost the entire landscape.

HRI is the same

Human-robot interaction is often placed near the top of robotics diagrams, next to speech, language, social interaction and intent understanding.

That is only part of it.

If a humanoid hands an object to a person, HRI starts at the fingertips. Grip force matters. Compliance matters. The exact moment of release matters.

Move outward and the robot needs to detect the person’s hand, understand whether they are reaching for the object, track body position and maintain a safe distance.

Move upward again and it needs to interpret speech, gestures, intent and context.

Then there is the behavioral side: should it approach? Should it wait? Has the person changed their mind? Is the robot blocking somebody’s path? Does the person understand what the robot is about to do?

This connects directly to something I discussed in A Humanoid Robot Is Judged Against a Human, Not a Machine. Human-like morphology creates expectations. If the robot has hands, eyes, a head and human-scale movement, people assume certain interaction capabilities. HRI therefore touches mechanical design, sensing, control, behavior and semantic reasoning at the same time.

Again, it is not one layer.

So how should we judge the intelligence of a humanoid?

This brings me back to the investment question that started this article.

I do not think asking “How intelligent is this humanoid?” gives us much.

Instead, I would ask:

  1. What decisions can the system actually make autonomously?

  2. What does each learned model output directly?

  3. At what frequency does each part operate?

  4. How much context does each decision consider?

  5. Which capabilities are learned, and which are classically controlled or programmed?

  6. What can adapt after deployment, and what is frozen after training?

  7. What happens when one layer fails or becomes uncertain?

  8. Which safety constraints sit outside the learned policy?

  9. How does information move between semantic reasoning, skill execution and physical control?

Two humanoids can both advertise a “state-of-the-art VLA” and have completely different answers to those questions.

One may use the VLA mainly to select from predefined skills. Another may generate end-effector actions. Another may output whole-body joint targets. Another may connect vision and language almost directly to full-body control but still rely on a faster learned or model-based layer underneath.

The model name tells you surprisingly little without the rest of the architecture.

Perhaps the stack is disappearing, but the timescales are not

There is one counterargument to everything I have written here.

What if end-to-end learning eventually absorbs most of these layers?

I think that is entirely possible.

We may gradually replace explicit perception modules, planners, state machines, locomotion controllers and skill libraries with larger learned systems. We are already moving in that direction.

But I do not think the physical problem itself becomes flat.

A motor still needs a fast response.

Balance still needs to react much faster than long-horizon task planning.

A fingertip contact contains different information from a spoken instruction.

A fall cannot wait for a large semantic model to finish reasoning.

A task planner does not need to operate at 1 kHz.

So even if the software boundaries disappear, the timescales, scopes and physical constraints remain. The hierarchy may become implicit inside learned models instead of being written manually by engineers.

Figure’s move from Helix to Helix 02 is a good example. The architecture becomes more unified, yet the published system explicitly contains a 1 kHz physical layer, a 200 Hz visuomotor layer and a slower semantic layer.

The layers did not disappear.

They learned to talk to each other differently.

A humanoid does not have one intelligence

So I am no longer convinced that “high-level intelligence” and “low-level intelligence” are enough to describe a humanoid.

A better picture is a landscape.

On one axis is time, from millisecond physical reactions to long-horizon task decisions.

On the other is scope, from one joint or contact point to the whole body, environment, task and eventually multiple agents.

Then we can place control, perception, planning and learned policies across that landscape. Learning becomes another property layered on top: some parts are programmed, some optimized, some pretrained, some learned from demonstration, some trained through reinforcement, and perhaps some continuously adapt from experience.

Safety cuts vertically across it.

HRI cuts across it too.

And the robot we eventually call “intelligent” is the result of all of them working together.

A humanoid with a brilliant language model but poor contact control is not an intelligent physical system. A humanoid with extraordinary balance and locomotion but no ability to understand its task is not one either.

So maybe the question is not:

How intelligent is this humanoid?

Maybe the better questions are:

What can it decide? At what speed? Over what scope? What does it learn? And what happens when reality does something the model did not expect?

That, to me, tells us much more about the intelligence of a robot than the name of the model sitting somewhere inside it.



Next
Next

Humanoid Safety: What We Know, and What We Still Don't