A Humanoid Robot Is Judged Against a Human, Not a Machine

How human-like design raises capability expectations, creates an expectation gap, and risks overpromising before the robot can deliver.

How human-like design raises capability expectations, creates an expectation gap, and risks overpromising before the robot can deliver.

A question has haunted me for over a decade. Although there have been studies and surveys around it, I am still not convinced that we have a complete answer:

Does making a robot more human-like make it appear more technically capable, while also raising users’ expectations of what it should be able to do?

The answer, unsurprisingly, is yes and no.

These are actually two related but different questions. A human-like appearance may make a robot seem more capable before it does anything. At the same time, that appearance may raise the benchmark against which users judge its actual performance.

The first effect is an opportunity. The second is where humanoid companies can get trapped.

So, what does this mean for companies designing general-purpose humanoids?

This post is not a research paper or a systematic literature review. It is a mixture of research findings, personal experience, reasoning logic, and some practical advice for designing the next generation of humanoids.

Human-likeness is a capability signal

A robot’s appearance is not just styling. It communicates what the machine can probably do.

A wheeled box with a simple gripper suggests a relatively narrow range of abilities. It can move objects, transport things, or perform repetitive tasks.

A humanoid with two arms, five-fingered hands, legs, eyes, and a face suggests something much broader. It appears able to navigate human environments, use our tools, manipulate unfamiliar objects, understand gestures, communicate naturally, and adapt to new situations.

Even before the robot moves, we have already formed a mental model of its capabilities.

Research in human-robot interaction suggests that both a robot’s physical appearance and its behaviour influence these mental models. Human-like cues can cause people to generalize from their understanding of humans, sometimes creating expectations that exceed the robot’s real capabilities.[1]

This is why a dexterous five-fingered hand may look more intelligent than a simple industrial gripper, even when both are connected to exactly the same AI stack.

In one study, 160 participants evaluated images of 73 different robot-hand designs on characteristics including intelligence and capability. The visual features of the hands explained a large proportion of the variation in those ratings. Hand size, fingertip shape, surface design, segmentation, and colour influenced how capable or intelligent the hands appeared.[2]

Interestingly, simply adding more fingers was not necessarily the most important factor.

The hardware becomes a proxy for the mind.

From tool to agent

Anthropomorphism does something deeper than making a product friendly or attractive. It changes how we categorize it.

It turns a tool into an agent.

Think about Transformers (The movies, not the AI method!). As cars or trucks, they are impressive machines, but their capabilities appear relatively narrow. Once they transform into humanoids, they suddenly seem capable of fighting, planning, communicating, improvising, manipulating objects, and making independent decisions.

Of course, they are fictional. But the perception mechanism is real.

A car implies transportation. A humanoid implies general agency.

This is one reason the humanoid form is so attractive to general-purpose robotics companies. A human-shaped body is an extremely powerful visual metaphor for adaptability. It communicates that the robot is not designed for only one fixed operation.

That can be a major opportunity. It can also become a serious trap.

A human-like robot is judged against humans

A conventional industrial robot is judged against other machines.

Is it accurate? Is it fast? Is it reliable? Is it safe?

A highly human-like robot is more likely to be judged against humans.

Can it walk naturally? Can it recover when something unexpected happens? Can it use unfamiliar tools? Can it understand what I mean rather than only what I say? Can it notice social cues? Can it move its hands smoothly? Can it respond quickly?

The closer the appearance gets to a human, the more demanding the comparison becomes.

This creates what I would call an expectation gap: the distance between the capabilities implied by the robot’s appearance and the capabilities it can reliably demonstrate.

If that gap becomes too large, the robot may feel disappointing, unreliable, or even deceptive.

A basic-looking gripper that successfully completes nine out of ten tasks may appear highly competent. A beautiful five-fingered robotic hand completing the same nine tasks may be judged more harshly because people expected much more from it.

The objective performance is the same. The psychological benchmark is not.

Honest design

Dieter Rams famously argued that good design is honest. It should not make a product appear more powerful or valuable than it really is, or make promises that cannot be kept.[6]

That principle is highly relevant to humanoid robotics.

The robot’s proportions, form factor, materials, coverings, hands, face, motion, voice, and interaction style should be congruent with its real technical capabilities.

This does not mean humanoid companies should deliberately make robots look primitive. It means every human-like feature should communicate something that the robot can actually support.

Human-like proportions may be justified because the robot needs to use stairs, tools, workstations, vehicles, doors, and shelves built for humans.

Five-fingered hands may be justified when the robot has the sensing, control, accuracy, and manipulation policies needed to use them effectively.

A realistic face is harder to justify if the robot cannot maintain eye contact, understand conversation, read social signals, or produce appropriate expressions.

Visible mechanical components can sometimes be an advantage. They remind users that they are interacting with a machine and help calibrate expectations. A robot does not need silicone skin to demonstrate general-purpose capability.

The goal should be:

Make the robot look as capable as it actually is, not as capable as the company hopes it will eventually become.

Is this the uncanny valley?

Yes and no.

The uncanny valley usually describes the discomfort people experience when something appears almost human but contains noticeable nonhuman inconsistencies.

Appearance-motion mismatch is one possible cause. In a well-known neuroscience study, participants watched a human, a mechanical robot, and a human-looking android perform similar actions. The android combined human-like appearance with visibly mechanical movement, producing stronger prediction-error responses in the brain’s action-perception system.[4]

Its appearance suggested one kind of movement. Its actual behaviour delivered another.

However, the uncanny valley is not simply a universal curve where adding more human-likeness always leads to discomfort. A major review found inconsistent support for the simplest version of the theory, while finding stronger evidence for perceptual mismatch as one route to uncanniness.[5]

More importantly, not every appearance-capability mismatch is uncanny.

A robot may not feel creepy at all. It may simply feel clumsy, unintelligent, frustrating, or overmarketed.

The uncanny valley is therefore one possible consequence of incongruent design, but it is not the whole problem.

What about context?

Context matters, but “design for the task” is not a sufficient answer for general-purpose humanoids.

Research has shown that perceptions and preferences can depend on the combination of robot appearance and the task assigned to it.[3] A machine-like robot, animal-like robot, and human-like robot may not be equally appropriate for teaching, entertainment, guiding, security, or other roles.

A factory robot may benefit from human-like reach, manipulation, and body proportions while gaining little from a realistic face.

An elderly-care robot may benefit from gaze, speech, readable gestures, and an approachable appearance.

But a general-purpose humanoid does not have one fixed application. It is supposed to move between tasks, environments, and social situations.

The relevant target is therefore not task-specific congruence. It is capability-envelope congruence.

The robot should visually communicate the broad class of activities it can reliably perform.

For a general-purpose humanoid, this may mean an adult-sized, relatively neutral, clearly robotic body with human-compatible proportions. It should have enough anthropomorphism to make its actions understandable, but not so much realism that users assume human-level physical or social intelligence.

Human-likeness should be functional rather than decorative.

Practical principles for humanoid companies

First, design for human compatibility, not maximum human resemblance. Human proportions, reach, joints, and hands can be valuable because our physical world is designed around the human body.

Second, treat every anthropomorphic feature as a promise.

A face promises social responsiveness.

Eyes promise attention.

Hands promise dexterity.

Legs promise mobility.

A voice promises understanding.

Companies should ask whether the robot can reliably keep each of those promises.

Third, evaluate perception alongside engineering performance. Do not only measure task-success rates, walking speed, payload, or manipulation accuracy. Ask users what they believe the robot can do before seeing it operate, and then ask again after the demonstration.

The difference between these answers may be as important as the technical benchmark itself.

Fourth, reduce the expectation gap progressively. As the robot becomes more capable, its appearance can become more expressive, refined, or human-like. The visual promise should grow with demonstrated capability.

Finally, remember that greater anthropomorphism does not automatically mean better design. Human-likeness can improve intuitiveness, social interaction, and perceived versatility. But it can also increase expectations faster than engineering capability improves.

The real design problem

Humanoid companies are not only designing machines.

They are designing expectations.

The body is the first interface, the first product demonstration, and the first capability claim. Before the robot says a word or performs a task, its appearance has already told the user a story.

The safest strategy is not to avoid human-likeness. The humanoid form is powerful precisely because it communicates mobility, manipulation, communication, and general agency.

But that power needs discipline.

A human-like robot is judged against humans, not machines.

So the goal should be functional human-likeness, not maximum human-likeness.

The best humanoid may not be the one that looks most human. It may be the one whose appearance, movement, intelligence, and real-world performance tell the same honest story.

Cheers

References

[1] Kwon, M., Jung, M. F., & Knepper, R. A. (2016). “Human Expectations of Social Robots.” 2016 11th ACM/IEEE International Conference on Human-Robot Interaction, 463–464.
https://doi.org/10.1109/HRI.2016.7451807

[2] Seifi, H., Vasquez, S. A., Kim, H., & Fazli, P. (2022). “Charting Visual Impression of Robot Hands.” arXiv preprint.
https://doi.org/10.48550/arXiv.2211.09397

[3] Li, D., Rau, P. L. P., & Li, Y. (2010). “A Cross-Cultural Study: Effect of Robot Appearance and Task.” International Journal of Social Robotics, 2, 175–186.
https://doi.org/10.1007/s12369-010-0056-9

[4] Saygin, A. P., Chaminade, T., Ishiguro, H., Driver, J., & Frith, C. (2012). “The Thing That Should Not Be: Predictive Coding and the Uncanny Valley in Perceiving Human and Humanoid Robot Actions.” Social Cognitive and Affective Neuroscience, 7(4), 413–422.
https://doi.org/10.1093/scan/nsr025

[5] Kätsyri, J., Förger, K., Mäkäräinen, M., & Takala, T. (2015). “A Review of Empirical Evidence on Different Uncanny Valley Hypotheses: Support for Perceptual Mismatch as One Road to the Valley of Eeriness.” Frontiers in Psychology, 6, Article 390.
https://doi.org/10.3389/fpsyg.2015.00390

[6] Rams, D. “Ten Principles for Good Design.” Vitsœ.
https://www.vitsoe.com/about/good-design

Next
Next

The State of Data Collection for Robotic Manipulation