AI Generation

What Is the Uncanny Valley?

Also called uncanny valley effect, almost human, near-human

The uncanny valley is the dip in comfort that appears when a synthetic face or body becomes almost human but not quite. Viewers accept a stylised character and accept a real person, yet find the near-perfect version between them unsettling. It is the failure mode most AI-generated humans land in.

Where the dip comes from

Plot how comfortable people are against how human something looks and you do not get a straight line. Comfort rises through stylised faces, keeps rising through cartoon and puppet territory, then falls off a cliff just before photorealism, and only recovers once the thing is indistinguishable from a person. That collapse is the uncanny valley, and the word valley is doing real work: there is a climb on the far side, and most synthetic humans never reach it.

The usual explanation is category confusion. Your brain runs cheap general object recognition on things it has filed as objects, and runs a dedicated, ruthless face and body pipeline on things it has filed as people. Stylisation keeps a character in the object bucket. Photorealism moves it into the person bucket, where a different and far pickier set of checks applies. Nothing about the image got worse. The standard it is being held to did.

The cues that actually trigger it

They are not evenly weighted, and knowing the order saves a lot of re-rolling.

  • Eyes. Pupils that do not adjust, a gaze that does not converge on anything, absent or metronomic blinking. This is the single strongest cue.
  • Mouth and teeth. A jaw that moves while the tooth line stays rigid, lips with no wetness, speech shapes that arrive slightly before or after the sound.
  • Skin. No pores, no fine hair, no light bleeding through the ear or nostril. Skin that reads as painted plastic.
  • Motion. Head turns with no settle at the end, breathing that never disturbs the shoulders, movement that is too even in speed. Video exposes this within a second.
  • Asymmetry. Real faces are lopsided. A perfectly mirrored face is unnerving even when every other cue is right.

Note that four of those five are about imperfection being missing rather than something being added. That is the practical lesson.

Why AI-generated humans land there so often

Image and video models are trained toward the average of what they have seen and toward whatever their aesthetic tuning rewards, and both pressures push away from the specific defects that make a face read as real. You get symmetric features, unblemished skin, even lighting, and a neutral pleasant expression, which is the exact centre of the valley. The models are not failing at realism so much as succeeding at idealisation.

Video adds a second cause. Temporal smoothing is what keeps a clip from flickering, and it also sands off the micro-jitter of real human motion. A face can be flawless in every single frame and still feel wrong the moment it moves, because the motion has no noise in it.

Prompting and shooting around it

Two strategies work, and they pull in opposite directions. Pick one per project rather than splitting the difference, because the middle is the valley.

Retreat toward stylisation. Ask for illustration, cel shading, a graphic-novel look, a puppet, visible brush texture. You lose photoreal credibility and you get a character the audience never audits as a person.

Or push for specificity instead of quality. This is the useful move when you need photoreal output. Name the imperfections you want: visible skin texture and pores, slight asymmetry in the eyes, one strand of hair out of place, natural blink. Add a real optical context, because lens and light language correlates with un-idealised training images: 50mm at f/2, soft window light, documentary realism. Avoid the words that summon the centre of the valley, which are the flattering ones: beautiful, perfect skin, flawless, 8k hyperrealistic.

For clips with dialogue, the fastest single fix is usually not in the prompt at all. Generate the shot slightly wider than you need. Faces read as uncanny in close-up because the cues are large enough to inspect, and a medium shot hides most of what a model gets wrong about eyes and teeth while keeping the performance intact.

The prompt for this

A starting point that reliably produces the effect. Adjust the subject and setting; keep the technical clauses.

Close-up portrait of a woman speaking to camera, soft window light from frame left, visible skin texture and pores, slight asymmetry in the eyes, one strand of hair out of place, natural blink, shot on 50mm at f/2, documentary realism

Try Uncanny Valley yourself

Open the generator with a starting point already filled in.

Frequently asked questions

Who came up with the uncanny valley?
Masahiro Mori, a Japanese robotics professor, described it in 1970 while writing about prosthetic hands and humanoid robots. He plotted comfort against human likeness and found a sharp dip just before the curve reaches a real person. The English name came later, from the 1978 translation of his essay.
Why does something almost human feel worse than something obviously fake?
Because likeness changes which perceptual system you use. A cartoon is read as a drawing, so nothing about it is checked against a real face. Something photoreal is read as a person, which switches on the face and motion machinery you use on actual humans, and that machinery is extremely sensitive to small errors.
Which cues trigger it most reliably?
Eyes and mouth first, then motion. Dead or mistimed eyes, a blink that never happens, teeth that stay identical while the jaw moves, and skin with no pores or no subsurface glow are the usual culprits. In video, motion that is too smooth reads as wrong faster than a still image ever does.
Can you get out of the valley by making things more realistic?
Sometimes, but it is the expensive direction. Climbing the far wall means getting every cue right at once, because one wrong cue drags the whole shot back down. Retreating toward stylisation is cheaper and more reliable: a deliberately illustrated or slightly graphic look sits on the safe side of the dip.
Does the uncanny valley apply to voices?
Yes, and the mechanism is the same. A synthetic voice with correct phonemes but flat prosody, breaths in the wrong places, or no mouth noise sits in an audible version of the same dip. Listeners usually describe it as sounding hollow rather than robotic.
Is it cultural or universal?
The dip shows up across studies, but its depth and location move. Familiarity with animation, expectation set by context, and even how a character is introduced all shift how tolerant a viewer is. There is no fixed line you can design against, which is why testing on real viewers beats reasoning about it.

Related terms