Where the dip comes from
Plot how comfortable people are against how human something looks and you do not get a straight line. Comfort rises through stylised faces, keeps rising through cartoon and puppet territory, then falls off a cliff just before photorealism, and only recovers once the thing is indistinguishable from a person. That collapse is the uncanny valley, and the word valley is doing real work: there is a climb on the far side, and most synthetic humans never reach it.
The usual explanation is category confusion. Your brain runs cheap general object recognition on things it has filed as objects, and runs a dedicated, ruthless face and body pipeline on things it has filed as people. Stylisation keeps a character in the object bucket. Photorealism moves it into the person bucket, where a different and far pickier set of checks applies. Nothing about the image got worse. The standard it is being held to did.
The cues that actually trigger it
They are not evenly weighted, and knowing the order saves a lot of re-rolling.
- Eyes. Pupils that do not adjust, a gaze that does not converge on anything, absent or metronomic blinking. This is the single strongest cue.
- Mouth and teeth. A jaw that moves while the tooth line stays rigid, lips with no wetness, speech shapes that arrive slightly before or after the sound.
- Skin. No pores, no fine hair, no light bleeding through the ear or nostril. Skin that reads as painted plastic.
- Motion. Head turns with no settle at the end, breathing that never disturbs the shoulders, movement that is too even in speed. Video exposes this within a second.
- Asymmetry. Real faces are lopsided. A perfectly mirrored face is unnerving even when every other cue is right.
Note that four of those five are about imperfection being missing rather than something being added. That is the practical lesson.
Why AI-generated humans land there so often
Image and video models are trained toward the average of what they have seen and toward whatever their aesthetic tuning rewards, and both pressures push away from the specific defects that make a face read as real. You get symmetric features, unblemished skin, even lighting, and a neutral pleasant expression, which is the exact centre of the valley. The models are not failing at realism so much as succeeding at idealisation.
Video adds a second cause. Temporal smoothing is what keeps a clip from flickering, and it also sands off the micro-jitter of real human motion. A face can be flawless in every single frame and still feel wrong the moment it moves, because the motion has no noise in it.
Prompting and shooting around it
Two strategies work, and they pull in opposite directions. Pick one per project rather than splitting the difference, because the middle is the valley.
Retreat toward stylisation. Ask for illustration, cel shading, a graphic-novel look, a puppet, visible brush texture. You lose photoreal credibility and you get a character the audience never audits as a person.
Or push for specificity instead of quality. This is the useful move when you need photoreal output. Name the imperfections you want: visible skin texture and pores, slight asymmetry in the eyes, one strand of hair out of place, natural blink. Add a real optical context, because lens and light language correlates with un-idealised training images: 50mm at f/2, soft window light, documentary realism. Avoid the words that summon the centre of the valley, which are the flattering ones: beautiful, perfect skin, flawless, 8k hyperrealistic.
For clips with dialogue, the fastest single fix is usually not in the prompt at all. Generate the shot slightly wider than you need. Faces read as uncanny in close-up because the cues are large enough to inspect, and a medium shot hides most of what a model gets wrong about eyes and teeth while keeping the performance intact.