Editing & Sound

What Is the Kuleshov Effect?

Also called kuleshov experiment, montage effect, juxtaposition effect

The Kuleshov effect is the tendency of viewers to read emotion into a neutral face based on the shot cut next to it. The same unchanged close-up appears hungry, grieving or aroused depending on what precedes it, which demonstrates that meaning in film is produced by juxtaposition rather than by performance alone.

What the experiment established

Kuleshov's demonstration is one of the few pieces of film theory that is also an empirical result. He took a single close-up of an actor with a blank expression and cut it against three different shots. Audiences reported watching a subtle, varied performance, and described emotions the actor had never played.

The conclusion is structural rather than psychological: a shot does not carry a fixed meaning. It carries a meaning produced in combination with its neighbours, and the editor is therefore a co-author of the performance rather than someone arranging finished units.

Modern replications complicate the story without overturning it. The effect is real, but it depends on the face being genuinely ambiguous. Give a viewer an unmistakable expression and they read the expression, not the context. That constraint is where the practical craft lives.

Consequences for performance and editing

Three things follow, and all three are counterintuitive if you think of a scene as a sequence of complete moments.

  • Less is more on camera. An expression that reads clearly in isolation frequently reads as too much in the cut, because the surrounding shots are already supplying the emotion. Film performance is quieter than stage performance for this reason rather than for microphone reasons.
  • The reaction shot is where meaning is decided. What a character is looking at determines what their face means, so choosing the object shot is choosing the emotion.
  • Order changes content. The same three shots in a different order tell a different story, which is why an edit can be rescued or ruined without any new material.

Using it with generated footage

There is a happy alignment here between a theoretical insight and a practical weakness of current models.

Video models are unreliable at specific emotion. Asking for grief, dawning realisation or suppressed anger tends to produce either an exaggerated theatrical expression or something in the uncanny range, because subtle emotion in a human face is exactly the hardest thing to synthesise. Asking for a neutral face is much easier, and the results are far more often usable.

The Kuleshov effect says you do not need the emotion in the face. So generate the neutral close-up, generate the context shot, and let the cut do the work.

Practical notes. Prompt for neutrality with more force than feels necessary, because models drift toward expressiveness: a completely neutral expression, no discernible emotion, relaxed mouth. Hold the shot long enough to cut into and out of, since a two second clip gives an editor no choice about timing. Keep the light even and unmotivated, because dramatic lighting is itself a context cue that will fight whatever the neighbouring shot is trying to say. And generate the context shot separately with no continuity requirement at all, which is the one thing this technique does not need.

The prompt for this

A starting point that reliably produces the effect. Adjust the subject and setting; keep the technical clauses.

Close-up of a middle-aged man looking slightly off-lens with a completely neutral expression, no discernible emotion, relaxed mouth, even soft light, static camera, held for four seconds, 85mm

Try Kuleshov Effect yourself

Open the generator with a starting point already filled in.

Frequently asked questions

What was the original experiment?
Lev Kuleshov, working in Soviet cinema in the 1910s and 1920s, intercut the same neutral close-up of an actor with different shots: a bowl of soup, a girl in a coffin, a woman on a divan. Audiences praised the actor's range, describing hunger, grief and desire in a performance that had never changed.
Does the effect actually hold up?
Modern replications find it real but weaker and more conditional than the original story suggests. It works best when the face is genuinely ambiguous and the context shot is unambiguous. A face with a clear expression overrides the context rather than absorbing it, which is a useful practical limit.
How is it different from montage generally?
It is the narrowest and most testable case of montage theory. Montage covers all meaning generated by the arrangement of shots. This effect isolates one variable: hold the face constant, change only the neighbour, and observe that the reading changes. It is the proof rather than the theory.
How do you use it deliberately?
Shoot or generate faces with less expression than the scene requires, then let the surrounding shots supply the emotion. This is why film performances often look underplayed on set and land correctly in the cut, and why an actor who pushes an expression can make a sequence feel overstated.
Why does it matter for AI-generated shots?
Because models default to legible, slightly exaggerated expressions, which is exactly what defeats the effect. Prompting for a genuinely neutral face gives an editor something the context can act on. It is also a workaround for a real weakness: a neutral face is easier to generate convincingly than a specific emotion.

Related terms