Why ADR exists
Dialogue is the one element of a soundtrack that has almost no tolerance for noise. Audiences forgive a wrong footstep and will not forgive an unintelligible line, so any production sound that cannot be cleaned has to be replaced. Common causes are mundane: aircraft, traffic, wind, wet weather cover, a generator, a costume rustling into the lavalier, or a camera the mixer could not isolate.
The second driver is the edit. Lines get rewritten, characters get renamed, a joke lands badly in a test screening, and a legal note requires a brand not be spoken. None of those problems can be solved on set because the set is gone.
The third is performance. Sometimes the take is technically clean and simply the wrong reading, and a director wants the line delivered smaller, later, or with different weight.
How a session runs
An ADR session puts the actor at a microphone watching their own picture, usually with three beeps counting them in to the line. They perform the line repeatedly against the loop until the timing and the energy match. Modern sessions lean less on the beeps and more on the actor watching their mouth and simply going again, which tends to produce a freer reading.
An ADR editor then does the real alignment work: cutting the chosen take into place, stretching or nudging syllables so the lip sync holds, and matching perspective so the replaced line sits in the same acoustic space as the surrounding production audio. Reverb, EQ, and worldizing (playing the line into a real room and re-recording it) are all part of making a booth recording believable in a corridor or a car.
What makes replacement audio convincing
- Match the mic and distance. A wide shot recorded on a boom sounds nothing like a close lavalier. Replacing one with the other is instantly audible.
- Match the room. A booth is dead. The scene is not. Some reflection has to be added back, either with reverb or by re-recording the line in a space.
- Match the body. If the character is walking, moving, or out of breath, the actor has to be too. Standing still and imitating breathlessness rarely lands.
- Keep the neighbours. A single replaced line inside a scene of production audio is the hardest case, because the cut in and out of the ADR is exposed. Replacing the surrounding lines too is often easier than blending one.
Where AI voice tools actually help
Two capabilities matter here. Voice cloning can produce a line in an actor's voice from a reference recording, and text to speech can produce a placeholder or a background voice from scratch. Both change the economics of small fixes.
Realistic uses today:
- Word-level repairs. A misspoken name, a changed number, a place that got renamed. Splicing a generated word into a real take is the highest hit rate available.
- Temp lines. Filling the edit so a cut can be shown before the actor is available. This alone removes a scheduling bottleneck.
- Crowd and background voices. Nobody is doing loop group for a corridor of passing extras if a model can produce twenty separate voices.
- Other languages. Cloning a lead's voice for a localised version, which is dubbing work rather than same-language replacement.
The limits are honest ones. Models handle plain readings well and struggle with a specific intent held across a long line, especially overlapping breath, held emotion, and the tail of a shout. They also produce clean, dry, roomless audio, which means every generated line still needs the perspective and room treatment described above. And if the replacement has to fit a mouth that is already on screen, a lip sync pass is a separate step, not something the voice model solves.
Permission is not a technical question. Using a performer's cloned voice is governed by their contract, and in several territories by law, so it belongs in the deal rather than in the sound department's workflow.