What the model changes, and what it leaves alone
A lip sync pass is a local edit. The model detects the face, isolates a region around the mouth and jaw, and regenerates only that region for each frame, then blends it back into the original picture. Identity, hair, background, and camera movement all come from your source, untouched.
That scope explains the artifacts. Because the region is composited back, a mismatch in skin tone or sharpness at the blend boundary shows up as a faint patch around the chin. Because the region is small, teeth and tongue get very few pixels. And because head pose comes from the source, the model cannot fix a performance where the head is doing something the new audio does not support.
Choosing a source clip
Most bad results are decided before the job runs. What matters, roughly in order:
- Face size in pixels. A 4K frame with a tiny distant face is worse than a 720p closeup. If the face is small, crop to it, run the sync, then composite back.
- Mouth visibility. No hand, microphone, cigarette, or hair across the mouth. Occlusion in even a few frames produces a visible tear.
- Angle. Near frontal to three-quarter is reliable. Full profile is not, since half the mouth is not there to reconstruct.
- Motion blur and compression. Both destroy mouth detail. A clean, well-lit, moderately still take beats a dynamic one.
- Original mouth state. A source where the subject is already speaking gives the model plausible jaw geometry to work from. A locked closed mouth for ten seconds is harder.
Preparing the audio
Clean the audio before you sync, not after. Strip music and background dialogue, tame reverb, normalize to a consistent level, and trim leading silence so the first phoneme is where you think it is. If you generated the voice, keep it dry: reverb should be added after the sync, on the final mix.
For long takes, cut into segments of roughly ten to twenty seconds at natural sentence breaks and run them separately. Timing accuracy tends to degrade over a long single pass, and short segments also let you rerun one bad line instead of the whole scene.
Failure modes and their causes
- Mushy or blurred mouth on fast speech. Not enough pixels, or audio with smeared consonants.
- Jaw jitter or chattering. Frame-to-frame instability, common on noisy or low-light sources.
- Visible patch around the chin. Blend boundary mismatch, usually from a heavily graded or grainy source.
- Right mouth, wrong feeling. The head and eyebrows still perform the original line. This is the ceiling of the technique.
Lip sync versus talking head generation
They are different tasks that get confused constantly. Lip sync edits an existing shot and preserves everything except the mouth. A talking head model creates the whole performance, including head motion and expression, from one portrait plus audio. Choose lip sync when you have footage you want to keep, and full performance synthesis when you have only a photo and need it to feel alive.