What the shape describes
Put two clips on a timeline and extend the first clip's audio to the right, past where its picture ends. Video stops, sound continues underneath the next shot, and the pair traces an L. The name comes from tape and film assembly and survived into every modern edit suite, where the operation is usually a matter of dragging one track.
Underneath the terminology it is simply a split edit: picture and sound cut at different frames. That should be the norm. When every transition cuts both at once, an edit acquires a metronomic quality that audiences read as amateur long before they can explain it.
Why editors reach for it constantly
Dialogue coverage is the main case. A scene between two people has more useful pictures than it has speaking moments, because the listener is often doing something more interesting than the speaker. Holding one character's voice while the picture moves to the other lets you have both. That single move is doing most of the work in almost every conversation you have watched.
It solves three problems at once:
- Reactions. You can watch a line land on the person receiving it.
- Bad takes. If the best delivery of a line came from a take where the picture was unusable, hold the audio and cut away.
- Rhythm. Because picture and sound break at different moments, the scene stops falling into a pattern of one line, one shot.
The same device works outside dialogue. A car engine or a room's ambience held for a second past the cut ties two spaces together. Music carried over a cut to a new location makes the second place feel like a continuation rather than a new beginning.
Getting the length right
The judgement is when the trailing sound stops belonging to the new picture. Practically, an L cut used for smoothing runs a few frames to half a second and nobody notices it. Used expressively it can run several seconds, at which point the outgoing audio has become a layer over the new scene, which is a strong effect and a deliberate one.
Two habits help. Fade rather than hard cut when the sound is ambience, since a room tone stopping abruptly announces the edit you just hid. And check that the trailing audio is not colliding with something the incoming shot needs, which is the most common reason a smooth-looking edit still sounds cluttered.
With generated material
Split edits are one of the easiest techniques to apply to AI footage, precisely because sound and picture were never joined in the first place. Generated clips usually arrive silent, so every audio decision is already yours and the edit points are free.
What matters in practice:
- Build dialogue as separate audio. Whether it comes from recorded voice, a cloned voice, or text to speech, keep each line as its own clip. Then holding a voice over a cut is a drag, not a re-render.
- Generate listening shots on purpose. The reason this technique pays off is that you have somewhere to go during a held line, so generate reaction clips as deliberately as you generate the speaker. A quiet close-up with small movement is cheap and reusable.
- Lay one continuous ambience under the whole scene. Independently generated shots have no shared acoustic space, and a single room tone running under all of them is what makes them feel like one location. It also hides the seams at every cut.
- Do not bake sound into clips. If a model produces audio with the video, mute it and rebuild. Baked audio ends exactly where its picture ends, which is the one thing this technique needs to avoid.
Alternating L cuts and J cuts across a scene is what stops a sequence of generated shots feeling like a list. The pictures may have been made one at a time, but the soundtrack can be continuous, and continuity of sound is what an audience actually uses to decide whether a scene is one place.