Why sound crossing a cut changes the cut
A picture cut is instantaneous. One frame is a bedroom, the next is a street, and there is no intermediate state. Sound does not work like that, and the mismatch is what a bridge exploits.
When an audio layer continues unbroken across a picture cut, the viewer is given something continuous to hold on to during a discontinuous moment. The join stops being a break and becomes a change of view within something ongoing. That is why a bridge is the default device for showing time passing, for moving between locations, and for connecting two things thematically without stating the connection.
The direction of the carry changes the meaning. Sound held over from the previous scene reads as memory, or as a thought that has not finished. Sound arriving early from the next scene reads as anticipation, or as the future intruding. Both are the same technique pointed in opposite directions, and choosing consciously between them is most of the craft.
The variants worth naming
- Trailing. Outgoing audio continues over the incoming picture. Reflective, lingering.
- Leading. Incoming audio starts before its picture. Anticipatory, propulsive, and the more common of the two in contemporary editing.
- Neutral third layer. Score, rain, or an abstract texture belonging to neither scene, laid across the join. Useful when both scenes have specific sound that would clash.
- Match on sound. Two different sources that sound similar, so the cut feels like a single continuing sound: a kettle becoming a siren, applause becoming rain. This is the most satisfying version and the hardest to find material for.
What all of these need is audio without strong internal structure. A continuous hum, a crowd, or a sustained note bridges well. A phrase of dialogue or a piece of music with a clear phrase ending does not, because the sound announces its own edges and the viewer hears the seam anyway.
Building one when the picture is generated
There is a specific trap here. Video models that produce a native soundtrack give you one fused file per clip, and that audio cannot extend beyond its own shot. A bridge by definition has to cross a shot boundary, so native audio makes the device unavailable.
The workable approach is to keep audio as its own layer from the start.
Generate the picture silently, or discard the native track. Then generate the bridging audio as a single long file, longer than either shot, using a text to audio model: a continuous bed, no events, explicitly excluding music and voices unless you want them.
Lay that bed across both clips so it is genuinely unbroken, then place each scene's specific effects on separate tracks above it. The specific layers can cut with the picture while the bed does not, which is exactly the structure the effect requires.
One prompt detail matters on the picture side too. The shots on either side of the join should have a little dead air in them, a second or two where nothing happens, because a bridge needs somewhere to sit. A clip that is full of action from the first frame to the last leaves the audio nowhere to breathe, and the transition ends up feeling crowded rather than connected.