What it is doing that a recording cannot
Location sound gives you dialogue, some traffic, and whatever else happened to be nearby. It does not give you a place, because real acoustic reality is thin, inconsistent between takes, and full of material you do not want.
Sound design replaces it with something constructed. The result is not more realistic than the location recording; it is more legible. A door in a finished film is louder, cleaner and more specific than any real door, and the audience accepts it as a door because it matches what a door means rather than what one measures.
That is the underlying principle. Every choice is about legibility and meaning rather than accuracy, which is why an invented sound for something that does not exist follows exactly the same craft as a footstep.
The layers
The work is organised into layers so that a mix can rebalance them later without re-editing.
- Room tone or ambience. A continuous bed. Establishes the space and, more importantly, prevents silence.
- Backgrounds. Traffic, birds, distant crowd, machinery. Non-specific, keeps the world running past the frame.
- Foley. Synchronous human sounds: steps, cloth, handling. Performed to picture.
- Hard effects. Specific loud events: doors, gunshots, crashes. Placed to the frame.
- Designed elements. Sounds with no source in reality, built from unrelated recordings and processing.
- Space. Reverb and filtering that place every other layer in the same room.
The instinct of anyone new to this is to start at the top, with the loud specific things. The professional order is the reverse. Get the bed right and the rest gets easier, because you are now adding sounds to a place rather than to a void.
Assembling one for generated footage
Generated video usually arrives silent or with a single fused soundtrack you cannot rebalance, which puts the whole job in front of you. The route through it follows the layers, and it maps neatly onto what generation can and cannot do.
Start with the bed, and generate it. Continuous ambience is the ideal case for a text to audio model, because there is nothing to synchronise: low HVAC hum with a faint electrical buzz from strip lights, thirty seconds, consistent throughout. Ask for length and consistency explicitly, and exclude everything else with negative statements, since models add music and voices as texture unless told not to.
Then handle synchronous events, where generation is weaker. Either use a video to audio model, which watches the picture and places sounds at the events it detects, or generate isolated effects and place them by hand. Manual placement is slower and more accurate, and for anything the audience will notice, accuracy wins.
Two habits make the difference between a workable mix and a bad one. Generate each layer as a separate file, never as one combined request, because a single stereo file cannot be rebalanced against dialogue. And mix everything under the dialogue track from the beginning, since effects that sound right alone almost always overwhelm speech.