Editing & Sound

What Is Sound Design?

Also called sfx design, audio design, sound editing

Sound design is the construction of everything a film sounds like apart from dialogue and score: effects, ambience, textures and the invented sounds that have no real-world source. It is built in layers rather than recorded, and its job is to make a picture feel like a place.

What it is doing that a recording cannot

Location sound gives you dialogue, some traffic, and whatever else happened to be nearby. It does not give you a place, because real acoustic reality is thin, inconsistent between takes, and full of material you do not want.

Sound design replaces it with something constructed. The result is not more realistic than the location recording; it is more legible. A door in a finished film is louder, cleaner and more specific than any real door, and the audience accepts it as a door because it matches what a door means rather than what one measures.

That is the underlying principle. Every choice is about legibility and meaning rather than accuracy, which is why an invented sound for something that does not exist follows exactly the same craft as a footstep.

The layers

The work is organised into layers so that a mix can rebalance them later without re-editing.

  • Room tone or ambience. A continuous bed. Establishes the space and, more importantly, prevents silence.
  • Backgrounds. Traffic, birds, distant crowd, machinery. Non-specific, keeps the world running past the frame.
  • Foley. Synchronous human sounds: steps, cloth, handling. Performed to picture.
  • Hard effects. Specific loud events: doors, gunshots, crashes. Placed to the frame.
  • Designed elements. Sounds with no source in reality, built from unrelated recordings and processing.
  • Space. Reverb and filtering that place every other layer in the same room.

The instinct of anyone new to this is to start at the top, with the loud specific things. The professional order is the reverse. Get the bed right and the rest gets easier, because you are now adding sounds to a place rather than to a void.

Assembling one for generated footage

Generated video usually arrives silent or with a single fused soundtrack you cannot rebalance, which puts the whole job in front of you. The route through it follows the layers, and it maps neatly onto what generation can and cannot do.

Start with the bed, and generate it. Continuous ambience is the ideal case for a text to audio model, because there is nothing to synchronise: low HVAC hum with a faint electrical buzz from strip lights, thirty seconds, consistent throughout. Ask for length and consistency explicitly, and exclude everything else with negative statements, since models add music and voices as texture unless told not to.

Then handle synchronous events, where generation is weaker. Either use a video to audio model, which watches the picture and places sounds at the events it detects, or generate isolated effects and place them by hand. Manual placement is slower and more accurate, and for anything the audience will notice, accuracy wins.

Two habits make the difference between a workable mix and a bad one. Generate each layer as a separate file, never as one combined request, because a single stereo file cannot be rebalanced against dialogue. And mix everything under the dialogue track from the beginning, since effects that sound right alone almost always overwhelm speech.

The prompt for this

A starting point that reliably produces the effect. Adjust the subject and setting; keep the technical clauses.

Continuous room tone for an empty office at night, low HVAC hum with a faint electrical buzz from strip lights, no music, no voices, no footsteps, thirty seconds, consistent throughout

Try Sound Design yourself

Open the generator with a starting point already filled in.

Frequently asked questions

What is the difference between sound design and foley?
Foley is a specific discipline within it: performing everyday synchronous sounds in a studio while watching picture, mainly footsteps, cloth and handled objects. Sound design covers everything else as well, including ambience, effects, transitions and sounds that were never real, such as a spaceship or a monster.
What are the standard layers?
Working from the bottom up: room tone or ambience, background effects, foley, hard effects, designed or invented elements, and then whatever the mix does with reverb and space. Each layer is edited separately so the mix can rebalance them against dialogue without redoing the work.
Which layer matters most?
The ambience bed, and it is the one beginners skip. A scene with no continuous background sounds dead and makes every effect feel pasted on. Adding a quiet, consistent bed under everything is usually the single biggest improvement available to an amateur mix.
How loud should effects be?
Quieter than instinct suggests, once dialogue is present. Almost every effect that sounds correct in isolation is too loud under speech. The working method is to mix effects against the dialogue track from the start rather than balancing them on their own and discovering the conflict later.
Can generated audio do this work?
Partly, and the split is predictable. Beds and background textures are well suited to generation because nothing has to hit a frame. Synchronous events such as footsteps and impacts need either video-to-audio, which reads the picture for timing, or manual placement. Anything carrying a performance beat should stay under manual control.

Related terms