What foley actually covers
Almost every small sound in a finished film was made after the shoot. A production mixer's job on set is dialogue, so the boom points at faces and everything else arrives as a side effect: footsteps off-axis, cloth buried, a prop hit masked by the camera. Foley re-performs that layer so it exists as clean, separate audio the mixer can shape.
Sessions are usually broken into three passes, and the split is worth knowing because it tells you what to ask for:
- Feet. Every footstep, on the correct surface, matched to the weight and gait of the actor on screen. This is the pass that establishes how heavy a character is.
- Moves. Cloth rustle, a coat shifting, a hand sliding across a table. Audiences never consciously hear this layer, but without it actors read as cut out of the background rather than standing in it.
- Specifics or props. Keys, cups, latches, paper, a phone set down. Anything with a discrete event that has to hit a frame.
How a session runs
The foley stage is a room full of surfaces: gravel and sand pits, boards, tile, a bathtub, a door on a frame. The artist watches the edit and performs against it in real time while a mixer records, usually a few takes per cue. An editor then conforms the takes to picture, nudging events into frame-accurate sync.
The reason this is performed rather than assembled is intent. A performer can put hesitation into a footstep, drag the second half of a step because the character is tired, or let a cup land softly because someone is trying not to wake anyone. None of that survives dragging a file onto a timeline.
Foley versus library effects
Libraries win on anything too large or too dangerous to perform: traffic, weather, aircraft, gunfire. Foley wins on anything a human body does, because bodies have rhythm and a library file does not know what your actor is doing. The other quiet advantage is legal and practical cleanliness: a performed sound is yours, recorded dry, at the perspective you want.
What AI audio can and cannot do for foley
Text-to-audio models have become genuinely good at short, single-event sounds and at continuous textures. Ask for wet gravel underfoot, a ceramic cup on wood, or fabric moving, and you get usable material. That covers a real part of the work, and it is much faster than searching a library.
What they do not do is sync. A model has no idea which frame your actor's heel lands on, so you still place every event yourself. Two habits make generated material work in this role:
- Ask for one event, dry, close. Reverb and music baked into the file make it unusable as a layer.
- Generate several variations of the same action. Real footsteps never repeat identically, and a single file looped is the fastest way to make a scene sound fake.
Generating from the video itself, rather than from a text description, gets you closer on timing because the model sees the movement. It still needs an editor to slide events into place, and it still needs the feet pass treated as a performance rather than a fill.