Why the absence of sound is a problem
Record a conversation in a kitchen and the file contains more than the conversation. It contains the fridge, the building, the road outside, air moving, and the noise floor of the microphone itself. That composite is the sound of that kitchen, and the listener has learned it within the first second.
Now cut out a stumble in the dialogue and leave the gap empty. What follows is not quiet, it is nothing, and the drop from the kitchen's noise floor to absolute zero is one of the most conspicuous things you can put in a soundtrack. The listener does not hear an edit; they hear a fault.
Room tone exists to prevent that. It is the same composite, recorded with nothing happening, so an editor has an unlimited supply of the sound of that place with no dialogue in it. Laid under a scene it makes the whole track continuous, and continuity is what lets cuts disappear.
Capturing it properly
The requirements are strict, and they are the reason it has to be recorded on the day rather than approximated later.
- Same microphone, same position, same settings. A bed recorded on a different mic does not match the dialogue's own noise floor, which defeats the purpose.
- At least thirty seconds, preferably sixty. The material gets looped, and short loops reveal themselves through repeated details.
- Nobody moves. Standard set practice is total stillness, because a single shifting foot makes a section unusable.
- One per setup. If the lighting changed, a fan came on, or a window was opened, it is a different space acoustically and needs its own recording.
Skipping it is the classic low-budget mistake, and it cannot be fully fixed afterwards. What can be done is extraction: harvesting quiet stretches from the dialogue takes themselves, which works but yields short, awkward fragments.
Generating one
Continuous ambience is close to the ideal task for a text to audio model, because the whole point is that nothing happens. There is no timing to hit, no event to place, and no performance to match, which removes every weakness generated audio normally has.
The prompt should describe the components of the noise floor rather than an atmosphere. Very low air conditioning rumble, faint hum from a computer fan, distant muffled traffic through closed glass gives the model a stack of specific continuous sources. Asking for a mood, such as an eerie empty office, tends to return something with events and music in it.
Then constrain the shape explicitly. Ask for length, ask for consistency, and exclude everything you do not want: sixty seconds, perfectly consistent, no voices, no music, no events. Models add texture unprompted, and a bed with a distant door slam in it cannot be looped.
For generated video this stops being a substitute and becomes the correct method. There is no location, so there is nothing to match, and the noise floor of the scene is whatever you decide it is. Building the bed first and then placing effects on top of it reproduces the layer order a real mix uses, and it is the single cheapest way to make silent generated footage stop sounding synthetic.