Inside the world or outside it
Every sound on a finished soundtrack falls on one side of a single question: could a character in the scene hear this? Diegetic sound is the yes side. It exists in the world of the film, has a physical source somewhere in or near the frame, and behaves accordingly.
The no side is addressed to the audience alone. Score, a narrator, a stylised whoosh on a title card, and the sudden bass drop under a reveal are all outside the story world. The characters live their lives without ever noticing them.
Examples of each
Inside the story world:
- Dialogue between characters
- Footsteps, doors, cloth, cutlery, the whole foley layer
- Traffic, rain, wind, a room's air conditioning
- A band on a stage, a radio, a phone ringing, a TV in the background
Outside it:
- Orchestral or synth score
- Voiceover narration and inner monologue
- Editorial effects: risers, stings, transitions
- A song dropped over a sequence with no source in the world
The grey areas, which are the interesting part
Three cases sit deliberately between the two, and they are where sound design does its most expressive work.
Source music becoming score. A track begins on a practical speaker, then the room acoustics fall away and it swells to full width. The audience feels the scene lift out of literal space. Reversing this, bringing a score down into a car radio, is just as effective and often funnier.
Internal diegetic sound. A character's heartbeat, a ringing in their ears after a blast, a remembered voice. The character hears it and nobody else does, which puts it inside the story world but only for one person. Almost every subjective sequence in modern film runs on this.
Sound that comments. An effect that is technically motivated but exaggerated far past realism, such as a clock that gets louder than physics allows while someone waits. It stays inside the world and starts working as score.
Why it matters when you generate audio separately
AI audio tools give you three separate streams: sound effects, music, and speech. Each arrives clean, dry, and at full level, which is exactly right for score and exactly wrong for anything meant to exist in the room.
If a generated sound is supposed to sit inside the story world, it has to be placed in that space after generation:
- Perspective. A radio in the next room is muffled and quiet. Generating it as a clean, close recording and then simply lowering the fader does not sound like distance, it sounds like a quiet clean recording. Ask for the muffling and the wall in the prompt, then add room reverb that matches the shot.
- Reaction to picture. Source music has to change when a door shuts or the camera moves outside. That is an automation pass, not a generation setting.
- One consistent room. Effects generated in separate runs carry different implied spaces. Adding a single shared reverb across all of them is what makes them agree.
Score is the easy case, and it is worth reaching for AI music generation there first: it does not need to obey the space, so a clean generated track can sit under a sequence at a constant level. The moment a piece of music is supposed to come out of something visible on screen, it becomes source music, and the work moves from generating it to placing it.