What the model is watching for
A v2a model runs two jobs at once. It classifies what is in the frame well enough to guess what things would sound like, and it locates the moments where something happens: an impact, a footfall, a door meeting its jamb, a hand landing on a surface. The second job is the one that decides whether the result is usable.
That ordering runs against intuition. It seems as though the difficulty would be in synthesising a convincing door slam, but convincing door slams are cheap now. What is expensive is deciding that the door meets the frame on frame 71 rather than frame 74. Sound and picture are judged together by a viewer who will accept a wooden door that sounds slightly metallic and will not accept a door that closes before it shuts.
This is also why v2a output tends to be good at continuous sound and weaker at discrete events. Rain, traffic, room tone and wind have no attack to align, so they land convincingly almost every time. Footsteps, impacts and handling noise are where errors are audible.
Prompting it
The video already carries the timing, so the prompt should carry the material, and it should exclude everything you intend to supply yourself.
- Name the surface, not just the action.
footsteps on wet gravelproduces something usable.footstepsproduces a generic thud that will not match your picture. - Anchor the one event that matters.
a heavy wooden door closing on the third secondgives the model a target to align to rather than leaving it to find every event on its own. - Ask for the layer, not the mix. Generating footsteps, ambience and effects in one request gives you a single stereo file you cannot rebalance. Separate passes cost more and are far more useful in an edit.
- Exclude aggressively.
no music, no voicesis nearly always worth including. Models add both as texture and both will fight your real mix. - Keep clips short. Accuracy degrades over length as the event list grows, and eight seconds of correctly synced audio is worth more than thirty seconds that drift.
Native audio versus v2a
Many current video models can produce sound in the same generation as the picture, exposed as a switch rather than as a separate step. When that option exists it is usually the better one, because sound and image are decided together and synchronisation is not a matching problem at all.
The split is simply about what already exists. If you are generating the shot, turn native audio on and treat v2a as the fallback for clips that came out silent or where the generated sound was wrong. If the footage is filmed, or came from a model without audio, or you deliberately want to keep the picture and replace the sound, v2a is the only route.
One caveat about native audio worth knowing before you rely on it: it tends to include music and voices whether or not you asked, and you cannot separate them from the effects afterwards, because it is one mix. For anything that has to sit under real dialogue, generating the picture silently and running v2a for effects gives you control that native audio does not.
Where it fits against real foley
Traditional foley exists because recorded location sound is thin, inconsistent, and full of things you do not want. A foley artist performs the sounds in sync while watching picture, which produces material that is clean, exaggerated in the right places, and separated into layers.
Generated audio covers the least interesting part of that job well: ambience beds, background layers, and effects nobody will inspect. It is weakest exactly where a foley artist adds value, which is performance. A footstep that gets slightly heavier as a character hesitates is an acting choice, and a model watching the picture does not make choices. Use v2a to fill the ninety percent of a soundtrack that only needs to be present, and keep the moments that carry meaning under manual control.