The three jobs it does
Calling this material supplementary undersells it, because a piece assembled only from its main footage usually does not work. B-roll is doing specific structural work, and knowing which job you need decides what to shoot.
Illustration. Someone describes a place, a process or an object, and you show it. This is the obvious use, and the easiest to over-serve: literal illustration of every noun becomes exhausting quickly.
Concealment. A talking head has pauses, repetitions and stumbles you want gone, and cutting them creates jump cuts. Laying supplementary footage over the join hides it entirely, which is why interview editing is largely the craft of deciding where to cut away.
Pace. A sequence with no room in it reads as relentless. A few seconds of something quiet lets an idea land before the next one arrives, and this is the use most people underestimate.
What to shoot, and how much
The quantity guidance is unintuitive: plan for three to five times the length of the final piece. The reason is that you cannot know which moments need covering until the edit exists, and by then the shoot is over.
What makes a shot usable is not how good it looks.
- Long enough to trim. Five to ten seconds of a single continuous idea. Short clips can only land in one place.
- Static or very slow. A locked shot can be cut in anywhere. A whip pan can be cut in almost nowhere.
- One subject. A frame with three things happening in it belongs to one specific moment in the edit.
- Ambiguous about time. Anything with a clock, a screen, or a distinctive lighting change constrains where it can go.
- No lip movement. Visible speech ties the shot to specific dialogue and makes it useless as cover.
Generating it
This is one of the clearest cases where generative video is genuinely appropriate rather than a compromise, and the reasons are structural. These shots are short, they contain no dialogue, they do not need to match another angle of the same space, and a large share of them are pure texture rather than narrative. Every weakness of generated video, which shows up in continuity, long takes, faces and speech, is absent from the brief.
Practical approach. Write the shot as one still idea rather than as an event: steam rising from a coffee cup on a workshop bench. Add the optical language that makes it read as photographed rather than illustrated, since that is what has to match your real footage: shallow depth of field, 50mm at f/2, morning light through a dirty window.
Ask for stillness explicitly. Slow static shot prevents the model from inventing a camera move you did not want, and a static clip is more usable in an edit anyway.
Exclude people unless you need them. No people removes the two things generated video handles worst, faces and hands, from a shot that almost never requires either.
Finally, generate variations rather than perfecting one. Four takes of the same idea at slightly different framings gives an editor choices, which is exactly what this material is for.