Editing & Sound

What Is B-Roll?

Also called b roll footage, supplementary footage, overlay footage, coverage

B-roll is supplementary footage cut alongside the main material to illustrate it, cover edits, or give a sequence room to breathe. The name comes from the second reel of film that ran beside the primary one, and the primary material is still called A-roll by contrast.

The three jobs it does

Calling this material supplementary undersells it, because a piece assembled only from its main footage usually does not work. B-roll is doing specific structural work, and knowing which job you need decides what to shoot.

Illustration. Someone describes a place, a process or an object, and you show it. This is the obvious use, and the easiest to over-serve: literal illustration of every noun becomes exhausting quickly.

Concealment. A talking head has pauses, repetitions and stumbles you want gone, and cutting them creates jump cuts. Laying supplementary footage over the join hides it entirely, which is why interview editing is largely the craft of deciding where to cut away.

Pace. A sequence with no room in it reads as relentless. A few seconds of something quiet lets an idea land before the next one arrives, and this is the use most people underestimate.

What to shoot, and how much

The quantity guidance is unintuitive: plan for three to five times the length of the final piece. The reason is that you cannot know which moments need covering until the edit exists, and by then the shoot is over.

What makes a shot usable is not how good it looks.

  • Long enough to trim. Five to ten seconds of a single continuous idea. Short clips can only land in one place.
  • Static or very slow. A locked shot can be cut in anywhere. A whip pan can be cut in almost nowhere.
  • One subject. A frame with three things happening in it belongs to one specific moment in the edit.
  • Ambiguous about time. Anything with a clock, a screen, or a distinctive lighting change constrains where it can go.
  • No lip movement. Visible speech ties the shot to specific dialogue and makes it useless as cover.

Generating it

This is one of the clearest cases where generative video is genuinely appropriate rather than a compromise, and the reasons are structural. These shots are short, they contain no dialogue, they do not need to match another angle of the same space, and a large share of them are pure texture rather than narrative. Every weakness of generated video, which shows up in continuity, long takes, faces and speech, is absent from the brief.

Practical approach. Write the shot as one still idea rather than as an event: steam rising from a coffee cup on a workshop bench. Add the optical language that makes it read as photographed rather than illustrated, since that is what has to match your real footage: shallow depth of field, 50mm at f/2, morning light through a dirty window.

Ask for stillness explicitly. Slow static shot prevents the model from inventing a camera move you did not want, and a static clip is more usable in an edit anyway.

Exclude people unless you need them. No people removes the two things generated video handles worst, faces and hands, from a shot that almost never requires either.

Finally, generate variations rather than perfecting one. Four takes of the same idea at slightly different framings gives an editor choices, which is exactly what this material is for.

The prompt for this

A starting point that reliably produces the effect. Adjust the subject and setting; keep the technical clauses.

Slow static shot of steam rising from a coffee cup on a workshop bench, shallow depth of field, dust in the air, morning light through a dirty window, no people, 50mm at f/2, five seconds

Try B-Roll yourself

Open the generator with a starting point already filled in.

Frequently asked questions

Where does the name come from?
From physical film assembly, where the main footage ran on the A-roll and a second reel of supplementary material ran alongside it on the B-roll, allowing dissolves and cutaways to be printed. The terminology outlived the practice, which is why editors still say it about digital files.
What is it actually for?
Three jobs. It illustrates what someone is describing, so a voice talking about a workshop can play over the workshop. It hides edits, since cutting away from a talking head lets you remove a pause invisibly. And it controls pace, giving a sequence somewhere to slow down.
How much do you need?
More than feels reasonable. A common working figure is three to five times the length of the finished piece, because you cannot predict which pauses will need covering until the edit exists. Running out is the most common reason an otherwise finished cut has visible jump cuts in it.
What makes a shot useful rather than just pretty?
Length, stillness and a clear single subject. A five to ten second static shot of one thing can be trimmed anywhere and cut against anything. A three second moving shot with three subjects in it can only be used at one specific moment, which is why beautiful footage often turns out to be unusable.
Is generated footage suitable for this?
It is one of the strongest uses for it. Supplementary footage is short, has no dialogue, needs no continuity with other shots, and is frequently just texture: steam, hands, traffic, light on a wall. Those are exactly the conditions generative video handles well, and the conditions where its weaknesses do not show.

Related terms