The zone, not the point
Focus is a plane, but sharpness is a range. Only one distance is truly in focus; either side of it, detail degrades gradually, and the span where the softening is small enough to pass as sharp is what the term names. It is not a fixed number of centimetres. It shifts with how large the image will be shown and how close the viewer sits, which is why a shot that looked fine on a laptop can fall apart on a cinema screen.
The zone is also asymmetric. Roughly a third of it sits in front of the focus plane and two thirds behind, and the imbalance grows with distance. Practically: when focusing on a face in a group, focus on the front row, not the middle, because the sharpness extends backward more generously than forward.
The four levers
- Aperture. The most used control. Opening up from f/8 to f/2 collapses the sharp zone dramatically. It also changes exposure, so on set this is where ND filters come in.
- Focal length. A 135mm lens at f/4 gives a much narrower zone than a 24mm at f/4. Long lenses are how you isolate a subject you cannot walk up to.
- Subject distance. The strongest lever of all, and the cheapest. Halving the distance to the subject shrinks the sharp zone far more than opening the aperture one stop.
- Sensor size. Bigger sensors need longer lenses for the same framing, so they deliver less depth of field. This is the whole reason large-format digital cinema cameras look the way they do.
Deep focus is a choice too
The default assumption is that blur equals quality, and that is a modern habit rather than a rule. Deep focus lets a director stage meaning at several distances in one shot: a face in the foreground, a reaction at the far wall, an object between them, all legible, and the audience decides what to watch. Wide-lens comedy needs it, since the joke often lives in the background. Documentary needs it, because you cannot pull focus on something you did not expect.
Blur is the right call when the background is uncontrollable, when you need to hide a location, or when you want the viewer's attention pinned rather than free.
Getting either look from a model
Models do not simulate optics. They reproduce the statistical look of images described a certain way, so the aperture number in a prompt behaves as a style cue rather than a calculation. That still works, because the correlation in the training data is strong.
- ✅
f/1.8, 85mm, background dissolvedfor the narrow look - ✅
f/11, 28mm, deep focus, sharp from the table in front to the doorway at the backfor the wide look - ⚠️
bokeh, cinematicgives blur you did not specify the amount of - ⚠️ Naming an aperture without a focal length is a weak instruction, since the two only mean something together
The harder direction is deep focus, and it is worth knowing why. Anything captioned cinematic in a training set is overwhelmingly shot with the background destroyed, so the model treats blur as a synonym for quality. To beat that, name specific background content you want readable: a clock on the far wall, a sign, a second person's expression. Giving the model a thing to keep sharp works better than giving it an f-number.
What to check in the result
Look at where the sharp zone begins and ends. Generated frames often place it inconsistently: a subject's near shoulder soft while their far ear is crisp, or a sharp band that follows the subject outline rather than a distance. Also check for a plane of blur that runs vertically across the whole frame regardless of what is in it, which means the blur was painted as a gradient rather than derived from the scene.