The shapes worth knowing
- 16:9 (1.78:1). Screens, streaming, most web players. The 16:9 aspect ratio is where you land when you have no reason to land anywhere else.
- 9:16. The same shape rotated for phones. Full-screen vertical feeds crop or letterbox anything else, so the 9:16 aspect ratio is not a stylistic choice there, it is the container.
- 2.39:1. The wide cinema shape, still commonly called cinemascope. The 2.39:1 aspect ratio reads as film immediately, mostly because nothing else looks like it.
- 1.85:1. Slightly wider than 16:9, the other standard theatrical shape. Rarely worth requesting from a model, since 16:9 is close enough.
- 4:3 (1.33:1). Boxy. Now used deliberately for period, archival, and intimacy, since a narrower frame crowds the subject.
- 1:1. Square. Survives the most feed placements and flatters the fewest compositions.
- 2:1. A compromise shape popularised by streaming originals: wide enough to feel cinematic, safe enough to crop to 16:9.
What the shape does to the composition
A wide frame gives you horizontal room and takes away vertical room. That sounds neutral and is not. Two characters can sit at opposite ends of a 2.39:1 frame with air between them, which is why wide shapes are good at landscape, isolation and confrontation. The same two characters in 9:16 have to be stacked, one behind the other or one above the other, which is why vertical is good at faces, hands and single objects and bad at scale.
Vertical also changes where the eye moves. A tall frame invites scanning up and down, so vertical compositions lean on foreground and background layering instead of left and right placement. This is why a shot that works beautifully in landscape often reads as empty when cropped to vertical: the composition lived in the width you just removed.
Pick it before you generate, not after
The frame shape is an input to almost every image and video model, and it is one of the few parameters models respect exactly, because it is a size argument rather than a description. Use it. Two habits that save real time:
- Match the delivery format at generation time. If the piece ships vertical, generate vertical. Cropping a wide generation to 9:16 throws away most of the frame and usually cuts a limb off the subject.
- If you need both, generate the wide version and shoot a second pass for vertical. Reframing costs less than regenerating, but a purpose-built vertical composition beats a crop every time, and viewers can tell.
How models behave with each shape
Extreme shapes push a model outside its training distribution, and it shows. Requesting a very wide frame often produces duplicated subjects near the edges, because the model has fewer wide examples to draw on and fills the extra width by repeating what it knows. Very tall frames do the same thing vertically, most visibly with stacked or doubled figures.
Two practical countermeasures. First, describe what occupies the extra space, so the model has content for it instead of inventing a second copy of your subject. Second, keep the total pixel count near what the model was trained at: an unusual shape at an unusual size is two problems at once.
Square requests have the opposite failure. They pull everything to dead center, since centering is the cheapest solution in a square frame. If you want an off-center composition, a wider shape will fight you less.