99 terms
AI & Filmmaking Glossary
Every term you need to describe a shot, and every term you need to generate one. Each entry pairs a plain definition with a real example produced on ZOOOP and the exact prompt behind it.
A
ADR
Editing & SoundADR, short for automated dialogue replacement, is the process of re-recording dialogue in a studio and syncing it to picture after the shoot. It is used when the location audio is unusable, when a line has to change in the edit, or when a performance needs to be reshaped.
automated dialogue replacement · additional dialogue recording · looping · post-sync dialogue
AI Hallucination
Models & ParametersAn AI hallucination is output a model presents confidently even though it is wrong or invented. In text that means a fabricated fact; in images and video it means six-fingered hands, unreadable text on signs, physics that does not hold, and detail an upscaler adds that was never in the source.
hallucination · ai hallucinations · model hallucination · hallucinate
Anamorphic Lens
Look & CompositionAn anamorphic lens uses cylindrical glass to squeeze a wider field of view onto the sensor along one axis, then the image is stretched back out in post to give a wide 2.39:1 frame. The squeeze is why the look comes with oval highlights, horizontal streak flares and stretched edge distortion.
anamorphic · anamorphic look · squeeze factor
Aspect Ratio
Look & CompositionAspect ratio is the proportion of a frame's width to its height, written as width:height. It fixes the shape of the image before anything is placed inside it, which constrains composition more than any other single setting. 16:9 is the widescreen default, 9:16 is vertical, and 2.39:1 is the wide cinema shape.
16:9 aspect ratio · 9:16 aspect ratio · 2.39:1 aspect ratio · cinemascope · frame ratio
Autoregressive Model
Models & ParametersAn autoregressive model produces its output one element at a time, each new element conditioned on everything already emitted. Language models work this way, and so does a growing set of image and video models that predict picture tokens or frames in sequence rather than denoising a whole canvas in parallel.
ar model · autoregressive · next token prediction · autoregressive image model
B
Back Light
LightingA back light sits behind the subject relative to camera and points toward it, drawing a bright edge that lifts the subject off the background. Push it further and the subject goes to silhouette; add haze or rain and it becomes the only light that lets you see the air in a scene.
backlight · hair light · kicker
Bird's Eye View Shot
Camera & ShotsA bird's eye view shot looks straight down at a scene from high above, with the lens roughly perpendicular to the ground. Depth collapses, so figures read as shapes moving across a flat plane rather than as people we are standing among.
overhead shot · top-down shot · god's eye view
Blue Hour
LightingBlue hour is the window before sunrise and after sunset when the sun is below the horizon and the only remaining daylight is scattered skylight, which arrives cool, soft and almost directionless. Its value is balance: for a short time the sky and the artificial lights in a scene sit at the same exposure.
twilight · civil twilight
Bokeh
Look & CompositionBokeh is the aesthetic character of the out-of-focus areas in an image, not the amount of blur. It covers the shape of defocused highlights, how smoothly the blur falls off, and whether the background dissolves cleanly or breaks into busy edges. Lens design produces it; aperture size only decides how much of it you see.
bokeh balls · background blur · out-of-focus highlights
C
Camera Angles
Camera & ShotsCamera angles describe where the camera sits relative to the subject's eyeline, and therefore what the audience is invited to feel about that subject. The angle never changes what is in the frame, only the vantage point on it, which is why the same performance reads as powerful from below and vulnerable from above.
types of camera angles · camera angle · shooting angles
Camera Movements
Camera & ShotsCamera movements are the ways a camera changes position or orientation during a shot, from a pan that rotates it in place to a dolly that carries it through space. Movement adds information across time rather than within a frame, which is why a moving shot reveals, follows or reframes instead of simply showing.
camera movement · types of camera movement · camera moves
CFG Scale
Models & ParametersCFG scale is the strength of classifier free guidance, the setting that decides how far a diffusion model is pushed toward your prompt and away from what it would have drawn unprompted. Low values give loose, soft results, high values give literal, contrasty ones, and every model family has its own working range.
what is cfg scale · guidance scale · classifier free guidance · cfg · prompt strength
Chiaroscuro
LightingChiaroscuro is the use of strong contrast between light and dark to model form, usually built from a single hard source with little or no fill. The name is Italian for light-dark, and the technique moved from Renaissance painting into film noir and horror largely unchanged.
chiaroscuro lighting · light and shadow contrast
Cinematic Lighting
LightingCinematic lighting is not a single technique but a set of habits: one dominant direction, contrast that is allowed to go dark, sources that appear to come from somewhere in the scene, and light shaped so only part of the frame is lit. Crews rarely use the phrase; they name the specific technique instead.
lighting techniques · film lighting terms · filmic lighting
CLIP
Models & ParametersA CLIP model is a pair of encoders, one for text and one for images, trained so that a caption and its picture land in the same place in a shared embedding space. Image generators use its text encoder to turn a prompt into the numbers that steer generation.
clip · text encoder · contrastive language-image pre-training · clip score · t5 encoder
Close Up Shot
Camera & ShotsA close up shot fills the frame with one thing, most often a face from the shoulders up or a single object. Because context is cropped out, the audience has nowhere else to look, which is why this is the size that carries emotion, decision and detail.
close-up · extreme close up · big close up · insert
Consistency Model
Models & ParametersA consistency model is trained so that any point along a noise-to-image trajectory maps straight to the same endpoint, which lets it finish a generation in one to four steps instead of dozens. It is the distillation technique behind every fast, turbo, flash and lightning tier you see on a model catalog.
model distillation · latent consistency model · lcm · step distillation · turbo model
Continuity Editing
Editing & SoundContinuity editing is the set of conventions that make a cut feel invisible, so a scene assembled from many separate shots reads as one continuous event. Its core rules are the 180 degree rule, eyeline match, consistent screen direction, and cutting on action.
eyeline match · 180 degree rule · invisible editing · continuity system · classical continuity
ControlNet
Models & ParametersControlNet is an add-on network that conditions a diffusion model on a structural map extracted from a reference image, such as a pose skeleton, a depth map, or an edge outline. The prompt still decides the style and content, but the layout is pinned to the map you supply.
what is controlnet · ip adapter · control net · openpose controlnet · depth map conditioning
Crane Shot
Camera & ShotsA crane shot is taken from a camera mounted on a crane or jib arm, so the camera physically rises, descends, or arcs through the air during the take. The vertical travel is what defines it, and it is what separates the move from a dolly, which stays on the ground, or a tilt, which only rotates.
jib shot · boom shot · crane up
Cross Cutting
Editing & SoundCross cutting alternates between two or more actions in different places, usually happening at the same time, so the audience follows both at once. It is the standard way editors build suspense, draw a comparison, or make separate lines of story feel like a single event.
crosscutting · parallel editing · intercutting · parallel action
D
Denoising Strength
Models & ParametersDenoising strength decides how far your input image is pushed back into noise before a diffusion model starts to denoise it again, which sets how much of the original survives. Sampling steps decide how many passes that rebuild takes. Together they are the two dials behind every image to image edit.
denoising strength · sampling steps · denoise strength · image strength · steps
Depth Map
AI GenerationA depth map is a grayscale image where each pixel encodes distance from the camera instead of color, conventionally with near surfaces bright and far ones dark. Generation pipelines use a depth map as a structural control signal, because it carries the layout and volume of a scene without carrying its style, so you can change how an image looks while keeping where everything is.
depth estimation · monocular depth · depth pass · z-depth
Depth of Field
Look & CompositionDepth of field is the distance between the nearest and farthest points in a scene that look acceptably sharp. A narrow zone isolates a subject against blur, a wide zone keeps foreground and background legible at once, and four things move it: aperture, focal length, subject distance, and sensor size.
DOF · focus range · deep focus
Diegetic Sound
Editing & SoundDiegetic sound is any sound that exists inside the world of the story, so the characters could hear it: dialogue, footsteps, a car engine, a radio playing in the room. Anything the characters cannot hear, such as the score or a narrator, is non-diegetic.
non-diegetic sound · source music · actual sound · diegesis
Diffusion Model
Models & ParametersA diffusion model generates by starting from random noise and removing a little of it at a time, predicting at each step what the image would look like with less noise, guided by your prompt. Training teaches it to reverse a gradual noising process, and that step-by-step structure is what makes the output steerable.
what is a diffusion model · how do diffusion models work · latent diffusion · video diffusion · diffusion vs gan · denoising diffusion
Diffusion Transformer (DiT)
Models & ParametersA diffusion transformer is a diffusion model whose denoising backbone is a transformer rather than a convolutional UNet. The latent is cut into patches, each patch becomes a token, and every token attends to every other one at every layer. That single change is what gave current models legible text, long prompts that hold, and video measured in seconds instead of frames.
dit · dit model · mmdit · transformer diffusion backbone
Dolly Zoom
Camera & ShotsA dolly zoom moves the camera toward or away from a subject while zooming the lens the opposite way, so the subject stays the same size while the background appears to stretch or compress. The subject holds still and the world behind it changes shape, which is why the effect reads as vertigo or dawning realisation.
vertigo effect · Hitchcock zoom · zolly · contra-zoom · trombone shot
Dutch Angle
Camera & ShotsA dutch angle is a shot where the camera is rolled sideways so the horizon sits at a slant instead of level. The tilt reads as instability, which is why it signals unease, disorientation, or a character losing control.
canted angle · oblique angle · dutch tilt
E
F
Fill Light
LightingA fill light raises exposure on the shadow side of a subject without creating shadows of its own, which makes it the main control over contrast in a setup. It is measured against the key as a ratio, and taking fill away with black flags (negative fill) is as much a technique as adding it.
fill · bounce fill
Film Grain
Look & CompositionFilm grain is the visible texture produced by clumps of light-sensitive silver crystals in photographic emulsion. Faster stocks have larger crystals and coarser grain. Digital sensors have no grain at all, so the look is now added deliberately to soften digital cleanliness and to signal a period or a format.
grain · grain structure · film emulation
Fine-Tuning
Models & ParametersFine tuning AI models means continuing training a pretrained model on your own images so it learns a subject, style, or product it did not know before. It changes the weights, which is what separates it from prompting and from reference images, and it is the last resort rather than the first.
dreambooth · textual inversion · checkpoint model · stable diffusion checkpoint · custom model training
Flow Matching
Models & ParametersFlow matching is a training objective that teaches a model a velocity field carrying noise to data along nearly straight paths. Rectified flow is the straight-line version that current image and video families are built on. Because the path is straight, sampling needs far fewer steps than a classic diffusion schedule to reach the same quality.
rectified flow · flow matching model · velocity prediction · rf
Foley
Editing & SoundFoley is the practice of performing everyday sounds by hand in a studio, in sync with the picture, so they can be recorded cleanly and mixed at the right level. Footsteps, cloth movement, keys, cups, and door handles in a finished film are almost always foley rather than location audio.
foley sound · foley artist · foley effects · foley sound effects
Foundation Model
Models & ParametersA foundation model is a large model pretrained on broad data and then adapted to many downstream jobs rather than built for one. In generative media it is the base checkpoint an entire ecosystem hangs off, since fine-tunes, LoRAs, control adapters, distilled fast variants and task specific versions are all built against a particular base.
base model · foundation models · pretrained model · base checkpoint
Frame Interpolation
AI GenerationFrame interpolation synthesizes new frames between existing ones, either to raise a clip's frame rate or to stretch it into slow motion. Modern methods estimate how every pixel moves between two frames and render the in-between along that motion, rather than cross-fading the two images.
video frame interpolation · keyframe interpolation · interpolation · frame blending · fps boost
Frame Within a Frame
Look & CompositionA frame within a frame is a composition where something in the scene, such as a doorway, window, mirror, archway or overhanging branches, encloses the subject and forms a second border inside the edges of the shot. It adds depth, directs attention, and can imply confinement or observation.
framing device · natural framing · foreground framing · doorway shot
G
GAN
Models & ParametersA GAN (Generative Adversarial Network) is a pair of neural networks trained against each other, a generator that produces candidate images and a discriminator that judges whether they look real. The contest drives quality up in a single forward pass, which is why GANs remain standard for upscaling and face restoration even though diffusion replaced them for open-ended synthesis.
gans · generative adversarial network · generative adversarial networks · adversarial network · esrgan · stylegan
Generative AI
Models & ParametersGenerative AI is any model that produces new content (images, video, audio, text) by sampling from a probability distribution it learned during training, rather than retrieving or editing an existing file. Because every run draws a fresh sample, the output changes even when your input does not.
generative ai · gen ai · genai · generative artificial intelligence
Golden Hour
LightingGolden hour is the stretch just after sunrise and just before sunset when the sun sits low enough that its light travels through far more atmosphere, arriving warm, soft and strongly directional. Long shadows plus an open sky filling the shadow side are what make it read as flattering.
magic hour · golden light
Golden Ratio Composition
Look & CompositionGolden ratio composition places the subject using the proportion 1:1.618, either as a phi grid whose lines sit closer to the centre than a thirds grid, or as a spiral that the eye is meant to follow inward. It is a slightly tighter alternative to thirds, and the spiral is the part that has practical value.
golden ratio · phi grid · golden spiral · Fibonacci spiral · composition techniques photography
H
I
Image Segmentation
AI GenerationImage segmentation divides a picture into labeled regions at the pixel level, producing a mask rather than a bounding box. Inside a generation pipeline it is the step that decides which pixels an edit is allowed to touch, which puts it underneath background removal, object removal, masked inpainting, and any per-subject adjustment.
semantic segmentation · instance segmentation · masking · cutout
Image-to-Image (i2i)
AI GenerationImage to image, usually shortened to i2i, is the generation task where an input image conditions the output. The model rewrites an existing picture instead of starting from noise alone, so composition, color, and pose carry over to whatever degree the strength setting allows.
img2img · image to image translation · what is image to image · reference image
Image-to-Video (i2v)
AI GenerationImage to video, usually shortened to i2v, is the generation task where a still image becomes the starting frame of a clip and the model only has to invent motion. Appearance is locked by the picture you supply, which removes the largest source of randomness in video generation.
img2vid · image2video · what is image to video · still to video
Inpainting
AI GenerationInpainting regenerates the pixels inside a mask you draw and leaves everything outside it untouched. It is how you remove an object, swap a detail, or repair a mistake without re-rendering the whole picture, and the quality of the result depends far more on the mask than on the prompt.
what is inpainting · ai inpainting · video inpainting · generative fill · masked editing
Insert Shot
Camera & ShotsAn insert shot is a tight shot of a detail that exists inside the scene, such as a hand on a door handle, a watch face, a note, or a gun on a table, cut into the wider coverage. It carries information the wide framing cannot, and it gives the editor a place to compress time or hide a join.
cutaway shot · detail shot · insert
J
J-Cut
Editing & SoundA J cut is an edit where the audio from the next shot starts before the picture changes, so you hear the new scene while still looking at the old one. It is the standard way to make a transition feel led rather than abrupt, and it is why most dialogue scenes never cut on picture and sound together.
J-cut · audio lead · pre-lap · split edit
Jump Cut
Editing & SoundA jump cut removes a piece of time from the middle of a continuous shot, so the subject appears to jump forward while the framing stays roughly the same. It breaks continuity on purpose, which is why it reads either as restlessness and time slipping, or, in talking-head video, as plain compression.
jump cuts · jump cutting
K
Key Light
LightingThe key light is the dominant source in a setup, the one that sets exposure and decides where every shadow falls. Its position controls the shape of a face, its size controls how hard the shadow edges are, and everything else in the setup is a response to it.
main light · key
Keyframe
AI GenerationA keyframe is a frame you fix in advance so the model has to pass through it. In AI video the two that matter are the first frame and the last frame, and supplying either one converts an open-ended generation into a constrained one. It is the single biggest lever on consistency.
first frame · last frame · start frame · end frame · first-last frame
L
L-Cut
Editing & SoundAn L cut is an edit where the audio from the outgoing shot continues after the picture has already changed, so you keep hearing one scene while looking at the next. In dialogue it is what lets an editor cut to a listener without breaking the line being spoken.
L-cut · split edit · audio hangover · overlapping audio
Latent Space
Models & ParametersLatent space is the compressed coordinate system a generative model works inside, where an image is a few thousand numbers instead of millions of pixels. Generation is a path through that space, decoded back to pixels at the end, which is why a seed is reproducible and why two nearby points look like relatives.
what is latent space · latent · latent representation · latent vector · latents
Leading Lines
Look & CompositionLeading lines are linear elements in a frame, such as a road, a railing, a shadow edge or a row of columns, that pull the viewer's eye along a path toward the subject. They work because the eye follows continuity automatically, so the line decides what gets looked at first and in what order.
converging lines · diagonal lines · lines in composition
Lens Flare
Look & CompositionA lens flare is stray light that scatters and reflects between the glass elements of a lens instead of forming part of the image, showing up as streaks, coloured ghost blobs, or a milky wash over the frame. It happens when a strong light source is in the shot or just outside it.
flare · anamorphic flare · veiling glare · ghosting
Lip Sync
AI GenerationLip sync is the task of driving a face in a photo or video so its mouth matches a supplied audio track. The model repaints the mouth, jaw, and usually the whole lower face frame by frame so the visemes line up with the sound, while the rest of the shot is left as it was.
lipsync · lip syncing · audio driven face · mouth sync
Long Take
Camera & ShotsA long take is a single uninterrupted shot that runs far longer than the average cut around it, often a minute or more, so a whole scene plays without an edit. Duration is the only requirement: the camera can be locked off on a tripod or travelling through a building.
oner · continuous take · sequence shot
LoRA
Models & ParametersA LoRA (Low-Rank Adaptation) is a small add-on file that shifts a base model's behaviour toward one specific subject, style or concept without retraining the model itself. It is typically 10 to 300 MB against a multi-gigabyte checkpoint, and it is the standard way to lock a recurring character or house look across many shots.
what is a lora · lora model · lora training · lora fine tuning · what is lora ai · lora vs checkpoint · low-rank adaptation
Low Angle Shot
Camera & ShotsA low angle shot places the camera below the subject's eyeline so the lens looks upward. The vantage point makes the subject loom over the viewer, which reads as power, threat or heroism, and it replaces the floor in the frame with ceiling or sky.
low angle · upward angle · worm's eye view
M
Master Shot
Camera & ShotsA master shot is one continuous take that covers a whole scene from start to finish, framed wide enough to hold every actor and every important action. It is the take an editor can always cut back to, and the reference that keeps the closer angles consistent.
master angle · master scene technique
Match Cut
Editing & SoundA match cut joins two shots that share a shape, a movement, a sound, or an idea, so the transition reads as a deliberate link instead of a break. It is the standard way to move between scenes, times, or places without reaching for a dissolve or a title card.
matching cut · graphic match · match cuts · audio match cut
Match on Action
Editing & SoundA match on action cut lands in the middle of a movement, so the same gesture starts in one shot and finishes in the next. Because the eye is busy following the motion, the change of angle registers as continuous rather than as an interruption, which makes it the most reliable way to hide a cut.
cutting on action · action match · matching action · action cut
Medium Shot
Camera & ShotsA medium shot frames a person from roughly the waist up, close enough to read facial expression and still wide enough to show gesture and posture. It is the default size for dialogue because it carries performance and body language in the same frame without forcing the audience to choose.
mid shot · waist shot · medium two shot
Montage
Editing & SoundA montage is a sequence of short shots cut together to compress time, show a process, or build an idea that no single shot could carry on its own. The meaning comes from the relationship between the shots rather than from what happens inside any one of them.
montage sequence · montage editing · film montage
Motion Transfer
AI GenerationMotion transfer copies the movement of a subject in a driving video onto a different character, so a still image or a new identity performs the same action. The pipeline extracts a pose track from the driver, retargets it to the target's proportions, then generates frames conditioned on that track plus your reference identity, which is why it is also called pose transfer.
pose transfer · motion reference · character animation · driving video
Multimodal AI
Models & ParametersMultimodal AI describes models that read or write more than one kind of data, usually some mix of text, images, video and audio held in a shared representation. In creative tools it is the reason you can hand a model a reference frame plus a written note and have it understand both at once instead of one after the other.
multimodal model · multimodal models · omni model · vision language model
N
Negative Prompt
Models & ParametersA negative prompt is a second piece of text a diffusion model is steered away from, the mirror image of the prompt it is steered toward. It reliably suppresses styles, materials, and recurring artefacts, and it is unreliable at removing specific objects from specific places.
negative prompting · negative prompt examples · undesired prompt · exclude prompt
Negative Space
Look & CompositionNegative space is the empty area around and between the subjects in a frame. It is not leftover room: the size and shape of the emptiness is what gives the subject scale, direction and emotional weight, which is why deliberately empty frames read as calm, lonely, or tense rather than unfinished.
empty space · white space · breathing room
O
Optical Flow
AI GenerationOptical flow is a per-pixel map of how content moved between two consecutive frames, stored as a 2D displacement vector for every pixel. It is the measurement that interpolation, video-to-video, stabilization, and video upscaling pipelines rely on to know what should stay the same across frames, which makes it the machinery behind temporal consistency.
flow field · motion estimation · temporal consistency · flow map
Outpainting
AI GenerationOutpainting extends an image beyond its original borders, generating new pixels outside the frame so they continue the existing scene. It is the same masked fill mechanism as inpainting, just pointed outward: you enlarge the canvas, mark the empty area as the region to fill, and the model paints what would plausibly have been there.
generative fill · image extension · uncrop · canvas expand
Over the Shoulder Shot
Camera & ShotsAn over the shoulder shot frames one character past the shoulder of another, so the audience sees the speaker from roughly the listener's position. The near shoulder anchors the frame in a real place, which is what makes a conversation feel shared rather than assembled from separate portraits.
OTS · over-the-shoulder · dirty single · reverse angle
P
Pan and Tilt
Camera & ShotsPan and tilt are the two ways a camera rotates from a fixed position: a pan turns it horizontally, a tilt turns it vertically. Neither changes where the camera is, which is what separates them from a tracking or crane move, where the camera itself travels through space.
panning · tilting · zoom in zoom out
POV Shot
Camera & ShotsA POV shot places the camera where a character's eyes are, so the audience sees exactly what that character sees. It only reads as subjective when the surrounding edit establishes whose eyes they are, which is why it almost always arrives between a shot of someone looking and a shot of their reaction.
point of view shot · first-person shot · subjective shot
Practical Light
LightingA practical light is a working light source that appears in the shot itself, such as a table lamp, a candle, a neon sign, a television or a car headlight. It gives the audience a visible reason for the light in the frame, and on modern sets it often does much of the lighting as well.
practicals · motivated source
Prompt Engineering
Models & ParametersPrompt engineering is the practice of writing and structuring model input so the output lands where you want it. For image and video models it is mostly about word order, specificity and weighting, because these models do not follow instructions the way a chat model does; they match a description.
what is prompt engineering · prompt template · few shot prompting · system prompt · prompt weighting · prompting
R
Rack Focus
Camera & ShotsA rack focus shifts the plane of focus from one subject to another during a single take, so what the audience is looking at changes without a cut and without the camera moving. It needs shallow depth of field and two subjects at clearly different distances from the lens.
focus pull · pulling focus · snap focus
Rim Light
LightingA rim light is the thin bright outline along the edge of a subject, produced by a source placed behind and slightly to the side. It only exists as a contrast effect, so it reads clearly against a dark background and disappears entirely against a bright one.
edge light · rimlight
Rotoscoping
AI GenerationRotoscoping is the practice of isolating a moving subject from its background frame by frame, producing an alpha matte that can be composited over something else. It started as literal tracing over live-action footage and is now mostly done by segmentation and matting models that propose the mask automatically.
alpha matte · green screen · roto · ai rotoscoping
Rule of Thirds
Look & CompositionThe rule of thirds divides the frame into a 3x3 grid and places the subject or the horizon on one of the four intersections or on the lines themselves, rather than dead center. Offsetting the subject leaves unequal space around it, which gives the frame a direction and keeps the eye moving.
thirds grid · off-center composition
S
Seed
Models & ParametersA seed number is the integer that initialises the random noise a diffusion model starts from. Hold the seed and every other setting fixed and you get the same output again, which is what turns generation from a slot machine into a controlled experiment where one variable changes at a time.
image seed · random seed · seed value · fixed seed
Shallow Depth of Field
Look & CompositionA shallow depth of field means only a thin slice of the scene is sharp while everything nearer and farther falls into blur. It is the standard way to isolate a subject from a background, produced by combining a wide aperture, a longer lens and a short distance to the subject.
shallow focus · subject isolation · blurred background look
Shot Types
Camera & ShotsShot types are the standard names for how a camera frames a subject. Each name fixes three things: how much of the subject the frame holds (shot size), where the camera sits relative to the subject (angle), and whether the camera moves during the take. Naming all three is what makes a shot repeatable.
types of camera shots · camera shots and angles · camera terms · shot list vocabulary
Style Transfer
AI GenerationStyle transfer takes the look of one image, its palette, mark-making, texture, and light quality, and applies it to the content of another. Classical neural style transfer optimized a single output against two loss targets. Diffusion pipelines now do the same job from a style reference image or a prompt, which is faster and more flexible but less literal about copying texture.
neural style transfer · style reference · stylization · style matching
T
Text-to-Image (T2I)
AI GenerationText to image is the generation task where a model turns a written prompt into a still picture with no reference image. A text to image model samples from what it learned rather than retrieving anything, so the same words with a different seed produce a different but equally valid image.
t2i · txt2img · text2img · what is text to image
Text-to-Video (T2V)
AI GenerationText to video (T2V) is a generation task where a model turns a written prompt into a video clip with no reference image. The model invents the subject, the framing, and the motion at once, which makes it the least controllable and most exploratory of the video tasks.
t2v · txt2vid · text2video · text to video model
Three Point Lighting
LightingThree point lighting is the standard setup built from a key light that establishes exposure and shadow direction, a fill light that controls how dark the shadow side goes, and a back light that separates the subject from the background. It is a starting scaffold rather than a rule, and plenty of finished work uses only one or two of the three.
3 point lighting · classic lighting setup
Tilt Shift
Look & CompositionA tilt shift lens can move independently of the sensor in two ways. Tilt swings the plane of focus so it no longer sits parallel to the sensor, and shift slides the lens sideways or up to keep vertical lines parallel. Tilt is what produces the famous miniature effect; shift is what fixes leaning buildings.
tilt shift lens · miniature effect · perspective control lens
Tokenizer
Models & ParametersA tokenizer is the component that splits text into the discrete units a model can process, then maps each unit to an id that becomes a vector (an embedding). It sets the hard limit on how much of your prompt is read at all, and it decides how unusual words get broken apart, which is why rare names often generate something unrelated.
embedding ai · tokenization · tokens · text encoder · clip tokenizer
Tracking Shot
Camera & ShotsA tracking shot is a shot in which the camera physically travels through space with a moving subject, holding them in roughly the same part of the frame while the background slides past. The movement of the camera itself, not a rotation or a zoom, is what defines it.
trucking shot · travelling shot · follow shot
Transformer Model
Models & ParametersA transformer model is a neural network that treats its input as a set of tokens and uses an attention mechanism to decide which tokens influence which. It powers large language models, and since the shift to diffusion transformers it powers most current image and video generators too.
attention mechanism · self-attention · transformer architecture · dit · diffusion transformer
U
UNet
Models & ParametersA UNet is the convolutional encoder and decoder, joined by skip connections, that predicts the noise to remove at each denoising step of a diffusion model. It was the standard image generation backbone from 2022 to 2024, and its shape is why LoRA and ControlNet exist in the form they do.
u-net · unet architecture · denoising unet · unet diffusion
Upscaling
AI GenerationUpscaling increases the pixel dimensions of an image or video. Classic resampling enlarges what is already there, while AI upscaling (super resolution) uses a model to synthesize detail that the source never contained, which is why it can look sharper than the original and also why it can quietly change a face.
super resolution · super resolution ai · ai upscaling · image enlargement · what is upscaling
V
VAE
Models & ParametersA VAE (variational autoencoder) is the compressor bolted to both ends of a diffusion model. Its encoder turns pixels into a small latent grid cheap enough to denoise, and its decoder expands the finished latent back into an image or into video frames. Most of the detail flaws you notice at 100% zoom are made here, not by the model that did the generating.
vae stable diffusion · variational autoencoder · vae decoder · latent encoder
Video-to-Video (vid2vid)
AI GenerationVideo to video ai takes an existing clip as its main input and re-renders it, restyled, relit, or with objects changed, while the original motion and timing are preserved. The hard part is not the new look, it is holding that look identical across every frame.
vid2vid · v2v · video restyle · video style transfer
Voice Cloning
AI GenerationVoice cloning builds a synthetic voice from a recording of a real speaker, then reads new text in that voice. Current zero-shot systems need seconds rather than hours of audio: the model extracts a speaker embedding from your reference clip and conditions speech synthesis on it, so nothing has to be retrained for each new voice.
voice clone · speaker cloning · zero-shot tts · voice replication
Volumetric Lighting
LightingVolumetric lighting is light you can see travelling through the air, appearing as beams or shafts because particles of haze, smoke or dust scatter it back toward the camera. It needs three things at once: a source with a defined edge, something in the air to scatter, and a dark area for the beam to read against.
god rays · light shafts · atmospheric lighting
W
Whip Pan
Camera & ShotsA whip pan is a pan executed so fast that the image smears into horizontal streaks. It is used either inside a shot, as a violent redirection of attention, or across two shots, where the blur at the peak of the move hides a cut and the two clips read as one continuous camera move.
swish pan · flick pan · whip transition
Wide Shot
Camera & ShotsA wide shot frames a subject in full with visible space above and below, so the audience reads the subject and the environment at the same time. It is the size that explains where a scene takes place and how the people in it are arranged, which is why scenes are usually built outward from one.
long shot · full shot · extreme wide