Why motion reference beats a text prompt here
Describing a dance in words is one of the least reliable things you can ask a video model to do — timing, weight and beat are exactly what language is bad at specifying. Binding a fixed reference clip instead removes the whole problem: the motion is already correct, and the model's only job is to make your animal perform it. That is why this template produces consistent results where a prompt-driven attempt at the same thing usually does not.
Photos that map cleanly
Full body, facing camera, on a background that separates from the animal. The reference needs to find a head, a torso and four limbs to drive; a face close-up gives it one of those, and a pet curled asleep gives it none. Good separation from the background also keeps the coat pattern intact through the movement, which is the part that makes it read as your pet rather than a generic one.