AI video models respond best to prompts that read like a shot list, not a wish list. This guide covers the structure that works across the models on WandVideo.
Describe the shot, not the story
Models generate a few seconds at a time. Describe one camera setup, one subject and one action:
A golden retriever shakes water off its coat on a wooden dock at sunset, slow motion, low angle, backlit.
Use the four ingredients
- Subject – who or what, with two or three concrete details.
- Action – a single verb phrase that can happen in 5 to 10 seconds.
- Camera – angle, movement (static, slow pan, dolly in), lens feel.
- Look – lighting, time of day, film stock or style.
Image to video needs less text
When you upload a start image, the model already knows the subject and composition. Describe only the motion: "the camera slowly pushes in while leaves drift across the frame".
Things that usually fail
- More than one scene change.
- Text that must be legible.
- Precise counts ("exactly seven birds").
- Real people by name (blocked by our safety policy anyway).
Start with a 5 second, 720p test to check composition, then re-run at higher resolution once the prompt is right.