The picture is the first frame, not the complete video
After uploading the picture, what really needs to be told to the video model is not "what is in this picture", but what happens in the next few seconds. The picture has already determined the character, composition, light and general style, and the text should mainly describe the action, environmental changes and shots.
A picture is responsible for the starting point, and the prompts is responsible for the time.
Method 1: One first frame plus one main action
The easiest way is to just upload a first frame and write an action. Don’t ask your character to turn, walk, wave, have the camera wrap around, and explode in the background all at the same time. The more movement there is, the easier it is to squeeze together in just a few seconds.
- Character: The camera is fixed, the character looks up out of the window, his hair is gently blown by the breeze.
- Product: The camera moves forward slowly, the bottle remains stationary, and the background light moves from left to right.
- Illustration: The clouds drift slowly to the right, slight ripples appear on the water surface, and the picture maintains a hand-painted texture.
Runway's guide to converting images to video also recommends focusing your text on subject movement, environmental movement, camera movement, speed, and direction, starting with the most important movement and working your way up.
Method 2: Use the first and last frames to control the end point
If you already know where the video will end up, you can provide both the first and last frames. The tool will try to compensate for intermediate changes, which is suitable for products changing from packaging to display, images transitioning from day to night, or a shot moving from a distant shot to a close shot.
Adobe Firefly and Google Flow currently both offer first and last frame methods. The two pictures should have a composition that can change continuously; if the character position, camera angle and background are completely inconsistent, the model can only be forcibly deformed or suddenly switched.
First using the same picture and making slight changes to get the final frame is usually more stable than taking two unrelated pictures and directly connecting them.
Method 3: Continue to bring reference materials for multiple videos
When making multiple videos, only relying on the last frame of the previous one to generate them, the characters and products may still become further and further away. Some tools allow pictures of people, products, or scenes to be used as ongoing references. Google Flow calls this type of material ingredients and recommends using a clean, simple background for your main body or product reference.
The reference picture is not a lock button. Faces, clothes, packaging text, buttons and colors still need to be checked after continuous generation. If an important product is deformed, it is better to shorten the shot and reduce the movement, or keep the product still and only let the background and camera move slightly.
The five most likely places to fail
- The character's face suddenly changes: the face in the input image is too small, blurry, or has errors; first change to a clear image, and then reduce large head turns and occlusions.
- Product packaging deformation: The action and shots are too intense; keep the product still and only move closer, with light and shadow or background movement.
- The camera is running around on its own: the prompt only describes the atmosphere, not the shot; clearly use a fixed lens, slowly zoom in, or pan to the left.
- A sudden cut in the middle: multiple scenes are crammed into a short period of time; only one continuous shot is made at a time, and multiple shots are generated separately before being cut.
- Want to be still but keep moving: Video models are inherently prone to motion; write down the only objects allowed to move in the picture, and use editing software to stabilize them after generation.
Also check for motion cues in the input image. Runway's guide gives the example of a car with dust and motion blur that makes it difficult to stay still on cue. The input image is already "rushing forward", but the text is required to not move at all, and the two signals will fight with each other.
A more cost-effective production sequence
- First revise the input image: the face, hands, product text and edges themselves must be correct.
- Only generate the shortest possible clip, verifying one action and one shot first.
- Get the action right, then add speed, environmental changes, and a second shot. Don't write it all at once.
- Generate several more versions, keeping only a few stable seconds in between.
- Finally, go to the editing software to crop, stabilize, add subtitles, sounds and transitions.
Converting pictures to videos is more like producing footage, rather than generating a complete film at once. Controlling each segment to one action, and changing the picture and movement requirements first if it fails, is usually more useful than constantly adding "high quality, cinematic feel".
official information
The first frame, last frame and motion prompt description of this article come from:



