How to Create a Realistic 10-Second AI Water Drinking Video from One Image
Artificial Intelligence has made it possible to turn a single photograph into a realistic short video with natural movement. One of the most interesting examples is creating a video where a person in a still image appears to continue drinking water naturally, while keeping the same face, clothing, body appearance, background, lighting, and overall composition.
In this tutorial, we will explore a 10-second image-to-video AI prompt designed for exactly this type of animation. The goal is to take an existing image of a man drinking water from a clear plastic bottle and animate only the drinking action. The result should look natural and realistic while preserving the original appearance of the person and the unusual water effect around his waist.
This type of prompt can be useful for AI video creators, social media content creators, bloggers, digital artists, and anyone interested in transforming still images into short cinematic videos.
What Is Image-to-Video AI?
Image-to-video AI is a technology that takes an existing image and generates a moving video from it using a text prompt. The original image acts as the visual starting point, while the prompt explains what should happen during the video.
According to Runway's official image-to-video prompting guide, the input image establishes important information such as the composition, subject matter, lighting, and visual style. The text prompt should primarily describe the desired motion, camera movement, timing, and progression of the scene.
This is different from traditional text-to-video generation. With text-to-video, the AI has to create the entire visual scene from a written description. With image-to-video, you already provide the visual starting point, making it easier to maintain the original appearance of the scene.
For this project, the uploaded photograph provides the man, water bottle, clothing, background, lighting, and unusual water splash effect. The prompt simply tells the AI how the man should move.
The Main Idea Behind This Video
The concept is very simple.
A man is standing in front of a gray studio background while drinking water from a transparent plastic bottle. A dramatic stream of water appears around his waist, creating a surreal visual effect.
Instead of changing the image into a completely different scene, the goal is to animate the existing image.
The man should continue drinking water naturally. He should slowly tilt the bottle, take several sips, swallow naturally, and breathe subtly.
At the same time, the rest of the scene should remain stable.
The face should remain the same. The hairstyle should remain the same. The shirt should remain the same. The trousers, belt, wristwatch, water bottle, background, lighting, and overall composition should remain consistent.
This is important because the purpose of image-to-video animation is not always to create a completely new scene. Sometimes the goal is simply to bring an existing photograph to life.
The Complete 10-Second Prompt
Here is the complete prompt used for this AI video:
Use the uploaded image as the exact reference and preserve everything exactly as shown. The man naturally continues drinking water from the clear plastic bottle, keeping the exact same face, facial features, hairstyle, skin tone, body shape, navy blue shirt, beige trousers, brown belt, wristwatch, pose, water splash effect, gray background, lighting, camera angle, and overall composition. He slowly tilts the bottle and takes several natural sips of water, with realistic swallowing and subtle breathing movements. The water bottle remains in his hand and stays aligned naturally with his mouth. The water splash around his waist remains visually consistent and realistic without changing its shape dramatically. Static camera with only very subtle cinematic movement, photorealistic, ultra-detailed 4K, natural motion, realistic physics, smooth animation, no face change, no clothing change, no body transformation, no background change, no new objects, no text, no watermark, maintain the exact identity and appearance of the reference image throughout the entire 10 seconds.
Why the Reference Image Matters
The reference image is one of the most important parts of the entire process.
A high-quality starting image gives the AI a clear understanding of the person and the environment. If the image contains blurry facial features, distorted hands, strange objects, or other visual problems, those problems can sometimes become more noticeable when the image is animated.
Runway's official guide recommends using a high-quality input image that is free from visual artifacts because problems in the original image can become more visible during video generation.
For this reason, it is better to begin with a sharp photograph where the face, hands, bottle, clothing, and background are clearly visible.
Why the Prompt Focuses on Motion
One of the most important principles of image-to-video prompting is to focus on movement rather than repeatedly describing everything that is already visible in the image.
Runway explains that image-to-video prompts should focus almost exclusively on motion. Important motion components include subject action, environmental motion, camera motion, timing, direction, and speed.
That principle is especially useful for this water-drinking video.
The image already shows the man drinking water, so the prompt does not need to spend most of its words explaining what a man looks like. Instead, it tells the AI what should happen next.
For example, the prompt specifies that he should:
- Continue drinking water.
- Slowly tilt the bottle.
- Take several natural sips.
- Swallow naturally.
- Breathe subtly.
- Keep the bottle aligned with his mouth.
- Maintain the original pose and appearance.
These instructions give the AI a clear animation target.
The Importance of Natural Drinking Motion
Drinking water may appear to be a simple action, but realistic animation requires several small movements.
The man's hand needs to hold the bottle naturally. The bottle should remain connected visually to his mouth. His head should move slightly as he drinks. His throat can show a subtle swallowing movement, while his body continues with natural breathing.
The motion should not be exaggerated.
If the bottle suddenly moves away from his mouth or his head turns too quickly, the video may look artificial. A slow and controlled movement is therefore more appropriate for this scene.
The phrase “slowly tilts the bottle and takes several natural sips” is designed to communicate this idea clearly.
Maintaining Face and Clothing Consistency
Another major goal of this prompt is maintaining the appearance of the original person.
The prompt specifically asks the AI to preserve:
- Facial features
- Hairstyle
- Skin tone
- Body shape
- Shirt
- Trousers
- Belt
- Wristwatch
- Pose
- Background
- Lighting
- Camera angle
This is particularly useful when you want the final video to look like a natural animation of the original photograph rather than a completely new AI-generated character.
Character consistency can still vary depending on the model being used, so it is important to understand that a prompt cannot guarantee perfect identity preservation in every generation.
However, providing a clear reference image and keeping the requested motion simple can improve the chances of maintaining visual consistency.
The Water Splash Effect
The unusual water splash around the man's waist is one of the most visually interesting parts of the image.
Because this effect is already present in the original image, the prompt asks the AI to keep it visually consistent rather than dramatically changing it.
The instruction:
“The water splash around his waist remains visually consistent and realistic without changing its shape dramatically.”
helps communicate that the splash should remain part of the scene.
The AI can add subtle movement to individual water droplets while avoiding a major transformation of the entire effect.
This is useful because large changes in a complex visual effect could make the scene look inconsistent from one frame to another.
Camera Movement
The prompt uses a mostly static camera with only subtle cinematic movement.
This is intentional.
A dramatic camera rotation or fast zoom would make the scene more complicated and could increase the chance of unwanted changes to the person's appearance.
A stable camera allows the viewer to focus on the drinking action.
Runway's prompting guidance identifies camera movement as one of the core motion components that can be controlled in image-to-video generation.
For this particular concept, a static or very gently moving camera is appropriate because the main subject action is already interesting.
Suggested 10-Second Timeline
The video can be imagined as a simple sequence.
0–2 seconds: The man holds the bottle to his mouth and begins drinking naturally.
2–5 seconds: He slowly tilts the bottle and takes several sips. His head and hand make small natural movements.
5–7 seconds: He continues drinking while breathing naturally. The water effect around his waist remains consistent.
7–9 seconds: He takes another gentle sip while maintaining the same posture and appearance.
9–10 seconds: He finishes the drinking motion while the camera remains stable, creating a clean ending.
Sequential prompting can also be used when more precise timing is required. Runway's official guide explains that prompts can describe events in sequence or use rough timestamps to control the order of actions.
Why Simplicity Is Important
A common mistake when creating AI videos is trying to include too many actions in a short clip.
For a 10-second video, one main action is usually enough.
In this case, the main action is drinking water.
Adding unnecessary movements such as walking, dancing, changing clothes, changing locations, spinning around, or dramatically changing the camera could make the animation less consistent.
Runway recommends starting with a simple prompt focused on the most important motion and adding details gradually when refinement is needed.
This means you can first generate a simple drinking animation and then improve it if necessary.
What to Do If the First Result Is Not Perfect
AI video generation is an iterative process.
The first result may not always be perfect. Perhaps the bottle moves strangely, the person's face changes slightly, the water effect becomes distorted, or the drinking motion is too fast.
Instead of changing the entire prompt, try changing only the part that causes the problem.
For example, if the bottle moves too quickly, strengthen the motion instruction by saying:
“The man slowly and naturally tilts the bottle toward his mouth.”
If the camera moves too much, you can emphasize:
“The camera remains stable with an extremely subtle cinematic push-in.”
If the character changes appearance, reinforce the importance of maintaining the original visual identity.
Runway describes iteration as a normal part of prompting: you generate a result, review how the model interpreted the instruction, and then refine the prompt.
Using the Prompt for Social Media
Once you have generated the 10-second video, there are many places where you can use this type of content.
You can potentially publish it as:
- YouTube Shorts
- TikTok videos
- Facebook Reels
- Instagram Reels
- WhatsApp Status
- AI video demonstrations
- Blog content
- Creative visual experiments
Short videos are especially useful for demonstrating what AI image-to-video technology can do.
You can also create a series using the same concept. For example, one video could show the person drinking water, another could show him putting the bottle down, and another could show a different visual effect while keeping the same character.
Adding the Video to Blogger
If you want to publish this tutorial on Blogger, you can create a new blog post and add the article, reference image, and generated video.
Google's official Blogger Help explains that you can add images and videos directly to blog posts. Videos can be uploaded from your computer or selected from YouTube.
Blogger Help – Add images and videos
Blogger also allows you to create a new post, preview it, save it as a draft, or publish it when you are satisfied with the result. You can also use labels to organize your content.
Blogger Help – Create, edit, manage, or delete a post
Conclusion
Creating a realistic 10-second AI video from a single photograph can be surprisingly simple when the prompt is designed correctly.
The main objective of this particular prompt is to animate one natural action: the man drinking water. Instead of changing the entire scene, the AI is instructed to preserve the original face, clothing, body, water bottle, water splash, lighting, background, camera angle, and composition.
The most important principle is to let the image establish the visual appearance while the text prompt focuses on movement. This matches the guidance provided by Runway for image-to-video prompting.
A high-quality reference image, a simple action, realistic timing, stable camera movement, and careful iteration can all help produce a better result.
If the first generation is not perfect, don't be discouraged. Adjust the prompt gradually and test different versions until the drinking movement looks natural and the character remains consistent.
This approach can be used for many other creative projects. You can animate people drinking, walking, waving, smiling, talking, sitting, or performing other simple actions while preserving the visual identity of the original image.
The key is to start simple, focus on the motion you actually want, and refine the prompt based on the result. With practice, image-to-video prompting can become a powerful technique for creating engaging AI videos for blogs, social media, and creative projects.

0 Comments