MiniMax H3 Max image to video prompts work differently from ordinary text-to-video prompts. When the starting image already defines the person, product, location, composition, lighting, and style, the prompt doesn’t need to describe that scene again — its real job is to explain what happens after the first frame.
A product photo may only need a controlled orbit and moving light. A portrait may need a blink, a head turn, moving hair, and a gentle push-in. A landscape may come alive through drifting clouds and shifting light without the main subject moving at all.
This guide covers how to structure MiniMax H3 Max image to video prompts, then gives 15 copy-ready examples for portraits, dialogue, products, vehicles, food, landscapes, illustrations, architecture, vertical video, and first-to-last-frame transitions.
The Most Important Rule for MiniMax H3 Max Image to Video
The source image should answer:
What does the scene look like?
The prompt should answer:
What happens next?
Imagine uploading a photo of a woman standing beside a window. The image already tells the model what she looks like, what she’s wearing, where she’s standing, how the shot is composed, and how it’s lit.
A weak prompt repeats all of that:
Weak prompt:
A beautiful woman with long brown hair stands beside a large window in a cinematic room with warm light.
A more useful prompt tells H3 Max how the still image should develop:
Better prompt:
She slowly turns from the window toward the camera, blinks once, and gives a restrained smile. A light breeze moves several strands of hair while the curtain shifts gently behind her. The camera makes a very slow, gentle push-in. Preserve her facial features, clothing, and the original room composition. Audio: soft room tone and distant city ambience.
The second version adds information the image doesn’t already contain: motion, camera behavior, environmental response, preservation instructions, and sound.
A Simple MiniMax H3 Max Image to Video Prompt Formula
Subject Motion + Camera Motion + Environmental Motion + Timing + Audio + Constraints
Not every prompt needs all six parts — the goal is clarity, not length.
1. Subject Motion
Lead with the most important action: turns toward the camera, takes two slow steps forward, lifts the cup, blinks and smiles, opens the box, begins driving, looks over her shoulder. Give the subject one main action rather than several unrelated ones.
2. Camera Motion
Describe camera movement in plain, concrete language rather than a mood word: pull back / zoom out, push in / zoom in, pan left / right, tilt up / down, tracking shot, handheld, macro, orbit, static camera.
This is the same style fal uses in its own H3 Max documentation, where the example prompt reads: “The camera slowly pulls back from the scene, revealing the full landscape as clouds drift overhead and light shifts across the terrain.” No specialized vocabulary is required — one clear direction is easier to control than three stacked together (avoid combining a zoom, a pan, and a tilt in the same prompt).
3. Environmental Motion
Secondary motion makes a still image feel alive without forcing the subject to move: hair in the wind, rising steam, shifting curtains, drifting clouds, moving reflections, falling rain, background traffic.
4. Timing
For prompts with several events, sequence them: 0–3s she looks outside, 3–6s she turns toward the camera, 6–8s she smiles as the camera moves closer. Useful whenever actions need a specific order.
5. Audio
H3 Max generates sound alongside the video, so describe it directly: quiet room tone, footsteps, distant traffic, engine sound, natural dialogue, restaurant ambience, or explicitly “no music.” Treat sound as part of the scene, not an afterthought.
6. Constraints
Add these when something must stay stable: preserve the original face, keep the product shape unchanged, keep the logo readable, maintain the original composition, no sudden camera movement, no scene change, keep the object fixed in place. Especially useful for product shots and portraits.
15 MiniMax H3 Max Image to Video Prompts
Adapt these rather than copying them blindly — swap in details that match what’s actually visible in your uploaded image.
1. Subtle Portrait Animation
Best for a close-up or medium portrait where the source image already looks correct.
Prompt:
She maintains her original pose, then slowly shifts her eyes toward the camera and blinks naturally. A faint smile develops without changing her facial structure. Several loose strands of hair move gently in a soft breeze. Use a very slow, almost imperceptible push-in with shallow depth of field. Preserve her face, hairstyle, clothing, skin texture, and the original background. Natural room ambience only; no music.
Why It Works
A few readable changes keep portraits stable — stacking walking, speech, dramatic expression, and camera movement all at once is usually what makes them drift.
2. Talking Portrait With Dialogue
Use this MiniMax H3 Max dialogue prompt when the uploaded image already provides the character and framing.
Prompt:
The woman looks directly into the camera and says, “I finally found the one detail everyone was missing.” Keep the delivery conversational and confident, with natural mouth and facial movement as she speaks. She gives a small nod near the end. Keep the camera nearly static, with at most a very subtle push-in — her delivery should carry the shot, not the camera. Preserve her identity, clothing, hairstyle, and background. Audio: clear natural voice, quiet indoor ambience, no music.
Tip
Keep spoken lines short enough to fit the clip comfortably. H3 Max generates natural-looking mouth movement as part of its combined audio-video pass, but its documentation doesn’t specify frame-accurate lip sync — review dialogue shots before relying on it for tight close-ups.
3. Cinematic Character Turn
Works well for fashion portraits, character concepts, and atmospheric photography.
Prompt:
The character remains still for the opening moment, then slowly turns his head toward the light coming from frame left. His coat shifts slightly in the wind as fine rain falls through the background. The camera makes a slow orbiting push-in, arcing around his right side while moving marginally closer and maintaining the same medium framing. Keep his face, clothing, body proportions, and the original nighttime environment consistent. Audio: gentle rainfall, distant traffic and soft footsteps from off screen; no music.
“Cinematic” isn’t the instruction — naming what actually creates the feeling is.
4. Luxury Product Orbit
MiniMax H3 Max product video is one of the strongest image-to-video use cases, since the design is already visible in the source image.
Prompt:
Keep the watch fixed in its original position while the camera performs a slow clockwise orbit around it. A narrow soft highlight travels gradually across the brushed metal case and crystal. Shift focus gently from the crown toward the center of the dial near the end. Preserve the exact watch shape, dial design, proportions, materials, markings, and background surface. Premium studio lighting, controlled movement, no sudden rotation. Audio: subtle mechanical ticking and quiet studio ambience.
The camera moves around the product — the product itself doesn’t deform or spin.
5. Product Macro Push-In
Good for cosmetics, electronics, jewelry, packaging, and other detail-heavy products.
Prompt:
Begin from the exact source composition. The camera makes a very slow macro push-in toward the product while focus gradually shifts from the front edge to the logo. A soft band of light passes across the surface, revealing material texture and small reflections. Keep the product completely stationary and preserve its shape, logo, typography, color, and packaging details. Shallow depth of field, controlled commercial lighting. Audio: subtle room tone with one soft tactile product sound; no music.
Visual identity matters more than dramatic movement here — say so explicitly.
6. Fashion Editorial Motion
Prompt:
The model takes one controlled step toward the camera and then pauses. Her coat and hair respond naturally to a light crosswind while maintaining the original styling. The camera tracks backward slowly at the same pace, keeping her centered in the frame. Preserve her face, outfit design, proportions, accessories, and the architecture behind her. Editorial fashion-film pacing, realistic fabric physics, no sudden pose changes. Audio: footsteps, soft wind and distant city ambience.
One character action plus one camera action reads more clearly than a full runway sequence crammed into 5–15 seconds.
7. Car Photography to Driving Shot
Prompt:
The car begins moving forward smoothly from its original position and accelerates along the wet street. The camera holds a low front three-quarter tracking shot, matching the car’s speed while keeping the body design consistent. Water sprays naturally from the tires and reflections move across the paint as streetlights pass overhead. Preserve the exact vehicle shape, wheel design, headlights, paint color, and surrounding environment. Audio: deep engine tone, wet tire noise, light rain and distant city ambience.
More precise than simply writing “make the car drive cinematically.”
8. Landscape With Environmental Motion
Not every prompt needs a moving subject. This is close to the pattern fal’s own H3 Max example demonstrates: a slow pullback revealing a landscape as clouds drift and light shifts across the terrain.
Prompt:
Keep the mountains and foreground composition fixed while low clouds drift slowly across the valley. Wind moves through the grass in gentle waves and a thin layer of mist travels between the distant trees. The camera performs a slow pullback, gradually revealing more of the surrounding landscape. Maintain realistic atmospheric perspective and natural movement. Audio: soft wind, distant birds and subtle movement through grass; no music.
The landscape stays recognizable while the environment supplies most of the motion.
9. Anime or Illustration Animation
Prompt:
Preserve the original illustration style, linework, colors, character design, and facial proportions. The character blinks once, looks slightly toward frame right, and gently raises her head. Hair ribbons and several strands of hair move in a light breeze while small particles drift through the background. Use a slow, minimal camera push-in without changing the drawing style or adding realistic 3D texture. Audio: soft wind and distant environmental ambience.
Avoid contradictory instructions like “keep the anime style” alongside “make it photorealistic.”
10. Food Photography Animation
Prompt:
Keep the bowl, ingredients, tableware, and original composition unchanged. Fresh steam rises naturally from the dish while the broth’s surface moves subtly. A pair of chopsticks enters from frame right, lifts a small portion of noodles slowly, and pauses above the bowl. Use a gentle macro camera drift and shallow depth of field. Preserve the food texture and color. Audio: quiet restaurant ambience, subtle ceramic movement and soft food sounds; no music.
A clear beginning, action, and payoff — without changing the location.
11. Architecture and Interior Reveal
Prompt:
Preserve the architecture, furniture placement, materials, geometry, and original room layout. The camera begins with a slow forward dolly through the space, moving through the doorway, while afternoon sunlight shifts gently across the floor. Thin curtains move slightly near the windows and tree shadows outside create subtle movement across the walls. Keep all structural lines stable and realistic. Audio: quiet interior room tone, soft outdoor birds and distant wind.
“Keep structural lines stable” matters here — warped walls read as artificial fast.
12. Dramatic Reveal From a Character Image
Prompt:
The character remains facing away from the camera for the first moment. He then slowly turns his head over his shoulder as a bright light appears in the distance behind him. Dust begins moving across the ground while his coat reacts to a sudden gust of wind. The camera pushes closer from behind, arcing slightly as it approaches, ending on a tight over-the-shoulder composition. Preserve the character design and location. Audio: rising wind, distant mechanical rumble and fabric movement; no dialogue.
Stillness → reaction → reveal → closer framing.
13. First and Last Frame Transformation
Some H3 Max image-to-video workflows support an opening image plus an optional ending image. With two frames, don’t spend the prompt describing the destination — the model already has it. Describe how the transition happens.
Prompt:
Begin exactly from the first image. The camera slowly advances toward the building as daylight gradually changes into blue-hour evening. Window lights turn on one area at a time while the sky darkens naturally and reflections become visible on the wet pavement. Transition continuously toward the supplied final frame without abrupt morphing or a hard cut. Preserve the building geometry and camera direction throughout. Audio: light city ambience gradually becoming more active as evening arrives.
14. 9:16 Social Video From a Portrait Image
Vertical source images work well for TikTok, Reels, and Shorts.
Prompt:
Keep the vertical composition and subject identity unchanged. The creator looks into the camera immediately, raises the product into frame, and says, “This is the detail I wish I had noticed sooner.” She then turns the product slightly toward the light and smiles. Use a subtle handheld creator-style camera feel without aggressive shaking. Preserve her face and the product appearance. Audio: clear close voice, quiet room ambience and one subtle product handling sound; no background music.
Start from a source image that already has strong 9:16 composition — output follows the input’s shape.
15. Cinematic Rack Focus Shot
Useful when the source image has both a foreground and background subject.
Prompt:
Begin with focus held on the foreground object exactly as shown in the source image. After a short pause, shift focus smoothly toward the person in the background as they slowly raise their eyes toward the camera. Keep the camera nearly static with only a very subtle drift, letting the focus shift carry the shot. Preserve the original composition, subject identities, lighting, and spatial relationship between foreground and background. Audio: quiet room ambience and a distant environmental sound; no music.
The focus transition carries the shot, not large physical movement.
How Long Should These Prompts Be?
A simple image needs little:
Short prompt:
The camera pushes in slowly as she turns her eyes toward the window. A soft breeze moves her hair and the curtain. Preserve her face and clothing. Quiet indoor ambience.
A more complex scene benefits from sequencing:
Detailed prompt:
For the first three seconds, keep the character nearly still while distant traffic moves through the background. She then turns toward the camera and takes one slow step forward. At the same time, the camera tracks backward slowly at her pace. Wind moves her coat naturally but does not obscure her face. Maintain the original lighting, identity, clothing, and street layout. Audio: footsteps on wet pavement, distant traffic and light wind; no music.
Specific is more useful than verbose.
Best Settings for MiniMax H3 Max Image to Video
Setting | Practical Starting Point |
|---|---|
Resolution | 768p for final evaluation |
Duration | 5s for one simple action |
Duration | 10s for action + camera movement |
Duration | 15s for longer sequences |
Start Image | Required for image-led workflow |
End Image | Optional for first-to-last-frame transitions |
Aspect Ratio | Follows the uploaded starting image |
Prompt Expansion | Balanced is the current API default |
Camera Movement | One clear direction per prompt — avoid stacking several |
A 15-second generation doesn’t automatically beat a 5-second one for a simple product orbit — match duration to the amount of action.
Why MiniMax H3 Max Image to Video Prompts Fail
These are the failure modes specific to working from a source image, where preserving what’s already there matters as much as directing what changes.
The face changes. Usually from combining an aggressive head turn, large body motion, dramatic expression, and strong camera movement at once. Fix: “Preserve her facial structure and hairstyle. She blinks once and turns her eyes toward frame left. Keep the head movement minimal.”
The product changes shape. Don’t ask the object to move in physically impossible ways. Instead of “the perfume bottle spins rapidly while transforming under dramatic lighting,” try: “Keep the perfume bottle fixed. The camera performs a slow orbit around it while a soft highlight travels across the glass. Preserve the bottle shape, label, typography, cap, and liquid color.” Let the camera and light create the movement.
Too much of the scene moves. If face, hair, clothes, camera, furniture, background, and lighting all shift at once, the result gets unstable. Set a hierarchy — one primary motion, one secondary motion, one camera movement, everything else stays put.
The camera direction gets ignored. “Cinematic dynamic camera” describes a feeling; “the camera performs one slow clockwise orbit around the product while keeping it centered” describes an action the model can execute.
The background morphs. Add preservation language: “Maintain the original room geometry, window placement, furniture, and camera perspective. No scene transition.” Especially useful for interiors and product shots.
Pacing feels off — unintentional slow motion, or dialogue that sounds rushed. State pace explicitly (“real-time movement, not slow motion”) and keep spoken lines short enough to fit the clip without being compressed.
Image-to-Video vs Text-to-Video Prompts
Text-to-video describes the world. Image-to-video develops an existing world. A text-to-video prompt has to establish character, clothing, environment, composition, lighting, camera, movement, and sound all at once. With image-to-video, the source image already makes most of those decisions, so the prompt can focus on movement, timing, camera behavior, environmental response, sound, and preservation instead — which is why copying a text-to-video prompt straight into an image-to-video workflow usually wastes the advantage the image already gives you.
Prompt Template
[Main subject] performs [main action]. The camera [pulls back / pushes in / tracks / pans / tilts / orbits / holds static]. [Secondary environmental motion] happens naturally. Preserve [important visual elements]. Keep [constraint]. Audio: [voice / ambience / foley / music direction].
Example: The woman slowly turns toward the camera and gives a restrained smile. The camera makes a gentle push-in while her hair and the curtain move in a light breeze. Preserve her face, clothing, room layout, and original lighting. Keep movement natural and avoid dramatic pose changes. Audio: quiet room ambience and distant city traffic; no music.
FAQ
What should I write in a MiniMax H3 Max image to video prompt?
Describe what happens after the starting frame — subject motion, camera motion, environmental changes, timing, sound, and anything from the source image that must stay consistent.
Should I describe the uploaded image again?
Usually not in detail. Repeat only details that must be preserved or need special attention.
Can MiniMax H3 Max animate a product image?
Yes. Camera or lighting motion — a slow orbit, a macro push-in, a focus shift, a moving highlight — usually works better than making the product itself move aggressively.
Can MiniMax H3 Max animate portraits?
Yes — MiniMax H3 Max portrait animation works best when you start subtle: blinking, a small head turn, a gaze change, hair movement. Too many simultaneous movements make identity harder to preserve.
How specific should camera direction be?
Use plain language — pull back, push in, pan, tilt, tracking, orbit, static — and stick to one direction per prompt. This mirrors fal’s own H3 Max example prompt: “The camera slowly pulls back from the scene, revealing the full landscape as clouds drift overhead and light shifts across the terrain.”
Can I use a first and last frame?
Current H3 Max image-to-video implementations can accept an optional ending image alongside the opening frame, for first-to-last-frame generation describing how the scene moves between the two states.
Does image-to-video support audio?
Yes — H3 Max generates synchronized audio with the video, so prompts can include dialogue, ambience, foley, or music direction.
Does MiniMax H3 Max support lip sync?
It generates audio and video together, and spoken lines typically come with natural-looking mouth movement — but its documentation doesn’t specify frame-accurate, phoneme-level lip sync. Treat it as expressive talking-head animation and review dialogue shots before relying on exact sync.
What’s the best MiniMax H3 Max image to video prompt?
There’s no universal one. A strong prompt separates what the source image already defines from what the video needs to add: one main action, one clear camera direction, one layer of environmental motion, and clear preservation instructions are usually enough.
Final Thoughts
The best MiniMax H3 Max image to video prompts don’t try to recreate the starting image in words — they give the still image something specific to do.
Decide what must stay unchanged, then define one main action, one clear camera direction, and any environmental motion that supports the scene. Add audio when it contributes, and use preservation language when faces, products, or compositions need to hold steady.
Preserve the image → define the action → direct the camera → add secondary motion → describe the sound.
Use the 15 examples above as starting points, then change only what matters for your own image.