MiniMax H3 Max logoMiniMax H3 Max
Loading

A good H3 Max Prompt Guide comes down to one idea: writing a prompt is really about controlling five things — the subject, the action, the camera, the audio, and the timing. Get those five right and H3 Max has a much clearer production brief to follow; leave them vague and you’re giving the model more room to guess.

Because H3 Max generates synchronized audio alongside the video, prompting should cover both what the viewer sees and what they hear. This guide walks through the core prompt formula, camera movement prompts, dialogue and sound direction, text-to-video, image-to-video, and reference-to-video prompts, first-and-last-frame prompting, ready-to-copy templates, real examples with a breakdown of why each one works, and a troubleshooting checklist — everything you need to write better AI video prompts and get consistent results out of H3 Max.

Try the H3 Max Generator

What Is MiniMax H3 Max?

H3 Max is designed for short video generation from text, a starting image, or a set of reference files, with synchronized audio produced in the same generation pass rather than added afterward. H3 Max focuses on fast 480p/768p generation, while standard MiniMax H3 offers a broader feature set, including 2K output and instruction-based video editing. For the full model background, see what H3 Max actually is.

H3 Max at a Glance

H3 Max Specification

Details

Resolution

480p or 768p (768p is the default)

Duration

5–15 seconds per generation

Frame rate

24 FPS

Aspect ratios

21:9, 16:9, 4:3, 1:1, 3:4, 9:16 in text-to-video; follows the source image in image-to-video

Audio

Native, synchronized stereo generated with the video

Generation modes

Text-to-video, image-to-video (with an optional end frame), and reference-to-video

Generation speed

Optimized for fast iteration

Because audio is generated in the same pass as the picture, a well-written H3 Max prompt should direct the soundtrack as deliberately as it directs the shot — we’ll cover exactly how later in this guide.

What Makes a Good H3 Max Prompt?

The most useful principle in this H3 Max Prompt Guide is simple: describe what should happen, how it should look, how the camera should move, and what the audience should hear.

A weak prompt might say:

A cinematic woman walking through a beautiful city.

This gives H3 Max plenty of freedom, but not much control. A stronger H3 Max prompt could say:

A young woman in a dark green coat walks through a rain-soaked downtown street at night. Start with a medium rear tracking shot, then slowly move to her left side as she looks toward a glowing storefront. Reflections shimmer on the pavement. Natural footsteps, distant traffic, light rain, and soft city ambience. No background music. End with a close-up as she smiles toward the camera.

The second version establishes the subject, the action, the environment, the camera, the lighting, the sound, and the ending. That is the basic philosophy behind this H3 Max Prompt Guide.

The H3 Max Prompt Formula

A reliable H3 Max prompt can follow this structure:

Subject
+ Action
+ Environment
+ Camera
+ Visual Style
+ Lighting
+ Timing
+ Dialogue
+ Sound
+ Constraints
+ Ending

You don’t need every element for every generation — this formula is a framework to fall back on when a simple prompt isn’t giving you enough control.

Subject

Clearly identify the main person, character, object, or product.

Instead of: A man Try: A middle-aged Japanese chef wearing a white apron and dark blue shirt

Action

Describe what the subject actually does, in the order it happens.

Instead of: A chef in a kitchen Try: The chef slices fresh vegetables, pauses, looks toward the camera, and raises one eyebrow.

Action is especially important in an H3 Max prompt because video generation needs temporal direction, not a static image description.

Environment

Give the subject a physical context.

Example: A compact open kitchen with warm wooden cabinets, stainless-steel counters, soft morning sunlight, and steam rising from a pot.

Camera

Tell H3 Max how the scene should be filmed.

Example: Start with a medium shot and slowly push toward the chef with subtle handheld movement.

Specific camera instructions are almost always more useful than simply saying “make it cinematic.”

Visual Style

Describe the intended visual language: photorealistic commercial, documentary, vintage film, luxury advertisement, anime-inspired animation, stop-motion clay animation, handheld street documentary, macro product photography.

Lighting, Timing, and Ending

Lighting sets mood (“warm morning sunlight,” “cold blue neon”). Timing controls pacing on longer clips (covered in the next section). And the ending matters more than it seems — don’t forget it.

Example ending: End on a close-up of the finished dish as steam rises from the plate. An explicit ending instruction can make a short video feel far more intentional than letting the model guess where to stop.

How to Structure an H3 Max Prompt

A useful H3 Max Prompt Guide should distinguish between a description and a sequence.

Single-Action Prompts for 5-Second Clips

For a 5-second clip, one primary action is often enough:

A ceramic coffee cup rotates slowly on a black stone table while warm sunlight moves across the surface. Slow macro push-in. Soft room ambience and the faint sound of coffee being poured. End with the cup centered in frame.

Multi-Beat Prompts with Timestamps

For a longer clip, introduce several beats with explicit timing:

0–3 seconds:
A woman walks toward the train station.

3–7 seconds:
She stops as a train arrives and turns toward the platform.

7–10 seconds:
The camera moves closer as she smiles and steps forward.

Audio:
Footsteps, distant station announcements, train brakes, and light crowd ambience.

This turns the prompt into a simple shot plan, and timestamps work just as well for cueing a sound effect or a music swell at a precise moment.

H3 Max Prompt Expansion Mode: Disabled vs Balanced vs Quality

Separate from the prompt itself, H3 Max offers a Prompt Expansion Mode setting that controls how much the model rewrites your wording before generating.

Mode

Best For

Disabled

When you want H3 Max to follow your exact wording

Balanced

The default, and the best starting point for most prompts

Quality

More extensive prompt expansion and refinement, at the cost of extra processing time

For prompt testing, keep the expansion mode unchanged between generations. If you’re comparing two versions of a prompt, changing the expansion mode at the same time makes it harder to tell whether the prompt edit or the expansion mode caused the different result. Quality mode can add up to roughly 30 seconds of prompt-rewriting time before generation even starts, so it’s better suited to a final render than to rapid back-to-back testing.

H3 Max Camera and Motion Prompts

Camera direction is one of the most important parts of an effective H3 Max prompt. Instead of writing “dynamic camera,” specify the movement.

Useful Camera Instructions

  • Push-in — Slowly push the camera toward the subject.

  • Pull-out — Gradually pull the camera backward to reveal the environment.

  • Tracking shot — Track alongside the character as she walks through the market.

  • Orbit — Slowly orbit around the product while keeping it centered.

  • Pan — Pan from the character toward the city skyline.

  • Handheld — Use subtle handheld movement with natural micro-shakes.

  • Macro — Use a macro close-up with a slow focus transition from the foreground to the product.

A strong H3 Max prompt usually benefits from one dominant camera movement rather than several unrelated movements happening at once.

Weak vs Strong Camera Prompt

Weak: Make the camera cinematic and dynamic.

Strong: Begin with a medium close-up and slowly dolly forward as the character raises the glass. Keep the camera at eye level and maintain shallow depth of field.

The second instruction gives H3 Max a much clearer description of the intended movement and framing.

H3 Max Motion Prompting

Motion should be described chronologically — list what happens first, second, and third, the same way you’d write a shot list, rather than piling every movement into one sentence.

The 1+1+1 Rule for H3 Max Prompts

When a shot starts to feel overloaded, fall back on a simple constraint: one primary action, one primary camera movement, one clear ending.

A red sports car waits at a wet intersection. The headlights turn on. The car accelerates through the intersection while the camera tracks from a low front three-quarter angle. Water sprays from the tires. End with the car disappearing into the fog.

This is far more controllable than “a cool sports car drives fast with lots of cinematic camera movements.” For short H3 Max clips, the 1+1+1 rule helps prevent a prompt from becoming overloaded — you can still add lighting, sound, dialogue, and environmental detail around the core shot, just without stacking multiple competing actions or camera movements on top of it. If the subject is walking, talking, turning, opening a door, picking up an object, and running within five seconds, the result becomes less predictable. Keep this rule in your back pocket for any prompt that starts to feel crowded — camera, dialogue, or product shots included.

H3 Max Dialogue Prompts

Dialogue deserves its own structure. A basic H3 Max dialogue prompt looks like this:

A woman stands beside a kitchen counter and looks directly at the camera.

She says naturally:
"Here is the easiest way to make this recipe."

Keep the delivery warm, confident, and conversational.
The dialogue should be clearly audible and naturally synchronized with her lips.

For multiple characters, identify the speakers clearly:

S1 is the woman in the red jacket.
S2 is the man wearing glasses.

S1 says:
"Are we late?"

S2 replies:
"Not yet. The train is still here."

S1 smiles and turns toward the platform.

This is far more useful than “two people have a conversation.” The more important the dialogue is to the scene, the more explicitly the prompt should identify the speaker, the exact words, and the delivery.

H3 Max Audio and Sound Prompts

One of the strongest reasons to make sound part of an H3 Max prompt is that the model generates audio alongside the video, rather than requiring you to treat sound as a separate production step.

Four Sound Categories to Direct

  • DialogueThe woman speaks clearly in a calm voice.

  • Ambient soundLight rain, distant traffic, and soft city ambience.

  • Physical sound effectsFootsteps on wet pavement and the metallic sound of the door opening.

  • MusicA subtle minimalist piano score begins as the character enters the room.

You can also explicitly exclude unwanted music: Natural environmental sound only. No background music. That’s a clearer instruction than simply writing “realistic sound.”

Using Timestamps for Precise Audio Timing

When a sound effect needs to land exactly on an action, use the timestamp format from earlier in this guide:

0.0–2.0s: Footsteps on gravel, ambient wind. 2.0–3.5s: A door creaks open, footsteps stop. 3.5–5.0s: Soft piano note as the door fully opens.

Writing Non-English Dialogue

For non-English dialogue, write the spoken line directly in the target language and specify the intended delivery, rather than describing the line in English — this gives the model the actual words to match pronunciation and lip movement to, instead of leaving translation and delivery to guesswork.

H3 Max Image-to-Video Prompts

H3 Max can generate video from a starting image as well as from text, including an optional end frame. When using an image, the prompt should focus on what happens after the image — don’t simply re-describe everything already visible in it.

Instead, write: Use the uploaded image as the opening frame. Keep the character’s face, hairstyle, clothing, and overall composition consistent. She slowly turns toward the camera, smiles, and raises her hand. The camera makes a subtle push-in. Preserve the original lighting and background.

This gives H3 Max a clear description of what should remain stable and what should change.

Character Reference Prompt

Preserve the character’s facial identity, short black hair, red jacket, silver earrings, and body proportions throughout the entire shot. Do not change the clothing or hairstyle.

Product Reference Prompt

Preserve the exact shape, proportions, color, label placement, and material appearance of the uploaded product. Rotate the product slowly on the table without changing its design.

Environment Reference Prompt

Preserve the architecture, color palette, lighting direction, and major objects from the reference image while adding subtle movement from people and environmental elements.

The key idea across all three is distinguishing between locked details and moving details — the same discipline carries over to full reference-to-video prompts below.

H3 Max Reference-to-Video Prompts

H3 Max also supports reference-to-video generation, combining reference images, video clips, and audio into a single generation — up to 12 files total, with the output ratio set to Adaptive rather than something you choose yourself. Reference video must be MP4/MOV and at least 2 seconds per clip, with combined video reference material capped at 15 seconds; you need at least one image or video reference to generate. Current file, format, and credit limits are listed on the H3 Max generator and pricing pages.

Because you can mix several files, the prompt should assign each reference a clear role rather than leaving H3 Max to guess which file controls what. A useful structure:

Reference roles + Subject + Action + Camera + Environment + Audio + Constraints + Ending

For example:

Reference Image 1: character appearance (face, hair, outfit)
Reference Image 2: product design (shape, color, logo)
Reference Video: movement reference (walking pace, gesture style)
Reference Audio: voice or sound reference (tone, accent)

The character from Reference Image 1 walks toward the product from
Reference Image 2, moving the way Reference Video shows. Medium
tracking shot, soft daylight. She picks up the product and turns it
toward the camera. Audio: match the voice in Reference Audio for her
one line of dialogue, plus quiet room ambience.

Keep the character's face and outfit from Reference Image 1 unchanged.
Keep the product's shape and branding from Reference Image 2 unchanged.
End with the product facing the camera, label visible.

This mirrors the locked-details discipline from image-to-video prompts, just spread across multiple files instead of one starting image. If you’re porting a reference-heavy prompt over from a different H3 workflow, check the file-count and duration limits above first — they’re specific to how H3 Max’s reference-to-video endpoint is configured today and may not match another platform’s limits.

H3 Max First and Last Frame Prompts

When using a first frame and an optional ending frame, describe the journey between them:

Starting frame → movement → transformation → ending frame

The video begins with the closed luxury watch centered on a marble table. Over the next eight seconds, the camera slowly moves closer as the watch lid opens. Warm reflections travel across the metal surface. The internal gears begin rotating. End exactly on the uploaded final frame showing the open watch from a slightly closer perspective.

Avoid writing “start with Image 1, then Image 2” — that doesn’t explain what should happen in between, and H3 Max is left guessing how to fill the middle of the shot.

Matching Aspect Ratio to Your Scene

H3 Max supports six aspect ratios in text-to-video mode, and picking the right one before you write your prompt saves a re-generation later.

Aspect Ratio

Best For

9:16

TikTok, Reels, Shorts

1:1

Social feed posts and square ads

4:3

Classic framing and product-focused shots

3:4

Portrait-oriented content

16:9

YouTube, websites, presentations

21:9

Cinematic and ultra-wide shots

If you’re animating an existing image, image-to-video mode inherits the aspect ratio of your source image automatically — choose the ratio at the photo stage, not the prompt stage. Reference-to-video works the same way in spirit: the generator sets it to Adaptive rather than a manual dropdown, so the output ratio follows your reference material instead of something you specify in the prompt.

Best H3 Max Prompts by Use Case

The best H3 Max prompt depends on the type of video you're creating. The examples below cover common use cases including cinematic scenes, product advertisements, character-driven clips, action sequences, 9:16 social videos, dialogue, and image-to-video generation — each with a short note on why the prompt is written the way it is, not just what it says.

Cinematic Video Prompt

A lone explorer walks across a frozen black-sand beach at blue hour, carrying a small lantern. Begin with a wide shot from behind and slowly track forward. Wind moves the explorer's coat while distant waves crash against dark volcanic rocks. Cold blue lighting contrasts with the warm lantern glow. Audio: strong coastal wind, footsteps on wet sand, distant waves, and subtle fabric movement. No music. End with the explorer stopping at the waterline.

Why this works: the camera move, the lighting contrast, and a full sound layer are all named explicitly, so there are fewer important elements left for the model to infer.

Product Advertisement Prompt

A premium black smartwatch rests on a polished stone surface. Start with an extreme macro shot of the metal edge and slowly orbit around the watch as soft studio lights create controlled reflections. The screen activates and displays a clean fitness dashboard. Keep the product geometry, screen proportions, and branding consistent. Audio: subtle electronic activation sound with a quiet premium ambience. End with the complete watch centered in frame.

Why this works: the orbit movement is named directly, the product's geometry is explicitly locked so it doesn't warp mid-rotation, and the audio stays minimal and premium instead of generic "exciting" music.

Character Prompt

A young woman wearing a cream sweater sits beside a large window on a rainy afternoon. She looks down at a notebook, writes one sentence, then looks outside and smiles softly. Use a slow medium-to-close push-in. Natural overcast lighting, realistic skin texture, shallow depth of field. Audio: rain against the window, pencil movement, quiet room ambience. No music.

Why this works: one quiet, continuous action — write, look up, smile — plays to H3 Max's strength at a single clear beat, and the ambient audio reinforces the mood instead of competing with it.

Action Prompt

A cyclist races down a narrow mountain road at sunrise. Start with a low tracking shot beside the bicycle. The cyclist leans into a sharp turn, accelerates downhill, and passes a row of pine trees. The camera follows smoothly without cutting. Audio: tire movement, wind, chain vibration, and distant birds. End with a wide shot revealing the mountain valley.

Why this works: the camera is told to move with the subject rather than cut between angles, which gives the model a clearer continuous-motion instruction and can make a fast-moving scene easier to control.

9:16 Social Video Prompt

Vertical 9:16 fashion video featuring a young designer wearing a black oversized jacket in a modern studio. Start with a waist-up shot, then slowly tilt downward as she turns toward the camera and adjusts the jacket. Maintain the subject in the central vertical composition. Bright softbox lighting, clean background, high-end fashion commercial aesthetic. Add subtle fabric movement and natural room ambience. End on a confident close-up.

Why this works: the composition instruction does the real work — most vertical-video failures come from a horizontally-composed shot simply being cropped, and naming the central vertical composition heads that off directly.

Dialogue Video Prompt

A barista behind a coffee counter looks up as a customer approaches. She says naturally, "Good morning — the usual oat latte?" in a warm, familiar tone. The customer nods and smiles. Medium two-shot, static camera at counter height. Audio: espresso machine hissing quietly in the background, light café chatter, no music.

Why this works: the line is quoted exactly with a tone cue attached, so H3 Max has both the words and the delivery to match lip movement and audio to, instead of inventing a line on its own.

Image-to-Video Prompt

Use the uploaded product photo as the starting frame. Keep the bottle's shape, label, and lighting exactly as shown. The bottle slowly rotates 180 degrees on the pedestal while soft studio light sweeps across the label. Camera holds static, centered. Audio: quiet ambient studio tone, no music. End with the label facing the camera.

Why this works: it describes only what should change — the rotation and the light sweep — and explicitly locks what shouldn't, which is the core discipline behind every H3 Max image-to-video prompt in this guide.

H3 Max Prompt Templates

These video prompt templates cover the four most common workflows — copy and fill in the brackets for a fast starting point.

Text-to-video template:

Create a [style] video featuring [subject] in [environment].
The subject [primary action].
The camera [shot type + camera movement].
Lighting: [lighting description].
Visual style: [style].

From [time] to [time], [action].
Then [next action].

Audio: [dialogue / ambience / sound effects / music].
Keep [important details] consistent.
End with [final composition or action].

Image-to-video template:

Use the uploaded image as the starting frame.
Preserve [character/product/environment details].
The subject then [action].

Camera: [camera movement].
Lighting: [lighting].
Motion: [specific movement].
Audio: [sound design].

Do not change [locked details].
End with [final action or composition].

Reference-to-video template:

Reference Image 1: [what this reference controls]
Reference Image 2: [what this reference controls]
Reference Video: [what this reference controls]
Reference Audio: [what this reference controls]

Create [shot description] using the references above.
Keep [locked details] unchanged from their reference source.
Camera: [camera movement].
Audio: [ambient sound / effects / dialogue].
End with [final action or composition].

Product video template:

Create a premium product advertisement featuring [product].
Keep the product shape, color, proportions, materials,
and visible branding consistent.

Start with [opening shot].
The camera [movement].
The product [action/transformation].

Lighting: [lighting].
Environment: [environment].
Audio: [sound effects / ambience / music].
End with the product centered in a clean hero shot.

Dialogue template:

[Character description] stands in [environment].

S1 says:
"[Exact dialogue]"

S1 speaks with a [delivery style] voice.
The character [action while speaking].
Camera: [shot + movement].

Audio includes [ambience / effects / music].
Keep the dialogue clear and synchronized.
End with [reaction or final action].

Weak vs Strong H3 Max Prompts

Weak: A beautiful cinematic coffee commercial with a woman drinking coffee.

Strong: A young woman sits alone at a wooden café table beside a large rain-covered window. Start with a medium side profile as she lifts a white ceramic cup and takes a slow sip. The camera gently pushes closer while warm interior light contrasts with the cool blue street outside. Steam rises naturally from the coffee. Audio: soft rain, distant café conversation, ceramic cup movement, and a quiet espresso machine. No vocals. End on a close-up of the cup as she places it back on the table.

The strong version gives H3 Max a sequence, camera direction, environmental context, physical motion, and sound design — not just adjectives.

Why Fast Iteration Matters for H3 Max Prompts

H3 Max is designed for fast generation, so treat your first prompt as a draft rather than a final answer. Generate it, identify what the model got wrong, and adjust one variable at a time — camera angle, then lighting, then action — instead of trying to perfect every detail before your first generation.

A Fast Testing Workflow for H3 Max Prompts

A five-step loop built around that speed advantage:

  1. Write a baseline prompt using the core formula above.

  2. When available, test the prompt at 480p first to validate composition, motion, and timing before spending credits on a higher-resolution render.

  3. Change exactly one element — camera move, lighting, or dialogue tone — and regenerate.

  4. Compare the two clips side by side and keep whichever version follows the prompt more accurately.

  5. Once the structure is locked, regenerate the winning prompt at 768p for your final clip.

This workflow makes each round of prompt iteration easier to evaluate than trying to perfect one long, over-specified prompt from the start.

Start Free on H3 Max

Common H3 Max Prompt Problems

Problem

Fix

H3 Max changes the character

Add what must remain consistent: “Preserve the character’s face, hairstyle, clothing, accessories, and body proportions throughout the entire shot.”

H3 Max ignores the camera movement

Replace vague terms with a specific movement: “Slowly dolly forward from a medium shot to a close-up” instead of “use a dynamic cinematic camera.”

H3 Max generates too much motion

Apply the 1+1+1 rule and simplify to one clear action: “The character remains seated and slowly turns her head toward the window.”

H3 Max gets dialogue wrong

Specify the speaker and exact line: S1 says: “This is the final design.” Avoid paraphrasing important dialogue.

H3 Max adds unwanted music

Be explicit: “Natural environmental audio only. No background music.”

H3 Max changes the product

List what must stay fixed: “Preserve the exact logo position, product color, dimensions, material, and button placement.”

H3 Max doesn’t reach the desired ending

Describe the final composition directly: “End with the product centered in a clean close-up, fully visible and facing the camera.”

How to Improve an H3 Max Prompt Step by Step

Don’t rewrite the entire prompt every time a generation misses one detail. Use this workflow instead:

  1. Create the basic shot — define the subject, environment, and main action.

  2. Fix the composition and motion — add shot size, camera movement, and describe the action chronologically.

  3. Add audio — specify dialogue, ambience, physical effects, and music.

  4. Lock important details — identify the character, product, clothing, or environment details that shouldn’t change.

  5. Define the ending — tell H3 Max exactly where the shot should finish.

  6. Change one variable at a time — if the character is right but the camera movement is wrong, change only the camera instruction instead of rewriting everything.

This makes each round of prompt iteration much easier to evaluate.

H3 Max vs Standard MiniMax H3: Which Should You Use?

Once your prompting technique is solid, the other lever worth getting right is which model you’re running it on. See what H3 Max actually is for the full background. As a rule of thumb: use H3 Max when fast iteration, a 480p/768p finish, or reference-driven generation fits the workflow, and reach for standard MiniMax H3 when a project specifically needs 2K output, instruction-based video editing, or an open-weight/self-hosted deployment.

Use Case

Better Choice

Fast prompt testing

H3 Max

Rapid variations

H3 Max

480p/768p workflow

H3 Max

Reference-to-video (up to 12 files)

H3 Max

2K output

Standard MiniMax H3

Instruction-based video editing

Standard MiniMax H3

Open-weight / self-hosted workflows

Standard MiniMax H3

For example, choose H3 Max when you’re testing a social ad concept and need to see five camera-angle variations of the same product shot in a few minutes; choose standard H3 once you know the exact shot and need to push it to 2K with a text-based edit, like changing a jacket color without regenerating the whole scene.

For the full feature-by-feature breakdown, see the complete H3 Max vs MiniMax H3 comparison.

Testing H3 Max Prompts on a Budget

H3 Max is priced in credits: 480p generation uses 1 credit per second while 768p uses 2 credits per second, so a 5-second 768p clip costs 10 credits at the current generation rate. Because 480p costs half as much as 768p, a practical workflow is to test several prompt variations at 480p first, then render the winning version at 768p. Reference-to-video costs more than a plain text- or image-to-video generation of the same length, since each reference image adds its own charge and reference video is billed per second on top of the output rate — budget for that before uploading a full 12-file reference set. Check the current H3 Max pricing page for the latest new-account allowance and plan details.

H3 Max Prompt Checklist

Run through this before every generation:

  • Is the main subject clearly identified?

  • Is the primary action specific?

  • Is the environment defined?

  • Is the camera direction clear, with one dominant movement?

  • Is the lighting described?

  • Are actions written in chronological order?

  • Is dialogue assigned to the correct speaker, with the exact line included?

  • Are sound effects, ambience, and music each specified (or explicitly excluded)?

  • Have you added timestamps for any cue that needs precise timing?

  • Are important character or product details locked for consistency?

  • Does your chosen aspect ratio match the platform you’re publishing to?

  • Is the ending clearly described?

  • Is the prompt written for the right mode — text-to-video, image-to-video, or reference-to-video?

  • Have you tested at 480p before spending credits on a 768p final render?

Frequently Asked Questions

What is the best H3 Max prompt format?

A practical H3 Max prompt can follow Subject + Action + Environment + Camera + Style + Audio + Timing + Ending. You don’t need every element in every prompt, but this structure makes complex shots much easier to control.

How long should an H3 Max prompt be?

There’s no fixed ideal prompt length. Use enough detail to define the subject, action, camera, sound, and ending without adding conflicting instructions — a short, specific prompt works well for a simple shot, while more complex scenes benefit from additional camera, timing, dialogue, and sound instructions.

What is the 1+1+1 rule for H3 Max prompts?

It’s a simple constraint for keeping a prompt controllable: one primary action, one primary camera movement, and one clear ending. Add lighting, sound, and dialogue around that core structure instead of stacking multiple competing actions or camera moves into a single short clip.

Can H3 Max generate dialogue and sound?

Yes. H3 Max can generate synchronized audio with the video, including spoken dialogue and environmental sound when specified in the prompt — so sound should be described directly inside the prompt rather than treated as a post-production afterthought.

Is H3 Max free to use?

Yes — new accounts currently get free trial credits, enough to test the prompt formula in this guide on a real clip before buying more. Check the H3 Max pricing page for the current allowance and paid plans.

How do I keep characters consistent in H3 Max?

Clearly identify the details that must remain stable — face, hairstyle, clothing, accessories, body proportions. When using a starting image, explicitly tell H3 Max which visual characteristics to preserve, as shown in the reference prompt examples above.

How do I write an H3 Max prompt for vertical 9:16 video?

Mention the vertical composition and keep the main subject inside the central visual area. Describe framing specifically for a vertical screen instead of simply adding “9:16” to an otherwise horizontal composition.

What are the best H3 Max prompts for image-to-video?

Focus the prompt on what happens after the source image, not on what’s already visible in it — describe the motion, the camera move, and the audio, and explicitly state which visual details (face, product shape, lighting) should stay locked. The image-to-video template earlier in this guide is a ready-to-copy starting point.

What are the best H3 Max prompts for product videos?

Lock the product’s shape, color, proportions, and branding first, then describe the camera movement, lighting, and any sound design. Explicitly fixing the product’s geometry is what keeps it from warping during a rotation or close-up shot — see the product video template above.

Can I combine reference images, video, and audio in one H3 Max prompt?

Yes. H3 Max’s reference-to-video mode accepts up to 12 files across images, video, and audio. Give each reference file a clear role in the prompt — which one controls character, product, movement, or voice — rather than leaving H3 Max to infer it.

Why does H3 Max ignore part of my prompt?

Overloaded prompts often contain competing instructions. Simplify the scene, prioritize the single most important action, use clear chronological order, and make constraints explicit — a focused production brief works better than a pile of unrelated adjectives.

What is the best H3 Max Prompt Expansion mode?

Balanced is the best starting point for most H3 Max prompts. Use Disabled when you want the model to work from your exact wording, and Quality when you want more extensive prompt expansion. When testing prompt variations, keep the expansion mode unchanged so you can isolate the effect of your prompt edits.

Final Thoughts

The best H3 Max Prompt Guide principle is simple: clarity beats keyword stuffing. Start with the subject and action, add the environment and camera, then define motion, dialogue, sound, important constraints, and the final frame. The same video prompt structure works whether you’re writing cinematic prompts, product ads, or dialogue-driven scenes. For image-based workflows, separate what should remain unchanged from what should move — and treat prompt iteration as cheap and fast rather than precious, since that discipline is what actually separates strong AI video prompts from generic ones.

Generate Your First Clip