MiniMax H3 Max logoMiniMax H3 Max
Loading

MiniMax H3 Max AI Video Generator

The post-trained version of MiniMax H3 that ranks first for image to video, and returns a finished 5-second clip faster than you can watch it.

Controls how much effort is used to rewrite the prompt before generation. Disabled skips expansion, Balanced returns in about a second, and Quality can spend up to about 30 seconds creating a richer prompt.

MiniMax H3 Max cinematic demo

Your generated video will appear here.

What is MiniMax H3 Max?

An AI video model that turns a written prompt or a single still into a short clip with sound already in it.

MiniMax H3 Max was post-trained by fal on top of the open-weight MiniMax H3 model. Post-training means fal took the base model and kept training it: new data aimed at two things, following the prompt more closely and looking better, with a large share of the compute spent on tasks where the result could be checked automatically. fal also built the model's architecture around its own inference engine, which the team has spent four years tuning for diffusion models.

That last part is why MiniMax H3 Max is fast rather than merely good. A 5-second clip at 768p comes back in under 3 seconds — roughly 35 times the throughput of the official H3 endpoint.

One thing to get straight, because search results currently get it wrong: MiniMax H3 Max is not the same model as MiniMax H3 — the base model — and it does not do 2K. The MiniMax H3 Max resolution ceiling is 768p. The base model is the one that goes to 2K. Module 6 lays out the whole split.

The alternative to a model like this is the old route: shoot or license footage, cut it, then find music and foley and line them up. That is what an MiniMax H3 Max AI video generator replaces — one prompt in, a finished clip out.

MiniMax H3 Max Key Features

Seven things MiniMax H3 Max does in one pass that normally take a timeline, a sound library, and an afternoon.

MiniMax H3 Max feature: Sound generated with the picture

Sound generated with the picture

Every MiniMax H3 Max generation returns with audio already in it — room tone, foley, music, ambience, cut to what is on screen. Describe the sound in the same sentence as the shot and it arrives in the same pass. There is no separate audio step and nothing to line up.

MiniMax H3 Max feature: Camera moves you can direct

Camera moves you can direct

Ask for a low tracking shot chasing a scooter downhill and you get that move, not something near it. The camera holds the subject, the background streaks at the speed the shot implies, and the lean carries through the corner.

MiniMax H3 Max feature: The same character across every shot

The same character across every shot

One request, several locations, one person: hair, jacket, proportions, and face hold as the light changes. Keeping a character stable across cuts is the hard part of animated work, and MiniMax H3 Max does it inside a single generation.

MiniMax H3 Max feature: Two stills become one continuous shot

Two stills become one continuous shot

Give MiniMax H3 Max an opening frame, and optionally a closing one, and it animates the whole journey between them. On the image-to-video endpoint this is an optional end image — one parameter, no extra call.

MiniMax H3 Max feature: It hits the beats in the order you wrote them

It hits the beats in the order you wrote them

Prompt adherence is what fal's post-training targeted first. Name the beats and they arrive in sequence. Words you ask for on screen come back legible and correctly set rather than approximated.

MiniMax H3 Max feature: One look, held to the last frame

One look, held to the last frame

Give MiniMax H3 Max a visual language and it keeps it. A single 15-second generation can cut between six shots without breaking style — same palette, same linework, same lettering from first frame to last.

MiniMax H3 Max feature: 5 seconds of video in under 3 seconds

5 seconds of video in under 3 seconds

Faster than real time. The API returns a timings.inference field with the actual render time on the backend, which lands around 2.5 seconds for a 5-second 768p clip. A 15-second clip takes about 15 seconds.

Specs at a glance

MiniMax H3 Max resolution
480p or 768p. At 16:9, 768p works out to 1344×768 at 24 FPS.
MiniMax H3 Max length
5 to 15 seconds per generation.
MiniMax H3 Max duration limit
15 seconds is the longest single request.
Aspect ratios
21:9, 16:9, 4:3, 1:1, 3:4, 9:16 on text to video; on image to video the output follows your input picture.
Audio
included in every generation, cut to the picture.

What Can You Create with MiniMax H3 Max?

Product shots, social clips, pieces to camera, animation, game-style footage, and slow motion — each from one prompt.

MiniMax H3 Max creation: Product and macro

Product and macro

Espresso pulling into a glass, a watch face catching light, fabric under a raking key. Shallow focus and slow drift read correctly, and the pour, the hiss, and the room come back with the picture.

Short-form social

Vertical at 9:16, 5 to 15 seconds, sound included — the length and shape a feed actually takes.

  • 9:16
  • 5–15 seconds
  • Native audio
MiniMax H3 Max creation: Short-form social
MiniMax H3 Max creation: A line to camera, lip synced

A line to camera, lip synced

Give MiniMax H3 Max a speaker, a setting, and the sentence. The mouth matches the words and the voice sits close and clear.

MiniMax H3 Max creation: Animation and stop motion

Animation and stop motion

Claymation with visible fingerprints, stepped 12fps motion, hand-set lettering, painted linework. Style holds across the whole clip.

MiniMax H3 Max creation: Game-style footage

Game-style footage

Pixel-art side-scrollers with parallax layers, a running cycle, a score counter, and chiptune audio under it.

MiniMax H3 Max creation: Slow motion and nature

Slow motion and nature

Wings at a blur, pollen in a backlit garden, a long lens and very shallow focus, wing hum and birdsong underneath.

MiniMax H3 Max Prompt Examples You Can Copy

Four prompts written for MiniMax H3 Max, with the shot, the beats, and the sound in each one.

Product — mechanical watch on basalt

What works here is the ordered reveal: subject first, a camera path with a clear focus change, controlled light, then exact sound cues tied to visible beats.

A brushed-steel mechanical watch rests on a slab of wet black basalt. Begin on an extreme macro of the crown as a bead of water travels along its knurled edge. The camera slides left and slowly pulls focus through the sapphire crystal to reveal the second hand moving, then arcs into a clean three-quarter hero view. Use one narrow white strip light above, a faint cool rim behind, deep charcoal shadows, and crisp reflections without blown highlights. Keep the engraved brand mark sharp and unchanged. End with the full watch centered while a final droplet lands beside it. Audio: quiet studio room tone, a close metallic tick every second, the soft drag of water over stone, and one precise droplet impact at the end; no music and no voice.

Dialogue — late-night laundromat

What works here is the short quoted line, the pause before and after it, and a restrained camera move that leaves the model room to preserve lip sync and expression.

Inside a nearly empty laundromat at midnight, a tired radio host in a dark green coat sits on a molded plastic chair between two turning dryers. Frame a medium close-up at eye level and make a very slow push toward her face. Blue street light flickers through the front window while warm dryer light rolls across the wall behind her. She looks directly into the lens, waits for one beat, and says clearly: ‘Some nights, the machines are the only things still listening.’ Her mouth must match every word and her expression should move from guarded to quietly amused. Hold for two seconds after the line as she glances toward the spinning clothes. Audio: intimate natural voice, low dryer rumble, a coin clicking inside one drum, distant rain against glass; no score.

Animation — paper observatory

What works here is a strict material and color system, numbered visual beats, and sound effects that reinforce the deliberately stepped stop-motion rhythm.

Handmade paper-cut stop-motion animation at 12 frames per second. A small rust-red fox climbs a spiral staircase inside a midnight-blue cardboard observatory. Start with a wide side view showing layered paper hills through the open dome, then cut to the fox turning a brass paper crank, then to a close-up of a telescope assembled from folded cream card and tiny inked rivets. The dome opens in three distinct stepped movements. A field of punched-paper stars rotates into view and one silver star unfolds into the words ‘LOOK CLOSER’ in neat hand-set lettering. Keep visible paper fibers, scissor-cut edges, slight frame-to-frame jitter, the same fox proportions, and the same limited red, navy, cream, and silver palette throughout. Audio: wooden clicks, paper rustle, a gentle crank squeak, soft night wind, and three bright celesta notes when the silver star opens.

Camera move — market to tram

What works here is a continuous route built from precise height, direction, speed, and subject-lock instructions, with audio perspective changing at the tram door.

One unbroken cinematic shot at dawn in a rain-soaked hillside market. Begin inches above a gutter, tracking backward ahead of a bicycle wheel as it cuts through a shallow puddle. Rise smoothly to waist height without cutting, pivot around the rider as she passes fruit stalls opening their striped awnings, then accelerate beside her down the narrow lane. Keep her red scarf and silver bicycle centered while foreground baskets sweep past with strong parallax. At the intersection, crane above the crowd, cross over a blue tram arriving from the right, and descend through its open rear door just as the rider stops outside. Finish looking back through the rain-streaked window as the tram pulls away. Overcast silver light with warm shop bulbs; realistic wet reflections. Audio moves with the camera: tire spray close, vendors calling farther back, awning snaps, tram bell, door hiss, then a muffled interior motor hum.

Who is MiniMax H3 Max Built For?

Social creators, marketing teams, indie animators, and developers putting video generation inside their own product.

For creators posting every day

You need vertical clips with sound, and you need them today. MiniMax H3 Max returns a 5-second clip in under 3 seconds, so trying eight versions of an idea costs minutes rather than an evening. The audio arrives with the picture, which removes the step most people find slowest.

For marketing teams testing ad variants

Run the same brief through MiniMax H3 Max with five different openings and see which one holds. Because a 15-second clip costs a fraction of what the comparable models charge, testing widely stops being a budget question.

For indie animators and small game teams

Character consistency across shots and a style that holds to the last frame are the two things that usually break when a model animates. MiniMax H3 Max keeps both inside one generation, which is what makes a six-shot sequence possible from a single request.

For developers building video into a product

There are two endpoints, MiniMax H3 Max text to video and MiniMax H3 Max image to video, both callable from your own app. Responses carry a `timings.inference` field so you can show real progress instead of a spinner. There are no GPUs to provision.

MiniMax H3 Max vs MiniMax H3

Same family, different jobs — MiniMax H3 Max trades resolution and input modes for speed and price.

MiniMax H3 MaxMiniMax H3 (base)
Maximum resolution480p or 768p2K
Render time, 5s clipUnder 3 seconds~35× longer on the official endpoint
Clip length5–15 seconds5–15 seconds
Native audioYesYes
Text to videoYesYes
Image to videoYes, with optional end frameYes
Reference to videoNot at launchYes
Video editingNoYes
Open weightsNo — hosted onlyYes, downloadable
Design Arena Elo (image to video)1,3411,333
Price at 768p$0.06 / secondPriced per its own tiers

Where MiniMax H3 Max loses, plainly.

If you need 2K, MiniMax H3 Max cannot do it and the base model can. If your work depends on reference-to-video or on editing an existing clip, MiniMax H3 Max has neither at launch. And if you want to download weights and run the model on your own hardware, that is the base model — MiniMax H3 Max is hosted only and the weights are not released.

Where MiniMax H3 Max wins.

Speed is not a small margin here; it is the difference between iterating and waiting. At 768p it is also cheaper per second than the models it sits beside on the leaderboards. And on Design Arena's image-to-video board it scores 1,341 against the base model's 1,333 — a real but honest gap, not a landslide.

Pick by the job.

Delivering at 768p for social, ads, or previz, and want many takes: MiniMax H3 Max. Delivering at 2K, or you need reference and editing endpoints: the base model.

Why Choose MiniMax H3 Max

Three reasons that hold up against the whole field, not just against the model it came from.

It is the cheapest model at the top of the board

On Artificial Analysis' image-to-video leaderboard with audio, MiniMax H3 Max ranks first at an Elo of 1,201, and at $3.60 per minute it is also the cheapest listed model in that board's top fifteen. Being first and cheapest at the same time is unusual; normally you pick one.

The speed changes what you do, not just how long you wait

Most models make you commit to a prompt because a bad take costs real minutes. At under 3 seconds for a 5-second clip, MiniMax H3 Max makes the cost of a wrong guess almost nothing, so you write worse first drafts and get to a better shot faster.

Sound is not a second project

Most video models hand back silent footage. MiniMax H3 Max returns audio cut to the picture in the same pass, which removes the step that usually decides whether a clip ships today or next week.

Design Arena image-to-video Elo leaderboard with MiniMax H3 Max first at 1,341Artificial Analysis image-to-video leaderboard with audio and MiniMax H3 Max first

How To Use MiniMax H3 Max

Four steps, and the first clip is back before you finish reading step four.

1. Write the shot and the sound in one block

Name the subject, the camera move, the light, and then the audio, in that order. MiniMax H3 Max reads a long brief down to the beat, so being specific pays off — vague prompts get vague shots. If you are starting from a picture instead, upload it and describe only what should happen next.

2. Pick resolution, length, and shape

The MiniMax H3 Max resolution choices are 480p and 768p; stay at 768p, since it is the default and the one the model is tuned around. Pick a length between 5 and 15 seconds. On MiniMax H3 Max text to video, choose from 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16. On MiniMax H3 Max image to video the output follows the aspect ratio of the picture you uploaded.

3. Leave prompt expansion on balanced

Balanced decides per request how much to rewrite your prompt, which keeps total time close to render time. Fast returns in about a second; quality can spend up to 30 seconds rewriting alone. Balanced is the setting to leave alone unless you have a reason.

4. Generate, listen, then export

A 5-second clip is back in under 3 seconds, a 15-second clip in around 15. Turn the sound on before you judge it — half of what MiniMax H3 Max produced is audio, and a clip that looks average often works once you hear it. Download the MP4 and it is finished; there is nothing to sync.

MiniMax H3 Max Pricing

Pay for the seconds you generate — no subscription, no minimum.

What a clip costs. At 768p, MiniMax H3 Max lists at $0.06 per second, or $3.60 per minute. So a 5-second clip is $0.30, a 10-second clip is $0.60, and the longest single generation, 15 seconds, is $0.90.

Free

Included on sign-up
$0
10 credits included
One 5-second clip at 768P
Two 5-second clips at 480P
Included when you sign up

Enough for one 5-second clip at 768P, or two at 480P. It exists so you can hear what the audio sounds like before you decide, since that is the part of H3 Max that is hard to judge from someone else's sample.

Starter

One-time
$9.9
99 credits included
$0.10 per credit
Nine 5-second clips at 768P
One-time purchase

Nine 5-second clips at 768P, or nineteen at 480P. The right size if you have one project in front of you and no idea yet whether you will have another.

Most popular

Basic

One-time
$29.9
370 credits included
About $0.08 per credit
Thirty-seven 5-second clips at 768P
One-time purchase

Thirty-seven 5-second clips at 768P. Per credit it works out about 19% cheaper than Starter. This is the pack that makes sense once you are generating several versions of the same idea rather than one take and done.

Professional

One-time
$99.9
1,665 credits included
$0.06 per credit
A hundred and sixty-six 5-second clips at 768P
One-time purchase

A hundred and sixty-six 5-second clips at 768P, or 55 clips at the full 15-second length. At $0.06 per credit it is exactly 40% cheaper than Starter — the largest step down in the range.

MiniMax H3 Max FAQ

The questions people actually search before they generate anything.

What is MiniMax H3 Max?

MiniMax H3 Max is an AI video generation model that makes 5 to 15 second clips with synchronized audio from a text prompt or a starting image. It was post-trained by fal Research on top of the open-weight MiniMax H3 model, with new data aimed at prompt adherence and aesthetics. fal also designed the model's architecture around its own inference engine, which is why a 5-second clip comes back in under 3 seconds. Everything distinctive about the base model carries over, including audio and video generated together rather than in two passes. It is not a setting on MiniMax H3 — it is a separate model with its own endpoints.

What is the MiniMax H3 Max resolution, length, and duration?

MiniMax H3 Max generates at 480p or 768p, with 768p the default and the resolution the model is tuned around. At 16:9 that works out to 1344×768 at 24 FPS. Clips run from 5 to 15 seconds, and 15 seconds is the longest single generation. On text to video you can pick 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16; on image to video the output follows the aspect ratio of the picture you upload. If you need 2K, that is the base MiniMax H3 model, not MiniMax H3 Max — this is the detail most search results currently get wrong.

How fast is MiniMax H3 Max?

A 5-second clip at 768p renders in under 3 seconds, which is faster than real time. That is roughly 35 times the throughput of the official MiniMax H3 endpoint. Every response carries a `timings.inference` field — the actual denoising time on the backend — which lands around 2.5 seconds for a 5-second 768p generation. Longer clips scale from there, so 15 seconds takes about 15 seconds. The speed comes from co-designing the inference engine alongside the post-training rather than dropping new weights onto a generic serving path.

Is MiniMax H3 Max really ranked #1 for image to video?

It is first on both public image-to-video boards fal tracks. Design Arena scores MiniMax H3 Max at an Elo of 1,341, ahead of base MiniMax H3 at 1,333 and every other model on that board. Artificial Analysis ranks it first on its image-to-video leaderboard with audio, at an Elo of 1,201 with a 95% confidence interval of ±11 over 2,177 samples — where it is listed under its internal name, MiniMax H3 Turbo (768p). fal's own head-to-head human preference studies against twelve leading video models put it first on overall quality, prompt understanding, and aesthetics. Those are two independent boards plus one first-party study, so treat the third with the caution any vendor's own testing deserves.

How is MiniMax H3 Max different from MiniMax H3?

MiniMax H3 Max is fal's post-trained variant, tuned for prompt adherence and aesthetics and co-optimized with fal's inference stack. MiniMax H3 is a separate frontier model with its own endpoints: it generates at 2K, and it adds reference-to-video and video editing, neither of which MiniMax H3 Max has at launch. The base model also ships open weights you can download and run yourself; MiniMax H3 Max is hosted only. Reach for MiniMax H3 Max when you want speed, prompt adherence, and the lowest price at 768p. Reach for the base model when you need 2K, reference-to-video, or editing.

How much does MiniMax H3 Max cost?

MiniMax H3 Max lists at $0.06 per second at 768p, which is $3.60 per minute — the lowest listed API price of any model at the top of the Artificial Analysis image-to-video board. A 5-second clip is $0.30 and a 15-second clip is $0.90. Those are fal's published rates as of 2026-08-26.

Which MiniMax H3 Max endpoints are available?

Two at launch: text to video and image to video. The image-to-video endpoint also handles first-to-last keyframes through an optional end image, so you do not need a separate call to animate between two stills. Reference to video was announced as following later, so check before you build a workflow that depends on it. There are no GPUs to provision — both endpoints are serverless and callable from a few lines of Python or JavaScript.

Can I use MiniMax H3 Max videos commercially?

Yes. Content generated through the fal.ai API can be used in commercial projects, and fal's terms of service carry the full detail on usage rights and licensing. Read them before you put a generated clip in a paid campaign, particularly around recognisable people and trademarked material, which no model's terms fully solve for you.

Start Generating with MiniMax H3 Max

One prompt, one pass, sound included — the first clip takes about three seconds.

Ranked #1 for image to video on Design Arena and on Artificial Analysis' leaderboard with audio.