First, the bigger picture
If you have not been following AI video closely, three ideas explain most of what H3 Max is.
AI video generation models turn a written description or still image into moving footage. Current systems such as Veo, Kling, Seedance, Wan, and MiniMax H3 usually produce a few seconds at a time, and more of them now generate sound with the picture.
Open-weight models publish their trained parameters so others can run or modify them. Most frontier video systems are closed. MiniMax released H3 as an open-weights model, which made H3 Max possible.
Post-training continues training after a base model exists, using new data and preference feedback to push behavior toward goals such as literal instruction following or better-looking output.
H3 Max sits at the intersection: a post-trained variant of an open-weight AI video model. MiniMax built the base; fal did the additional training and serves the result.
Who actually built H3 Max?
MiniMax built and released MiniMax H3, a general-purpose model capable of up to 2K video with native stereo audio.
fal Research took the open weights and post-trained them into a separate variant called H3 Max. It is not an official MiniMax tier. On the Artificial Analysis leaderboard, it may appear under the internal name MiniMax H3 Turbo (768p).
The naming rule
MiniMax made H3. fal post-trained H3 Max. This site provides a ready-made interface to run the model.
The problem it was built to solve
AI video prompting is iterative. A clip that takes minutes rather than seconds stretches ten attempts into a long wait, breaks the creative loop, and makes every exploratory prompt a budget decision.
Silent output adds a second workflow: generate or source audio, then synchronize it in an editor. Prompt drift adds more retries when requested beats arrive in the wrong order or text becomes unreadable.
These issues reflect a familiar trade-off: quality was paid for with latency and cost. H3 Max matters because its published results attempt to move all three at once.
Why use H3 Max
Faster than real time
fal publishes a sub-three-second render for a five-second 768p clip, making iteration feel interactive rather than batched.
Sound with the picture
Room tone, foley, music, ambience, and dialogue can be generated in the same pass and returned synchronized.
Stronger brief adherence
fal says post-training targeted prompt adherence directly, including ordered beats and legible on-screen text.
Low published API price
At $3.60 per generated minute, it is positioned below comparable top-ranked models in the cited board.
Visual consistency
The model is designed to hold a character or visual language across cuts within one generation.
First-to-last-frame animation
An opening and closing image can define the endpoints of an image-to-video shot.
How H3 Max was built
Post-training
fal says it added substantial new data on top of the base weights and used its reinforcement-learning framework against real generation workloads and human preferences. Prompt adherence, audio-visual quality, and aesthetics were evaluated independently.
Inference co-design
The serving stack was optimized while the model was still being developed, allowing training and inference decisions to inform one another. Training and serving used NVIDIA GB200 NVL72 systems.
The quality constraint
Dropping precision or sampling work can make diffusion faster at the cost of output quality. fal states that an optimization shipped only when the resulting model held its position in internal quality evaluations.
How fast is it, really?
The published headline is under three seconds to render a five-second clip at 768p. API responses expose a timings.inference field, with fal reporting roughly 2.5 seconds of backend denoising for that case. A 15-second clip is described as taking about 15 seconds.
- Roughly 35× the throughput of the official MiniMax H3 endpoint
- About 15× faster than models of comparable quality, on average
- Design Arena describes more than 50× the speed of the base model
Is the number-one ranking real?
Two public boards placed it first after launch. Design Arena reported an image-to-video Elo of 1,341, eight points over base MiniMax H3. Artificial Analysis reported 1,201 with a ±11 confidence interval over 2,177 samples in its image-to-video-with-audio category.
The margin matters. Eight Elo points is a narrow lead. Your own prompts and footage should carry more weight than a leaderboard. The clearer advantage is reaching comparable quality much faster and at a lower listed price.
H3 Max vs standard MiniMax H3
They are different models for different jobs. Choose speed and unit economics at 768p, or choose the broader capability set and 2K path.
| Capability | H3 Max | MiniMax H3 |
|---|---|---|
| Built by | fal, post-trained | MiniMax, open weights |
| Max resolution | 768p | 2K |
| Duration | 5–15s | 4–15s |
| Frame rate | 24 FPS | 24 FPS |
| Modes | Text and image to video | Text, image, reference, editing |
| Reference generation | Not at launch | Up to nine images plus media clips |
| Native audio | Yes | Yes |
| 5-second render | Under 3s, published | ~30s+, published comparison |
| Best for | Adherence, aesthetics, speed | Maximum quality and capability breadth |
Why does H3 Max stop at 768p? The published open weights generate at 768p. MiniMax H3's 2K output uses a separate upscaling stage, so a variant trained from the released base does not inherit that complete 2K path.
What H3 Max costs
The published list price is $0.06 per generated second at 768p, or $3.60 per minute. The launch promotion lists $0.03 per second for its first 14 days; because that offer is time-limited, confirm the current rate before purchase.
| Model | Tier | 15-second clip |
|---|---|---|
| H3 Max | 768p | $0.90 |
| Wan 3.0 | 720p | $1.50 |
| Kling v3 Standard with audio | 720p | $1.89 |
| FLUX 3 | 720p | $2.55 |
| Seedance 2.5 | 720p | $7.10 |
Closest available tier to 768p with audio. Rates can change; verify each provider before budgeting.
The trade-offs
- No 2K. 768p is the ceiling. Use standard MiniMax H3 when the deliverable requires 2K.
- Fifteen seconds maximum. Longer sequences must be assembled from multiple clips.
- No reference-to-video at launch. Text-to-video and image-to-video shipped first.
- Image-to-video follows the input ratio. The six-ratio menu applies to text-to-video.
- Quality prompt expansion adds latency. It can spend up to 30 seconds rewriting a prompt before rendering; balanced is the better default for speed.
- The leaderboard lead over H3 is thin. Test both when absolute output quality matters more than time.
Settings that give the best results
- Stay at 768p. It is the resolution the model is tuned around.
- Generate five to ten seconds. Longer is supported, but this range is the published reliability recommendation.
- Leave prompt expansion on balanced. Fast minimizes rewriting; quality can add substantial delay.
- Describe the sound. Name room tone, foley, music, and cue timing so the audio half of the model has direction.
Slow dolly forward down a rain-slick night-market alley, steam pouring off a noodle cart under a buzzing red neon sign. 35mm anamorphic, practical lights only. Sound: wok sizzle and clang, rain ticking on tin awnings, neon hum.
What we measured ourselves
First-party benchmark pending
We have not published a 20-run first-party benchmark yet, so this page does not substitute fal's published numbers into an “our results” column. The planned test will record render time, end-to-end wall time on balanced and quality modes, ordered-prompt success, and usable audio without retry. Results and a timing screenshot will be added only after the runs are complete.
Who H3 Max is for
A good fit
- Fast idea iteration
- Social and ad content at 768p
- API products that need high throughput
- Shots where synchronized sound matters
Look elsewhere
- 2K masters
- Multi-image reference-driven characters
- Single generations longer than 15 seconds
Generate your first H3 Max video
You can call fal's API or use a ready-made interface. If you want the video rather than the integration, the interface is the shorter path.
- Open the generator. No SDK, API key, or GPU setup is required.
- Choose Text to Video or Image to Video.
- Write the shot in order: subject, action, camera, lighting, then sound.
- Choose 768p, five to ten seconds, and the target aspect ratio.
- Keep Prompt Expansion on Balanced.
- Generate, review, then change one variable at a time for the next attempt.
For application integration, use minimax/h3-max/text-to-video or minimax/h3-max/image-to-video and consult fal's API documentation for current parameters.
H3 Max FAQ
What is H3 Max?+
H3 Max is a post-trained variant of MiniMax H3 built by fal Research. It generates 5-to-15-second video at up to 768p with synchronized stereo audio; fal publishes an under-three-second render time for a five-second clip.
Is H3 Max made by MiniMax?+
No. MiniMax built and released the base H3 model. fal post-trained those open weights into the separate H3 Max variant and serves it.
What does post-trained mean here?+
It means training continued from MiniMax's published weights using new data and preference-based reinforcement learning, targeting prompt adherence, audio-visual quality, and aesthetics.
What resolution does H3 Max support?+
480p or 768p, with 768p the default. For 2K output, use standard MiniMax H3.
How long can an H3 Max video be?+
Each generation can be 5 to 15 seconds. Five to ten seconds is the recommended range.
Does H3 Max generate sound?+
Yes. Stereo audio is generated in the same pass as the video and arrives synchronized with the picture.
How much does H3 Max cost?+
fal's published list price is $0.06 per generated second at 768p, or $0.90 for a 15-second clip. Check current pricing before production use.
Which H3 Max endpoints are available?+
Text-to-video and image-to-video are available. Image-to-video can optionally animate from a first frame to an end frame.
Sources and editorial note
Written by the MiniMax H3 Max editorial team. Published and reviewed August 28, 2026. Product specifications and benchmark claims are attributed to the sources below; the first-party measurement section remains explicitly pending.