A mechanic repairs a brass bird, whispers, “Let’s see if you still remember the sky,” and watches it fly out of his workshop. MiniMax H3 fits the whole scene into 15 seconds.
In 15 seconds, H3 keeps the mechanic, workshop, spoken line, bird’s activation, and departure in one sequence. It also accepts supplied endpoint frames and mixed image, video, and audio references. The model has enough room for a setup, an action, a spoken line, and a payoff. It can also begin and end on supplied frames or take images, video, and audio as creative references.
MiniMax launched H3, also called Hailuo H3, on July 31 as a general-purpose multimodal video model with native stereo sound and 2K output. I tested its narrative, reference-driven, educational, abstract, and continuation workflows. Across the batch, H3 kept a setting, character, dialogue, and ending action coherent in one shot. It also produced an anatomically absurd heart and an origami boat whose geometry collapsed while moving.
H3 can stage a complete 15-second scene
The clockwork bird was a clear result. The prompt described a repair, the bird’s activation, one line from the mechanic, and the bird’s flight through an open doorway. H3 preserved the mechanic and workshop through the camera move, gave the bird a readable change of state, and reached the requested ending.
The night-market reunion asked for more simultaneous action: two recurring people, two speakers, a moving crowd, rain, steam, physical contact, and an embrace. Both lines appeared with the correct speaker split, while the wet market and surrounding crowd stayed coherent enough for the exchange to read immediately.
The rooftop parkour run gave the camera a harder motion test. The runner sprints, clears an alley, lands, recovers, and stops at sunset without losing her identity.
Seven prompts requested speech or narration. Local transcription recovered the requested words in all seven. That is a useful small-sample result, though it is not a lip-sync benchmark. One of those seven clips was the rejected heart animation, which proves another point: accurate speech says nothing about the accuracy of the pictures behind it.
Write H3 prompts around visible change
H3 responded best when the prompt explained what should change and where the scene should end. The clockwork bird had to leave the workshop. The night-market characters had to find and embrace each other. A magnetic-ink prompt asked two droplets to merge, stretch into ribbons, form a reflective structure, and collapse into a sphere.
The same pattern worked in a generated volcano cutaway. H3 rotated an intact mountain into a cross-section, moved magma through a conduit, erupted at the surface, and settled into a dark lava layer. It also spoke the requested narration exactly.
A Patagonian storm clearing into sunlight, a dancer changing from charcoal to watercolor and oil paint, an underwater rescue, and a museum installation moving through a day-night transformation followed the same principle. Their prompts described motion through time instead of relying on a subject and a visual mood.
The stronger runs used prompts that named a starting state, visible actions, and a final state.. A sequence of visible actions gives the model a chance to build a scene.
The same structure holds across quieter scenes: a cartographer’s river leaves the page, a lunar flower blooms under glass, and a potter’s lesson ends with a finished blue cup.
Reference mode accepts image, video, and audio
The MiniMax video guide documents text-to-video, first-frame and last-frame control, and reference generation through one MiniMax-H3 model. On the direct V2 API, duration accepts every whole number from 4 through 15 seconds for duration, including six-second and eleven-second requests. Reference mode accepts for input up to nine images, three video clips, and three audio clips, with a 12-file cap. Video and audio references can each total up to 15 seconds. An audio reference must accompany an image or video.
The V2 API reference treats reference generation and endpoint control as separate modes. A reference request cannot also set first_frame or last_frame. A first-frame or last-frame request cannot include reference images, video, or audio. Each generation therefore chooses between multimodal references and fixed endpoint images.
The mode boundary shaped the overview clip below. Its middle shot uses the article cover and a supplied endpoint frame. Its final shot uses excerpts from the rooftop, cartographer, and pottery videos, plus greenhouse and observatory images. The two shots had to be generated separately.
Longer reference attempts damaged several narrated words. The finished clip keeps H3’s generated rain, wing, water, and landscape sounds beneath the clean narration. Visually, the bird carries the sequence from the cover into the rooftop leap, parchment river, blue glaze, lunar flower, and mountain observatory.
In this test, the references supplied subjects, voice, composition, movement, style, and rhythm. The prompt defined the action; it still has to turn those ingredients into action.
First and last frames can carry a scene across clips
H3’s frame controls offer a way past the 15-second limit. I created three matching observatory keyframes in Reve 2.1, then used H3 to generate two adjacent clips. The first transforms a frozen observatory into a building wrapped in suspended water. The second settles the water into a pool while snow recedes and alpine plants grow.
The shared middle frame kept the building, telescope, copper surface, mountain horizon, and comet aligned closely enough to join the clips into one 30-second sequence. The generated soundtracks still needed a softened transition at the edit.
The same technique failed on a paper boat. Its shape broke inside both clips, and the second clip began with implausible motion. Endpoint images helped preserve composition in the observatory scene. Motion state and object geometry still needed separate review.
H3’s polish can hide factual errors
The volcano clip looks clean enough to place in an explainer. That appearance is not evidence of scientific accuracy. The heart test made the risk impossible to miss: H3 delivered the requested narration while drawing blood flow and anatomy so badly that the clip had to be discarded.
Other polished outputs raised similar questions. A tesseract animation looked geometric without proving mathematical correctness. A mechanical watch looked convincing without proving that its parts could function. A cicada emergence looked plausible without an entomologist’s review.
Speech fidelity, visual coherence, and subject accuracy are separate variables. H3 can score well on the first two while failing the third. Technical, medical, scientific, and educational footage needs review from someone who knows the domain.
MiniMax and fal sell access to H3 differently
MiniMax exposes H3 through one multimodal V2 endpoint. fal separates text-to-video, image-to-video, and reference-to-video into three routes.
MiniMax lists 2K output at $0.13 per second, or $1.95 for a 15-second clip. Its 768P option costs $0.09 per second and remains in closed beta. Audio references and the first five image references are free.
fal lists 2K H3 output at $0.26 per second, or $3.90 for 15 seconds. It charged the same rate for the output side of the reference route. fal completed every request in the larger batch and made it easy to queue each generation mode through a dedicated wrapper.
The matched image-and-narration control cost an expected $1.30 through MiniMax and $2.60 at fal’s list price. Each provider produced one stochastic sample, so the visual difference between those two clips is not a quality benchmark. The test established access, price, and input handling.
A pay-as-you-go MiniMax key worked on the first submission. The operator’s Subscription Key returned error 2013 before task creation because H3 is outside current Token Plan coverage. MiniMax’s Token Plan pricing page states that exclusion, while its FAQ still says Token Plan supports all API Platform models and tells customers to use a pay-as-you-go key outside plan coverage.
The Max plan’s advertised daily video allowance did work with the older MiniMax-Hailuo-2.3 model through the v1 endpoint. H3 requires the v2 endpoint and pay-as-you-go access under the tested account.
fal also requires a privacy decision before uploading sensitive references. Its retention documentation says uploaded inputs and generated media are served through CDN URLs, while request payloads are stored for 30 days by default. fal provides headers to disable payload storage, set media expiration, and choose the initial file ACL.
Seedance 2.5 sets the next comparison
ByteDance launched Seedance 2.5 on the same day as H3. Its published envelope is larger: up to 30 seconds per generation, 30 images, 10 video clips, 10 audio clips, multi-round extension, and timestamp-controlled editing.
Dreamina’s official Seedance 2.5 pages now promote the model, although access language and account availability remain inconsistent. ByteDance says API access through BytePlus ModelArk is coming soon.
Those specifications make Seedance the relevant comparison for H3’s scene-building pitch. H3 has a documented API now, a 15-second ceiling, and demonstrated first-plus-last-frame control. Seedance promises twice the single-generation duration and a much larger reference envelope. Output quality needs a matched test once stable API access allows both models to receive the same prompts, references, and review criteria.
H3 weights are available
MiniMax has released the H3 weights. Developers have started running the model locally on consumer hardware through early ComfyUI support. Published configurations differ in memory, quantization, resolution, and runtime, so they do not yet establish a hardware baseline or repeatable production setup.
This article reports hosted API tests completed before the weight release. A follow-up will examine the released model on local consumer hardware and what the community is building in this ecosystem.
Where MiniMax H3 fits now
H3 is ready for creators and developers who can review every output and discard convincing failures. It fits short narrative scenes, concept footage, visual transformations, stylized motion, dialogue, and frame-controlled transitions. Prompts work better when they describe a visible progression and a final state.
Anatomy, science, machinery, exact geometry, logos, text, and multi-clip continuity need stricter supervision. The model’s surface quality makes review more important because obvious ugliness is no longer a reliable warning.
MiniMax pay-as-you-go is the cheaper tested route. fal costs twice as much at 2K, but its dedicated wrappers and reliable batch completion make it a practical alternative. Token Plan subscribers should use a separate pay-as-you-go key for H3 until MiniMax changes the observed coverage or reconciles its documentation.
H3 gave me a convincing workshop, a reunion, and an impossible heart with the same visual confidence. It can direct a short synthetic scene. The creator still owns every claim the scene appears to make.



