Model Brief
MiniMax H3: One model for sight, sound, and motion
Explore MiniMax’s open general-purpose video model for multimodal direction, native stereo audio, and output up to 2K.
MiniMax H3 treats text, images, video, and audio as one creative context. Instead of separating visual generation from sound design, it can interpret mixed references and produce a synchronized audiovisual result—giving creators one workflow for subjects, movement, camera language, voices, ambience, and rhythm.
- 4–15s
- Output duration
- Up to 2K
- Maximum output resolution
- 24 FPS
- Video frame rate
- 32 kHz
- Native stereo audio
- 12 files
- Maximum mixed reference inputs
- 11
- Stably supported dialogue languages
From mixed references to one finished scene
H3-Context-IR
Reads text, images, video, and audio as one directed context.
H3-Base
Creates a synchronized 768p audiovisual sequence.
H3-Regenerate-2K
Revisits the original context and regenerates the selected result up to 2K.
What MiniMax H3 can do
Direct with multiple media types
Combine text instructions with image, video, and audio references. H3 can use them to understand characters, environments, motion, camera behavior, visual style, voices, sound effects, and editing rhythm.
Generate picture and sound together
Create 4-to-15-second video at 24 FPS with native 32 kHz stereo audio. Dialogue, ambience, effects, and visuals are generated as one synchronized result rather than assembled in separate passes.
Control the opening and ending frames
Use a first frame, a last frame, or both to anchor how a shot begins and resolves. This makes image-to-video work more deliberate and helps preserve composition across a transition.
Build from rich reference sets
The omni-reference variant accepts up to nine images, three video clips, and three audio clips, with no more than 12 files in total. References can establish identity, motion, composition, camera language, voice, sound, or rhythm instead of forcing one asset to carry every instruction.
Work across formats and resolutions
Create common landscape, square, and portrait aspect ratios at a 768-pixel default short side, with H3-Regenerate-2K available for higher-resolution output up to 2K.
Create multilingual dialogue
H3 provides stable dialogue support across 11 languages, including English, Chinese, Japanese, and Korean, for globally oriented audiovisual production.
Choose API or local deployment
Use the complete hosted workflow through the MiniMax API, or deploy the released BF16 H3-Base checkpoints for text-to-audio-video, first/last-frame generation, and multimodal reference generation at 768p.
Model Brief
How the H3 system works
A unified multimodal context
H3 encodes text, visual material, and audio with modality-specific encoders or VAEs, then packs them into a single multimodal sequence. Spatial and temporal relationships are carried into one dense, single-stream H3-Omni-Transformer, allowing the model to reason about what should appear, how it should move, and what should be heard as connected parts of the same scene.
Context-IR turns references into direction
H3-Context-IR is the hosted interpretation and orchestration layer. It parses instructions, associates information across media, understands timing, resolves relationships among references, and converts the result into a structured representation for generation. It may fill underspecified semantic details while preserving the creator’s intent, which is especially useful when many references have different roles.
H3-Base generates synchronized audio and video
H3-Base produces the 768p audiovisual sequence. Its visual and audio latent spaces are separate, but generation happens inside the same system: the temporally causal VisualVAE represents motion and image detail, while AudioVAE processes left and right channels independently before recombining them as stereo output.
2K is regenerated, not simply enlarged
H3-Regenerate-2K does more than conventional super-resolution. It feeds the 768p result and the original creative context back into H3, then regenerates the sequence at higher resolution. Because the model can revisit the source instructions and references, it can recover more faithful details than a purely pixel-based upscaler.
Two checkpoints serve different workflows
H3-Base-FL2VA covers text-to-video plus optional first-frame, last-frame, or first-and-last-frame control. H3-Base-Ref2VA handles multimodal references: up to nine images, three videos, and three audio clips. Video and audio references must each be 2–15 seconds, their total duration is capped at 15 seconds per type, and audio cannot be the only reference modality.
What “open” currently includes
MiniMax released the complete H3-Base model weights, including FL2VA and Ref2VA checkpoints, for deployment and further development such as fine-tuning. The hosted H3-Context-IR orchestration layer, H3-Regenerate-2K, and the native sparse-attention implementation are not part of the initial open-source release, so reproducing the full official 2K pipeline still requires MiniMax services.
Where it fits in production
H3 is suited to cinematic shots, product films, multilingual character scenes, music-led visuals, campaign adaptations, storyboard motion tests, and reference-driven variations. The practical advantage is not merely another text-to-video endpoint: it is the ability to assign different creative responsibilities to different media and generate one coherent audiovisual scene from them.
Direct one coherent audiovisual scene
- State the subject, action, camera movement, environment, and timing in a clear sequence.
- Assign each reference a specific role, such as character, motion, style, voice, or editing rhythm.
- Use first- and last-frame control when the composition at either end of the shot matters.
- Describe dialogue, ambience, and sound effects alongside the visuals so they serve the same dramatic beat.
- Keep video and audio reference clips between 2 and 15 seconds; audio must be paired with an image or video reference.
- Use 768p for rapid iteration, then move selected results through the 2K regeneration workflow.
- Check rights for every uploaded reference and generated likeness, voice, logo, or copyrighted element.
- Start with one focused scene before increasing the number of subjects, references, and transitions.