Model Brief

MiniMax H3: One model for sight, sound, and motion

Explore MiniMax’s open general-purpose video model for multimodal direction, native stereo audio, and output up to 2K.

MiniMax H3 treats text, images, video, and audio as one creative context. Instead of separating visual generation from sound design, it can interpret mixed references and produce a synchronized audiovisual result—giving creators one workflow for subjects, movement, camera language, voices, ambience, and rhythm.

4–15s
Output duration
Up to 2K
Maximum output resolution
24 FPS
Video frame rate
32 kHz
Native stereo audio
12 files
Maximum mixed reference inputs
11
Stably supported dialogue languages

From mixed references to one finished scene

01 · Understand

H3-Context-IR

Reads text, images, video, and audio as one directed context.

02 · Generate

H3-Base

Creates a synchronized 768p audiovisual sequence.

03 · Refine

H3-Regenerate-2K

Revisits the original context and regenerates the selected result up to 2K.

What MiniMax H3 can do

01

Direct with multiple media types

Combine text instructions with image, video, and audio references. H3 can use them to understand characters, environments, motion, camera behavior, visual style, voices, sound effects, and editing rhythm.

02

Generate picture and sound together

Create 4-to-15-second video at 24 FPS with native 32 kHz stereo audio. Dialogue, ambience, effects, and visuals are generated as one synchronized result rather than assembled in separate passes.

03

Control the opening and ending frames

Use a first frame, a last frame, or both to anchor how a shot begins and resolves. This makes image-to-video work more deliberate and helps preserve composition across a transition.

04

Build from rich reference sets

The omni-reference variant accepts up to nine images, three video clips, and three audio clips, with no more than 12 files in total. References can establish identity, motion, composition, camera language, voice, sound, or rhythm instead of forcing one asset to carry every instruction.

05

Work across formats and resolutions

Create common landscape, square, and portrait aspect ratios at a 768-pixel default short side, with H3-Regenerate-2K available for higher-resolution output up to 2K.

06

Create multilingual dialogue

H3 provides stable dialogue support across 11 languages, including English, Chinese, Japanese, and Korean, for globally oriented audiovisual production.

07

Choose API or local deployment

Use the complete hosted workflow through the MiniMax API, or deploy the released BF16 H3-Base checkpoints for text-to-audio-video, first/last-frame generation, and multimodal reference generation at 768p.

Model Brief

How the H3 system works

01

A unified multimodal context

H3 encodes text, visual material, and audio with modality-specific encoders or VAEs, then packs them into a single multimodal sequence. Spatial and temporal relationships are carried into one dense, single-stream H3-Omni-Transformer, allowing the model to reason about what should appear, how it should move, and what should be heard as connected parts of the same scene.

02

Context-IR turns references into direction

H3-Context-IR is the hosted interpretation and orchestration layer. It parses instructions, associates information across media, understands timing, resolves relationships among references, and converts the result into a structured representation for generation. It may fill underspecified semantic details while preserving the creator’s intent, which is especially useful when many references have different roles.

03

H3-Base generates synchronized audio and video

H3-Base produces the 768p audiovisual sequence. Its visual and audio latent spaces are separate, but generation happens inside the same system: the temporally causal VisualVAE represents motion and image detail, while AudioVAE processes left and right channels independently before recombining them as stereo output.

04

2K is regenerated, not simply enlarged

H3-Regenerate-2K does more than conventional super-resolution. It feeds the 768p result and the original creative context back into H3, then regenerates the sequence at higher resolution. Because the model can revisit the source instructions and references, it can recover more faithful details than a purely pixel-based upscaler.

05

Two checkpoints serve different workflows

H3-Base-FL2VA covers text-to-video plus optional first-frame, last-frame, or first-and-last-frame control. H3-Base-Ref2VA handles multimodal references: up to nine images, three videos, and three audio clips. Video and audio references must each be 2–15 seconds, their total duration is capped at 15 seconds per type, and audio cannot be the only reference modality.

06

What “open” currently includes

MiniMax released the complete H3-Base model weights, including FL2VA and Ref2VA checkpoints, for deployment and further development such as fine-tuning. The hosted H3-Context-IR orchestration layer, H3-Regenerate-2K, and the native sparse-attention implementation are not part of the initial open-source release, so reproducing the full official 2K pipeline still requires MiniMax services.

07

Where it fits in production

H3 is suited to cinematic shots, product films, multilingual character scenes, music-led visuals, campaign adaptations, storyboard motion tests, and reference-driven variations. The practical advantage is not merely another text-to-video endpoint: it is the ability to assign different creative responsibilities to different media and generate one coherent audiovisual scene from them.

Direct one coherent audiovisual scene

  • State the subject, action, camera movement, environment, and timing in a clear sequence.
  • Assign each reference a specific role, such as character, motion, style, voice, or editing rhythm.
  • Use first- and last-frame control when the composition at either end of the shot matters.
  • Describe dialogue, ambience, and sound effects alongside the visuals so they serve the same dramatic beat.
  • Keep video and audio reference clips between 2 and 15 seconds; audio must be paired with an image or video reference.
  • Use 768p for rapid iteration, then move selected results through the 2K regeneration workflow.
  • Check rights for every uploaded reference and generated likeness, voice, logo, or copyrighted element.
  • Start with one focused scene before increasing the number of subjects, references, and transitions.