Skip to main content
MiniMax H3 is MiniMax’s general-purpose, omni-modal generation model, available in ComfyUI through the MiniMax H3 API nodes. The API workflows run generation on MiniMax’s servers, so no model downloads or local GPU is required, and each second of generated video is billed to your Comfy API account. MiniMax H3 generates video with native stereo audio: voice, sound effects, and music are modeled together in a single forward pass instead of being layered on afterward. Output is up to 2K resolution at 5-15 seconds per clip.

What MiniMax H3 is good at

  • Native stereo audio: Voice, sound effects, and music are modeled together with the video in a single forward pass instead of being layered on afterward
  • High-resolution output: Up to 2K resolution at 5-15 seconds per clip
  • Text-to-video: Generates videos, with audio, from text prompts
  • First-last-frame video: Generates the motion between a first frame and an optional last frame image
  • Reference-conditioned generation: Generates videos conditioned on up to 9 reference images, 3 reference videos, and 3 reference audio clips

Example outputs

Text-to-video generation from a single prompt, with native stereo audio: First-last-frame generation, with the model creating the motion between two frames:

Use it in ComfyUI

MiniMax H3 workflows

Run the text-to-video, first-last-frame, and reference-to-video workflows in ComfyUI, locally or on Comfy Cloud