What MiniMax H3 is good at
- Native stereo audio: Voice, sound effects, and music are modeled together with the video in a single forward pass instead of being layered on afterward
- High-resolution output: Up to 2K resolution at 5-15 seconds per clip
- Text-to-video: Generates videos, with audio, from text prompts
- First-last-frame video: Generates the motion between a first frame and an optional last frame image
- Reference-conditioned generation: Generates videos conditioned on up to 9 reference images, 3 reference videos, and 3 reference audio clips
Example outputs
Text-to-video generation from a single prompt, with native stereo audio: First-last-frame generation, with the model creating the motion between two frames:Use it in ComfyUI
MiniMax H3 workflows
Run the text-to-video, first-last-frame, and reference-to-video workflows in ComfyUI, locally or on Comfy Cloud