Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Generate 2K clips with stereo sound using the comfyui minimax h3 workflow—text, image, and reference inputs with full ComfyUI node control.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator
Why Choose the comfyui minimax h3 Pipeline
The comfyui minimax h3 pipeline brings MiniMax's omni-modal generation model into ComfyUI as open weights. It processes text, images, video, and audio together in one context, producing clips with native stereo sound — dialogue, effects, and music generated in a single pass. Outputs scale up to 2K at 24fps for roughly 15 seconds, with complete node-level control over every setting.
- Built-In Stereo AudioVoice, sound effects, and music render together with the visuals in one MP4 — perfectly synced through the comfyui minimax h3 workflow.
- Full Open-Weight CustomizationRun the comfyui minimax h3 model on your own hardware with complete control over resolution, duration, and every diffusion parameter — no API constraints.
- Multimodal Reference MixingBlend text, images, video, and audio references into a single generation, locking in character, style, motion, camera movement, or voice with comfyui minimax h3 nodes.
Getting Started with comfyui minimax h3
Produce open-weight video with native audio in three simple steps using the comfyui minimax h3 pipeline.
Key Capabilities of the comfyui minimax h3 Pipeline
The comfyui minimax h3 workflow includes three native ComfyUI templates, open-weight multimodal generation, native stereo audio, reference-driven control, and Sage Attention acceleration — a complete on-premise video production stack.
Three Ready-Made Workflow Templates
The comfyui minimax h3 template pack provides text-to-video, image-to-video, and reference-to-video examples, covering every generation mode right out of the box.
Unified Multimodal Understanding
The comfyui minimax h3 model handles text, images, video, and audio in a shared context, letting you combine multiple reference types within a single generation.
Reference-Guided Creation
Preserve a character's identity, a visual style, a motion pattern, a camera move, or a voice from references — up to 9 images, 3 videos, and 3 audio clips through the comfyui minimax h3 R2V node.
Crisp Text & Brand Rendering
Spelled-out text and brand logos render with precision using the comfyui minimax h3 model, while instruction following naturally describes relationships between references.
Sage Attention Boost
Accelerate generation roughly 2x with negligible quality drop by inserting the Patch Sage Attention KJ node into the comfyui minimax h3 workflow.
Resolution & Duration Grid System
The comfyui minimax h3 Resolution Selector calculates width and height from aspect ratio and megapixels, snapped to the model's 32-multiple grid and 17-frame-per-block duration at 24fps.
comfyui minimax h3 — Common Questions
Everything you need to know about running the MiniMax H3 model inside ComfyUI.
What exactly is the comfyui minimax h3 workflow?
It's ComfyUI's native integration of the MiniMax H3 omni-modal model, released as open weights. The pipeline generates video with native stereo audio from text, image, video, and audio references in a single forward pass.
What video quality can I expect?
The comfyui minimax h3 workflow produces up to 2K resolution at 24fps for around 15 seconds. Its native canvas uses a 768px short edge, capped at 768x1344 pixels and rounded to a multiple of 32.
Which generation modes are available?
The comfyui minimax h3 template library includes three workflows: text-to-video (T2V), image-to-video (I2V) with optional first/last-frame control, and reference-to-video (R2V) that locks character, style, motion, camera, or voice.
Can it really generate audio?
Yes — the comfyui minimax h3 model creates native stereo audio including voice, sound effects, and music, all modeled together with the video in one pass and synced into a single MP4 file.
How do I start generating?
Update ComfyUI to version 0.30.0 or later, open Template Library > Video, select a comfyui minimax h3 workflow, and follow the pop-up to download models from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Is there a way to speed up rendering?
Yes — install SageAttention and the KJNodes custom nodes, then place a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to roughly double rendering speed.
Bring Your Ideas to Life with comfyui minimax h3
Run MiniMax H3 locally in ComfyUI with native stereo audio, open weights, and total parameter control — text-to-video, image-to-video, and reference-to-video workflows ready to use.
