MiniMax H3 has day-zero support in ComfyUI
ComfyUI natively supports the new open-weights MiniMax H3 video generation model, which takes text, images, video, or audio as input and generates video with native stereo sound up to 2K resolution and 15 seconds length. The model is optimized for consumer hardware (such as NVIDIA RTX 3060) through weight pruning, lookup tables, and int8 quantization.
What are the main capabilities of MiniMax H3?
MiniMax H3 supports text-to-video, image-to-video, first-and-last-frame control, and reference-based video generation using images, videos, or audio. It performs all these tasks in a single multimodal model with native stereo audio generation.
Why is MiniMax H3 significant as an open-weights model?
MiniMax H3 is the first generative video model from MiniMax released with open weights, meaning developers and researchers can freely use, study, and modify it, unlike proprietary models.
How is the model optimized to run on consumer hardware?
Engineers reduced the model's memory footprint by pruning modulation weights (about 40% of parameters) and replacing them with a functionally equivalent lookup table. The model also uses int8 convolution quantization and custom kernels to reduce peak VRAM usage.
- Lyria 3.5, Google's music generation model, now available in AI Studio — deepmind.google 76 % match
- Experiment with Gemma 4 as an offline translator — github.com 74 % match
- Naomi Bashkansky leaves OpenAI to develop thought-to-text AI at Conduit — naomibashkansky.com 73 % match
- MiniMax H3
- ComfyUI
- MiniMax