How to run MiniMax H3 on a single 12GB graphics card
ComfyUI's repackaged weights cut MiniMax H3 from 123.6GB to 42.5GB, but the local build leaves out the hosted 2K stage behind Hailuo's showreels.

Key takeaways
- MiniMax H3 is a 33B dense transformer released with open weights on Hugging Face on 3 August 2026.
- ComfyUI's repackaged weights cut H3's footprint from 123.6GB in full precision to 42.5GB through int8 quantisation and pruning.
- ComfyUI states the optimised build can run on a 12GB RTX 3060-class card, but the H3 guides say 16GB or 24GB is the realistic tier.
- The 2K regeneration stage and the Context-IR module stay on MiniMax's hosted API, so local clips will not match the Hailuo app.
MiniMax H3 is one of the few open-weight video models that writes picture and sound in the same pass, and since August it has run in ComfyUI on consumer hardware if you accept a few trade-offs. The 33B model's open weights posted to Hugging Face on 3 August 2026, and ComfyUI shipped same-day support with repackaged weights that cut the download by about two thirds. What the download is not is the full pipeline behind the clips Hailuo, MiniMax's app, puts out at 2K. This guide covers what to install, what to download, what card you need, and where the local build stops.
What MiniMax H3 is
H3 is a 33B dense single-stream transformer, MiniMax's third-generation video model after Hailuo 01 and Hailuo 02 and the first the company has released with open weights, per ComfyUI's day-0 announcement. The model card describes it as a general-purpose omni-modal system that treats text, images, video and audio as one context and generates video with native stereo audio at up to 2K and up to 15 seconds.
In the open release, output is 768p by default. ComfyUI's documentation calls the model's native canvas a 768px short edge, which is 1344x768 at 16:9, with resolutions rounded to a multiple of 32. Audio is 32kHz stereo, generated with the video rather than added afterward.
The model card lists two H3-Base checkpoints: one for text and first-or-last-frame to video (FL2VA), and one for reference-to-video (Ref2VA) with up to 9 images, 3 video clips and 3 audio clips. Both ship in BF16.
What you actually download
Skip the original weights unless you plan to fine-tune. ComfyUI users want the Comfy-Org/MiniMax-H3 repackage, which swaps the raw checkpoints for int8 "convrot" quantised files plus pruned variants.
The pruning is the interesting part. The ComfyUI team found that H3's modulation weights, about 40% of the model's parameters, could be replaced with a functionally equivalent lookup table, and combined that with int8 quantisation and custom kernels tuned to reduce peak VRAM. The result is a total memory footprint cut by 66%, from 123.6GB in full precision to 42.5GB, according to the day-0 announcement. The guide at AI Video Sensei notes that the 42.5GB has to live somewhere, so you want serious system RAM: it suggests 64GB.
What hardware you need
ComfyUI states that the 42.5GB build, combined with dynamic VRAM offloading, lets the model run on a GPU like the RTX 3060. That is the 12GB floor the optimised weights are advertised against, and it comes with heavy offloading, which means slower generation and reliance on system memory.
AI Video Sensei is blunter about the practical tier: it puts H3 in the "worth it if you own or rent 24GB plus" bracket and suggests skipping it on 8GB to 16GB cards when your work is silent b-roll. A comparison post at minimaxh3.app cites community measurements of about 20GB staged for the pruned int8 checkpoint and about 11.9GB for an NVFP4 build, and says 12GB works only with heavy offloading. Read the 12GB figure as a floor you can reach, not the card you want.
How to run it, step by step
- Update ComfyUI to version 0.30.0 or later, or use Comfy Cloud.
- Open the Template Library, go to Video, and pick a MiniMax H3 workflow. ComfyUI ships three base templates (text-to-video, image-to-video and reference-to-video) plus Multiframe Reference and Fun ControlNet Union examples.
- Let the template's model scan pull the files, or download them yourself from Comfy-Org/MiniMax-H3 and place them in the folders the workflow notes name.
- Set the output size. The shipped Resolution Selector defaults to 1344x768, the 16:9 native canvas. Keep the "multiple" setting at 32 to match H3's resolution grid, and do not push past 0.98 megapixels at 16:9: 1.0 megapixels lands at 1376x768, above the model's 768x1344 pixel-area cap.
- Set the duration. H3 snaps duration to a 17-frame-per-block grid at 24 frames per second, so the length you type is rounded up to a valid 17k+5 frame count, per the native-workflow docs.
- Run a short diagnostic clip first to check the length, audio and motion, then render the final.
An optional Sage Attention package roughly doubles generation speed at minor quality cost, though ComfyUI's docs warn you may see dtype fallback messages in the console.
What the local build leaves out
Two pieces of the full 2K pipeline stay on MiniMax's hosted API. The model card names them: H3-Context-IR and H3-Regenerate-2K. The full 2K workflow combines a local deployment with those hosted modules and MiniMax API credentials, which is how MiniMax reproduces the quality of its own direct 2K output. Without them, local clips top out at 768p. If you have watched a Hailuo showreel and expect the same sharpness from a local run, you will be disappointed.
There is a licensing catch too. The weights ship under the MiniMax H3 Community License, which per Fuser's licence comparison excludes the United States, the European Union, the United Kingdom and South Korea, and asks for separate authorisation above $20 million in yearly revenue. ComfyUI's own overview adds that commercial use of locally generated outputs requires a MiniMax commercial licence, available through Comfy as the licence's official reseller.
How it compares
Against Wan 2.2, H3 trades permissiveness for features. Wan 2.2 is Apache 2.0 with no territory limits, and its 5B model fits on a consumer card at 720p. H3 gives you stereo audio, reference conditioning and 15-second clips, but a higher hardware floor and a restricted licence. The minimaxh3.app comparison draws the same line: Wan for licence terms and iteration speed, H3 for audio, references and clip length.
What comes next
MiniMax has said the initial open release ships full attention only and that its sparse-attention implementation will follow in a future update, which would cut the cost of long sequences if it lands. For now, local H3 is a genuine dialogue-and-sound model on your own GPU, with a clear ceiling at 768p and a licence worth reading before you ship anything commercial.
Start in ComfyUI's Template Library under Video, and pull the repackaged weights from the Comfy-Org/MiniMax-H3 repository on Hugging Face.
Sources
- huggingface.co - model card: 33B dense transformer, 2K and 15s output, native stereo audio, H3-Context-IR and H3-Regenerate-2K modules
- blog.comfy.org - day-0 support, 123.6GB to 42.5GB, RTX 3060 claim, ComfyUI 0.30.0
- docs.comfy.org - output resolution and native canvas, licence via Comfy, Sage Attention
- docs.comfy.org - three base workflows, 17k+5 frame grid, reference limits
- huggingface.co - repackaged weights and int8_convrot files
- aivideosensei.com - pruning lookup table, 64GB RAM note, 24GB advice
- fuser.studio - community licence territory exclusions
- minimaxh3.app - release date, VRAM measurements, Wan 2.2 comparison