Slop TVNewsLatest
Explainers

How to run MiniMax H3 on an 8 GB graphics card

FlashML-org's FreeVideo runs an eight-step MiniMax H3 derivative on small cards by streaming weights between GPU, RAM and disk.

Illustration: How to run MiniMax H3 on an 8 GB graphics card
Illustration: SLOP TV News.

Key takeaways

  • FlashML-org's FreeVideo runs an eight-step MiniMax H3 derivative on as little as 8 GB of VRAM and 16 GB of RAM, per the project's README.
  • The project's own tests report a 10-second 1344x768 clip in 558 seconds on an RTX 4060 Ti and 122 seconds on an RTX 5090.
  • FreeVideo's code is Apache-2.0, but the H3 weights carry the MiniMax H3 Community License, which excludes the EU, UK, South Korea and the US.
  • FreeVideo runs the eight-step VDN-H3 checkpoint, which outputs at a 768-pixel short edge rather than the 2K the hosted MiniMax H3 model supports.

MiniMax H3, a heavy open video model that has usually meant a 24 GB card or a hosted API, can now run locally on a consumer GPU with 8 GB of VRAM. FreeVideo, a local inference engine that FlashML-org published on GitHub, fits an eight-step derivative of H3 onto hardware that small, on 16 GB of system RAM.

The repository appeared on GitHub on September 13, 2026, and the project went open source on October 2. Its README makes one promise, MiniMax H3 video on the computer you already own. Here is what that takes, what the memory trick is, and what you give up for it.

What FreeVideo runs

FreeVideo does not run the dense MiniMax H3 transformer. It runs OpenVDN's eight-step VDN-H3 checkpoint, a rebuild of H3's attention that replaces most of its long-range softmax attention with a linear branch. The Video DeltaNet paper reports denoising a 14.3-second, 768p clip in 6.70 seconds across eight B200 GPUs, which the authors put at 14.5 times faster than the 50-step dense baseline. The research comes from Haocheng Xi, a UC Berkeley researcher working with Impossible Inc. and UT Austin.

FreeVideo wraps that model in an execution planner and a creative workspace inside ComfyUI. The planner is the part that turns a heavy model into an 8 GB figure.

How FreeVideo fits a 50-block model into 8 GB

H3's transformer has 50 blocks, and the full model is far larger than any consumer card's memory. FreeVideo's execution-planning documentation explains how the planner decides where each block lives for every request. Some blocks stay resident in VRAM. Some sit pinned in system memory. The rest are re-read from disk on each sampling step.

In plain terms, FreeVideo packs as much of the model onto the card as will fit, then streams the remainder in from your RAM and your SSD while it builds the video. It measures free VRAM and available host memory before each request and sets that split accordingly. When the model will not fit in VRAM and RAM together, the leftover blocks come off the drive at every step, which is why an SSD matters more than raw compute on a small card. Chunked computation splits the feed-forward and projection layers into smaller pieces so no single operation blows the memory budget.

The trade is waiting time and drive traffic, not the shape of the output. The same checkpoint runs on a laptop and on a workstation; only the plan and the clock change.

What you need before you install

The stated minimum is 8 GB of VRAM and 16 GB of RAM. FreeVideo installs three ways: a Windows launcher that sets up ComfyUI and the right model pack, a ComfyUI plugin added under custom_nodes, or a Linux command-line path. The launcher's v0.3.0 release, dated October 7, 2026, switches GeForce and RTX 30-series cards to an int8 model and reports an RTX 3090 making a 5-second video about 2.3 times faster.

Inputs match the hosted model's range: text prompts, first and last frames, and image, video and audio references. Community MiniMax H3 LoRAs work in the same graph, and the workspace supports two-pass sampling and batch generation. The v0.2.0 release added four quality levels, Light, Medium, High and Max, where higher settings take longer.

How fast it is

The project's own end-to-end table covers Windows, a 1344x768 frame, 10 seconds and two-pass sampling. It reports 122 seconds on an RTX 5090, 486 seconds on an RTX 5060 Ti, 558 seconds on an RTX 4060 Ti, and 603 seconds for a community RTX 4060 Ti run at 8 GB of VRAM and 64 GB of RAM.

The project's community-results log shows 8 GB cards working but waiting longer. An RTX 5060 with 16 GB of RAM reported 1,367 seconds for a 15-second clip at the Light setting, later trimmed to about 1,160 seconds. An RTX 4070 laptop at 8 GB reported 1,031 seconds for an 8-second clip. Those are the project's and its users' measurements, not tests SLOP TV ran, and the timing table's capacity grid was measured on an H200 with imposed memory limits rather than on a consumer card.

What you give up against a hosted API

MiniMax describes H3 as 2K-capable, but that path runs through a hosted regeneration step; the open VDN-H3 checkpoint that FreeVideo carries outputs at a 768-pixel short edge, roughly one megapixel depending on aspect ratio. AlphaSignal puts the local output at about one megapixel and 24 frames per second, in clips of 5 to 15 seconds. A hosted call returns a clip in seconds. FreeVideo takes minutes, and the smallest cards take much longer than that. Local runs also skip H3's hosted context-preprocessing system, which MiniMax says is critical to final quality.

What you keep is the rest of the workflow. There is no per-second fee, no upload, and a graph you can edit and re-run. For a creator iterating on a shot at home, those three things decide whether a local run is worth the wait.

The licence, in plain terms

Two licences apply, and they are not the same. The FreeVideo code is released under Apache-2.0. The H3 weights are not. They carry the MiniMax H3 Community License, which the README says includes territorial and acceptable-use restrictions.

MiniMax's own licence Q&A explains the scope. The licence grants its rights only inside an "Applicable Territory" defined as worldwide minus four excluded places, the European Union, the United Kingdom, South Korea and the United States. Creators in those four places need separate permission from MiniMax first. The company says the limit reflects a more complex regulatory environment for video models than for text or code, not a wish to exclude developers. Read the licence before you ship anything commercial.

What comes next

FreeVideo is on a fast release cadence. In its first weeks it added a public gallery of shared clips, four quality levels, workflow-carrying videos, Mac support in preview, and community LoRA support. The stated minimum has not moved below 8 GB of VRAM, and the heavy disk streaming will stay the binding constraint on smaller cards until the weights shrink further. FreeVideo also cannot run the dense H3 model, only the eight-step VDN-H3 derivative; a reader who needs the fullest H3 quality still needs a hosted service or a much larger GPU.

FreeVideo is on GitHub at github.com/FlashML-org/FreeVideo. Download the Windows launcher or add the repository under ComfyUI's custom_nodes, then run ./freevideo plan --vram-gib 8 --ram-gib 16 to see the block placement for your card before you generate.

Sources

  1. github.com - FreeVideo README: 8 GB VRAM and 16 GB RAM minimum, features, releases, licences
  2. github.com - the planner, the 50-block residency scheme, and the end-to-end timing table
  3. github.com - community-reported times on 8 GB cards
  4. github.com - v0.3.0 release date, int8 model, prompt enhancement
  5. github.com - v0.2.0 four quality levels and two-pass sampling
  6. huggingface.co - VDN-H3 eight-step checkpoint, weight sizes, licence
  7. openvdn.github.io - Video DeltaNet paper page: 14.4s 768p in about 9.0s on 8xB200, authors
  8. arxiv.org - Video DeltaNet abstract: 14.3s 768p denoising in 6.70s, 14.5x
  9. huggingface.co - H3 768p base output and hosted 2K path, model size
  10. huggingface.co - licence territory scope: EU, UK, South Korea, US
  11. alphasignal.ai - local output of about one megapixel at 24 fps, clips of 5 to 15 seconds