Tencent Hunyuan's Prism renders 2K video with sound in one pass
The Fudan and Tencent Hunyuan model is MIT-licensed and open on GitHub, but native 2K output needs four or more 80 GB GPUs.

Key takeaways
- Tencent Hunyuan's Prism generates video and audio together in one pass at 720p, 1080p and 2K (2560x1440), and its code, paper and preview checkpoints are public.
- Prism is MIT-licensed and builds on the open MOVA base model; each of its two preview checkpoints is 65.3 GB in safetensors.
- 720p inference runs on one 80 GB GPU with CPU offload, but 1080p and 2K need four or more 80 GB GPUs, and native training starts at 32 to 64 of them.
Tencent's Hunyuan team, with Fudan University and Zhejiang University, has released Prism, a model that renders video and its soundtrack together in one pass at native 720p, 1080p and 2K. The code and two preview checkpoints are public under the MIT licence.
The project is on GitHub and Hugging Face, with a technical report on arXiv and a project page. The repository dates to 21 September 2026, its code was pushed on 24 September and its README was updated on 6 October; the paper landed on 4 October. Watch the maker's demo below.
Prism renders the picture and the sound in the same pass, so there is no separate dubbing step. A prompt splits into a video prompt and an audio prompt, and the audio side understands tags such as <music>, <sfx> and <speech>. Given a reference image, Prism animates that frame with matching sound.
The default clip is 205 frames, about 8.5 seconds at 24 fps, and the frame count must satisfy (n − 1) % 4 == 0. Recommended outputs are 720×1280, 1072×1920 and 1440×2560, which the project page calls 720p, 1080p and 2K (2560×1440).
The method is a sparse attention scheme. Prism cuts the token sequence into spatiotemporal macro-zones and gives each zone its own attention block shape, with finer blocks where the picture changes fast or the sound ties to what is on screen. In the paper and the README the authors report that Prism trains 2.5 times faster than full attention while matching or beating its quality.
Prism builds on MOVA, the open joint video-audio model from the OpenMOSS team, and the MOVA-360p base ships in the same Hugging Face repository. The MOVA paper describes a 32B-parameter mixture-of-experts model with 18B active at inference, and Prism inherits the MOVA DiT.
Two preview checkpoints are up: Prism-preview-alpha, the stable one, and Prism-preview-beta, which the repository says is tuned for motion. Each is 65.3 GB in safetensors, and both need the MOVA base model alongside them. A "Prism-pro" release sits on the project's to-do list.
The hardware is the honest part. The repository lists one 80 GB GPU for 720p inference with CPU offload, and four or more 80 GB GPUs for 1080p and 2K. Native training starts at 32 such GPUs for 720p and 64 for 1080p or 2K. There is no hosted API and no ComfyUI node yet, so Prism today is a repository to clone and a GPU environment to configure. Shuyuan Tu of Fudan University is the lead author and the repository's contact.
Prism is on GitHub and Hugging Face now. Clone the repo, pull one preview checkpoint and the MOVA-360p base, set the paths in prism_infer.sh, and start at 720p on one 80 GB card.
Sources
- github.com - repo: resolutions, GPU needs, quickstart, inference settings, checkpoint layout
- github.com - licence text: MIT, with MOVA-360p under Apache-2.0
- arxiv.org - paper abstract: 2.5x claim, MOVA inheritance, submitted 4 Oct 2026
- huggingface.co - model card, checkpoint layout and MOVA base
- huggingface.co - checkpoint file size 65.3 GB
- francis-rings.github.io - maker project page: 2K is 2560 x 1440, comparison clips
- youtu.be - maker demo video, verified via YouTube oEmbed
- arxiv.org - MOVA base model: 32B parameters, 18B active
- artrealmai.com - secondary summary: release dates, MOVA parts, output sizes
- dev.to - independent explainer