Kandinsky 6.0 Video ships open weights with 44 kHz audio
Kandinsky Lab released Lite (3B) and Pro (29B) free under MIT, with five-second clips and lip-sync that run on a 16 GB GPU.

Key takeaways
- Kandinsky Lab open-sourced Kandinsky 6.0 Video on October 6, 2026 under the MIT licence, with a 3B Lite model and a 29B Pro model.
- Both generate five-second clips with synchronized 44 kHz audio, including lip-sync, from text or an image, and a separate super-resolution model lifts output to 1920x1080.
- The maker's own device presets target GPUs down to a 16 GB RTX 5060 Ti, where a non-distilled Pro clip takes about 3,080 seconds.
- The technical report says Pro outperforms Kandinsky 5.0 Pro and stays competitive with leading audio-video models on speech quality, in the team's own human study.
Kandinsky Lab open-sourced Kandinsky 6.0 Video on October 6, a model family that generates five seconds of picture and 44 kHz sound together.
The release is announced in the lab's GitHub repository, which carries an October 6 note that Kandinsky 6.0 and its super-resolution companion are open source. The checkpoints live in a Hugging Face collection, and the technical report went up on arXiv on October 4.
Kandinsky 6.0 Video comes in two sizes. The repository states that Kandinsky 6.0 Video Lite runs 3 billion parameters and Kandinsky 6.0 Video Pro runs 29 billion, and that both make five-second clips with synchronized 44 kHz audio, lip-sync included. Text-to-audio-video and image-to-audio-video both work, so one still can be animated with a matching soundtrack.
The base render is 864x480. A separate plug-in super-resolution model raises it to Full HD (1920x1080), and a hosted API lists a 2x, 2.25x and 4x upscaler that works on the first 121 frames of a clip. Kandinsky Lab ships each size in full, 10-step distilled and pretrain checkpoints.
The line follows Kandinsky 5.0 Video, which made silent clips. In the team's own side-by-side human evaluation, the report says 6.0 Pro clearly beats Kandinsky 5.0 Video Pro and stays competitive with leading audio-video models, especially on speech quality. That is a first-party study, and independent leaderboards had not listed the release when this was written.
Running it locally is the point. The repository's device presets include a 16 GB RTX 5060 Ti, and its timing table lists a non-distilled Lite clip at about 1,310 seconds on that card and a Pro clip at about 3,080 seconds. On a 24 GB RTX 4090 the figures drop to roughly 437 and 936 seconds. AICoder, citing the report's deployment notes, says block offload cut Pro peak memory from 72.8 GiB to 21.7 GiB.
The nearest open rival is LTX 2.5, which also generates audio. The report says 6.0 Pro beats it on most VABench metrics and is preferred on speech, trades wins with Google's closed Veo 3.1 Fast, and trails MiniMax H3 and Seedance 2.0 on visuals.
For a creator, the ceiling is the clip itself. Five seconds is one shot, and Full HD arrives through upscaling rather than native generation, so this is a shot-builder to stitch in an editor rather than a scene-maker. MIT terms mean the models, your fine-tunes and your finished clips can go into commercial work with no territory carve-out.
Kandinsky 6.0 Video is free to download now from the Kandinsky Lab repository and Hugging Face; only the full Pro checkpoint asks for a one-click terms acceptance, while the Lite, distilled and pretrain checkpoints download directly, and a free Hugging Face Space runs the Pro distill checkpoint without a local GPU.
Sources
- github.com - maker's repo: 3B/29B, 5s clips, 44 kHz audio, MIT, Oct 6 open-source note, per-GPU timings, 16 GB device preset
- arxiv.org - technical report: abstract, architecture, MIT licence, submitted Oct 4
- huggingface.co - checkpoints and MIT licence field
- huggingface.co - Lite model card: usage, 480p render, 44 kHz audio sample rate
- x.com - maker's X post announcing the open-source release
- flixly.ai - third-party host: 480p native, upscaling, VSR scale factors
- explainx.ai - independent write-up of the release and licence
- aicoder.com - 16 GB preset and offload memory figures
- segmind.com - VSR 2x/2.25x/4x scale factors and 121-frame limit