Slop TVNewsLatest
Lists

5 few-step AI video releases, 3 you can run right now

Two days brought a wave of distillation papers, and three of the five releases hand you a file you can load onto a GPU today.

Illustration: 5 few-step AI video releases, 3 you can run right now
Illustration: AI-generated for SLOP TV News with GPT Image 2

Key takeaways

  • Elastic Forcing's 1.3B model raises VBench Total from 83.80 to 84.64 by matching reference videos in a frozen embedding space instead of training a critic model; checkpoints are on Hugging Face under liuyueyi-8.
  • DuoMatching ships an Apache-2.0 checkpoint built on Wan2.1-T2V-1.3B that runs 4-step first-latent and 2-step subsequent denoising at 832x480, available at huggingface.co/JohnZhan/DuoMatching.
  • PDMD's project page links both code and weights: a GitHub repository at github.com/ZeamoxWang/pdmd and a 4-NFE checkpoint published on Hugging Face as pdmd2026/pdmd_4NFE_full under Apache-2.0.
  • MILD has no public checkpoint and FlowAct-R2's Space is a static project page, so of the five releases only three can be run today.

Two days, several papers, one theme: making video diffusion models fast enough to run without a data center. September 27 and 28, 2026 brought a cluster of distillation releases, the technique that trains a lightweight student model to match a slow teacher's output in a handful of steps instead of dozens. That is the difference between a clip that takes minutes and one that takes seconds, and sometimes the difference between renting an H100 and using the card already in your tower. Three of the five below give you a file to run today; two are worth reading rather than loading. The technique they share is explained in today's explainer on few-step distillation.

Elastic Forcing (Tsinghua and collaborators)

Most few-step distillation leans on Distribution Matching Distillation, which needs a frozen diffusion teacher plus a second "critic" model trained online. Elastic Forcing, posted to arXiv on September 28, throws the critic out and matches the student's rollouts against real reference videos inside a frozen self-supervised embedding space, using maximum mean discrepancy as the training signal.

The payoff is concrete. Using the same architecture and initialization as Self-Forcing, the 1.3B model lifts VBench Total from 83.80 to 84.64 while holding 17 FPS. Dropping the critic also freed enough memory that the team post-trained a 14B model on a single node of eight H200 GPUs, in 23.2 hours over 80 updates, according to the project page. Two checkpoints are live on Hugging Face: liuyueyi-8/Elastic-forcing-14B, a generator-only FP32 release, and liuyueyi-8/Elastic-forcing-1.3B-three-encoder, released at step 150. Neither model card states a licence, and both expect the project's own runtime rather than a plain Diffusers pipeline, so budget time for setup.

Card

If you would rather stay inside the Wan ecosystem with a licence you can read, DuoMatching is the friendlier stop.

DuoMatching (JohnZhan)

DuoMatching keeps the joint-distribution matching DMD already does across frames, then adds a second, frame-level matching objective supervised by an image model instead of a video teacher. A LatentBridge adapter translates between the video student's latents and the image teacher's latent space so that supervision can cross over at all. It is built on the Wan2.1-T2V-1.3B backbone and denoises the first latent in four steps, then every following latent in two, at 832x480, according to the model card and code repo, both posted September 28.

What is actually on Hugging Face is an EMA FP32 checkpoint from training step 1020, plus the frozen Qwen-Image LatentBridge weights used during training rather than inference. The release code carries an Apache-2.0 licence, though the card notes that the separately downloaded Wan2.1 and Qwen-Image weights keep their own upstream terms. The repo's own release notes also say the environment that packaged the weights had no CUDA device, so end-to-end generation was not re-verified there before upload. Its model card is refreshingly blunt about the numbers: there is no checkpoint-specific benchmark table in the release.

Card

  • Tool: DuoMatching
  • Maker: JohnZhan
  • Backbone: Wan2.1-T2V-1.3B
  • Denoising: 4 steps for the first latent, 2 for the rest
  • Resolution: 832x480
  • Licence: Apache-2.0 code; upstream terms on the downloaded weights
  • Get it: https://huggingface.co/JohnZhan/DuoMatching

Not every distillation fix needs an extra network at all.

PDMD (ByteDance and UC San Diego)

Projected Distribution Matching Distillation targets a specific failure mode: DMD samples drifting into oversaturation and artifacts as training goes on. The paper, posted to arXiv on September 28 under a CC-BY-4.0 licence, traces the drift to critic errors that compound across student updates, then removes them with a projection against an estimate of that error. The change is small: one line, no extra loss, network, data pass or training stage.

The numbers on paper: on Wan2.1, PDMD reports a VBench total of 83.73 at 4 NFE, 1.03 points above matched DMD. On MiniMax-H3 joint video-and-audio generation, it reports a VideoGen-Eval visual score of 83.17 and the best results among compared 4-NFE models on all six audio metrics. The release is real this time: the project page links a GitHub repository and a weights release on Hugging Face, pdmd2026/pdmd_4NFE_full, which the Hub lists under Apache-2.0 with a Diffusers config and its weights split across seven safetensors shards. That is the cleanest licence of the three runnable picks here.

Card

MILD (arXiv 2609.34371)

MILD is short for Motion-Preserving Image-to-Video Latent Distillation, and it is not a speed paper in the same sense as the other four. It transfers an image model's aesthetic and OCR strengths into a video diffusion model through a learnable latent connector, plus an optical-flow motion reward that keeps the transferred fixes from wrecking motion quality. The arXiv listing, posted September 28, describes results across multiple video-student backbones outperforming video-teacher on-policy distillation baselines. No checkpoint, code repository or project page turned up in this search, so treat MILD as a method to read until that changes.

FlowAct-R2 (ByteDance Intelligent Creation)

FlowAct-R2 aims at interactive humanoid streaming rather than clip generation: a Streaming Multimodal Reference Diffusion Transformer built on the pretrained Seedance 2.0 Mini reference-to-video backbone, paired with a Proactive Interaction Agent that plans a persona and an agenda in advance, then handles live audience interruptions. The arXiv paper, posted September 28, claims real-time 720p generation and hour-scale streaming. There is a Hugging Face Space at ProAudience/FlowAct-R2, but its own README describes a "Project Clean Export" running a static index.html bundle, not an interactive generator you feed a prompt. The backbone is closed, so what is open here is a description of the interaction framework rather than a model.

Card

  • Tool: FlowAct-R2
  • Maker: ByteDance Intelligent Creation
  • Backbone: Seedance 2.0 Mini (closed)
  • Claimed capability: real-time 720p, hour-scale streaming
  • Availability: static project-page Space, no downloadable model
  • Try it: https://huggingface.co/spaces/ProAudience/FlowAct-R2

Elastic Forcing, DuoMatching and PDMD all hand you something to run this week; PDMD is the one with the licence you can read at a glance. MILD is a paper without a release, and FlowAct-R2 is worth reading, not running.

Sources

  1. arxiv.org - Elastic Forcing paper, VBench figures, 14B training claim
  2. huggingface.co - 14B checkpoint file and precision
  3. huggingface.co - 1.3B checkpoint file and encoders
  4. video-examples-m8r2v6.pages.dev - Elastic Forcing project page, training hardware and benchmark table
  5. huggingface.co - DuoMatching model card, checkpoint spec, licence
  6. github.com - DuoMatching code repo and licence
  7. arxiv.org - PDMD abstract, VBench and VideoGen-Eval figures, licence
  8. pdmd2026.github.io - PDMD project page and weights status
  9. arxiv.org - MILD abstract and method
  10. arxiv.org - FlowAct-R2 abstract, backbone and claims