Few-step distillation explains this week's open video models
Four papers posted on September 28, 2026, Elastic Forcing, PDMD, DuoMatching and MILD, each attack the same training bottleneck in video generation.

Key takeaways
- Elastic Forcing, described in a paper posted to arXiv on September 28, 2026, removes the critic model from DMD-style video distillation and raised VBench Total from 83.80 to 84.64 on a 1.3B model at 17 FPS.
- PDMD reports a Wan2.1 VBench total of 83.73 at 4 denoising steps, 1.03 points above matched DMD, from a one-line change to DMD's training code, per its arXiv paper posted the same day.
- DuoMatching's Hugging Face release runs a Wan2.1 1.3B model on 4 denoising steps for the first latent and 2 for the rest, but its model card says the checkpoint ships with no benchmark table.
- MILD, from a paper also posted September 28, 2026, has no public checkpoint; it describes a method for pulling image-model skill into video models.
Four papers landed on arXiv and Hugging Face on September 28, 2026, and none of them announced a new video model. Elastic Forcing, PDMD, DuoMatching and MILD are not rival products. They are four attempts at the same training problem: how to teach a video diffusion model to render a clip in two to four passes instead of the dozens a full model normally needs, without the picture drifting into mush. That shared problem is why so many open video releases this week can run in a handful of steps on modest hardware.
What a step actually is
A video diffusion model starts from random noise and cleans it up over a series of passes through the network. Each pass is a denoising step, or NFE (number of function evaluations) in the papers here. A model built to run 50 steps calls the network 50 times per clip; a model built for 4 steps calls it four times. Fewer calls means faster generation and less GPU time, which is the whole appeal of a few-step model to anyone generating video on their own hardware rather than in a data center.
The problem is that a video model trained to expect 50 gradual corrections does not automatically know how to do the same job in four large jumps. Skip most of the steps and the output tends to blur, oversaturate or lose the motion it is meant to describe. Getting a model to work in two to four steps takes a second training pass, called distillation, that teaches a student model to match what a full, many-step model would have produced.
Why a critic makes that training expensive
The standard method for this second pass is Distribution Matching Distillation, or DMD. DMD does not just compare the student's finished video to a target. It trains an extra model, often called a critic or fake-score model, that scores how far the student's in-progress output has drifted from the distribution of video the teacher would produce, and feeds that score back into training. As Elastic Forcing's paper, From Scores to Samples, puts it, standard DMD requires "a bidirectional diffusion teacher and an online fake-score model." That second model adds its own memory footprint, its own training instability and its own errors, which is the exact problem all four of this week's papers are trying to solve, from different angles.
Elastic Forcing removes the critic entirely
Elastic Forcing's answer is to stop scoring the student against a critic and instead compare it directly to real reference videos, inside a frozen, pretrained video-embedding space built for a different job (self-supervised representation learning, not generation). The paper reports the change on a 1.3B model built on the same architecture as an existing model called Self-Forcing: VBench Total rises from 83.80 to 84.64, while the model still runs at 17 frames per second. That is a specific improvement on one 1.3B model, not a claim about every few-step video model.
The bigger practical effect, according to the paper, is what removing the critic frees up. Without a second model taking GPU memory, the authors say they were able to post-train a 14B model on eight H200 GPUs, a scale the extra score model would have made harder to reach. Two checkpoints are already up on Hugging Face: Elastic-forcing-14B, a generator-only release, and Elastic-forcing-1.3B-three-encoder, trained with three separate video-embedding models (DINOv3, V-JEPA2 and VideoMAE) and released at step 150. Both pages describe generator weights only, meant to be loaded into an existing Elastic-Forcing or Self-Forcing runtime rather than run as a standalone pipeline.
PDMD keeps the critic, but filters its mistakes
PDMD, short for Projected Distribution Matching Distillation, takes the opposite approach: keep the critic, but stop its errors from accumulating. The paper traces the oversaturation and artifacts that few-step DMD models develop over training back to "critic errors, which enter successive student updates and accumulate over time." Its fix is a single mathematical projection step that removes the part of the training update that lines up with the critic's own error, which the authors describe as "only a one-line code change to DMD, with no extra loss, network, data, model pass, or multi-stage training."
On Wan2.1, the paper reports a VBench total of 83.73 at 4 NFE, 1.03 points above matched DMD trained the ordinary way. On a separate model, MiniMax-H3, tested on joint video-and-audio generation, PDMD reports a VideoGen-Eval visual score of 83.17 and says it beat every compared 4-NFE model on all six audio metrics measured. Those are two different backbones tested for two different things: a video-only score on Wan2.1, and a joint video-audio score on MiniMax-H3. The paper is released under a CC BY 4.0 licence and its abstract says code and models are available at the project site, which links a code repository on GitHub and a Hugging Face release of the 4-NFE weights.
DuoMatching brings in an image teacher
DuoMatching, uploaded to Hugging Face under the handle JohnZhan, attacks the problem from a third direction: pull in a teacher that already knows what a good single frame looks like. It pairs standard joint video distribution matching with frame-level supervision from an image model, Qwen-Image, passed through an adapter the authors call LatentBridge. The released checkpoint fine-tunes a Wan2.1 text-to-video 1.3B model to generate 832x480 video using four denoising steps for the first latent and two for each one after that. The release is under an Apache-2.0 licence.
DuoMatching's own model card is unusually direct about what it does not show: "There is no checkpoint-specific benchmark table in this release. Do not infer numerical performance from the project page's comparisons." No VBench or other quality number for this checkpoint could be confirmed, so none is reported here.
MILD has no checkpoint at all
The fourth paper, titled From Static to Dynamic, introduces MILD, for Motion-Preserving Image-to-Video Latent Distillation. Like DuoMatching, it moves expertise from image models into video models, using a learnable connector to align a video student's internal states with an image expert's, plus an optical-flow motion reward meant to stop the video losing its sense of movement as it borrows image-level polish. The paper says the method "consistently outperforms video-teacher OPD baselines" across the image experts and video backbones it tested. No Hugging Face checkpoint or code repository could be found for it, and none is claimed in the abstract.
What this changes for you
None of these four papers claims to be the best few-step method available, and this article does not either. What they add up to is a wave of open experiments in one specific bottleneck: making a critic model cheaper to run, cheaper to trust, or unnecessary altogether. If a critic no longer eats a large share of training memory, more labs can afford to distill bigger models, which is the direct argument Elastic Forcing makes for its 14B checkpoint. If a critic's errors can be filtered out with a one-line change, as PDMD claims, existing DMD pipelines could improve without a rebuild. And if an image model can stand in for part of a critic's job, as DuoMatching and MILD both try, a video model's few-step training could piggyback on the much larger ecosystem of image models.
For a creator, the practical payoff shows up further downstream, in models that run in two to four steps on a single consumer GPU instead of dozens of steps on a rented cluster. Two of this week's four releases, Elastic Forcing's 1.3B and 14B weights and DuoMatching's Wan2.1 1.3B weights, are already on Hugging Face for testing today. PDMD's 4-NFE weights are posted under pdmd2026 on Hugging Face, with its code in a GitHub repository, and MILD remains a paper without a release.
Elastic Forcing's checkpoints are on Hugging Face under liuyueyi-8; DuoMatching's Apache-2.0 checkpoint is live under JohnZhan. PDMD's code and 4-NFE weights are linked from its project page, and MILD has no public release to try.
Sources
- arxiv.org - Elastic Forcing paper, VBench figures, hardware claim
- huggingface.co - 14B checkpoint details
- huggingface.co - 1.3B checkpoint details
- arxiv.org - PDMD paper, benchmark figures, licence
- pdmd2026.github.io - PDMD project page and release status
- huggingface.co - DuoMatching model card, specs, licence
- arxiv.org - MILD paper and method description