Slop TVNewsLatest
News

Tsinghua's SGF+ generates 24-hour video from 5-second clips

The method doubles generator parameters to 2.8 billion to separate the two roles, and its weights and code ship under Apache-2.0.

Illustration: Tsinghua's SGF+ generates 24-hour video from 5-second clips
Illustration: SLOP TV News. Tsinghua University seal via Wikimedia Commons (public domain); JD.com mark from a photograph by N509FZ, CC BY-SA 4.0. Logos are trademarks of their owners.

Key takeaways

  • Researchers at Tsinghua University, JD's Joy Future Academy and the Chinese University of Hong Kong released SGF+, which trains on 5-second video windows yet generates one continuous video for up to 24 hours without long-video fine-tuning.
  • SGF+ gives context writing and frame denoising separate parameters after the team measured a mean gradient separation of 104.2 degrees in attention and 106.3 degrees in the feed-forward layers, with all 512 paired cosine similarities negative.
  • The parameter split doubles the generator from 1.4 billion to 2.8 billion parameters and raises training time per step from 11.79 to 12.76 seconds, about 8.2 percent, while inference latency stays nearly flat.
  • The code and model checkpoints are released under Apache-2.0, and the student model runs on Wan2.1-T2V-1.3B with Wan2.1-T2V-14B as teacher.

A Tsinghua University team with JD and the Chinese University of Hong Kong generates 24 hours of video from 5-second training clips.

The method, Self Gradient Forcing Plus (SGF+), is described in a paper posted to arXiv on 7 October 2026, with code and model checkpoints released on GitHub and Hugging Face under Apache-2.0.

Autoregressive video models do two jobs at once. Denoising turns a noisy frame into a clean one, and context writing encodes that clean frame into the key-value memory that later frames read. SGF+ gives the two roles separate parameters, keeping them linked through causal attention while routing each role's training updates to its own weight set.

The team measured why the split matters. In the earlier Self Gradient Forcing baseline, across 128 prompts and four denoising timesteps, the mean angle between context-writing and denoising gradients was 104.2 degrees in attention and 106.3 degrees in the feed-forward layers, and all 512 paired cosine similarities were negative in both module families. Gradients that point in opposite directions partly cancel when they share weights.

SGF+ is the third step in that line of work. Self Forcing trains on self-generated history but blocks gradients into the context path, and Self Gradient Forcing restores them on shared weights. SGF+ separates them. The split doubles the generator from 1.4 billion to 2.8 billion parameters, and the trained student model sits on Wan2.1-T2V-1.3B with a frozen Wan2.1-T2V-14B teacher.

The extra parameters cost little at inference. On an 81-frame workload, latency is 4.969 seconds versus 4.962 for SGF, inference memory rises from 24.85 to 27.96 GB, and training time per step goes from 11.79 to 12.76 seconds, about 8.2 percent. On VBench-Long at 60 seconds and MovieGen-128 at 240 seconds, generation horizons of 12 and 48 times the training window, SGF+ scores above both baselines on subject consistency, background consistency, flickering, motion smoothness and aesthetics in every setting, and on imaging in all but the 60-second framewise setting. Dynamic degree is the one metric where the baselines score higher; the authors attribute higher motion scores elsewhere to scene jumps and deformation.

The project page shows continuous rollouts labelled up to 24 hours, including clips titled "White Cat in a Basket" and "Gwen Reading." The aggregator Made by Agents lists the checkpoints at 11.4 GB and confirms the Apache-2.0 licence, and AI Weekly notes the paper publishes no per-hour quality curve inside its 24-hour rollouts. This is a research result on a 1.3-billion-parameter base model, not a product, and no 24-hour film was made with it.

Code, weights and checkpoints are on GitHub and Hugging Face under Apache-2.0; running them means downloading the Wan2.1 base models separately.

Sources

  1. arxiv.org - paper: 104.2 and 106.3 degree separations, 512 negative pairs, 1.4B to 2.8B parameters, 11.79 to 12.76 s per step, 5s to 24h, VBench tables
  2. github.com - Apache-2.0 licence, 7 October 2026 release of paper, checkpoints and code
  3. zihan-su.github.io - project page: method, labelled 24-hour demo clips, VBench scores
  4. huggingface.co - inference checkpoints and weight downloads
  5. aiweekly.co - independent write-up: timings, memory, no per-hour quality curve
  6. madebyagents.com - independent write-up: Apache-2.0, base model, 11.4 GB checkpoints