Slop TVNewsLatest
News

WanPE boosts AI video preference 50.86 points at 30 seconds

A 397B prompt enhancer trained on 1.05 million real videos beats raw user requests by a far wider margin at 30 seconds than on shorter clips, and the Wan team has published no weights.

Illustration: WanPE boosts AI video preference 50.86 points at 30 seconds
AI-generated illustration by SLOP TV News. Wan logo from the WanPE project page (wan-pe.github.io). Logos are trademarks of their owners.

Key takeaways

  • WanPE, a 397B-parameter prompt enhancement model trained on 1.05 million real-world videos, raised human preference over raw user prompts by 10.66 to 18.84 points on five to fifteen second clips.
  • On 30-second clips the same model raised preference by 50.86 points, tested on Wan 3.0's video generator.
  • The paper says WanPE leads all evaluated commercial offerings at five to fifteen seconds and is competitive with Seedance 2.5 at 30 seconds.
  • No code and no weights have been released; the project page's code link is empty.

A prompt enhancement model built by the Wan team raised human preference over users' own prompts by 50.86 points on 30-second clips, between about 2.7 and 4.8 times the gain it delivered on shorter ones.

WanPE is a 397-billion-parameter model trained on 1.05 million real-world videos, and the paper that describes it went up on arXiv on September 24, 2026. Its job is not to expand adjectives into more adjectives. It writes shot-level cinematic plans, deciding how action, camera trajectories, lighting and sound unfold across a sequence of shots, which is the same thing a director does before anyone picks up a camera.

The headline numbers come from WanPEval, a human-annotated testbed the team built to cover clips from five to 30 seconds and roughly 11,000 blind pairwise comparisons. Run as the prompt layer for Wan 3.0's video generator, WanPE beats the raw user prompt by different amounts depending on how long the clip is.

Clip length WanPE's gain over the raw user prompt
5 to 15 seconds +10.66 to +18.84 points
30 seconds +50.86 points

At 30 seconds the comparison on the project page is with Seedance 2.5 as a complete system, and the table there has Wan 3.0 with WanPE at 60.24 overall against Seedance 2.5's 59.76. The abstract's summary of the same work says WanPE "leads all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds." The project page does not list which offerings those are.

Two design choices separate WanPE from forward prompt rewriting, the approach the paper tests it against. The first is reverse construction, where the model reconstructs a shot-level plan from real video rather than rewriting a short request forward. The paper reports that reverse construction "demonstrates clear superiority over forward rewriting" in its ablations. The second is Semantic-Consistency GRPO, a training method intended to stop a longer prompt from drifting away from what the user actually asked for across shots.

That second problem is the one creators will recognise. A rewriter that turns six words into two hundred can invent specifics, and over a sequence of shots those inventions can compound. The project page shows the output format: numbered shots with timecodes, camera notes and sound direction, the shape of a screenplay rather than a paragraph.

The caveat is large. The code link on the project page is empty, no weights have been published, and, as Fervor Creative AI's briefing notes, a 397B model would not run on a desktop anyway. WanPE is evidence about where the leverage sits in a generation pipeline, not a tool anyone can install.

If the prompt layer decides more of the outcome than the generator does, a better prompt may buy more than a better model.

The paper is on arXiv as 2609.30221. The project page at wan-pe.github.io carries the WanPEval tables and sample shot plans.

Sources

  1. arxiv.org
  2. wan-pe.github.io
  3. fervorcreativeai.com