KAIST's GRACE makes Wan2.1 video 11x faster on an A100
Two-stage latent compression cuts Wan2.1-I2V-14B's token count nearly 8x while matching the model's VBench quality on one A100.

Key takeaways
- GRACE renders a 480x832, 81-frame Wan2.1-I2V-14B clip in 78 seconds, down from 863 seconds on the uncompressed model, an 11.1x speed-up.
- The method compresses Wan2.1's latent from 32,800 tokens per 480x832 clip to 4,300, nearly 8x fewer, by squeezing both the spatial and temporal axes.
- At 736x1280 with 81 frames GRACE runs 15.5x faster, cutting the same clip from 3,397 seconds to 219 seconds on one A100 80GB.
- GRACE scores 85.81 on VBench-T2V at 480p against Wan2.1-14B's 83.93, and 87.90 on VBench-I2V against 87.92.
A KAIST AI and Kakao Corp. team has published GRACE, a compression method that renders a 480x832, 81-frame clip with Wan2.1-I2V-14B in 78 seconds. The uncompressed model needs 863 seconds for the same clip on the same card.
GRACE stands for Generation-Aware Latent Compression for Efficient Video Generation. The paper went up on arXiv on October 7 and was submitted to Hugging Face Papers the same week. The project page notes the work "was done while the first three authors were interns at Kakao Corp."
The project page puts the speed-up at 11.1x at 480x832 with 81 frames, and 15.5x at 736x1280 with 81 frames. It states the timings are end to end, autoencoder included, measured on one A100 80GB at 50 steps, CFG 5, batch 1 and bf16.
How GRACE's latency compares
| Model | Latent tokens | 480x832x81 | 736x1280x81 |
|---|---|---|---|
| Wan2.1-14B | 32.8k | 863.2s | 3,396.8s |
| LTX-Video 0.9.7 | 4.3k | 104.1s | 274.6s |
| GRACE | 4.3k | 77.7s | 218.8s |
Image-to-video latency on VBench, single A100, per the project page.
The 863-second baseline comes from that same table: Wan2.1-14B takes 863.2 seconds per 480x832 image-to-video clip and GRACE takes 77.7 seconds, which the page rounds to its headline 78. So 11.1x is the figure the paper supports at 480x832, and the "15x" in some write-ups is the 15.5x figure at 736x1280. Both are the team's own measurements on one A100 80GB, not on consumer hardware.
The compression moves Wan2.1's latent from f8t4 to f16t8, 16x on the spatial axis and 8x on the temporal one, cutting a clip from 32,800 latent tokens to 4,300. GRACE keeps a frozen base latent from the pretrained encoder and learns a residual latent for what the heavier compression drops. During training it matches the compressed latent to the original inside the frozen transformer's feature space, and at inference it denoises the base ahead of the residual by a fixed offset.
GRACE is built on Wan2.1, Alibaba's 14-billion-parameter model from 2025, rather than the newer Wan releases. ArtRealmAI's write-up notes there is no ComfyUI node yet and that the training code and a hosted demo are listed as coming soon. AI Weekly reports the retrofit cost 38.5 H200 GPU-days, 8.5 for the autoencoder stage and 30 for adapting the transformer, and that reconstruction quality on Panda-70M falls, with PSNR dropping from 35.15 to 32.63. The generation scores hold: at 480p GRACE reaches 85.81 on VBench-T2V against Wan2.1-14B's 83.93, and 87.90 on VBench-I2V against 87.92.
For a creator, the number that changes how you iterate is the 78 seconds. A 14B open model that used to take a quarter of an hour per clip now returns one in a little over a minute, on the same card, before any step distillation is stacked on top.
Run it yourself: the GRACE checkpoints are on Hugging Face and the inference code for text-to-video and image-to-video is on GitHub; you still need the official Wan2.1 base weights.
Sources
- huggingface.co - paper page and abstract carrying the 8x token and 11.1x latency claims
- arxiv.org - arXiv record with authors, submission date and full abstract
- cvlab-kaist.github.io - project page with the 78s headline, benchmark tables and the A100 measurement note
- github.com - official code repo listing the released checkpoints and method summary
- huggingface.co - released GRACE checkpoints
- artrealmai.com - secondary write-up carrying the 863s baseline and per-resolution timings
- aiweekly.co - secondary write-up on retrofit cost and VBench parity