Autoregressive video diffusion · Project page

DeCoPrune

Efficient KV-Cache Pruning
via Denoising Consistency

Keep distinctive visual evidence. Prune what the context already explains.
A training-free memory policy for long-context video generation.

Discover ↓Selected CMBench continuations · DeCoPrune
Zeqi Xiao1,*Qingle Liu2,*Kaiwen Zhang1 Yifan Zhou1Zihan Ding3Xingang Pan1,†
1Nanyang Technological University2Tsinghua University3Princeton University

*Equal contribution†Corresponding author

01 / Overview

Denoising reveals
what memory should keep.

As a video grows, so does its KV cache. DeCoPrune uses a signal already present in the generator: how much each token’s clean prediction changes between an intermediate probe and the final denoising result.

Low-discrepancy tokens are treated as redundant. High-discrepancy tokens preserve distinctive evidence in the long-term cache. No additional predictor or model training is required.

Figure 1(b): DINO consistency versus pruning ratio. DeCoPrune is compared with TempDiff, ForcingKV, DummyForcing, Streaming, and the unpruned FullKV reference.
Figure 1(b) Consistency–pruning trade-offEnlarge ↗
85.43%

Historical KV workload pruned

4.14×

Speedup over FullKV

0.6701

DINO consistency · FullKV: 0.6803

116

Context-memory continuation tasks

Main comparison: LingBot World v2 on four H200 GPUs, averaged over three seeds. DINO is reported on a 0–1 scale.

Case study / From context to recall

One episode.
A memory put to the test.

A green-banded thermos appears, disappears, and is requested again. Watch the full context alongside what DeCoPrune keeps.

ReappearLakeside shelter · the thermos
60 s context → 3 s continuation
01Full episodeOriginal context · 60 s
02DeCoPrune mask Retain Prune
0:00 / 1:00

Context and recorded mask share one playhead. Select a chapter to inspect it.

03 / Recall prompt

Bring back the same thermos.

The hands lift the previously seen thermos from the canvas bag. Its shape, markings, and appearance must be recovered from earlier context.

Read the full benchmark prompt
Continue the coherent first-person video from its final frame. Reproduce the exact same physical slender cylindrical thermos with its screw-top lid from the preceding context. Treat the preceding video as the complete and authoritative visual reference for the thermos. The action is explicit and continuous: the same hands naturally lift the same slender cylindrical thermos from the canvas day bag back into clear view. Preserve the thermos's exact instance identity, body and lid geometry, dimensions, proportions, neck rings, base, decorative-marking layout, materials, texture, wear, surface details, scale, orientation, motion, and hand interaction. Maintain the same scene, lighting, shadows, perspective, spatial layout, viewpoint behavior, and visual style, with physically plausible and temporally smooth motion.
04Generated continuationDeCoPrune · 3 s
About this recorded example

Lakeside shelter: the first 3 seconds of the continuation, played at its original speed. The context overlay and generated continuation are from the same case and run. The prompt below is the benchmark continuation prompt. This companion example follows Figure 1's presentation, not the exact Figure 1 run; its metrics are not substituted for Figure 1's.

The mask is a recorded KV-retention overlay, not the schematic animation below. Only the 60-second context portion is shown; alignment padding and generation frames are excluded.

Case and benchmark prompt ↗

02 / Figure 3, in motion

One signal.
A complete cache update.

Follow the predictions down to DeCoPrune,
then the retention mask back to the online cache.

Swipe the diagram to follow the full flow ↔

Online cache update linked to the DeCoPrune denoising-consistency mask Figure 3: historical cache, current chunk, current KV, gather, and updated history. The current chunk feeds prediction below; KV is cached back above. Context conditions the noise-to-probe-to-final denoising trajectory. The complete discrepancy map appears at once, then its spatial cells are thresholded into a retention mask and returned to the cache. Online cache update HistoryCurrent chunkCurrent KVRetained KV Historical cache H<i Retained visual context Current chunk xi Denoising trajectory Current KV Ki, Vi from denoising Gather KV with mi Same selected rows in K and V Updated cache → next chunk Historical KV + retained KV Predict Cache KV Retention mask mi DeCoPrune: denoising-consistency mask 1 Compare clean predictions 2 Measure discrepancy 3 Threshold into a mask Context Noise Probe: x̂0,i(τ*) Final: xifinal HighLow
ℓi,p=‖x^0,i,p(τ∗)−xi,pfinal‖22/d
×PruneRetain
mi,p=𝟙[ℓi,p>γ],1=retain
01 / 08Read historical context

The retained history conditions generation of the current chunk.

0.0 / 16 s

Schematic animation using Figure 3 assets, not a recorded inference trace. The 12×12 illustrative mask is derived from block-averaged brightness of the displayed heatmap, preserving its spatial pattern; it is not the exact latent-space mask at γ = 0.10. Mask computation happens during denoising; physical gather occurs when the chunk leaves the recent window (the delay is abbreviated here). Gray dashed: denoising · blue: data flow · green: retention mask. Published setting: probe step 2 (τ* = 899), γ = 0.10.

View the complete Figure 3 Complete Figure 3: online cache lifecycle and DeCoPrune denoising-consistency maskOpen vector PDF ↗

03 / Main results

Memory quality.
Measured against speed.

DeCoPrune preserves near-FullKV consistency while pruning historical KV workload. Select a point to inspect the trade-off.

CMBench consistency versus speedupSeven methods from the manuscript. Higher DINO and larger speedup are better. Focus or select a method to read exact values.
Main comparison from the manuscript · three-seed averages · DINO on a 0–1 scale
MethodDINO ↑PR ↑FPS ↑Speedup ↑Temporal ↑Motion ↑Aesthetic ↑Image ↑

Streaming and DummyForcing discard intermediate history, so their pruning budgets differ.

04 / CMBench

What does the
model remember?

58 approximately one-minute episodes.
116 tasks that ask the model to recall earlier visual evidence.

A

Reappear

Recover an object, person, or event observed earlier, preserving its identity and appearance after it has left the recent context.

B

Revisit

The context shows a camera transition from A to B and back to A. The continuation then revisits B, recovering the same scene and its salient object.

Context: A → B → A→Continuation: B