🤖 AI Summary
This work addresses the susceptibility of video variational autoencoders (VAEs) to inter-channel frequency aliasing and round-trip dynamics degradation during spectral editing in latent space, which often leads to reconstruction artifacts. To mitigate these issues, the authors propose Latent Frequency Validity (LFV), a novel approach that, for the first time, models channel mixing capacity as a tunable resource. LFV learns VAE-specific compact spectral responses and integrates a diagonal calibrator (C1), a channel mixing operator (CM), and a validation-driven path selection mechanism to enhance reconstruction fidelity while preventing round-trip drift. Evaluated across 544 VAE-editing configurations, the method yields 423 high-efficiency operators (99%+ pass rate under hold-out evaluation), generalizes to CogVideoX and HunyuanVideo without fine-tuning, and achieves editing speeds three times faster than pixel-domain filtering followed by re-encoding.
📝 Abstract
Direct spectral editing in video-VAE latents can control noise, flicker, smoothness, and frequency content without a decode--filter--reencode pass. However, video VAEs may redistribute pixel-space frequency bands across latent channels, and latent edits can disrupt VAE round-trip dynamics. We introduce \emph{latent-frequency validity} (LFV), which learns a compact VAE-specific spectral response and deploys it only when it improves decoded-target fidelity without worsening round-trip drift. LFV follows a validation-selected path from a diagonal per-frequency calibrator (C1) to full channel mixing (CM), making cross-channel capacity a controllable per-edit resource. Across 544 VAE--edit cells spanning six spectral families, LFV emits 423 cheap operators: 277 are handled by C1, while 146 (34.5\% of emitted operators) require channel mixing. On the primary 120-cell radial sweep, 99/100 emitted operators pass source-video-grouped held-out evaluation. Across five additional filter families, all 323 emitted operators pass held-out evaluation. Fully frozen OpenVid-fitted operators, including the validation-selected path coefficient, pass all 20 tested CogVideoX and HunyuanVideo generated-domain cells without adaptation. The selected response matches direct latent-filter latency and is about $3\times$ faster than pixel filter--reencode. The resulting maps reveal distinct VAE regimes, including strongly channel-coupled CogVideoX responses and a sharp Open-Sora high-band stability frontier.