Institution profile

Descript Inc.

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation

May 19, 2026

Long-term video generation often suffers from visual degradation and temporal inconsistency due to frame-by-frame autoregressive modeling. This work proposes Anchor Tree Sampling (ATS), a training-free inference scheduling method that introduces hierarchical tree-structured generation to video-to-video tasks for the first time. ATS leverages sparse anchor initialization, recursive refinement, and leaf-node interpolation, operating under a static camera assumption to support diverse multimodal conditions—including inpainting, pose, and depth maps. By confining temporal drift within anchor intervals, ATS substantially shortens the critical generation path. Experiments demonstrate that ATS outperforms existing autoregressive baselines on Wan 2.1 + VACE across five conditioning modalities, consistently improving both output quality and drift resistance, and enables stable generation of high-quality videos exceeding 40 minutes in duration on LTX-2.3.

0 citationsRead paper

PoDAR: Power-Disentangled Audio Representation for Generative Modeling

May 11, 2026

This work addresses the challenge in audio latent diffusion models where signal power and semantic content are tightly coupled in the latent space, complicating effective modeling. To resolve this, the authors propose the first explicit decoupling of these two factors by introducing stochastic power augmentation and a latent consistency objective, thereby constructing a structurally cleaner and more tractable latent representation. This approach not only enhances training efficiency and generation quality but also enables classifier-free guidance (CFG) to be applied solely to semantic content, significantly improving stability under high guidance scales. Evaluated on the LibriSpeech-PC dataset, the method achieves approximately 2× faster convergence compared to the baseline, along with a 0.055 improvement in speaker similarity and a 0.22 gain in UTMOS score.

0 citationsRead paper
Recent publications

Latest Papers

Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation

May 19, 2026

Long-term video generation often suffers from visual degradation and temporal inconsistency due to frame-by-frame autoregressive modeling. This work proposes Anchor Tree Sampling (ATS), a training-free inference scheduling method that introduces hierarchical tree-structured generation to video-to-video tasks for the first time. ATS leverages sparse anchor initialization, recursive refinement, and leaf-node interpolation, operating under a static camera assumption to support diverse multimodal conditions—including inpainting, pose, and depth maps. By confining temporal drift within anchor intervals, ATS substantially shortens the critical generation path. Experiments demonstrate that ATS outperforms existing autoregressive baselines on Wan 2.1 + VACE across five conditioning modalities, consistently improving both output quality and drift resistance, and enables stable generation of high-quality videos exceeding 40 minutes in duration on LTX-2.3.

0 citationsRead paper

PoDAR: Power-Disentangled Audio Representation for Generative Modeling

May 11, 2026

This work addresses the challenge in audio latent diffusion models where signal power and semantic content are tightly coupled in the latent space, complicating effective modeling. To resolve this, the authors propose the first explicit decoupling of these two factors by introducing stochastic power augmentation and a latent consistency objective, thereby constructing a structurally cleaner and more tractable latent representation. This approach not only enhances training efficiency and generation quality but also enables classifier-free guidance (CFG) to be applied solely to semantic content, significantly improving stability under high guidance scales. Evaluated on the LibriSpeech-PC dataset, the method achieves approximately 2× faster convergence compared to the baseline, along with a 0.055 improvement in speaker similarity and a 0.22 gain in UTMOS score.

0 citationsRead paper