Institution profile

Center for Robotics

Academic institution
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

StrAD: A Streaming Method and Benchmark for Audio Description Generation for Long-form Videos

Aug 12, 2026

This work addresses the limitations of existing audio description generation methods, which struggle with long-form videos and rely on manually annotated timestamps, thereby failing to meet the accessibility needs of visually impaired users at scale. The paper introduces the first streaming framework for audio description generation tailored to long videos, leveraging a sliding window mechanism to enable real-time caption insertion without ground-truth timestamps. The framework supports both fine-tuning (StrAD-FT) and zero-shot prompting with vision-language models (StrAD-Zero). Additionally, the authors construct StrAD, a diverse benchmark of long videos, to standardize full-video-level evaluation. Experiments show that the proposed method achieves a CIDEr score of 36.3 on CMD-AD—outperforming prior work by 10.0—and reaches 51.0 CIDEr on the StrAD benchmark, with a streaming task SODA score of 2.4, substantially exceeding the zero-shot baseline of 1.1.

0 citationsRead paper
Recent publications

Latest Papers

StrAD: A Streaming Method and Benchmark for Audio Description Generation for Long-form Videos

Aug 12, 2026

This work addresses the limitations of existing audio description generation methods, which struggle with long-form videos and rely on manually annotated timestamps, thereby failing to meet the accessibility needs of visually impaired users at scale. The paper introduces the first streaming framework for audio description generation tailored to long videos, leveraging a sliding window mechanism to enable real-time caption insertion without ground-truth timestamps. The framework supports both fine-tuning (StrAD-FT) and zero-shot prompting with vision-language models (StrAD-Zero). Additionally, the authors construct StrAD, a diverse benchmark of long videos, to standardize full-video-level evaluation. Experiments show that the proposed method achieves a CIDEr score of 36.3 on CMD-AD—outperforming prior work by 10.0—and reaches 51.0 CIDEr on the StrAD benchmark, with a streaming task SODA score of 2.4, substantially exceeding the zero-shot baseline of 1.1.

0 citationsRead paper