MUSE: Unlocking Timestep as Native Task Steering for One-Step Dense Prediction
This work addresses the computational redundancy in existing single-step diffusion models for multitask dense prediction, which typically rely on parameter-heavy adapters or learnable task tokens. The study is the first to reveal and exploit the fixed sinusoidal timestep embeddings inherent in diffusion models as endogenous task-conditioning signals, proposing a unified multitask learning paradigm that requires no additional parameters. Built upon pretrained diffusion models, the method leverages timestep embeddings for task guidance and incorporates manifold disentanglement to enable task-specific generation, compatible with both U-Net and DiT architectures. Experiments across ten datasets demonstrate that the approach achieves performance on par with state-of-the-art methods in monocular depth and surface normal estimation, confirming its effectiveness and broad applicability.