Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving
为解决自动驾驶中长期预测与即时决策难以兼顾的问题,提出Drive-HWM框架,通过慢-快层级模型在不同时间尺度上进行未来场景预测和动作生成。
为解决自动驾驶中长期预测与即时决策难以兼顾的问题,提出Drive-HWM框架,通过慢-快层级模型在不同时间尺度上进行未来场景预测和动作生成。
为解决机器人操作中动作表示与控制频率纠缠或依赖固定时间参数化的问题,提出CAT框架,通过编码动作轨迹为连续潜在标记,并加入频率感知位置编码,提高了不同控制频率下的表现。
This work addresses the limited dynamic depth selection in conventional Transformer residual architectures, which stems from insufficient historical information exchange across multiple parallel streams. To overcome this, we propose a reciprocal cross-stream addressing mechanism that enables bidirectional historical retrieval within a multi-stream framework: each stream computes depth weights based on the state of its counterpart stream and applies these weights to its own historical values. Our approach uniquely integrates inter-stream interaction into historical retrieval, preserving inter-layer representational diversity through cross-stream depth selection while mitigating redundancy and functional imbalance. Key components include reciprocal cross-attention, normalized state weighting, constrained gated writing, and block-level history storage. Experiments demonstrate consistent and significant improvements over standard residual Transformers and Attention Residuals across dense models (0.1B–1B) and a 7B sparse MoE model, with ablation studies confirming that performance gains arise from cross-stream interaction rather than additional parameters or projections.
This work addresses a key limitation in existing robotic manipulation approaches, which often neglect the dependence of action semantics on environmental context, resulting in noisy, redundant, and poorly structured control trajectories. To overcome this, the paper introduces EDAR—a novel framework that explicitly models the coupling between actions and their surrounding context. EDAR constructs environment-dependent action tokens by jointly embedding executable control commands with their visual outcomes in specific scenes. This representation enables the action space to capture interaction semantics rather than merely encoding command patterns. Evaluated in both simulated and real-world robotic manipulation tasks, EDAR significantly enhances downstream policy learning performance, demonstrating particularly strong gains in long-horizon manipulation scenarios.
Existing real-time streaming methods for virtual human video generation struggle to simultaneously maintain long-term visual temporal consistency and accurately perceive user intent. This work proposes a real-time framework capable of generating videos of unlimited duration, leveraging autoregressive distillation to enhance inference efficiency. It introduces an innovative fusion of short- and long-term visual memory mechanisms with a reasoning-and-response module, complemented by a state-recurrent strategy and a cache-switching mechanism. This design enables high visual consistency while effectively aligning with complex user intentions. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches across diverse scenarios, achieving—for the first time in real-time streaming generation—concurrent long-term visual coherence and responsive interactive intent alignment.
为解决自动驾驶中长期预测与即时决策难以兼顾的问题,提出Drive-HWM框架,通过慢-快层级模型在不同时间尺度上进行未来场景预测和动作生成。
为解决机器人操作中动作表示与控制频率纠缠或依赖固定时间参数化的问题,提出CAT框架,通过编码动作轨迹为连续潜在标记,并加入频率感知位置编码,提高了不同控制频率下的表现。
This work addresses the limited dynamic depth selection in conventional Transformer residual architectures, which stems from insufficient historical information exchange across multiple parallel streams. To overcome this, we propose a reciprocal cross-stream addressing mechanism that enables bidirectional historical retrieval within a multi-stream framework: each stream computes depth weights based on the state of its counterpart stream and applies these weights to its own historical values. Our approach uniquely integrates inter-stream interaction into historical retrieval, preserving inter-layer representational diversity through cross-stream depth selection while mitigating redundancy and functional imbalance. Key components include reciprocal cross-attention, normalized state weighting, constrained gated writing, and block-level history storage. Experiments demonstrate consistent and significant improvements over standard residual Transformers and Attention Residuals across dense models (0.1B–1B) and a 7B sparse MoE model, with ablation studies confirming that performance gains arise from cross-stream interaction rather than additional parameters or projections.
This work addresses a key limitation in existing robotic manipulation approaches, which often neglect the dependence of action semantics on environmental context, resulting in noisy, redundant, and poorly structured control trajectories. To overcome this, the paper introduces EDAR—a novel framework that explicitly models the coupling between actions and their surrounding context. EDAR constructs environment-dependent action tokens by jointly embedding executable control commands with their visual outcomes in specific scenes. This representation enables the action space to capture interaction semantics rather than merely encoding command patterns. Evaluated in both simulated and real-world robotic manipulation tasks, EDAR significantly enhances downstream policy learning performance, demonstrating particularly strong gains in long-horizon manipulation scenarios.
Existing real-time streaming methods for virtual human video generation struggle to simultaneously maintain long-term visual temporal consistency and accurately perceive user intent. This work proposes a real-time framework capable of generating videos of unlimited duration, leveraging autoregressive distillation to enhance inference efficiency. It introduces an innovative fusion of short- and long-term visual memory mechanisms with a reasoning-and-response module, complemented by a state-recurrent strategy and a cache-switching mechanism. This design enables high visual consistency while effectively aligning with complex user intentions. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches across diverse scenarios, achieving—for the first time in real-time streaming generation—concurrent long-term visual coherence and responsive interactive intent alignment.