Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving
为解决自动驾驶中长期预测与即时决策难以兼顾的问题,提出Drive-HWM框架,通过慢-快层级模型在不同时间尺度上进行未来场景预测和动作生成。
为解决自动驾驶中长期预测与即时决策难以兼顾的问题,提出Drive-HWM框架,通过慢-快层级模型在不同时间尺度上进行未来场景预测和动作生成。
本文提出SMTrap,一种基于SMT冲突指导的方法,生成计算密集型CSP查询以低成本实现对大型推理模型的DoS攻击。
This work addresses the limitations of traditional stereo matching methods, which model disparity estimation as deterministic regression and consequently struggle with multimodal distributions and ambiguous regions, often yielding mean bias. To overcome this, the paper introduces the first generative stereo matching framework that unifies deterministic regression with probabilistic distribution modeling. The approach leverages a two-stage progressive cascaded network, a frequency-decoupled StereoDiT pixel diffusion Transformer, and a novel Transition Flow Matching optimization objective to enable efficient few-step generation. The proposed method achieves state-of-the-art performance across multiple benchmarks—including Scene Flow, KITTI, ETH3D, and Middlebury—and demonstrates exceptional zero-shot generalization, significantly improving detail fidelity and consistency in geometrically discontinuous and textureless regions.
This work addresses the escalating human supervision cost in robotic autonomous data collection, where repeated human corrections for recurring failure modes lead to linearly increasing labor demands over task duration. To overcome this limitation, the authors introduce PhysClaw-0, a human–robot symbiotic agent that enables cross-episode memory and reuse of natural language corrections for the first time. By leveraging a large language model to parse human instructions into structured adjustment policies, integrating a vision-language model as a verification module, and incorporating an autonomous reset mechanism, PhysClaw-0 forms a closed-loop data collection system that solicits human intervention only after exhausting retry attempts. Evaluated on a real-world tabletop clearing task, the approach reduces human effort to 16% of baseline levels, boosts single-attempt success rates from 12.5% to 47.5%, and achieves fine-tuning performance comparable to teleoperation-based training, substantially lowering long-term human labor costs.
This work addresses the challenge of end-effector misalignment or insertion failure in high-precision robotic manipulation caused by calibration inaccuracies, perception errors, and contact dynamics within world-action (WA) models. To overcome these limitations, the authors propose a hybrid attention-based latent-variable-guided online reinforcement learning framework. The approach integrates latent features and action priors from the WA model generation process through a lightweight actor-critic adapter, augmented with a hybrid attention mechanism that jointly captures task-relevant information from visual context and end-effector correction demands while preserving temporal consistency in actions. Evaluated on four real-world high-precision tasks, the method achieves an average success rate of 87.1%, surpassing the strongest baseline by 19.2 percentage points, with only 45–75 minutes of online training per task.
为解决自动驾驶中长期预测与即时决策难以兼顾的问题,提出Drive-HWM框架,通过慢-快层级模型在不同时间尺度上进行未来场景预测和动作生成。
本文提出SMTrap,一种基于SMT冲突指导的方法,生成计算密集型CSP查询以低成本实现对大型推理模型的DoS攻击。
This work addresses the limitations of traditional stereo matching methods, which model disparity estimation as deterministic regression and consequently struggle with multimodal distributions and ambiguous regions, often yielding mean bias. To overcome this, the paper introduces the first generative stereo matching framework that unifies deterministic regression with probabilistic distribution modeling. The approach leverages a two-stage progressive cascaded network, a frequency-decoupled StereoDiT pixel diffusion Transformer, and a novel Transition Flow Matching optimization objective to enable efficient few-step generation. The proposed method achieves state-of-the-art performance across multiple benchmarks—including Scene Flow, KITTI, ETH3D, and Middlebury—and demonstrates exceptional zero-shot generalization, significantly improving detail fidelity and consistency in geometrically discontinuous and textureless regions.
This work addresses the escalating human supervision cost in robotic autonomous data collection, where repeated human corrections for recurring failure modes lead to linearly increasing labor demands over task duration. To overcome this limitation, the authors introduce PhysClaw-0, a human–robot symbiotic agent that enables cross-episode memory and reuse of natural language corrections for the first time. By leveraging a large language model to parse human instructions into structured adjustment policies, integrating a vision-language model as a verification module, and incorporating an autonomous reset mechanism, PhysClaw-0 forms a closed-loop data collection system that solicits human intervention only after exhausting retry attempts. Evaluated on a real-world tabletop clearing task, the approach reduces human effort to 16% of baseline levels, boosts single-attempt success rates from 12.5% to 47.5%, and achieves fine-tuning performance comparable to teleoperation-based training, substantially lowering long-term human labor costs.
This work addresses the challenge of end-effector misalignment or insertion failure in high-precision robotic manipulation caused by calibration inaccuracies, perception errors, and contact dynamics within world-action (WA) models. To overcome these limitations, the authors propose a hybrid attention-based latent-variable-guided online reinforcement learning framework. The approach integrates latent features and action priors from the WA model generation process through a lightweight actor-critic adapter, augmented with a hybrid attention mechanism that jointly captures task-relevant information from visual context and end-effector correction demands while preserving temporal consistency in actions. Evaluated on four real-world high-precision tasks, the method achieves an average success rate of 87.1%, surpassing the strongest baseline by 19.2 percentage points, with only 45–75 minutes of online training per task.