G-ray: Ray-Level Relative Geometric Position Encoding in Multi-View Vision Transformers under Camera Heterogeneity
研究针对多视角视觉Transformer在相机异质性下的相对位置编码问题,提出G-ray方法,通过基于光线角度的旋转相位参数化实现投影不变的位置一致性。
研究针对多视角视觉Transformer在相机异质性下的相对位置编码问题,提出G-ray方法,通过基于光线角度的旋转相位参数化实现投影不变的位置一致性。
本文提出EDGE框架,通过经验蒸馏和引导探索解决强化学习中经验复用问题,提升策略性能并减少对外部检索依赖。
This work addresses the limitations of World Action Models (WAMs) in complex embodied tasks—specifically, their inadequate planning, state maintenance, and failure recovery stemming from a disconnect between prediction and reasoning. To overcome this, we propose an agent-based framework that employs a vision-language model–driven task manager to maintain a structured scene belief and task graph. High-level semantic plans are projected into sequences of atomic skills that respect both task dependencies and robot capability constraints. An event-driven dual-timescale feedback mechanism, coupled with a lightweight progress estimator, enables a verifiable and recoverable execution loop. Our approach introduces, for the first time, external structured state maintenance and closed-loop decision-making, substantially enhancing WAMs’ global planning and local fault tolerance. Experiments show our method achieves a 59.6% end-to-end success rate (69.9% on subtasks) on RoboMemArena and 23.7% on RoboCerebra Ideal, significantly outperforming existing approaches.
This work addresses the challenges of data scarcity and misalignment between semantic representations and fixed acoustic supervision in end-to-end spoken dialogue modeling for low-resource Chinese dialects. To overcome these issues, the authors construct a scalable data pipeline for dialectal spoken dialogue synthesis and propose a two-stage post-training strategy incorporating a self-aligned speech supervision mechanism that dynamically aligns acoustic targets with the model’s evolving semantic representations. This approach achieves the first successful end-to-end spoken dialogue modeling for low-resource Chinese dialects, significantly outperforming existing baselines across multiple dialects. Substantial improvements are observed in dialect consistency, response quality, and speech intelligibility. The complete framework is open-sourced to facilitate future research in this underexplored domain.
This work addresses the limitation of existing research agents that oversimplify complex scientific projects into single tasks, resulting in ambiguous task boundaries, disorganized execution, and missing deliverables—challenges that hinder long-horizon, multi-objective, and dependency-sensitive research planning. To overcome this, the authors propose a graph-guided, project-level planning approach that explicitly decomposes a research project into executable task compositions with clearly attributed contributions and explicit dependencies, leveraging an innovative atomic representation and a directed provenance graph. A lightweight Bernoulli block model optimizes task selection, generating standardized task contracts that specify objectives, dependencies, and constraints, enabling seamless decoupled integration with arbitrary executors. Evaluated on ten scientific benchmarks, the method achieves an average quality score of 7.15, significantly outperforming baselines (4.58 and 5.31), and when integrated with AutoResearchClaw, boosts downstream task accuracy from 0.536 to 0.759.
研究针对多视角视觉Transformer在相机异质性下的相对位置编码问题,提出G-ray方法,通过基于光线角度的旋转相位参数化实现投影不变的位置一致性。
本文提出EDGE框架,通过经验蒸馏和引导探索解决强化学习中经验复用问题,提升策略性能并减少对外部检索依赖。
This work addresses the limitations of World Action Models (WAMs) in complex embodied tasks—specifically, their inadequate planning, state maintenance, and failure recovery stemming from a disconnect between prediction and reasoning. To overcome this, we propose an agent-based framework that employs a vision-language model–driven task manager to maintain a structured scene belief and task graph. High-level semantic plans are projected into sequences of atomic skills that respect both task dependencies and robot capability constraints. An event-driven dual-timescale feedback mechanism, coupled with a lightweight progress estimator, enables a verifiable and recoverable execution loop. Our approach introduces, for the first time, external structured state maintenance and closed-loop decision-making, substantially enhancing WAMs’ global planning and local fault tolerance. Experiments show our method achieves a 59.6% end-to-end success rate (69.9% on subtasks) on RoboMemArena and 23.7% on RoboCerebra Ideal, significantly outperforming existing approaches.
This work addresses the challenges of data scarcity and misalignment between semantic representations and fixed acoustic supervision in end-to-end spoken dialogue modeling for low-resource Chinese dialects. To overcome these issues, the authors construct a scalable data pipeline for dialectal spoken dialogue synthesis and propose a two-stage post-training strategy incorporating a self-aligned speech supervision mechanism that dynamically aligns acoustic targets with the model’s evolving semantic representations. This approach achieves the first successful end-to-end spoken dialogue modeling for low-resource Chinese dialects, significantly outperforming existing baselines across multiple dialects. Substantial improvements are observed in dialect consistency, response quality, and speech intelligibility. The complete framework is open-sourced to facilitate future research in this underexplored domain.
This work addresses the limitation of existing research agents that oversimplify complex scientific projects into single tasks, resulting in ambiguous task boundaries, disorganized execution, and missing deliverables—challenges that hinder long-horizon, multi-objective, and dependency-sensitive research planning. To overcome this, the authors propose a graph-guided, project-level planning approach that explicitly decomposes a research project into executable task compositions with clearly attributed contributions and explicit dependencies, leveraging an innovative atomic representation and a directed provenance graph. A lightweight Bernoulli block model optimizes task selection, generating standardized task contracts that specify objectives, dependencies, and constraints, enabling seamless decoupled integration with arbitrary executors. Evaluated on ten scientific benchmarks, the method achieves an average quality score of 7.15, significantly outperforming baselines (4.58 and 5.31), and when integrated with AutoResearchClaw, boosts downstream task accuracy from 0.536 to 0.759.