IL-ACT: Imitation Learning with Adaptive Cartesian Tracking Control for a 30-ton Excavator
为解决30吨级挖掘机的自主控制问题,提出了一种结合模仿学习与自适应笛卡尔跟踪控制的新框架IL-ACT,通过预训练的操作员演示生成关节速度,并利用自适应反馈和增益/偏置估计来修正命令。
为解决30吨级挖掘机的自主控制问题,提出了一种结合模仿学习与自适应笛卡尔跟踪控制的新框架IL-ACT,通过预训练的操作员演示生成关节速度,并利用自适应反馈和增益/偏置估计来修正命令。
本文提出了一种新的高维时间序列频谱差分网络分析方法,通过直接估计两个高维逆谱密度差异来研究不同条件下的网络变化。
研究探讨了在单轮心理健康问答中,选择性检索如何改善或影响回答质量,通过特定条件下的检索需求维度来优化检索策略。
This work addresses the challenges of view inconsistency and geometric distortion in single-image 3D reconstruction, which often arise from projection ambiguities during multi-view synthesis. To mitigate these issues, the authors propose a view-adaptive neural rendering framework that employs a shared feature backbone to capture global structure while enabling per-view independent correction of rendering errors. A lightweight self-attention fusion module is introduced to integrate multi-view information and enhance geometric consistency without relying on supervision from diffusion models such as SDS. The method optimizes solely with photometric loss, achieving near state-of-the-art reconstruction fidelity while maintaining computational efficiency and significantly improving view consistency and practical performance.
This work addresses the high false positive and false negative rates in cross-modal association between faces and voices, which stem from modality heterogeneity. To mitigate this issue, the authors propose a joint learning framework that integrates convex hull feature embedding with a cross-modal attention mechanism. By compactly aggregating cross-modal features of the same identity within a unified embedding space and incorporating deep metric learning, the method effectively narrows the semantic gap across modalities. Experimental results on the VoxCeleb dataset demonstrate that the proposed approach significantly outperforms state-of-the-art methods in cross-modal verification, matching, and retrieval tasks, achieving substantial reductions in error rates.
为解决30吨级挖掘机的自主控制问题,提出了一种结合模仿学习与自适应笛卡尔跟踪控制的新框架IL-ACT,通过预训练的操作员演示生成关节速度,并利用自适应反馈和增益/偏置估计来修正命令。
本文提出了一种新的高维时间序列频谱差分网络分析方法,通过直接估计两个高维逆谱密度差异来研究不同条件下的网络变化。
研究探讨了在单轮心理健康问答中,选择性检索如何改善或影响回答质量,通过特定条件下的检索需求维度来优化检索策略。
This work addresses the challenges of view inconsistency and geometric distortion in single-image 3D reconstruction, which often arise from projection ambiguities during multi-view synthesis. To mitigate these issues, the authors propose a view-adaptive neural rendering framework that employs a shared feature backbone to capture global structure while enabling per-view independent correction of rendering errors. A lightweight self-attention fusion module is introduced to integrate multi-view information and enhance geometric consistency without relying on supervision from diffusion models such as SDS. The method optimizes solely with photometric loss, achieving near state-of-the-art reconstruction fidelity while maintaining computational efficiency and significantly improving view consistency and practical performance.
This work addresses the high false positive and false negative rates in cross-modal association between faces and voices, which stem from modality heterogeneity. To mitigate this issue, the authors propose a joint learning framework that integrates convex hull feature embedding with a cross-modal attention mechanism. By compactly aggregating cross-modal features of the same identity within a unified embedding space and incorporating deep metric learning, the method effectively narrows the semantic gap across modalities. Experimental results on the VoxCeleb dataset demonstrate that the proposed approach significantly outperforms state-of-the-art methods in cross-modal verification, matching, and retrieval tasks, achieving substantial reductions in error rates.