A multimodal large language model for evidence-based autism spectrum disorder screening
为解决自闭症谱系障碍早期筛查瓶颈,提出ASDchat模型,采用视频、音频和对话作为输入,基于双分支架构进行筛查并生成行为证据。
为解决自闭症谱系障碍早期筛查瓶颈,提出ASDchat模型,采用视频、音频和对话作为输入,基于双分支架构进行筛查并生成行为证据。
Existing video object removal methods often suffer from inefficiency due to multi-step denoising or introduce noticeable artifacts when relying on single-step generation. This work proposes a draft-free, end-to-end single-step removal framework that distills the refinement capability of multi-step processes into a single-step diffusion model via a Prior-Privileged Consistency Distillation strategy. To guide inpainting without external drafts, the method introduces a Self-Guided Fast Planting module based on a Temporal Masked Transformer, which automatically generates temporally coherent pseudo-drafts. This approach achieves high-quality, draft-independent single-step video object removal for the first time, attaining state-of-the-art performance across multiple metrics while processing an entire video in approximately one second—significantly outperforming existing techniques.
This work addresses the significant gap between theory and practice in single-loop methods for bilevel optimization, such as Approximate Implicit Differentiation (AID) and Iterative Differentiation (ITD), whose existing convergence bounds are notoriously loose. To bridge this gap, the authors propose a Decoupled Norm Analysis (DNA) framework that enables a refined convergence analysis of single-loop AID and ITD. Theoretically, they improve the convergence rate of AID from 𝒪(κ⁶/K) to 𝒪(κ⁵/K) and establish that ITD achieves an asymptotic error of 𝒪(κ²)—matching the known lower bound and improving upon the previous 𝒪(κ³) guarantee. Extensive numerical experiments on both synthetic and real-world tasks corroborate the tightness and practical relevance of the derived bounds.
This work proposes EISAM, a novel optimizer that integrates extragradient methods with Sharpness-Aware Minimization (SAM) to address the tendency of conventional optimizers to converge to sharp minima, which harms model generalization. EISAM employs a prediction step to explore the geometry of the loss landscape and a perturbation step that jointly updates parameters with the base optimizer, enabling more efficient convergence to flat minima. This design substantially reduces sensitivity to the perturbation radius, enhancing robustness and simplifying hyperparameter tuning. Extensive experiments demonstrate that EISAM consistently outperforms SGD, Adam, and SAM across diverse models and benchmark datasets, achieving higher test accuracy and improved training efficiency. The study also provides practical guidelines for hyperparameter configuration.
Existing dual-arm manipulation approaches generally lack effective inter-arm interaction and dynamic role allocation mechanisms, hindering efficient collaboration. This work proposes the PA-BiCoop framework, which introduces—for the first time—a dynamic leader-follower role assignment mechanism within a single model. By employing a shared global feature encoder, role-specific decoders, and a follower-arm pose prediction module grounded in a relative coordinate system, the framework enables automatic role allocation and functional coordination. It further enhances inter-arm knowledge sharing and task synergy through heatmap-driven manipulability prediction for core tasks. Evaluated in the RLBench2 simulation environment, the method outperforms the current state-of-the-art by an average of 48%; in real-world tasks, it achieves performance gains exceeding 50%, demonstrating its effectiveness and strong generalization capability.
为解决自闭症谱系障碍早期筛查瓶颈,提出ASDchat模型,采用视频、音频和对话作为输入,基于双分支架构进行筛查并生成行为证据。
Existing video object removal methods often suffer from inefficiency due to multi-step denoising or introduce noticeable artifacts when relying on single-step generation. This work proposes a draft-free, end-to-end single-step removal framework that distills the refinement capability of multi-step processes into a single-step diffusion model via a Prior-Privileged Consistency Distillation strategy. To guide inpainting without external drafts, the method introduces a Self-Guided Fast Planting module based on a Temporal Masked Transformer, which automatically generates temporally coherent pseudo-drafts. This approach achieves high-quality, draft-independent single-step video object removal for the first time, attaining state-of-the-art performance across multiple metrics while processing an entire video in approximately one second—significantly outperforming existing techniques.
This work addresses the significant gap between theory and practice in single-loop methods for bilevel optimization, such as Approximate Implicit Differentiation (AID) and Iterative Differentiation (ITD), whose existing convergence bounds are notoriously loose. To bridge this gap, the authors propose a Decoupled Norm Analysis (DNA) framework that enables a refined convergence analysis of single-loop AID and ITD. Theoretically, they improve the convergence rate of AID from 𝒪(κ⁶/K) to 𝒪(κ⁵/K) and establish that ITD achieves an asymptotic error of 𝒪(κ²)—matching the known lower bound and improving upon the previous 𝒪(κ³) guarantee. Extensive numerical experiments on both synthetic and real-world tasks corroborate the tightness and practical relevance of the derived bounds.
This work proposes EISAM, a novel optimizer that integrates extragradient methods with Sharpness-Aware Minimization (SAM) to address the tendency of conventional optimizers to converge to sharp minima, which harms model generalization. EISAM employs a prediction step to explore the geometry of the loss landscape and a perturbation step that jointly updates parameters with the base optimizer, enabling more efficient convergence to flat minima. This design substantially reduces sensitivity to the perturbation radius, enhancing robustness and simplifying hyperparameter tuning. Extensive experiments demonstrate that EISAM consistently outperforms SGD, Adam, and SAM across diverse models and benchmark datasets, achieving higher test accuracy and improved training efficiency. The study also provides practical guidelines for hyperparameter configuration.
Existing dual-arm manipulation approaches generally lack effective inter-arm interaction and dynamic role allocation mechanisms, hindering efficient collaboration. This work proposes the PA-BiCoop framework, which introduces—for the first time—a dynamic leader-follower role assignment mechanism within a single model. By employing a shared global feature encoder, role-specific decoders, and a follower-arm pose prediction module grounded in a relative coordinate system, the framework enables automatic role allocation and functional coordination. It further enhances inter-arm knowledge sharing and task synergy through heatmap-driven manipulability prediction for core tasks. Evaluated in the RLBench2 simulation environment, the method outperforms the current state-of-the-art by an average of 48%; in real-world tasks, it achieves performance gains exceeding 50%, demonstrating its effectiveness and strong generalization capability.