Establishing a Dynamic Multimodal HRI Dataset for Engagement Analysis with a Humanoid Robot
本文通过构建包含生理信号、行为数据和自我报告的多模态数据集,以分析人机交互中的用户参与度,解决了以往研究中缺乏整合生理信号的问题。
本文通过构建包含生理信号、行为数据和自我报告的多模态数据集,以分析人机交互中的用户参与度,解决了以往研究中缺乏整合生理信号的问题。
This work addresses the challenge of reliably detecting surface scratches in semiconductor manufacturing, which are difficult to identify due to their irregular shapes, low contrast, and varying scales. To this end, the authors propose ScratNet, an end-to-end scratch segmentation framework built upon an enhanced Swin Transformer backbone and a custom-designed decoder. ScratNet introduces several key innovations: a multi-scale dilated aggregation module, a stem integration mechanism, and an anisotropic convolution-driven boundary refinement branch, collectively enabling stage-adaptive feature fusion and boundary-aware optimization. Experimental results demonstrate that ScratNet significantly outperforms existing methods under diverse and complex imaging conditions, achieving superior detection accuracy and robustness—particularly for fine and irregular scratches.
This work addresses the underexplored vulnerability of event-driven spiking neural network (SNN)–based object detection models to availability-oriented backdoor attacks. We propose the Event Burst Trigger (EBT) attack, which injects carefully crafted event-based triggers into training data to induce dense event streams during inference, thereby significantly increasing the computational overhead of non-maximum suppression (NMS) and degrading system availability. Notably, EBT requires no modifications to the model architecture, loss function, or inference pipeline, allowing it to evade existing detection mechanisms such as STRIP. Experimental results demonstrate that, with less than a 0.099 drop in mAP@0.5, the attack can increase NMS latency by up to 38%, effectively elevating resource consumption and reducing scheduling slack without producing conspicuous resource usage spikes.
This work addresses the semantic inconsistency arising from heterogeneous representations in multi-scale feature fusion, which limits object detection accuracy. To this end, the authors propose the FINE module, which leverages high-level contextual guidance to align low-level features through cross-level attention prior to fusion. An alignment-aware token sampling strategy is introduced to substantially reduce computational complexity. Furthermore, spatial-channel modulation combined with residual element-wise modulation is incorporated to enhance responses of semantically relevant pixels while preserving precise localization capabilities. The proposed method demonstrates consistent and significant performance gains across diverse detectors, achieving notable improvements in detection accuracy with negligible additional computational overhead.
Existing multimodal 3D CAD generation methods suffer from high computational costs and low training efficiency. This work proposes a lightweight prefix embedding mechanism that leverages a mapping network to transform image embeddings into textual prefixes, which guide a pretrained large language model to effectively fuse visual and textual information for predicting CAD modeling sequences. By introducing only a small number of trainable parameters, the approach substantially reduces resource consumption and accelerates training. Experimental results on a newly constructed text-image paired dataset demonstrate that the proposed method achieves comparable 3D CAD generation quality to baseline approaches while using approximately one-quarter of the parameters and training twice as fast.
本文通过构建包含生理信号、行为数据和自我报告的多模态数据集,以分析人机交互中的用户参与度,解决了以往研究中缺乏整合生理信号的问题。
This work addresses the challenge of reliably detecting surface scratches in semiconductor manufacturing, which are difficult to identify due to their irregular shapes, low contrast, and varying scales. To this end, the authors propose ScratNet, an end-to-end scratch segmentation framework built upon an enhanced Swin Transformer backbone and a custom-designed decoder. ScratNet introduces several key innovations: a multi-scale dilated aggregation module, a stem integration mechanism, and an anisotropic convolution-driven boundary refinement branch, collectively enabling stage-adaptive feature fusion and boundary-aware optimization. Experimental results demonstrate that ScratNet significantly outperforms existing methods under diverse and complex imaging conditions, achieving superior detection accuracy and robustness—particularly for fine and irregular scratches.
This work addresses the underexplored vulnerability of event-driven spiking neural network (SNN)–based object detection models to availability-oriented backdoor attacks. We propose the Event Burst Trigger (EBT) attack, which injects carefully crafted event-based triggers into training data to induce dense event streams during inference, thereby significantly increasing the computational overhead of non-maximum suppression (NMS) and degrading system availability. Notably, EBT requires no modifications to the model architecture, loss function, or inference pipeline, allowing it to evade existing detection mechanisms such as STRIP. Experimental results demonstrate that, with less than a 0.099 drop in mAP@0.5, the attack can increase NMS latency by up to 38%, effectively elevating resource consumption and reducing scheduling slack without producing conspicuous resource usage spikes.
This work addresses the semantic inconsistency arising from heterogeneous representations in multi-scale feature fusion, which limits object detection accuracy. To this end, the authors propose the FINE module, which leverages high-level contextual guidance to align low-level features through cross-level attention prior to fusion. An alignment-aware token sampling strategy is introduced to substantially reduce computational complexity. Furthermore, spatial-channel modulation combined with residual element-wise modulation is incorporated to enhance responses of semantically relevant pixels while preserving precise localization capabilities. The proposed method demonstrates consistent and significant performance gains across diverse detectors, achieving notable improvements in detection accuracy with negligible additional computational overhead.
Existing multimodal 3D CAD generation methods suffer from high computational costs and low training efficiency. This work proposes a lightweight prefix embedding mechanism that leverages a mapping network to transform image embeddings into textual prefixes, which guide a pretrained large language model to effectively fuse visual and textual information for predicting CAD modeling sequences. By introducing only a small number of trainable parameters, the approach substantially reduces resource consumption and accelerates training. Experimental results on a newly constructed text-image paired dataset demonstrate that the proposed method achieves comparable 3D CAD generation quality to baseline approaches while using approximately one-quarter of the parameters and training twice as fast.