Occlusion-Robust Multimodal Emotion Recognition in VR via Fusion of Facial Images and EMG
为解决VR中因头戴设备遮挡上半脸导致的情绪识别问题,通过融合下半脸视频与面部肌电图数据来分类七种情绪,提高识别准确性。
为解决VR中因头戴设备遮挡上半脸导致的情绪识别问题,通过融合下半脸视频与面部肌电图数据来分类七种情绪,提高识别准确性。
为解决3D高斯点绘场景文件体积过大的问题,提出KISS-GS模块化压缩方案,通过先进的剪枝技术和新颖的编码格式实现显著压缩效果。
This work addresses the challenge of simultaneously achieving substantial parameter reduction and performance preservation in large language model (LLM) compression by proposing an activation- and influence-aware low-rank approximation method. The approach uniquely integrates element-wise backward influence metrics into singular value decomposition (SVD)-based compression and employs a single closed-form alternating least squares (ALS) step to preserve functionality, offering both locality and monotonic descent properties. It is also orthogonal to end-to-end fine-tuning techniques. Experimental results demonstrate that with ≤60% of the original parameters retained, the method improves perplexity by over 18% compared to SVD-LLM(W). Moreover, it achieves comparable model quality using only ~10% calibration data while significantly reducing FLOPs, peak memory usage, and per-token latency.
This work addresses the critical challenge of silent failures in object detectors—such as missed pedestrian detections in safety-critical scenarios—that evade conventional out-of-distribution (OOD) detection methods. To this end, the authors propose KGFP, a knowledge-guided failure prediction framework that formulates detector failures as semantic inconsistencies between the detector’s internal features and embeddings from a vision foundation model. Leveraging a dual-encoder architecture and angular distance metrics, KGFP establishes a runtime selective prediction gating mechanism. Evaluated on pedestrian detection within the COCO benchmark, KGFP improves recall from 64.3% to 84.5% at a 5% false positive rate and consistently outperforms existing OOD detection approaches across six COCO-O domains.
This study addresses the challenge of high-precision 3D object localization for human-robot interaction by integrating monocular RGB images, natural language instructions, and robot state information. To this end, the authors propose an end-to-end framework built upon a pretrained vision-language model (VLM), enhanced with QLoRA-based efficient fine-tuning, a custom regression head, and a conditional routing mechanism. This design preserves the VLM’s general visual understanding capabilities while introducing dedicated 3D localization functionality. The work introduces a heterogeneous dataset comprising over 100,000 samples and demonstrates strong empirical performance, achieving a median absolute error of 13 mm—representing a fivefold improvement over the unmodified baseline. Notably, approximately 25% of predictions meet the accuracy threshold required for direct robotic manipulation.
为解决VR中因头戴设备遮挡上半脸导致的情绪识别问题,通过融合下半脸视频与面部肌电图数据来分类七种情绪,提高识别准确性。
为解决3D高斯点绘场景文件体积过大的问题,提出KISS-GS模块化压缩方案,通过先进的剪枝技术和新颖的编码格式实现显著压缩效果。
This work addresses the challenge of simultaneously achieving substantial parameter reduction and performance preservation in large language model (LLM) compression by proposing an activation- and influence-aware low-rank approximation method. The approach uniquely integrates element-wise backward influence metrics into singular value decomposition (SVD)-based compression and employs a single closed-form alternating least squares (ALS) step to preserve functionality, offering both locality and monotonic descent properties. It is also orthogonal to end-to-end fine-tuning techniques. Experimental results demonstrate that with ≤60% of the original parameters retained, the method improves perplexity by over 18% compared to SVD-LLM(W). Moreover, it achieves comparable model quality using only ~10% calibration data while significantly reducing FLOPs, peak memory usage, and per-token latency.
This work addresses the critical challenge of silent failures in object detectors—such as missed pedestrian detections in safety-critical scenarios—that evade conventional out-of-distribution (OOD) detection methods. To this end, the authors propose KGFP, a knowledge-guided failure prediction framework that formulates detector failures as semantic inconsistencies between the detector’s internal features and embeddings from a vision foundation model. Leveraging a dual-encoder architecture and angular distance metrics, KGFP establishes a runtime selective prediction gating mechanism. Evaluated on pedestrian detection within the COCO benchmark, KGFP improves recall from 64.3% to 84.5% at a 5% false positive rate and consistently outperforms existing OOD detection approaches across six COCO-O domains.
This study addresses the challenge of high-precision 3D object localization for human-robot interaction by integrating monocular RGB images, natural language instructions, and robot state information. To this end, the authors propose an end-to-end framework built upon a pretrained vision-language model (VLM), enhanced with QLoRA-based efficient fine-tuning, a custom regression head, and a conditional routing mechanism. This design preserves the VLM’s general visual understanding capabilities while introducing dedicated 3D localization functionality. The work introduces a heterogeneous dataset comprising over 100,000 samples and demonstrates strong empirical performance, achieving a median absolute error of 13 mm—representing a fivefold improvement over the unmodified baseline. Notably, approximately 25% of predictions meet the accuracy threshold required for direct robotic manipulation.