Motion-Aware Reasoning from Speech to Mask Tracks: Runner-up Solution for the MeViS-Audio Track of the 8th LSVOS Challenge 2026
本文提出Speech2MaskTrack方法,通过语音识别、运动中心时间定位、掩模跟踪等步骤,解决语音引导的视频对象分割问题。
本文提出Speech2MaskTrack方法,通过语音识别、运动中心时间定位、掩模跟踪等步骤,解决语音引导的视频对象分割问题。
Existing single-step generative visuomotor policies often suffer from spatial bias, loss of high-frequency details, and ambiguous action modes due to oversimplified inference. To address these limitations, this work proposes a high-fidelity single-step framework that integrates Recursive Consistent Action Flow (RCAF) to correct spatial truncation errors, Dual-Timestep Frequency Consistency (DTFC) to preserve high-frequency manipulation details, and Contrastive Flow Matching (CFM) to disentangle multimodal action distributions. Requiring only a single forward pass (1 NFE), the method achieves or surpasses the control performance of multi-step baselines—such as 10-step policies—on RoboTwin, Adroit, DexArt, and real robotic platforms, significantly reducing latency while enhancing both precision and action diversity.
To address the challenges of slow inference speed, low accuracy, and poor deployability of deep learning models for coal-gangue detection on industrial edge devices, this paper proposes a lightweight and efficient object detection framework. Methodologically: (i) we design ADown, a lightweight downsampling module; (ii) we construct C2PSA-TriAtt, a cross-layer feature aggregation architecture integrating Triplet Attention; and (iii) we introduce Inner-FocalerIoU, a novel loss function that jointly optimizes localization accuracy and convergence on hard samples. Built upon YOLOv11, the framework adopts ShuffleNetV2 as its backbone and incorporates all three innovations. Experiments demonstrate a state-of-the-art mAP of 99.10%, with a 38% reduction in model size, 41% fewer parameters, 40% lower computational cost (FLOPs), and a 1 ms decrease in per-image inference latency—significantly fulfilling real-time deployment requirements on coal-mine edge devices.
本文提出Speech2MaskTrack方法,通过语音识别、运动中心时间定位、掩模跟踪等步骤,解决语音引导的视频对象分割问题。
Existing single-step generative visuomotor policies often suffer from spatial bias, loss of high-frequency details, and ambiguous action modes due to oversimplified inference. To address these limitations, this work proposes a high-fidelity single-step framework that integrates Recursive Consistent Action Flow (RCAF) to correct spatial truncation errors, Dual-Timestep Frequency Consistency (DTFC) to preserve high-frequency manipulation details, and Contrastive Flow Matching (CFM) to disentangle multimodal action distributions. Requiring only a single forward pass (1 NFE), the method achieves or surpasses the control performance of multi-step baselines—such as 10-step policies—on RoboTwin, Adroit, DexArt, and real robotic platforms, significantly reducing latency while enhancing both precision and action diversity.
To address the challenges of slow inference speed, low accuracy, and poor deployability of deep learning models for coal-gangue detection on industrial edge devices, this paper proposes a lightweight and efficient object detection framework. Methodologically: (i) we design ADown, a lightweight downsampling module; (ii) we construct C2PSA-TriAtt, a cross-layer feature aggregation architecture integrating Triplet Attention; and (iii) we introduce Inner-FocalerIoU, a novel loss function that jointly optimizes localization accuracy and convergence on hard samples. Built upon YOLOv11, the framework adopts ShuffleNetV2 as its backbone and incorporates all three innovations. Experiments demonstrate a state-of-the-art mAP of 99.10%, with a 38% reduction in model size, 41% fewer parameters, 40% lower computational cost (FLOPs), and a 1 ms decrease in per-image inference latency—significantly fulfilling real-time deployment requirements on coal-mine edge devices.