From Prediction to Decision: World-Model-Guided Action Selection for Continuous Pile Excavation
研究解决了轮式装载机连续挖掘中的决策问题,通过世界-行动模型(WAM)预测和选择最佳挖掘动作,减少挖掘次数并提高效率。
研究解决了轮式装载机连续挖掘中的决策问题,通过世界-行动模型(WAM)预测和选择最佳挖掘动作,减少挖掘次数并提高效率。
该研究提出了一种增强隐私的联邦学习框架,通过结合动态差分隐私、轻量级同态加密和本地差分隐私技术,在保证数据隐私的同时支持异步环境下的分布式训练。
研究评估了深度扰动学习在机器遗忘中的三种可能角色,通过修正实现问题后发现其直接删除效果不佳,作为正则化器或热启动也表现欠佳。
Underwater object detection faces severe challenges including low-level feature degradation (e.g., texture, edge, and color distortion), noise interference, and class imbalance due to optical distortions inherent in aquatic environments. Method: This work conducts a systematic robustness evaluation of YOLOv8–v12 across six simulated underwater conditions using the DUO and Roboflow100 datasets (10,000 annotated images), employing cross-model and cross-environment benchmarking. It further proposes a noise-aware sample injection strategy and enhancement-domain fine-tuning to improve generalization under noise perturbations and real underwater domains. Contribution/Results: We identify—for the first time—the robustness bottlenecks of YOLO models underwater: although YOLOv12 achieves the highest overall accuracy, it exhibits extreme sensitivity to noise; detection performance is predominantly constrained by sample quantity and instance frequency. Our lightweight training paradigm and targeted image enhancement significantly boost robustness and domain adaptability, empirically validating their efficacy for underwater detection.
Underwater object detection (UOD) faces five core challenges: severe image degradation, small and highly deformable object scales, scarcity of annotated data, stringent real-time inference requirements, and poor generalization of existing models. To address these, this work systematically analyzes the challenges and surveys methodological evolution. Crucially, it pioneers the integration of large vision-language models (LVLMs) into UOD: leveraging DALL·E 3 to generate high-fidelity synthetic underwater imagery, and employing Florence-2 for multimodal fine-tuning and cross-domain transfer. Experiments demonstrate substantial improvements in detection accuracy and robustness under complex underwater conditions—particularly in realistic scene modeling—while highlighting persistent bottlenecks in small-object localization and dynamic scene adaptation. This study bridges a critical gap by establishing the first LVLM-based framework for UOD, offering a novel pathway toward data-efficient learning and lightweight deployment.
研究解决了轮式装载机连续挖掘中的决策问题,通过世界-行动模型(WAM)预测和选择最佳挖掘动作,减少挖掘次数并提高效率。
该研究提出了一种增强隐私的联邦学习框架,通过结合动态差分隐私、轻量级同态加密和本地差分隐私技术,在保证数据隐私的同时支持异步环境下的分布式训练。
研究评估了深度扰动学习在机器遗忘中的三种可能角色,通过修正实现问题后发现其直接删除效果不佳,作为正则化器或热启动也表现欠佳。
Underwater object detection faces severe challenges including low-level feature degradation (e.g., texture, edge, and color distortion), noise interference, and class imbalance due to optical distortions inherent in aquatic environments. Method: This work conducts a systematic robustness evaluation of YOLOv8–v12 across six simulated underwater conditions using the DUO and Roboflow100 datasets (10,000 annotated images), employing cross-model and cross-environment benchmarking. It further proposes a noise-aware sample injection strategy and enhancement-domain fine-tuning to improve generalization under noise perturbations and real underwater domains. Contribution/Results: We identify—for the first time—the robustness bottlenecks of YOLO models underwater: although YOLOv12 achieves the highest overall accuracy, it exhibits extreme sensitivity to noise; detection performance is predominantly constrained by sample quantity and instance frequency. Our lightweight training paradigm and targeted image enhancement significantly boost robustness and domain adaptability, empirically validating their efficacy for underwater detection.
Underwater object detection (UOD) faces five core challenges: severe image degradation, small and highly deformable object scales, scarcity of annotated data, stringent real-time inference requirements, and poor generalization of existing models. To address these, this work systematically analyzes the challenges and surveys methodological evolution. Crucially, it pioneers the integration of large vision-language models (LVLMs) into UOD: leveraging DALL·E 3 to generate high-fidelity synthetic underwater imagery, and employing Florence-2 for multimodal fine-tuning and cross-domain transfer. Experiments demonstrate substantial improvements in detection accuracy and robustness under complex underwater conditions—particularly in realistic scene modeling—while highlighting persistent bottlenecks in small-object localization and dynamic scene adaptation. This study bridges a critical gap by establishing the first LVLM-based framework for UOD, offering a novel pathway toward data-efficient learning and lightweight deployment.