Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation
研究针对多模态大语言模型在视频审核中对分布式隐性危害(DIH)的识别不足问题,通过开发多代理合成框架生成含DIH视频数据集,评估并揭示现有模型在此类问题上的缺陷。
研究针对多模态大语言模型在视频审核中对分布式隐性危害(DIH)的识别不足问题,通过开发多代理合成框架生成含DIH视频数据集,评估并揭示现有模型在此类问题上的缺陷。
This study addresses the challenges of task complexity, API orchestration, and cumbersome state management in natural language programming for social robots by proposing a multi-agent LLM framework. The system employs a dual-agent architecture to decouple physical and social interactions, utilizing task routing alongside sensor-skill binding and result reuse mechanisms to enable context-aware skill orchestration and efficient state management. Experimental results demonstrate superior performance in routing accuracy and scalability with low variance. Furthermore, user studies confirm the framework’s high usability and interaction quality, effectively enhancing both the efficiency and stability of robot programming.
This work addresses the challenge of evaluating and selecting discriminative regions for fine-grained image classification in the absence of ground-truth labels. The authors propose a unified, CLIP-guided, unsupervised region scoring framework that systematically compares regions generated by SAM segmentation masks against those from random crops. They introduce two label-free pseudo-labeling variants based on global and local embeddings, respectively. By integrating multiple scoring strategies—including cosine similarity, margin-based boundary scores, entropy, and Soft Negative Margin—with a top-k region aggregation mechanism, they find that Soft Negative Margin yields the best performance, producing pseudo-labels nearly on par with true labels. Furthermore, random crops combined with a small top-k consistently outperform SAM-derived masks across all datasets, and the optimal aggregation strategy varies depending on the region generation method.
This study addresses the gap between simulation-based research and real-world deployment in smart grid state estimation by presenting an experimental validation over a commercial 5G network. The authors develop a multi-node testbed integrating Raspberry Pi edge nodes with Typhoon hardware-in-the-loop (HIL) simulation, implementing end-to-end real-time state estimation and fault detection using an IEEE 4-bus feeder model, a phasor data concentrator (PDC), and key performance indicators (KPIs). Experimental results demonstrate that 5G achieves an average end-to-end latency approximately 6.5 times lower than LTE Cat-M, maintains high estimation accuracy under both steady-state and dynamic conditions, and enables fault detection with a latency as low as 0.80 seconds. This work provides the first empirical evidence of the real-time capability and reliability of smart grid state awareness in an operational 5G environment, effectively bridging the gap between simulation and practical implementation.
Traditional covariance struggles to effectively model second-order statistics on nonlinear Riemannian manifolds such as the space of symmetric positive-definite (SPD) matrices and Kendall’s shape space. This work proposes the first intrinsic and basepoint-invariant definition of Riemannian cross-covariance, achieved by parallel transporting local variations on the manifold to a common tangent space. The resulting estimator inherits desirable properties of Euclidean covariance and is supported by rigorous asymptotic theory. Efficient computational procedures are developed for spheres, SPD manifolds, and Kendall shape spaces. Numerical experiments and real-world analysis of heart valve shapes demonstrate the method’s effectiveness, establishing a foundational second-order statistical tool for representation learning with non-Euclidean data.
研究针对多模态大语言模型在视频审核中对分布式隐性危害(DIH)的识别不足问题,通过开发多代理合成框架生成含DIH视频数据集,评估并揭示现有模型在此类问题上的缺陷。
This study addresses the challenges of task complexity, API orchestration, and cumbersome state management in natural language programming for social robots by proposing a multi-agent LLM framework. The system employs a dual-agent architecture to decouple physical and social interactions, utilizing task routing alongside sensor-skill binding and result reuse mechanisms to enable context-aware skill orchestration and efficient state management. Experimental results demonstrate superior performance in routing accuracy and scalability with low variance. Furthermore, user studies confirm the framework’s high usability and interaction quality, effectively enhancing both the efficiency and stability of robot programming.
This work addresses the challenge of evaluating and selecting discriminative regions for fine-grained image classification in the absence of ground-truth labels. The authors propose a unified, CLIP-guided, unsupervised region scoring framework that systematically compares regions generated by SAM segmentation masks against those from random crops. They introduce two label-free pseudo-labeling variants based on global and local embeddings, respectively. By integrating multiple scoring strategies—including cosine similarity, margin-based boundary scores, entropy, and Soft Negative Margin—with a top-k region aggregation mechanism, they find that Soft Negative Margin yields the best performance, producing pseudo-labels nearly on par with true labels. Furthermore, random crops combined with a small top-k consistently outperform SAM-derived masks across all datasets, and the optimal aggregation strategy varies depending on the region generation method.
This study addresses the gap between simulation-based research and real-world deployment in smart grid state estimation by presenting an experimental validation over a commercial 5G network. The authors develop a multi-node testbed integrating Raspberry Pi edge nodes with Typhoon hardware-in-the-loop (HIL) simulation, implementing end-to-end real-time state estimation and fault detection using an IEEE 4-bus feeder model, a phasor data concentrator (PDC), and key performance indicators (KPIs). Experimental results demonstrate that 5G achieves an average end-to-end latency approximately 6.5 times lower than LTE Cat-M, maintains high estimation accuracy under both steady-state and dynamic conditions, and enables fault detection with a latency as low as 0.80 seconds. This work provides the first empirical evidence of the real-time capability and reliability of smart grid state awareness in an operational 5G environment, effectively bridging the gap between simulation and practical implementation.
Traditional covariance struggles to effectively model second-order statistics on nonlinear Riemannian manifolds such as the space of symmetric positive-definite (SPD) matrices and Kendall’s shape space. This work proposes the first intrinsic and basepoint-invariant definition of Riemannian cross-covariance, achieved by parallel transporting local variations on the manifold to a common tangent space. The resulting estimator inherits desirable properties of Euclidean covariance and is supported by rigorous asymptotic theory. Efficient computational procedures are developed for spheres, SPD manifolds, and Kendall shape spaces. Numerical experiments and real-world analysis of heart valve shapes demonstrate the method’s effectiveness, establishing a foundational second-order statistical tool for representation learning with non-Euclidean data.