Search-Based Metamorphic Testing of Vision-Language Models in Autonomous Underwater Robotic Software
为评估和提高视觉-语言模型在自主水下机器人软件中的适用性和可靠性,提出一种基于搜索的变形测试方法MetaVLM,通过最小化图像变换揭示模型错误。
为评估和提高视觉-语言模型在自主水下机器人软件中的适用性和可靠性,提出一种基于搜索的变形测试方法MetaVLM,通过最小化图像变换揭示模型错误。
研究用大型语言模型替代部分变异算子以修复Simulink-Stateflow模型中的错误,但结果表明这种方法降低了修复性能,揭示了直接集成LLMs的局限性。
This study addresses the challenge of reproducing failures in cyber-physical system (CPS) simulation testing, where non-deterministic behaviors often hinder consistent fault manifestation. To tackle this issue, the work extends delta debugging to stochastic CPS scenarios for the first time, introducing three novel delta debugging algorithms tailored for randomized environments. The proposed approach integrates statistical failure analysis, repeated execution, and environment-aware input minimization to identify a minimal triggering input while preserving fault semantics. Empirical evaluation on case studies involving elevator scheduling and autonomous mobile robots demonstrates that the method significantly enhances the stability of failure reproduction and substantially reduces debugging time, effectively mitigating execution flakiness inherent in such systems.
This work addresses the lack of reliable end-to-end testing methods for Retrieval-Augmented Generation (RAG) systems, which stems from complex interactions among their components. The authors propose RagTester, the first automated evaluation framework that introduces a coverage-guided test generation strategy. RagTester systematically synthesizes test documents, queries, and expected answers, and leverages large language models as judges to detect failure modes such as inaccurate retrieval, unsupported answers, and insufficient context utilization. Evaluated across 24 configurations with 72,000 tests, RagTester identified 21,633 faults—6.6% more than baseline methods—and outperformed baselines in 20 out of 24 configurations, substantially enhancing pre-deployment validation capabilities for RAG systems.
Existing static benchmarks struggle to effectively uncover sparse and clustered failures of vision-language-action (VLA) models in high-dimensional embodied spaces, leading to insufficient robustness evaluation. This work introduces active test generation into VLA assessment for the first time, proposing a failure-oriented dynamic testing framework that adaptively generates high-risk, diverse test cases through the synergy of diversity-guided scene exploration and an agent model trained on execution feedback. Experiments across four prominent VLA models demonstrate that this approach discovers 29.7% more failure cases on average—evidenced by, for instance, a drop in GR00T-N1.6’s success rate from 64.4% to 34.7%—significantly exposing model weaknesses and revealing richer failure modes, thereby advancing the evaluation paradigm toward active and dynamic methodologies.
为评估和提高视觉-语言模型在自主水下机器人软件中的适用性和可靠性,提出一种基于搜索的变形测试方法MetaVLM,通过最小化图像变换揭示模型错误。
研究用大型语言模型替代部分变异算子以修复Simulink-Stateflow模型中的错误,但结果表明这种方法降低了修复性能,揭示了直接集成LLMs的局限性。
This study addresses the challenge of reproducing failures in cyber-physical system (CPS) simulation testing, where non-deterministic behaviors often hinder consistent fault manifestation. To tackle this issue, the work extends delta debugging to stochastic CPS scenarios for the first time, introducing three novel delta debugging algorithms tailored for randomized environments. The proposed approach integrates statistical failure analysis, repeated execution, and environment-aware input minimization to identify a minimal triggering input while preserving fault semantics. Empirical evaluation on case studies involving elevator scheduling and autonomous mobile robots demonstrates that the method significantly enhances the stability of failure reproduction and substantially reduces debugging time, effectively mitigating execution flakiness inherent in such systems.
This work addresses the lack of reliable end-to-end testing methods for Retrieval-Augmented Generation (RAG) systems, which stems from complex interactions among their components. The authors propose RagTester, the first automated evaluation framework that introduces a coverage-guided test generation strategy. RagTester systematically synthesizes test documents, queries, and expected answers, and leverages large language models as judges to detect failure modes such as inaccurate retrieval, unsupported answers, and insufficient context utilization. Evaluated across 24 configurations with 72,000 tests, RagTester identified 21,633 faults—6.6% more than baseline methods—and outperformed baselines in 20 out of 24 configurations, substantially enhancing pre-deployment validation capabilities for RAG systems.
Existing static benchmarks struggle to effectively uncover sparse and clustered failures of vision-language-action (VLA) models in high-dimensional embodied spaces, leading to insufficient robustness evaluation. This work introduces active test generation into VLA assessment for the first time, proposing a failure-oriented dynamic testing framework that adaptively generates high-risk, diverse test cases through the synergy of diversity-guided scene exploration and an agent model trained on execution feedback. Experiments across four prominent VLA models demonstrate that this approach discovers 29.7% more failure cases on average—evidenced by, for instance, a drop in GR00T-N1.6’s success rate from 64.4% to 34.7%—significantly exposing model weaknesses and revealing richer failure modes, thereby advancing the evaluation paradigm toward active and dynamic methodologies.