Evaluating Automated Testing on an Open-Source Web Application Using Cypress
本文使用Cypress框架对开源Web应用进行端到端自动化测试评估,通过27个测试案例分析了执行速度、可靠性和可维护性,证明了Cypress的有效性。
本文使用Cypress框架对开源Web应用进行端到端自动化测试评估,通过27个测试案例分析了执行速度、可靠性和可维护性,证明了Cypress的有效性。
This work addresses the challenge of fine-grained retrieval of anomalous pedestrian behaviors from large-scale image collections based on natural language descriptions. To tackle this problem, we propose a robust cross-modal retrieval framework that integrates heterogeneous vision-language embeddings through score alignment and iterative ensemble strategies to effectively fuse multi-model representations. Furthermore, we introduce a discrepancy-aware re-ranking mechanism to handle semantically ambiguous queries. The proposed approach significantly enhances the robustness and accuracy of cross-modal matching in complex scenarios, achieving state-of-the-art performance on the PAB benchmark with 90.92% mAP, 85.13% Recall@1, 97.72% Recall@5, and 98.68% Recall@10, thereby demonstrating its effectiveness.
This work addresses the challenge of fine-grained matching in text-based pedestrian anomaly retrieval under synthetic-to-real (Sim2Real) scenarios. To this end, the authors propose an anchor-constrained coarse-to-fine retrieval framework that leverages multi-facet semantic decomposition and calibrated fusion. The approach integrates a heterogeneous vision-language retriever, a Qwen3-based reranker, and an anomaly-aware cloze-style verification module, complemented by an uncertainty-gated consensus mechanism operating over a small candidate pool to enable efficient fine-grained semantic reasoning. Innovatively, semantic facets serve as anchor constraints to jointly optimize recall and computational efficiency. Evaluated on the PAB benchmark, the method achieves 95.41% mAP@10, 94.44% R@1, and 99.09% R@5, significantly outperforming existing single-backbone models.
Long-term traffic scene prediction often suffers from temporal ghosting, geometric drift, and motion inconsistency, leading to unstable static structures and temporally incoherent outputs. To address this, this work proposes a training-free inference framework that enhances the outputs of pretrained video diffusion models through geometry-aware refinement during inference. The approach uniquely integrates multi-frame depth-layered rendering with vision-language model–guided view-conditioned routing, leveraging a frozen visual language model to achieve geometric stability and cross-view generalization without fine-tuning. Evaluated on the AI City Challenge Track 5, the method achieves state-of-the-art performance, significantly improving fidelity of static structures and consistency of low-level details.
This work addresses limitations in existing compound claim decomposition methods for automated fact-checking, which rely on lexical overlap metrics like Jaccard similarity and thus struggle to accurately assess the semantic fidelity of paraphrased atomic claims, while also lacking theoretical guarantees on the termination of repair processes. To overcome these issues, the authors propose CREDENCE, a framework that replaces lexical overlap with cosine similarity based on BGE-large embeddings to enable semantic-aware decomposition and self-repair. They formally prove, for the first time, the convergence of a hybrid repair pipeline combining symbolic rules and large language models. The study introduces Semantic-F1, a new evaluation metric validated across three cross-domain benchmarks—social media, encyclopedic texts, and news—demonstrating 15–32 percentage point improvements over Jaccard-F1, EPR scores of 0.94–1.00, and a 47%–100% reduction in atomicity violations via rule-based repair without compromising semantic fidelity.
本文使用Cypress框架对开源Web应用进行端到端自动化测试评估,通过27个测试案例分析了执行速度、可靠性和可维护性,证明了Cypress的有效性。
This work addresses the challenge of fine-grained retrieval of anomalous pedestrian behaviors from large-scale image collections based on natural language descriptions. To tackle this problem, we propose a robust cross-modal retrieval framework that integrates heterogeneous vision-language embeddings through score alignment and iterative ensemble strategies to effectively fuse multi-model representations. Furthermore, we introduce a discrepancy-aware re-ranking mechanism to handle semantically ambiguous queries. The proposed approach significantly enhances the robustness and accuracy of cross-modal matching in complex scenarios, achieving state-of-the-art performance on the PAB benchmark with 90.92% mAP, 85.13% Recall@1, 97.72% Recall@5, and 98.68% Recall@10, thereby demonstrating its effectiveness.
This work addresses the challenge of fine-grained matching in text-based pedestrian anomaly retrieval under synthetic-to-real (Sim2Real) scenarios. To this end, the authors propose an anchor-constrained coarse-to-fine retrieval framework that leverages multi-facet semantic decomposition and calibrated fusion. The approach integrates a heterogeneous vision-language retriever, a Qwen3-based reranker, and an anomaly-aware cloze-style verification module, complemented by an uncertainty-gated consensus mechanism operating over a small candidate pool to enable efficient fine-grained semantic reasoning. Innovatively, semantic facets serve as anchor constraints to jointly optimize recall and computational efficiency. Evaluated on the PAB benchmark, the method achieves 95.41% mAP@10, 94.44% R@1, and 99.09% R@5, significantly outperforming existing single-backbone models.
Long-term traffic scene prediction often suffers from temporal ghosting, geometric drift, and motion inconsistency, leading to unstable static structures and temporally incoherent outputs. To address this, this work proposes a training-free inference framework that enhances the outputs of pretrained video diffusion models through geometry-aware refinement during inference. The approach uniquely integrates multi-frame depth-layered rendering with vision-language model–guided view-conditioned routing, leveraging a frozen visual language model to achieve geometric stability and cross-view generalization without fine-tuning. Evaluated on the AI City Challenge Track 5, the method achieves state-of-the-art performance, significantly improving fidelity of static structures and consistency of low-level details.
This work addresses limitations in existing compound claim decomposition methods for automated fact-checking, which rely on lexical overlap metrics like Jaccard similarity and thus struggle to accurately assess the semantic fidelity of paraphrased atomic claims, while also lacking theoretical guarantees on the termination of repair processes. To overcome these issues, the authors propose CREDENCE, a framework that replaces lexical overlap with cosine similarity based on BGE-large embeddings to enable semantic-aware decomposition and self-repair. They formally prove, for the first time, the convergence of a hybrid repair pipeline combining symbolic rules and large language models. The study introduces Semantic-F1, a new evaluation metric validated across three cross-domain benchmarks—social media, encyclopedic texts, and news—demonstrating 15–32 percentage point improvements over Jaccard-F1, EPR scores of 0.94–1.00, and a 47%–100% reduction in atomicity violations via rule-based repair without compromising semantic fidelity.