RAFM-SER++: A Lightweight Multimodal Emotion Recognition Framework for Real-Time Behavioral Monitoring in Surveillance Systems
为解决计算成本高限制实时监控系统部署的问题,提出RAFM-SER++框架,采用轻量级单向残差注意力机制融合多模态信息,实现高效情绪识别。
为解决计算成本高限制实时监控系统部署的问题,提出RAFM-SER++框架,采用轻量级单向残差注意力机制融合多模态信息,实现高效情绪识别。
本文提出APIPilot框架,通过执行验证LLM推断的依赖关系并基于响应调整,生成有效的REST API测试序列,提高测试覆盖率和成功率。
This work addresses the challenge of fine-grained retrieval of anomalous pedestrian behaviors from large-scale image collections based on natural language descriptions. To tackle this problem, we propose a robust cross-modal retrieval framework that integrates heterogeneous vision-language embeddings through score alignment and iterative ensemble strategies to effectively fuse multi-model representations. Furthermore, we introduce a discrepancy-aware re-ranking mechanism to handle semantically ambiguous queries. The proposed approach significantly enhances the robustness and accuracy of cross-modal matching in complex scenarios, achieving state-of-the-art performance on the PAB benchmark with 90.92% mAP, 85.13% Recall@1, 97.72% Recall@5, and 98.68% Recall@10, thereby demonstrating its effectiveness.
This work addresses the challenges of scarce annotations and unreliable pseudo-labels in ambiguous gland regions for semi-supervised histopathology image segmentation. To this end, we propose a confidence-guided diffusion refinement mechanism built upon the Mean Teacher framework. In regions where the teacher model exhibits low prediction confidence, a conditional diffusion model is introduced to perform structure-aware refinement, while high-confidence predictions are leveraged to formulate a weighted consistency loss for training the student model. The proposed approach substantially enhances pseudo-label quality, achieving mDice scores of 88.09%/89.83% on the GlaS dataset and 89.19%/90.29% on the CRAG dataset using only 10% and 20% labeled data, respectively. Notably, the diffusion refinement module alone contributes a +6.36% mDice improvement, consistently outperforming current state-of-the-art methods.
This work addresses the challenge of fine-grained matching in text-based pedestrian anomaly retrieval under synthetic-to-real (Sim2Real) scenarios. To this end, the authors propose an anchor-constrained coarse-to-fine retrieval framework that leverages multi-facet semantic decomposition and calibrated fusion. The approach integrates a heterogeneous vision-language retriever, a Qwen3-based reranker, and an anomaly-aware cloze-style verification module, complemented by an uncertainty-gated consensus mechanism operating over a small candidate pool to enable efficient fine-grained semantic reasoning. Innovatively, semantic facets serve as anchor constraints to jointly optimize recall and computational efficiency. Evaluated on the PAB benchmark, the method achieves 95.41% mAP@10, 94.44% R@1, and 99.09% R@5, significantly outperforming existing single-backbone models.
为解决计算成本高限制实时监控系统部署的问题,提出RAFM-SER++框架,采用轻量级单向残差注意力机制融合多模态信息,实现高效情绪识别。
本文提出APIPilot框架,通过执行验证LLM推断的依赖关系并基于响应调整,生成有效的REST API测试序列,提高测试覆盖率和成功率。
This work addresses the challenge of fine-grained retrieval of anomalous pedestrian behaviors from large-scale image collections based on natural language descriptions. To tackle this problem, we propose a robust cross-modal retrieval framework that integrates heterogeneous vision-language embeddings through score alignment and iterative ensemble strategies to effectively fuse multi-model representations. Furthermore, we introduce a discrepancy-aware re-ranking mechanism to handle semantically ambiguous queries. The proposed approach significantly enhances the robustness and accuracy of cross-modal matching in complex scenarios, achieving state-of-the-art performance on the PAB benchmark with 90.92% mAP, 85.13% Recall@1, 97.72% Recall@5, and 98.68% Recall@10, thereby demonstrating its effectiveness.
This work addresses the challenges of scarce annotations and unreliable pseudo-labels in ambiguous gland regions for semi-supervised histopathology image segmentation. To this end, we propose a confidence-guided diffusion refinement mechanism built upon the Mean Teacher framework. In regions where the teacher model exhibits low prediction confidence, a conditional diffusion model is introduced to perform structure-aware refinement, while high-confidence predictions are leveraged to formulate a weighted consistency loss for training the student model. The proposed approach substantially enhances pseudo-label quality, achieving mDice scores of 88.09%/89.83% on the GlaS dataset and 89.19%/90.29% on the CRAG dataset using only 10% and 20% labeled data, respectively. Notably, the diffusion refinement module alone contributes a +6.36% mDice improvement, consistently outperforming current state-of-the-art methods.
This work addresses the challenge of fine-grained matching in text-based pedestrian anomaly retrieval under synthetic-to-real (Sim2Real) scenarios. To this end, the authors propose an anchor-constrained coarse-to-fine retrieval framework that leverages multi-facet semantic decomposition and calibrated fusion. The approach integrates a heterogeneous vision-language retriever, a Qwen3-based reranker, and an anomaly-aware cloze-style verification module, complemented by an uncertainty-gated consensus mechanism operating over a small candidate pool to enable efficient fine-grained semantic reasoning. Innovatively, semantic facets serve as anchor constraints to jointly optimize recall and computational efficiency. Evaluated on the PAB benchmark, the method achieves 95.41% mAP@10, 94.44% R@1, and 99.09% R@5, significantly outperforming existing single-backbone models.