Institution profile

People's Public Security University of China

Academic institutionasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

MIRA: Medical Image Reflection for Agentic Diagnosis

Aug 11, 2026

This work addresses the limitations of existing medical vision-language agents, which often invoke diagnostic tools indiscriminately, thereby introducing noisy or misleading evidence due to a lack of validation regarding tool necessity and evidential consistency. To overcome this, the authors propose a novel diagnostic framework endowed with autonomous evidence-seeking and reflective verification capabilities. The framework dynamically orchestrates image processing and web-search tools while evaluating the relevance and consistency of retrieved evidence. A two-stage training strategy is employed: first, high-quality fine-tuning trajectories are generated via tool-augmented Monte Carlo tree search; second, an online reflection-based evolutionary mechanism refines decision-making through self-correction and adaptation. Evaluated across nine medical visual reasoning benchmarks, the method achieves an average score of 64.73—outperforming Qwen3-VL-8B by 7.44 points—with a valid tool invocation rate of 73.8% and a harmful judgment rate reduced to 1.6%.

0 citationsRead paper

When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign Adapters

Feb 25, 2026

This work proposes MasqLoRA, a novel backdoor attack framework that exploits the open sharing mechanism of LoRA (Low-Rank Adaptation) modules, revealing a critical security vulnerability in parameter-efficient fine-tuning. By training a lightweight malicious LoRA adapter on only a few trigger-word–target-image pairs while keeping the base diffusion model frozen, MasqLoRA achieves highly effective (99.8% success rate) and targeted backdoor injection across modalities. The attack incurs minimal training overhead and exhibits perfect benign behavior in the absence of the trigger, rendering it highly stealthy. This study is the first to systematically demonstrate that LoRA adapters can serve as covert carriers for cross-modal backdoors, thereby exposing a new class of supply-chain security threats introduced by efficient adaptation techniques in AI models.

0 citationsRead paper

Multi-Head Spectral-Adaptive Graph Anomaly Detection

Dec 25, 2025

In graph-based anomaly detection—particularly for financial fraud—adversarial anomalies exhibit strong camouflage, and high-frequency discriminative signals are often smoothed out or lost by global graph filters. Method: We propose a spectral-fingerprint-driven dynamic graph learning framework tailored for financial fraud detection. First, we design a lightweight hypernetwork that dynamically generates multi-head Chebyshev filters per node based on its spectral fingerprint, enabling instance-adaptive spectral-domain modeling. Second, we introduce a multi-head decoupling mechanism combining Teacher–Student Contrastive (TSC) learning with Barlow Twins diversity regularization to enhance robust representation of anomaly-sensitive features. Results: Evaluated on four real-world heterogeneous graph datasets, our method consistently outperforms state-of-the-art approaches. Crucially, it preserves critical high-frequency anomaly signals under high heterogeneity, leading to improved detection accuracy and generalization.

0 citationsRead paper

A Novel Trustworthy Video Summarization Algorithm Through a Mixture of LoRA Experts

Mar 08, 2025

To address the challenges of efficient retrieval and browsing amid explosive growth in user-generated video content, existing methods (e.g., Video-LLaMA) suffer from fragmented spatiotemporal modeling, high computational overhead, and parameter redundancy. This paper proposes a lightweight and efficient spatiotemporal-cooperative video summarization framework. We introduce a novel hybrid LoRA-based mixture-of-experts architecture with dynamic expert routing; design a dual-path spatiotemporal adaptive module to jointly optimize temporal dynamics and spatial semantics; and incorporate video–language cross-modal alignment during training. Evaluated on VideoXum and ActivityNet, our method achieves state-of-the-art summarization quality while reducing model parameters by 62% and accelerating inference by 3.1× compared to prior approaches. These advances significantly enhance practicality and scalability for large-scale video understanding.

0 citationsRead paper
Recent publications

Latest Papers

MIRA: Medical Image Reflection for Agentic Diagnosis

Aug 11, 2026

This work addresses the limitations of existing medical vision-language agents, which often invoke diagnostic tools indiscriminately, thereby introducing noisy or misleading evidence due to a lack of validation regarding tool necessity and evidential consistency. To overcome this, the authors propose a novel diagnostic framework endowed with autonomous evidence-seeking and reflective verification capabilities. The framework dynamically orchestrates image processing and web-search tools while evaluating the relevance and consistency of retrieved evidence. A two-stage training strategy is employed: first, high-quality fine-tuning trajectories are generated via tool-augmented Monte Carlo tree search; second, an online reflection-based evolutionary mechanism refines decision-making through self-correction and adaptation. Evaluated across nine medical visual reasoning benchmarks, the method achieves an average score of 64.73—outperforming Qwen3-VL-8B by 7.44 points—with a valid tool invocation rate of 73.8% and a harmful judgment rate reduced to 1.6%.

0 citationsRead paper

When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign Adapters

Feb 25, 2026

This work proposes MasqLoRA, a novel backdoor attack framework that exploits the open sharing mechanism of LoRA (Low-Rank Adaptation) modules, revealing a critical security vulnerability in parameter-efficient fine-tuning. By training a lightweight malicious LoRA adapter on only a few trigger-word–target-image pairs while keeping the base diffusion model frozen, MasqLoRA achieves highly effective (99.8% success rate) and targeted backdoor injection across modalities. The attack incurs minimal training overhead and exhibits perfect benign behavior in the absence of the trigger, rendering it highly stealthy. This study is the first to systematically demonstrate that LoRA adapters can serve as covert carriers for cross-modal backdoors, thereby exposing a new class of supply-chain security threats introduced by efficient adaptation techniques in AI models.

0 citationsRead paper

Multi-Head Spectral-Adaptive Graph Anomaly Detection

Dec 25, 2025

In graph-based anomaly detection—particularly for financial fraud—adversarial anomalies exhibit strong camouflage, and high-frequency discriminative signals are often smoothed out or lost by global graph filters. Method: We propose a spectral-fingerprint-driven dynamic graph learning framework tailored for financial fraud detection. First, we design a lightweight hypernetwork that dynamically generates multi-head Chebyshev filters per node based on its spectral fingerprint, enabling instance-adaptive spectral-domain modeling. Second, we introduce a multi-head decoupling mechanism combining Teacher–Student Contrastive (TSC) learning with Barlow Twins diversity regularization to enhance robust representation of anomaly-sensitive features. Results: Evaluated on four real-world heterogeneous graph datasets, our method consistently outperforms state-of-the-art approaches. Crucially, it preserves critical high-frequency anomaly signals under high heterogeneity, leading to improved detection accuracy and generalization.

0 citationsRead paper

A Novel Trustworthy Video Summarization Algorithm Through a Mixture of LoRA Experts

Mar 08, 2025

To address the challenges of efficient retrieval and browsing amid explosive growth in user-generated video content, existing methods (e.g., Video-LLaMA) suffer from fragmented spatiotemporal modeling, high computational overhead, and parameter redundancy. This paper proposes a lightweight and efficient spatiotemporal-cooperative video summarization framework. We introduce a novel hybrid LoRA-based mixture-of-experts architecture with dynamic expert routing; design a dual-path spatiotemporal adaptive module to jointly optimize temporal dynamics and spatial semantics; and incorporate video–language cross-modal alignment during training. Evaluated on VideoXum and ActivityNet, our method achieves state-of-the-art summarization quality while reducing model parameters by 62% and accelerating inference by 3.1× compared to prior approaches. These advances significantly enhance practicality and scalability for large-scale video understanding.

0 citationsRead paper