Institution profile

Shanghai Polytechnic University

Academic institutionasia · cn
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

OSNIP: Breaking the Privacy-Utility-Efficiency Trilemma in LLM Inference via Obfuscated Semantic Null Space

Jan 30, 2026

This work addresses the trilemma in large language model (LLM) inference—balancing privacy preservation, model utility, and computational efficiency—by proposing a lightweight client-side encryption framework. It introduces, for the first time, a formal definition of the “obfuscated semantic nullspace,” into which input embeddings are projected to achieve privacy without requiring post-processing. The method integrates user-key-driven random perturbation trajectories with geometry-aware noise injection in the latent space, enabling personalized privacy guarantees while maintaining efficient inference. Evaluated across twelve generative and classification benchmarks, the approach achieves state-of-the-art performance, substantially reducing attack success rates while preserving high model utility under stringent security constraints.

0 citationsRead paper

Parasite: A Steganography-based Backdoor Attack Framework for Diffusion Models

Apr 08, 2025

Diffusion models for image-to-image (I2I) translation are vulnerable to backdoor attacks, yet existing approaches primarily target noise- or text-to-image generation and rely on explicit, static triggers—limiting stealth and adaptability. This paper introduces the first steganography-driven backdoor attack framework tailored to I2I tasks: it embeds target content as a covert trigger directly into the input image via steganographic encoding, reverse-sampling control, and joint embedding-reconstruction optimization, enabling content-aware, dynamic backdoor activation. Our method overcomes fundamental limitations of conventional trigger design, achieving zero detection rate under mainstream defenses—including fine-tuning, input purification, and anomaly detection. Ablation studies confirm that steganographic embedding strength enables precise, controllable trade-offs between attack success rate and output image fidelity, without perceptible degradation. The framework is model-agnostic, requires no access to model weights or training data, and preserves functional integrity during benign inference.

0 citationsRead paper

Gungnir: Exploiting Stylistic Features in Images for Backdoor Attacks on Diffusion Models

Feb 28, 2025

Diffusion models (DMs) are vulnerable to backdoor attacks, yet existing approaches rely on explicit, low-dimensional triggers that are readily detectable by mainstream defenses. This paper proposes the first implicit-style-feature-based backdoor attack paradigm: without modifying training data, it disentangles and injects stylistic features directly from input images as covert triggers, enabling end-to-end implicit backdoor injection in image-to-image translation tasks. Our method integrates Reconstruction-Adversarial Noise (RAN), Short-Term Trajectory Retention (STTR), and a novel style-feature disentanglement/injection mechanism. Extensive experiments demonstrate that the attack achieves a 0% detection rate across multiple state-of-the-art DM defense frameworks—including both trigger-detection and inverse-trigger-based methods—thereby fully evading existing defenses. It significantly enhances both stealthiness and robustness against defensive mitigation strategies.

0 citationsRead paper
Recent publications

Latest Papers

OSNIP: Breaking the Privacy-Utility-Efficiency Trilemma in LLM Inference via Obfuscated Semantic Null Space

Jan 30, 2026

This work addresses the trilemma in large language model (LLM) inference—balancing privacy preservation, model utility, and computational efficiency—by proposing a lightweight client-side encryption framework. It introduces, for the first time, a formal definition of the “obfuscated semantic nullspace,” into which input embeddings are projected to achieve privacy without requiring post-processing. The method integrates user-key-driven random perturbation trajectories with geometry-aware noise injection in the latent space, enabling personalized privacy guarantees while maintaining efficient inference. Evaluated across twelve generative and classification benchmarks, the approach achieves state-of-the-art performance, substantially reducing attack success rates while preserving high model utility under stringent security constraints.

0 citationsRead paper

Parasite: A Steganography-based Backdoor Attack Framework for Diffusion Models

Apr 08, 2025

Diffusion models for image-to-image (I2I) translation are vulnerable to backdoor attacks, yet existing approaches primarily target noise- or text-to-image generation and rely on explicit, static triggers—limiting stealth and adaptability. This paper introduces the first steganography-driven backdoor attack framework tailored to I2I tasks: it embeds target content as a covert trigger directly into the input image via steganographic encoding, reverse-sampling control, and joint embedding-reconstruction optimization, enabling content-aware, dynamic backdoor activation. Our method overcomes fundamental limitations of conventional trigger design, achieving zero detection rate under mainstream defenses—including fine-tuning, input purification, and anomaly detection. Ablation studies confirm that steganographic embedding strength enables precise, controllable trade-offs between attack success rate and output image fidelity, without perceptible degradation. The framework is model-agnostic, requires no access to model weights or training data, and preserves functional integrity during benign inference.

0 citationsRead paper

Gungnir: Exploiting Stylistic Features in Images for Backdoor Attacks on Diffusion Models

Feb 28, 2025

Diffusion models (DMs) are vulnerable to backdoor attacks, yet existing approaches rely on explicit, low-dimensional triggers that are readily detectable by mainstream defenses. This paper proposes the first implicit-style-feature-based backdoor attack paradigm: without modifying training data, it disentangles and injects stylistic features directly from input images as covert triggers, enabling end-to-end implicit backdoor injection in image-to-image translation tasks. Our method integrates Reconstruction-Adversarial Noise (RAN), Short-Term Trajectory Retention (STTR), and a novel style-feature disentanglement/injection mechanism. Extensive experiments demonstrate that the attack achieves a 0% detection rate across multiple state-of-the-art DM defense frameworks—including both trigger-detection and inverse-trigger-based methods—thereby fully evading existing defenses. It significantly enhances both stealthiness and robustness against defensive mitigation strategies.

0 citationsRead paper