Institution profile

Sangfor Technologies

Industry researchasia · cn
Official website
Research library21linked papers
Opportunities0open roles
Selected work

Representative Papers

Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production

Aug 09, 2026

Current static safety guardrails for large language models struggle to promptly counter emerging jailbreak attacks and harmful content, suffering from defensive lag. This work proposes SESG—the first self-evolving guardrail framework deployed in production—which employs a multi-agent architecture comprising monitoring, generation, validation, and routing agents. By analyzing live traffic to identify failure cases, SESG autonomously generates adversarial training data, dynamically rebalances the training distribution, and deploys updated models, enabling hour-scale threat response with minimal human intervention. Experiments show that a 1.7B-parameter guardrail autonomously mitigates new threats within 16–24 hours, requiring only ~2 hours of human effort over six evolution cycles. It significantly outperforms static and adaptive baselines (0.6B–9B parameters) across six emerging threat categories while preserving general screening capability, successfully remediating 14 of 15 novel threats within two months.

0 citationsRead paper

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

Aug 04, 2026

This work addresses the high computational cost, reliance on multi-stage training, and inclusion of control-irrelevant visual redundancy in existing World-Action Models (WAMs). To overcome these limitations, the authors propose LiLa-WAM, a lightweight architecture that jointly optimizes future state prediction and action generation within a compact latent space. Central to this approach is the introduction of language-agnostic Visual Transition Tokens (VTTs) as task representations, enabling fully end-to-end training on a single GPU. The method substantially reduces training costs while maintaining strong performance, achieving a 90.48% success rate across 50 tasks in RoboTwin 2.0, and demonstrating effectiveness on both LIBERO benchmarks and real-world robotic tasks—all trained on a single 24GB GPU.

0 citationsRead paper

Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation

Jul 30, 2026

This work addresses the critical privacy risks associated with releasing face training data, where existing methods struggle to disentangle identity information while preserving the class structure essential for recognition. To overcome this challenge, the authors propose a novel identity-disentangled and geometry-preserving face distillation framework that explicitly separates source identity semantics from proxy identity geometry. By enforcing orthogonal geometric preservation and aligning relational topologies, the method effectively eliminates linkability to original identities while retaining the hyperspherical proxy structure necessary for face recognition. Experimental results demonstrate that the proposed approach achieves a 3.94% improvement in TAR@FAR=1e-3 on the IJB-C surveillance benchmark, significantly outperforming baseline methods and offering a strong balance between privacy protection and model utility.

0 citationsRead paper

Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection

Jul 29, 2026

This work addresses the limitations of existing multimodal large language model–based methods for AI-generated image (AIGI) detection, which struggle with fine-grained visual anomaly perception and exhibit constrained generalization. To overcome these challenges, the authors propose a reliability-grounded authenticity reasoning framework that integrates three complementary perceptual capabilities: fine-grained visual details, semantic anomalies, and pixel-level discrepancies. The approach introduces two key innovations: Perception-oriented Representation Learning (PoRL), which replaces open-ended descriptive supervision with targeted perceptual guidance, and Value-aware Output Policy Distillation (VaOPD), a mechanism that prioritizes the transfer of high-value signals to strengthen the synergy between perception and reasoning. Experimental results demonstrate that the proposed method consistently achieves significant performance gains across standard, real-world, and emerging benchmarks, effectively bridging the perceptual gap while preserving original capabilities, thereby enabling robust and generalizable AIGI detection.

0 citationsRead paper
Recent publications

Latest Papers

Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production

Aug 09, 2026

Current static safety guardrails for large language models struggle to promptly counter emerging jailbreak attacks and harmful content, suffering from defensive lag. This work proposes SESG—the first self-evolving guardrail framework deployed in production—which employs a multi-agent architecture comprising monitoring, generation, validation, and routing agents. By analyzing live traffic to identify failure cases, SESG autonomously generates adversarial training data, dynamically rebalances the training distribution, and deploys updated models, enabling hour-scale threat response with minimal human intervention. Experiments show that a 1.7B-parameter guardrail autonomously mitigates new threats within 16–24 hours, requiring only ~2 hours of human effort over six evolution cycles. It significantly outperforms static and adaptive baselines (0.6B–9B parameters) across six emerging threat categories while preserving general screening capability, successfully remediating 14 of 15 novel threats within two months.

0 citationsRead paper

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

Aug 04, 2026

This work addresses the high computational cost, reliance on multi-stage training, and inclusion of control-irrelevant visual redundancy in existing World-Action Models (WAMs). To overcome these limitations, the authors propose LiLa-WAM, a lightweight architecture that jointly optimizes future state prediction and action generation within a compact latent space. Central to this approach is the introduction of language-agnostic Visual Transition Tokens (VTTs) as task representations, enabling fully end-to-end training on a single GPU. The method substantially reduces training costs while maintaining strong performance, achieving a 90.48% success rate across 50 tasks in RoboTwin 2.0, and demonstrating effectiveness on both LIBERO benchmarks and real-world robotic tasks—all trained on a single 24GB GPU.

0 citationsRead paper

Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation

Jul 30, 2026

This work addresses the critical privacy risks associated with releasing face training data, where existing methods struggle to disentangle identity information while preserving the class structure essential for recognition. To overcome this challenge, the authors propose a novel identity-disentangled and geometry-preserving face distillation framework that explicitly separates source identity semantics from proxy identity geometry. By enforcing orthogonal geometric preservation and aligning relational topologies, the method effectively eliminates linkability to original identities while retaining the hyperspherical proxy structure necessary for face recognition. Experimental results demonstrate that the proposed approach achieves a 3.94% improvement in TAR@FAR=1e-3 on the IJB-C surveillance benchmark, significantly outperforming baseline methods and offering a strong balance between privacy protection and model utility.

0 citationsRead paper

Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection

Jul 29, 2026

This work addresses the limitations of existing multimodal large language model–based methods for AI-generated image (AIGI) detection, which struggle with fine-grained visual anomaly perception and exhibit constrained generalization. To overcome these challenges, the authors propose a reliability-grounded authenticity reasoning framework that integrates three complementary perceptual capabilities: fine-grained visual details, semantic anomalies, and pixel-level discrepancies. The approach introduces two key innovations: Perception-oriented Representation Learning (PoRL), which replaces open-ended descriptive supervision with targeted perceptual guidance, and Value-aware Output Policy Distillation (VaOPD), a mechanism that prioritizes the transfer of high-value signals to strengthen the synergy between perception and reasoning. Experimental results demonstrate that the proposed method consistently achieves significant performance gains across standard, real-world, and emerging benchmarks, effectively bridging the perceptual gap while preserving original capabilities, thereby enabling robust and generalizable AIGI detection.

0 citationsRead paper