Institution profile

Education University of Hong Kong

Academic institutionasia · hk
Official website
Research library10linked papers
Opportunities0open roles
Selected work

Representative Papers

HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails

Aug 09, 2026

This work addresses the dilemma faced by current large language models, where safety alignment either distorts semantic representations through fine-tuning or incurs high inference costs. The authors propose a training-free geometric safety mechanism that freezes the pretrained encoder and maps text embeddings onto the unit hypersphere. Leveraging a precomputed library of topological anchor points, the method performs zero-shot safety judgments via Gibbs–Boltzmann free energy and a dual-timescale exponential moving average, effectively decoupling representation learning from inference. Requiring only a few fixed hyperparameters, the approach significantly enhances robustness against high-frequency perturbations and achieves state-of-the-art performance across eight benchmarks—e.g., AuthenHallu AUC = 1.0000 and HarmBench AUC = 0.9802—while offering sub-millisecond latency, zero cold-start overhead, and strong cross-lingual transferability, as demonstrated by CHIFRAUD AUC = 0.9758 on Chinese data.

0 citationsRead paper
Recent publications

Latest Papers

HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails

Aug 09, 2026

This work addresses the dilemma faced by current large language models, where safety alignment either distorts semantic representations through fine-tuning or incurs high inference costs. The authors propose a training-free geometric safety mechanism that freezes the pretrained encoder and maps text embeddings onto the unit hypersphere. Leveraging a precomputed library of topological anchor points, the method performs zero-shot safety judgments via Gibbs–Boltzmann free energy and a dual-timescale exponential moving average, effectively decoupling representation learning from inference. Requiring only a few fixed hyperparameters, the approach significantly enhances robustness against high-frequency perturbations and achieves state-of-the-art performance across eight benchmarks—e.g., AuthenHallu AUC = 1.0000 and HarmBench AUC = 0.9802—while offering sub-millisecond latency, zero cold-start overhead, and strong cross-lingual transferability, as demonstrated by CHIFRAUD AUC = 0.9758 on Chinese data.

0 citationsRead paper