Institution profile

The Chinese University of Hong Kong, Shen zhen

Academic institutionasia · hk
Official website
Research library117linked papers
Opportunities0open roles
Selected work

Representative Papers

Error-Aware Reverse Auction Mechanism for Large Language Model Routing

Aug 12, 2026

This work addresses the limitations of centralized prediction-based routing in large language models (LLMs), which often leads to misaligned information risks and scalability bottlenecks. To overcome these issues, the paper introduces, for the first time, a reverse auction mechanism into LLM routing and proposes the error-aware EA-RAM framework. In this framework, model providers autonomously bid their success rates and costs, while explicit modeling of dual sources of noise—arising from both prediction and evaluation—enables robust handling of uncertainty. The mechanism is shown to be Bayesian incentive-compatible and individually rational, with a provable upper bound on social welfare loss. Empirical results demonstrate that EA-RAM consistently outperforms centralized baselines across both simulated and real-world benchmarks, maintaining robustness under dual-error conditions and significantly advancing the cost-performance Pareto frontier.

0 citationsRead paper

ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval

Aug 12, 2026

Existing static memory retrieval mechanisms struggle to meet the demands of heterogeneous queries requiring diverse evidence construction. This work proposes ERSkill, a novel framework that models memory retrieval as a self-evolving and composable set of skills. ERSkill introduces a skill-routing co-evolution mechanism and a dual-frontier architecture to enable safe and efficient capability expansion. By integrating structured memory storage, experience-based Trie path recording, and large language model–driven dynamic routing, the framework substantially enhances retrieval customization. Evaluated across multiple agent memory benchmarks, ERSkill significantly outperforms current methods, achieving relative improvements of 31.3% and 28.1% in composite metrics on Qwen3 and GPT-5.4-nano, respectively.

0 citationsRead paper

An AoI-oriented Time-Frequency Distributed Access Mechanism in Wireless Sensor Networks with Spectrum Division

Aug 11, 2026

This work addresses the challenge of jointly achieving low communication overhead and high information freshness in large-scale randomly activated wireless sensor networks under spectrum partitioning. To this end, the paper proposes a deterministic time–frequency distributed access mechanism (D-TFDA) oriented toward minimizing the Age of Information (AoI). D-TFDA integrates centralized configuration with distributed execution, leveraging a token-based periodic time–frequency structure to provide conflict-free and predictable transmission opportunities. By uncovering structural properties of token allocation, the authors identify AoI-equivalent token clusters, transforming the optimal allocation problem into a linear program that drastically reduces the search space. A one-dimensional discrete-time Markov chain models the system’s steady-state behavior to analyze the long-term average AoI, enabling the design of a low-complexity auction-inspired heuristic algorithm. Simulations demonstrate that D-TFDA significantly outperforms optimized random-access baselines by reducing average AoI, eliminating collisions, and effectively exploiting heterogeneity in sensor-to-resource reliability.

0 citationsRead paper

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Aug 11, 2026

This work addresses the lack of coordinated visual expression in existing full-modal dialogue systems, which often generate speech responses without semantically aligned visual output, resulting in audio-visual misalignment. To bridge this gap, the authors propose a unified framework that jointly generates semantically consistent text, personalized speech, and conditionally synthesized video from multimodal queries, reference images, and audio inputs. Key innovations include a Visual Thought Planning (VTP) module for structured modeling of scenes, emotions, and actions; a shared multi-codebook acoustic unit serving as a unified interface to align speech and video generation; and a block-causal streaming student model integrated with a prefix-streaming mechanism to enable efficient incremental synthesis. Implemented in an end-to-end four-GPU pipeline, the system achieves 1.293× real-time inference at 400×720 or 720×400 resolution, striking a practical balance between generation quality and computational efficiency.

0 citationsRead paper

A Mechanistic Diagnostic of Rank Collapse in Post-Norm Decoder Transformers

Aug 10, 2026

This work addresses the instability and vanishing gradients in Post-Norm decoder-only Transformers caused by rank collapse, a phenomenon whose underlying mechanisms remain poorly understood. The authors model token similarity as a scalar state variable and, for the first time, disentangle the roles of causal attention—amplifying representation similarity during forward propagation—and the lack of corrective capacity during backward propagation in driving rank collapse, analyzing both initialization and training phases. By integrating theoretical analysis, modeling the backward dynamics of RMSNorm, and incorporating the branch effect of SwiGLU alongside a prefix-averaging operator approximation, they characterize the behavioral signatures of collapsing networks. Experiments on 48-layer decoders confirm that similarity grows at initialization, gradients contract upon collapse, and loss converges toward the prediction based on token frequency distribution.

0 citationsRead paper
Recent publications

Latest Papers

Error-Aware Reverse Auction Mechanism for Large Language Model Routing

Aug 12, 2026

This work addresses the limitations of centralized prediction-based routing in large language models (LLMs), which often leads to misaligned information risks and scalability bottlenecks. To overcome these issues, the paper introduces, for the first time, a reverse auction mechanism into LLM routing and proposes the error-aware EA-RAM framework. In this framework, model providers autonomously bid their success rates and costs, while explicit modeling of dual sources of noise—arising from both prediction and evaluation—enables robust handling of uncertainty. The mechanism is shown to be Bayesian incentive-compatible and individually rational, with a provable upper bound on social welfare loss. Empirical results demonstrate that EA-RAM consistently outperforms centralized baselines across both simulated and real-world benchmarks, maintaining robustness under dual-error conditions and significantly advancing the cost-performance Pareto frontier.

0 citationsRead paper

ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval

Aug 12, 2026

Existing static memory retrieval mechanisms struggle to meet the demands of heterogeneous queries requiring diverse evidence construction. This work proposes ERSkill, a novel framework that models memory retrieval as a self-evolving and composable set of skills. ERSkill introduces a skill-routing co-evolution mechanism and a dual-frontier architecture to enable safe and efficient capability expansion. By integrating structured memory storage, experience-based Trie path recording, and large language model–driven dynamic routing, the framework substantially enhances retrieval customization. Evaluated across multiple agent memory benchmarks, ERSkill significantly outperforms current methods, achieving relative improvements of 31.3% and 28.1% in composite metrics on Qwen3 and GPT-5.4-nano, respectively.

0 citationsRead paper

An AoI-oriented Time-Frequency Distributed Access Mechanism in Wireless Sensor Networks with Spectrum Division

Aug 11, 2026

This work addresses the challenge of jointly achieving low communication overhead and high information freshness in large-scale randomly activated wireless sensor networks under spectrum partitioning. To this end, the paper proposes a deterministic time–frequency distributed access mechanism (D-TFDA) oriented toward minimizing the Age of Information (AoI). D-TFDA integrates centralized configuration with distributed execution, leveraging a token-based periodic time–frequency structure to provide conflict-free and predictable transmission opportunities. By uncovering structural properties of token allocation, the authors identify AoI-equivalent token clusters, transforming the optimal allocation problem into a linear program that drastically reduces the search space. A one-dimensional discrete-time Markov chain models the system’s steady-state behavior to analyze the long-term average AoI, enabling the design of a low-complexity auction-inspired heuristic algorithm. Simulations demonstrate that D-TFDA significantly outperforms optimized random-access baselines by reducing average AoI, eliminating collisions, and effectively exploiting heterogeneity in sensor-to-resource reliability.

0 citationsRead paper

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Aug 11, 2026

This work addresses the lack of coordinated visual expression in existing full-modal dialogue systems, which often generate speech responses without semantically aligned visual output, resulting in audio-visual misalignment. To bridge this gap, the authors propose a unified framework that jointly generates semantically consistent text, personalized speech, and conditionally synthesized video from multimodal queries, reference images, and audio inputs. Key innovations include a Visual Thought Planning (VTP) module for structured modeling of scenes, emotions, and actions; a shared multi-codebook acoustic unit serving as a unified interface to align speech and video generation; and a block-causal streaming student model integrated with a prefix-streaming mechanism to enable efficient incremental synthesis. Implemented in an end-to-end four-GPU pipeline, the system achieves 1.293× real-time inference at 400×720 or 720×400 resolution, striking a practical balance between generation quality and computational efficiency.

0 citationsRead paper

A Mechanistic Diagnostic of Rank Collapse in Post-Norm Decoder Transformers

Aug 10, 2026

This work addresses the instability and vanishing gradients in Post-Norm decoder-only Transformers caused by rank collapse, a phenomenon whose underlying mechanisms remain poorly understood. The authors model token similarity as a scalar state variable and, for the first time, disentangle the roles of causal attention—amplifying representation similarity during forward propagation—and the lack of corrective capacity during backward propagation in driving rank collapse, analyzing both initialization and training phases. By integrating theoretical analysis, modeling the backward dynamics of RMSNorm, and incorporating the branch effect of SwiGLU alongside a prefix-averaging operator approximation, they characterize the behavioral signatures of collapsing networks. Experiments on 48-layer decoders confirm that similarity grows at initialization, gradients contract upon collapse, and loss converges toward the prediction based on token frequency distribution.

0 citationsRead paper