Institution profile

South-Central Minzu University

Academic institutionasia · cn
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

TRCGL-Net: A Long-Tailed Multi-Label Chest X-Ray Classification Framework with Generative Data Augmentation and Label Co-Occurrence Modeling

Jul 01, 2026

This work addresses the performance degradation in rare disease recognition caused by extreme long-tailed distributions in multi-label chest X-ray classification. To mitigate this challenge, the authors propose a novel framework that synergistically integrates text-guided generation with structured modeling. Specifically, a text-conditioned diffusion model synthesizes semantically consistent samples for tail classes, while channel re-weighting and a class-aware attention mechanism enhance lesion-related features. Furthermore, a label co-occurrence–based graph convolutional network facilitates inter-class information propagation. This approach represents the first unified integration of generative augmentation, feature recalibration, and graph-structured modeling to effectively alleviate class imbalance. Evaluated on the PadChest dataset, the method achieves a mean average precision (mAP) of 0.4904 on tail classes, an overall mAP of 0.4408, and a mean area under the ROC curve (mAUC) of 0.8989, outperforming current state-of-the-art methods.

0 citationsRead paper

Beyond Wireless Security: Covert Communications in Large Language Model-enabled Edge Networks

Jun 29, 2026

This work addresses the security challenges in LLM-enabled edge networks, where frequent wireless interactions and electromagnetic emissions render systems vulnerable to eavesdropping, interference, and prompt injection attacks, while existing defenses incur prohibitive overhead. To overcome this limitation, the paper introduces— for the first time—a lightweight security framework that synergistically integrates covert communication principles with computation. By jointly incorporating electromagnetic leakage suppression, low-overhead encryption, and secure scheduling strategies, the proposed approach simultaneously preserves the privacy of LLM tasks and significantly enhances execution efficiency. Experimental results demonstrate that under stringent security constraints, the method effectively reduces overall latency, achieving a co-optimized balance between security and performance and thereby transcending conventional high-overhead protection paradigms.

0 citationsRead paper

AgenticOS: An Intent-Oriented Secure Operating System Architecture for Autonomous AI Agents

Jun 19, 2026

This work addresses the structural security risks posed by large language model–driven autonomous agents to traditional operating systems, whose “resource exposure plus permission check” model proves inadequate—once compromised, attackers can abuse low-level resources to perform privilege-escalated operations. To mitigate this, the paper proposes AgenticOS, an intent-centric secure operating system architecture that treats structured agent intents as the entry point for system calls. The kernel synthesizes a least-privilege execution environment and enforces mandatory mediation, end-to-end auditing, and information flow control. Built upon a four-layer design—comprising the Ghost Kernel, Logic Shutter, Agent Capsule, and Semantic Boundary Gateway—and leveraging an Intent ABI, Manifest-Only Runtime, and Weaver capability mechanism, AgenticOS redefines the OS role from resource manager to intent filter, enabling semantic-level security governance of AI behaviors and substantially reducing the risk of resource misuse following agent hijacking.

0 citationsRead paper

CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval

May 01, 2026

This work addresses the challenge of users struggling to precisely articulate choreographic intent using natural language amid the surge of online dance content. To this end, we propose CustomDancer, a multimodal dance retrieval framework that jointly models textual semantics, musical rhythm, and full-body motion dynamics through an end-to-end cross-modal alignment architecture comprising a CLIP-based text encoder, dedicated music and motion encoders, and a fusion module. We also introduce TD-Data, the first large-scale, expert-annotated text-dance aligned dataset, enabling systematic evaluation of such systems. On TD-Data, CustomDancer achieves a Recall@1 of 10.23%, substantially outperforming existing methods, and user studies further confirm its superior recommendation quality.

0 citationsRead paper

MGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise Separable Module for Speaker Verification

May 06, 2025

Traditional speaker verification models (e.g., TDNN, ECAPA-TDNN) overly rely on long-term contextual information while neglecting fine-grained speaker characteristics, leading to limited discriminative power. To address this, we propose the Multi-granularity Feature Fusion TDNN (M-TDNN). Our method introduces three key innovations: (1) a 2D depthwise separable convolutional frontend to enhance local time-frequency modeling; (2) the first integration of phoneme-level feature pooling with TDNN, explicitly capturing phoneme-scale discriminative cues; and (3) a triple-domain fusion mechanism combining time-frequency, contextual, and fine-grained features. Evaluated on VoxCeleb1, M-TDNN achieves state-of-the-art performance—significantly reducing EER—while requiring fewer parameters and lower computational cost than standard TDNN and ECAPA-TDNN baselines.

0 citationsRead paper
Recent publications

Latest Papers

TRCGL-Net: A Long-Tailed Multi-Label Chest X-Ray Classification Framework with Generative Data Augmentation and Label Co-Occurrence Modeling

Jul 01, 2026

This work addresses the performance degradation in rare disease recognition caused by extreme long-tailed distributions in multi-label chest X-ray classification. To mitigate this challenge, the authors propose a novel framework that synergistically integrates text-guided generation with structured modeling. Specifically, a text-conditioned diffusion model synthesizes semantically consistent samples for tail classes, while channel re-weighting and a class-aware attention mechanism enhance lesion-related features. Furthermore, a label co-occurrence–based graph convolutional network facilitates inter-class information propagation. This approach represents the first unified integration of generative augmentation, feature recalibration, and graph-structured modeling to effectively alleviate class imbalance. Evaluated on the PadChest dataset, the method achieves a mean average precision (mAP) of 0.4904 on tail classes, an overall mAP of 0.4408, and a mean area under the ROC curve (mAUC) of 0.8989, outperforming current state-of-the-art methods.

0 citationsRead paper

Beyond Wireless Security: Covert Communications in Large Language Model-enabled Edge Networks

Jun 29, 2026

This work addresses the security challenges in LLM-enabled edge networks, where frequent wireless interactions and electromagnetic emissions render systems vulnerable to eavesdropping, interference, and prompt injection attacks, while existing defenses incur prohibitive overhead. To overcome this limitation, the paper introduces— for the first time—a lightweight security framework that synergistically integrates covert communication principles with computation. By jointly incorporating electromagnetic leakage suppression, low-overhead encryption, and secure scheduling strategies, the proposed approach simultaneously preserves the privacy of LLM tasks and significantly enhances execution efficiency. Experimental results demonstrate that under stringent security constraints, the method effectively reduces overall latency, achieving a co-optimized balance between security and performance and thereby transcending conventional high-overhead protection paradigms.

0 citationsRead paper

AgenticOS: An Intent-Oriented Secure Operating System Architecture for Autonomous AI Agents

Jun 19, 2026

This work addresses the structural security risks posed by large language model–driven autonomous agents to traditional operating systems, whose “resource exposure plus permission check” model proves inadequate—once compromised, attackers can abuse low-level resources to perform privilege-escalated operations. To mitigate this, the paper proposes AgenticOS, an intent-centric secure operating system architecture that treats structured agent intents as the entry point for system calls. The kernel synthesizes a least-privilege execution environment and enforces mandatory mediation, end-to-end auditing, and information flow control. Built upon a four-layer design—comprising the Ghost Kernel, Logic Shutter, Agent Capsule, and Semantic Boundary Gateway—and leveraging an Intent ABI, Manifest-Only Runtime, and Weaver capability mechanism, AgenticOS redefines the OS role from resource manager to intent filter, enabling semantic-level security governance of AI behaviors and substantially reducing the risk of resource misuse following agent hijacking.

0 citationsRead paper

CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval

May 01, 2026

This work addresses the challenge of users struggling to precisely articulate choreographic intent using natural language amid the surge of online dance content. To this end, we propose CustomDancer, a multimodal dance retrieval framework that jointly models textual semantics, musical rhythm, and full-body motion dynamics through an end-to-end cross-modal alignment architecture comprising a CLIP-based text encoder, dedicated music and motion encoders, and a fusion module. We also introduce TD-Data, the first large-scale, expert-annotated text-dance aligned dataset, enabling systematic evaluation of such systems. On TD-Data, CustomDancer achieves a Recall@1 of 10.23%, substantially outperforming existing methods, and user studies further confirm its superior recommendation quality.

0 citationsRead paper

MGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise Separable Module for Speaker Verification

May 06, 2025

Traditional speaker verification models (e.g., TDNN, ECAPA-TDNN) overly rely on long-term contextual information while neglecting fine-grained speaker characteristics, leading to limited discriminative power. To address this, we propose the Multi-granularity Feature Fusion TDNN (M-TDNN). Our method introduces three key innovations: (1) a 2D depthwise separable convolutional frontend to enhance local time-frequency modeling; (2) the first integration of phoneme-level feature pooling with TDNN, explicitly capturing phoneme-scale discriminative cues; and (3) a triple-domain fusion mechanism combining time-frequency, contextual, and fine-grained features. Evaluated on VoxCeleb1, M-TDNN achieves state-of-the-art performance—significantly reducing EER—while requiring fewer parameters and lower computational cost than standard TDNN and ECAPA-TDNN baselines.

0 citationsRead paper