Institution profile

Beijing National Research Center for Information Science and Technology

Academic institutionasia · cn
Research library66linked papers
Opportunities0open roles
Selected work

Representative Papers

Accurate Forgetting for Heterogeneous Federated Continual Learning

Feb 20, 2025International Conference on Learning Representations

To address statistical bias and noise interference arising from client data/task heterogeneity—or even adversarial behavior—in federated continual learning (FCL), this paper introduces the “Accurate Forgetting” (AF) paradigm: proactively identifying and discarding unreliable feature representations induced by skewed distributions and noise prior to knowledge reuse. Methodologically, we propose the first probability-based credibility assessment framework built upon normalized flows, enabling quantifiable, knowledge-granular filtering. Further, we integrate generative replay with selective knowledge inheritance to dynamically enhance global model robustness within the federated architecture. Evaluated on multiple heterogeneous FCL benchmarks, AF achieves an average accuracy improvement of 12.3%, significantly boosting generalization and noise resilience. Our approach provides a novel, interpretable, and computationally tractable pathway for bias mitigation in FCL.

5 citationsRead paper

A Consolidated Game Framework for Cooperative Defense Against Cross-Domain Cyber Attacks in Satellite-Enabled Internet of Things

Aug 11, 2026

This study addresses the challenge of cross-domain cyber threats in satellite-enabled Internet of Things (IoT) systems, where misaligned incentives between terrestrial and satellite operators and the difficulty of quantifying attack impacts hinder effective collaborative defense. To overcome these barriers, this work proposes the first tripartite security game framework that integrates both domains, aligning the defensive incentives of IoT operators and satellite service providers through a traffic pricing mechanism. Coupled with an efficient learning algorithm to optimize strategic decisions for both parties, the proposed approach mitigates incentive misalignment and significantly enhances cross-domain collaborative defense performance even under conditions of divergent interests. Experimental results demonstrate the framework’s effectiveness in improving system-wide security outcomes.

0 citationsRead paper

FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models

Aug 09, 2026

This work addresses the lack of systematic evaluation of multimodal large language models (MLLMs) in assessing exercise form quality (Action Quality Assessment, AQA), where existing benchmarks suffer from inconsistent error definitions and insufficient granularity. To bridge this gap, we introduce FitAQA, a novel benchmark comprising 2,219 videos across 30 bodyweight exercises and 5,512 question-answer pairs, grounded in a six-dimensional form error ontology collaboratively developed with sports science experts. FitAQA encompasses three core tasks—perception, judgment, and temporal localization—enabling, for the first time, fine-grained and interpretable AQA evaluation tailored to MLLMs. Experimental results reveal that current models still struggle with comprehensive assessment and precise error localization, with visual perception identified as the primary bottleneck; notably, incorporating real perceptual evidence substantially improves judgment accuracy.

0 citationsRead paper
Recent publications

Latest Papers

A Consolidated Game Framework for Cooperative Defense Against Cross-Domain Cyber Attacks in Satellite-Enabled Internet of Things

Aug 11, 2026

This study addresses the challenge of cross-domain cyber threats in satellite-enabled Internet of Things (IoT) systems, where misaligned incentives between terrestrial and satellite operators and the difficulty of quantifying attack impacts hinder effective collaborative defense. To overcome these barriers, this work proposes the first tripartite security game framework that integrates both domains, aligning the defensive incentives of IoT operators and satellite service providers through a traffic pricing mechanism. Coupled with an efficient learning algorithm to optimize strategic decisions for both parties, the proposed approach mitigates incentive misalignment and significantly enhances cross-domain collaborative defense performance even under conditions of divergent interests. Experimental results demonstrate the framework’s effectiveness in improving system-wide security outcomes.

0 citationsRead paper

FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models

Aug 09, 2026

This work addresses the lack of systematic evaluation of multimodal large language models (MLLMs) in assessing exercise form quality (Action Quality Assessment, AQA), where existing benchmarks suffer from inconsistent error definitions and insufficient granularity. To bridge this gap, we introduce FitAQA, a novel benchmark comprising 2,219 videos across 30 bodyweight exercises and 5,512 question-answer pairs, grounded in a six-dimensional form error ontology collaboratively developed with sports science experts. FitAQA encompasses three core tasks—perception, judgment, and temporal localization—enabling, for the first time, fine-grained and interpretable AQA evaluation tailored to MLLMs. Experimental results reveal that current models still struggle with comprehensive assessment and precise error localization, with visual perception identified as the primary bottleneck; notably, incorporating real perceptual evidence substantially improves judgment accuracy.

0 citationsRead paper

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution

Jul 29, 2026

This work addresses the high computational cost and representational mismatch between vision and language modalities in full-parameter fine-tuning of multimodal large language models for visual instruction tuning. To improve parameter efficiency, the authors propose a decoupled visual processing framework built upon LLaVA-1.5, which freezes the original decoder and explicitly separates visual and textual tokens at the top layer. A lightweight, trainable Transformer block is introduced exclusively for processing visual tokens, enabling distinct pathways for each modality. Despite updating only a minimal fraction of the model’s parameters, this approach achieves performance on par with full fine-tuning across multiple benchmarks—including MME, POPE, and ChartQA—demonstrating significantly enhanced parameter efficiency without compromising task performance.

0 citationsRead paper