Institution profile

Shanghai Normal University

Academic institutionasia · cn
Official website
Research library15linked papers
Opportunities0open roles
Selected work

Representative Papers

Multi-Relational Knowledge Graph Enhanced Embedding for Trajectory-User Linking

Aug 09, 2026

This work addresses the challenge of linking anonymized trajectories to user identities, a task hindered by the difficulty of effectively integrating multi-source heterogeneous mobility semantics and leveraging trajectory structural knowledge. To this end, the study introduces knowledge graph representation learning into the trajectory-user linkage problem for the first time, constructing a multi-relational mobility knowledge graph that models visit time, POI category, and transition speed as typed relations. These relations jointly constrain POI embeddings and incorporate higher-order co-occurrence patterns to enrich representations. A dual-branch classifier is further designed to fuse structural and sequential evidence. The proposed method achieves significant improvements in linkage accuracy, demonstrating particularly strong performance in scenarios involving sparse and overlapping trajectories.

0 citationsRead paper

Task-Specific Feature Fusion Method for Multi-Task Affective Behavior Analysis

Jul 15, 2026

This work addresses the performance limitations of unified models in multi-task affective behavior analysis, where discrepancies among valence-arousal regression, expression classification, and action unit detection hinder joint optimization. To overcome this, the authors propose a task-adaptive feature fusion strategy that leverages complementary frame-level features extracted from frozen DINOv2 ViT-L and DINOv3 ConvNeXt-base backbones. Each task is equipped with a dedicated prediction head and tailored fusion mechanism—such as gated or residual fusion—to avoid enforced architectural sharing. The framework further integrates temporal convolutions, LightGBM-based post-processing, and threshold calibration. Evaluated on the ABAW11 validation set, the method achieves an EXPR macro F1 of 0.4222, AU macro F1 of 0.5402, and VA mean CCC of 0.6717, yielding a total score of 1.6341 and demonstrating substantial improvement in multi-task collaborative performance.

0 citationsRead paper

Zipper-LoRA: Dynamic Parameter Decoupling for Speech-LLM based Multilingual Speech Recognition

Mar 18, 2026

This work addresses the stability-plasticity dilemma in multilingual speech recognition caused by data imbalance: fully shared parameters subject low-resource languages to negative interference, while entirely independent parameters impede cross-lingual knowledge transfer. To resolve this, the authors propose Zipper-LoRA, a novel framework featuring rank-level dynamic decoupling that enables fine-grained parameter sharing and disentanglement within the LoRA subspace via lightweight language-conditioned routers. The approach incorporates Static, Hard, and Soft variants alongside a two-stage training strategy—including an Initial-B warm start—to share parameters when compatible and decouple them when conflicting. Experiments across 12 languages with mixed resource levels demonstrate that Zipper-LoRA significantly outperforms both fully shared and fully independent baselines, with especially pronounced gains under extremely low-resource conditions, and maintains robust performance across both chunked and non-chunked encoder configurations.

0 citationsRead paper

A Two-Stage Dual-Modality Model for Facial Emotional Expression Recognition

Mar 12, 2026

This work addresses the challenges of frame-level eight-class facial expression recognition in unconstrained videos, including inaccurate localization, pose and scale variations, motion blur, and temporal inconsistency. To tackle these issues, a two-stage audio-visual multimodal approach is proposed. In the first stage, robust visual features are extracted using DINOv2, augmented with padding-aware data augmentation and processed by a mixture-of-experts classification head. The second stage integrates multi-scale re-cropped visual features with Wav2Vec 2.0 audio representations through a lightweight gated fusion module, further enhanced by temporal smoothing during inference to improve stability. The method achieves a Macro-F1 score of 0.5368 on the ABAW validation set and 0.5122 ± 0.0277 under five-fold cross-validation, significantly outperforming the official baseline.

0 citationsRead paper
Recent publications

Latest Papers

Multi-Relational Knowledge Graph Enhanced Embedding for Trajectory-User Linking

Aug 09, 2026

This work addresses the challenge of linking anonymized trajectories to user identities, a task hindered by the difficulty of effectively integrating multi-source heterogeneous mobility semantics and leveraging trajectory structural knowledge. To this end, the study introduces knowledge graph representation learning into the trajectory-user linkage problem for the first time, constructing a multi-relational mobility knowledge graph that models visit time, POI category, and transition speed as typed relations. These relations jointly constrain POI embeddings and incorporate higher-order co-occurrence patterns to enrich representations. A dual-branch classifier is further designed to fuse structural and sequential evidence. The proposed method achieves significant improvements in linkage accuracy, demonstrating particularly strong performance in scenarios involving sparse and overlapping trajectories.

0 citationsRead paper

Task-Specific Feature Fusion Method for Multi-Task Affective Behavior Analysis

Jul 15, 2026

This work addresses the performance limitations of unified models in multi-task affective behavior analysis, where discrepancies among valence-arousal regression, expression classification, and action unit detection hinder joint optimization. To overcome this, the authors propose a task-adaptive feature fusion strategy that leverages complementary frame-level features extracted from frozen DINOv2 ViT-L and DINOv3 ConvNeXt-base backbones. Each task is equipped with a dedicated prediction head and tailored fusion mechanism—such as gated or residual fusion—to avoid enforced architectural sharing. The framework further integrates temporal convolutions, LightGBM-based post-processing, and threshold calibration. Evaluated on the ABAW11 validation set, the method achieves an EXPR macro F1 of 0.4222, AU macro F1 of 0.5402, and VA mean CCC of 0.6717, yielding a total score of 1.6341 and demonstrating substantial improvement in multi-task collaborative performance.

0 citationsRead paper

Zipper-LoRA: Dynamic Parameter Decoupling for Speech-LLM based Multilingual Speech Recognition

Mar 18, 2026

This work addresses the stability-plasticity dilemma in multilingual speech recognition caused by data imbalance: fully shared parameters subject low-resource languages to negative interference, while entirely independent parameters impede cross-lingual knowledge transfer. To resolve this, the authors propose Zipper-LoRA, a novel framework featuring rank-level dynamic decoupling that enables fine-grained parameter sharing and disentanglement within the LoRA subspace via lightweight language-conditioned routers. The approach incorporates Static, Hard, and Soft variants alongside a two-stage training strategy—including an Initial-B warm start—to share parameters when compatible and decouple them when conflicting. Experiments across 12 languages with mixed resource levels demonstrate that Zipper-LoRA significantly outperforms both fully shared and fully independent baselines, with especially pronounced gains under extremely low-resource conditions, and maintains robust performance across both chunked and non-chunked encoder configurations.

0 citationsRead paper

A Two-Stage Dual-Modality Model for Facial Emotional Expression Recognition

Mar 12, 2026

This work addresses the challenges of frame-level eight-class facial expression recognition in unconstrained videos, including inaccurate localization, pose and scale variations, motion blur, and temporal inconsistency. To tackle these issues, a two-stage audio-visual multimodal approach is proposed. In the first stage, robust visual features are extracted using DINOv2, augmented with padding-aware data augmentation and processed by a mixture-of-experts classification head. The second stage integrates multi-scale re-cropped visual features with Wav2Vec 2.0 audio representations through a lightweight gated fusion module, further enhanced by temporal smoothing during inference to improve stability. The method achieves a Macro-F1 score of 0.5368 on the ABAW validation set and 0.5122 ± 0.0277 under five-fold cross-validation, significantly outperforming the official baseline.

0 citationsRead paper