Institution profile

Shanghai International Studies University

Academic institutionasia · cn
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

MAAM: Anchor-Preserving Compression and Contextual Calibration for Chinese Discriminatory Language Detection

Jun 08, 2026

Detecting discriminatory language in Chinese is highly challenging due to its implicit intent and strong context dependence. This work proposes MAAM, a lightweight, model-agnostic framework that introduces the novel Myopia–Astigmatism anchoring mechanism to preserve semantics relevant to discrimination judgments. MAAM further calibrates predictions by integrating three contextual priors: contextual tone, group identity, and stance polarity (C-I-S). The study also constructs ChLGBT, the first Chinese corpus focused on LGBT-related discriminatory language. Evaluated across various encoder architectures, MAAM consistently achieves substantial improvements in accuracy, F1 score, Brier score, and calibration performance under both zero-shot and few-shot settings, matching the effectiveness of large language models while offering superior compactness and stability.

0 citationsRead paper

QKCV Attention: Enhancing Time Series Forecasting with Static Categorical Embeddings for Both Lightweight and Pre-trained Foundation Models

Oct 21, 2025

To address the underutilization of static categorical information in time series forecasting, this paper proposes the QKCV attention mechanism: it extends the standard QKV framework by incorporating a static categorical embedding $C$, enabling explicit modeling of category-specific dynamic patterns. The mechanism is architecture-agnostic—compatible with Transformer, Informer, PatchTST, TFT, and others—and supports efficient transfer learning via fine-tuning only the categorical embedding $C$, drastically reducing computational overhead. Experiments across multiple real-world datasets demonstrate that QKCV significantly improves prediction accuracy in univariate time series forecasting. Notably, it delivers consistent performance gains in both lightweight models and fine-tuning scenarios of pretrained foundation models, achieving an optimal balance between accuracy and efficiency. QKCV establishes a general, scalable paradigm for categorical-aware time series modeling.

0 citationsRead paper

Red alert: Millions of "homeless" publications in Scopus should be resettled

Aug 18, 2025

Scopus contains millions of “homeless” publications—authored by researchers with complete institutional affiliations yet erroneously labeled “country-undefined”—undermining database reliability and compromising the accuracy and fairness of research evaluation. This study systematically identifies four primary root causes: incomplete address information, failure to recognize national name variants, typographical errors, and deficiencies in address parsing algorithms. Leveraging 124 years of Scopus metadata, we integrate bibliometric analysis, multilingual standardization of country names, and fine-grained data cleaning to classify and quantitatively trace these causes. Our findings yield a reproducible methodological framework for metadata quality enhancement, enabling institutional affiliation calibration, cross-national research performance assessment, and optimization of scholarly infrastructure. The approach advances best practices in bibliographic data curation and supports equitable, evidence-based science policy.

0 citationsRead paper
Recent publications

Latest Papers

MAAM: Anchor-Preserving Compression and Contextual Calibration for Chinese Discriminatory Language Detection

Jun 08, 2026

Detecting discriminatory language in Chinese is highly challenging due to its implicit intent and strong context dependence. This work proposes MAAM, a lightweight, model-agnostic framework that introduces the novel Myopia–Astigmatism anchoring mechanism to preserve semantics relevant to discrimination judgments. MAAM further calibrates predictions by integrating three contextual priors: contextual tone, group identity, and stance polarity (C-I-S). The study also constructs ChLGBT, the first Chinese corpus focused on LGBT-related discriminatory language. Evaluated across various encoder architectures, MAAM consistently achieves substantial improvements in accuracy, F1 score, Brier score, and calibration performance under both zero-shot and few-shot settings, matching the effectiveness of large language models while offering superior compactness and stability.

0 citationsRead paper

QKCV Attention: Enhancing Time Series Forecasting with Static Categorical Embeddings for Both Lightweight and Pre-trained Foundation Models

Oct 21, 2025

To address the underutilization of static categorical information in time series forecasting, this paper proposes the QKCV attention mechanism: it extends the standard QKV framework by incorporating a static categorical embedding $C$, enabling explicit modeling of category-specific dynamic patterns. The mechanism is architecture-agnostic—compatible with Transformer, Informer, PatchTST, TFT, and others—and supports efficient transfer learning via fine-tuning only the categorical embedding $C$, drastically reducing computational overhead. Experiments across multiple real-world datasets demonstrate that QKCV significantly improves prediction accuracy in univariate time series forecasting. Notably, it delivers consistent performance gains in both lightweight models and fine-tuning scenarios of pretrained foundation models, achieving an optimal balance between accuracy and efficiency. QKCV establishes a general, scalable paradigm for categorical-aware time series modeling.

0 citationsRead paper

Red alert: Millions of "homeless" publications in Scopus should be resettled

Aug 18, 2025

Scopus contains millions of “homeless” publications—authored by researchers with complete institutional affiliations yet erroneously labeled “country-undefined”—undermining database reliability and compromising the accuracy and fairness of research evaluation. This study systematically identifies four primary root causes: incomplete address information, failure to recognize national name variants, typographical errors, and deficiencies in address parsing algorithms. Leveraging 124 years of Scopus metadata, we integrate bibliometric analysis, multilingual standardization of country names, and fine-grained data cleaning to classify and quantitatively trace these causes. Our findings yield a reproducible methodological framework for metadata quality enhancement, enabling institutional affiliation calibration, cross-national research performance assessment, and optimization of scholarly infrastructure. The approach advances best practices in bibliographic data curation and supports equitable, evidence-based science policy.

0 citationsRead paper