Institution profile

Fast Accounting Co., Ltd.

Industry research
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

MESH: Memory-Efficient Sinkhorn Optimization for Mixture-of-Experts Training

Aug 04, 2026

This work addresses the failure of Sinkhorn-based optimization in Mixture-of-Experts (MoE) training, which stems from the expert routing matrix exhibiting a high and highly time-varying gradient condition number. The authors propose MESH, a novel method that identifies the expert matrix as the primary cause of performance degradation and introduces an implicit momentum mechanism. This mechanism provides first-moment signals along the temporal dimension within the gradient buffer’s lifetime, thereby eliminating the need to explicitly store expert-level AdamW optimizer states. MESH optionally incorporates block- or neuron-level inverse RMS preconditioning to further enhance stability. Compared to AdamW, MESH reduces optimizer state memory by 62.5% and peak CUDA memory by 12.6%, with only a minor increase in evaluation loss. Ablation studies confirm that temporal smoothing is the key factor underlying its effectiveness.

0 citationsRead paper

GHaLIB: A Multilingual Framework for Hope Speech Detection in Low-Resource Languages

Dec 27, 2025

This study addresses the scarcity of hope speech detection research for low-resource languages—particularly Urdu—and the limited cross-lingual generalization of existing models. We propose the first lightweight, multilingual adaptation framework for hope speech recognition. Methodologically, we systematically evaluate the cross-lingual transfer performance of pretrained multilingual models—including XLM-RoBERTa, mBERT, EuroBERT, and UrduBERT—on hope speech detection, employing simple text preprocessing and supervised fine-tuning for efficient binary/multiclass classification. On the PolyHope-M 2025 benchmark, our approach achieves 95.2% F1 for Urdu binary classification and 65.2% F1 for multiclass classification, with robust performance also observed for Spanish, German, and English. Our key contributions are threefold: (1) filling critical resource and methodological gaps in hope speech detection for low-resource languages; (2) providing the first empirical validation of multilingual Transformers’ generalization capability for positive discourse detection; and (3) introducing an extensible, lightweight adaptation paradigm.

0 citationsRead paper

Three Stage Narrative Analysis; Plot-Sentiment Breakdown, Structure Learning and Concept Detection

Nov 14, 2025

This study addresses the challenge of automated semantic analysis and multi-level narrative concept extraction from large-scale narrative texts, such as film screenplays. We propose a three-stage computational framework: (1) sentiment arc modeling using a VAD-semantic–customized lexicon integrating LabMTsimple and NRC-VAD; (2) plot structure learning via Ward hierarchical clustering; and (3) context-aware, joint detection of high- and low-level narrative concepts grounded in character semantics. The method balances interpretability with computational efficiency, significantly improving accuracy in sentiment trajectory identification and consistency in plot-pattern clustering. Experiments on multiple screenplay datasets demonstrate the framework’s effectiveness in supporting narrative quality assessment and personalized content recommendation. It provides a scalable, reproducible technical pathway for computational narratology.

0 citationsRead paper

HyRet-Change: A hybrid retentive network for remote sensing change detection

Jun 15, 2025

Remote sensing change detection faces three key challenges: severe false-change interference, difficulty in detecting subtle changes, and high computational complexity of Transformer-based self-attention mechanisms, which hinder effective local-detail modeling. To address these, this paper proposes a Siamese-style Hybrid Retention Network (HRNet). Its core contributions are: (1) a novel parallel Convolution–Retention module that jointly captures local texture patterns and long-range dependencies by computing feature differences; and (2) an adaptive local–global interactive context-aware mechanism enabling cross-scale feature mutual learning and discriminative enhancement. Extensive experiments on three benchmark datasets—LEVIR-CD, WHU-CD, and CDD—demonstrate state-of-the-art performance, with significant suppression of false changes and improved detection accuracy for fine-grained changes. The source code is publicly available.

0 citationsRead paper
Recent publications

Latest Papers

MESH: Memory-Efficient Sinkhorn Optimization for Mixture-of-Experts Training

Aug 04, 2026

This work addresses the failure of Sinkhorn-based optimization in Mixture-of-Experts (MoE) training, which stems from the expert routing matrix exhibiting a high and highly time-varying gradient condition number. The authors propose MESH, a novel method that identifies the expert matrix as the primary cause of performance degradation and introduces an implicit momentum mechanism. This mechanism provides first-moment signals along the temporal dimension within the gradient buffer’s lifetime, thereby eliminating the need to explicitly store expert-level AdamW optimizer states. MESH optionally incorporates block- or neuron-level inverse RMS preconditioning to further enhance stability. Compared to AdamW, MESH reduces optimizer state memory by 62.5% and peak CUDA memory by 12.6%, with only a minor increase in evaluation loss. Ablation studies confirm that temporal smoothing is the key factor underlying its effectiveness.

0 citationsRead paper

GHaLIB: A Multilingual Framework for Hope Speech Detection in Low-Resource Languages

Dec 27, 2025

This study addresses the scarcity of hope speech detection research for low-resource languages—particularly Urdu—and the limited cross-lingual generalization of existing models. We propose the first lightweight, multilingual adaptation framework for hope speech recognition. Methodologically, we systematically evaluate the cross-lingual transfer performance of pretrained multilingual models—including XLM-RoBERTa, mBERT, EuroBERT, and UrduBERT—on hope speech detection, employing simple text preprocessing and supervised fine-tuning for efficient binary/multiclass classification. On the PolyHope-M 2025 benchmark, our approach achieves 95.2% F1 for Urdu binary classification and 65.2% F1 for multiclass classification, with robust performance also observed for Spanish, German, and English. Our key contributions are threefold: (1) filling critical resource and methodological gaps in hope speech detection for low-resource languages; (2) providing the first empirical validation of multilingual Transformers’ generalization capability for positive discourse detection; and (3) introducing an extensible, lightweight adaptation paradigm.

0 citationsRead paper

Three Stage Narrative Analysis; Plot-Sentiment Breakdown, Structure Learning and Concept Detection

Nov 14, 2025

This study addresses the challenge of automated semantic analysis and multi-level narrative concept extraction from large-scale narrative texts, such as film screenplays. We propose a three-stage computational framework: (1) sentiment arc modeling using a VAD-semantic–customized lexicon integrating LabMTsimple and NRC-VAD; (2) plot structure learning via Ward hierarchical clustering; and (3) context-aware, joint detection of high- and low-level narrative concepts grounded in character semantics. The method balances interpretability with computational efficiency, significantly improving accuracy in sentiment trajectory identification and consistency in plot-pattern clustering. Experiments on multiple screenplay datasets demonstrate the framework’s effectiveness in supporting narrative quality assessment and personalized content recommendation. It provides a scalable, reproducible technical pathway for computational narratology.

0 citationsRead paper

HyRet-Change: A hybrid retentive network for remote sensing change detection

Jun 15, 2025

Remote sensing change detection faces three key challenges: severe false-change interference, difficulty in detecting subtle changes, and high computational complexity of Transformer-based self-attention mechanisms, which hinder effective local-detail modeling. To address these, this paper proposes a Siamese-style Hybrid Retention Network (HRNet). Its core contributions are: (1) a novel parallel Convolution–Retention module that jointly captures local texture patterns and long-range dependencies by computing feature differences; and (2) an adaptive local–global interactive context-aware mechanism enabling cross-scale feature mutual learning and discriminative enhancement. Extensive experiments on three benchmark datasets—LEVIR-CD, WHU-CD, and CDD—demonstrate state-of-the-art performance, with significant suppression of false changes and improved detection accuracy for fine-grained changes. The source code is publicly available.

0 citationsRead paper