AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP
研究通过构建AtlasNLP解决NLP数据集中国家代表性信息缺失问题,采用人工整理与大规模自动收集相结合的方法,揭示了数据集覆盖的不均衡性和地理不对称性。
研究通过构建AtlasNLP解决NLP数据集中国家代表性信息缺失问题,采用人工整理与大规模自动收集相结合的方法,揭示了数据集覆盖的不均衡性和地理不对称性。
To address the insufficient accuracy of story-point-based effort estimation in agile development, this paper proposes a hybrid modeling approach integrating LASSO and Elastic Net regression, optimized via grid search and validated using 5-fold cross-validation. Empirical evaluation on 21 real-world projects demonstrates that the proposed model significantly outperforms existing methods: LASSO achieves 100.0% prediction accuracy at thresholds of 8% and 25% (PRED(8%) and PRED(25%)), with a mean magnitude of relative error (MMRE) of 0.0491 and mean squared error (MSE) of 0.0007. The primary contribution lies in the first systematic empirical validation of sparse regularized regression for story-point-driven effort estimation—establishing a new statistical modeling paradigm that is highly accurate, interpretable, and deployable in agile practice.
To address insufficient semantic information exploitation, poor generalization of historical access patterns, and resource constraints on mobile devices in Web prefetching, this paper proposes a lightweight, multi-source collaborative prefetching framework. The framework supports plug-and-play integration of diverse prefetching strategies—including semantic graph modeling and access sequence prediction—without requiring modifications to underlying algorithms. It introduces a novel adaptive weighting mechanism that dynamically adjusts each strategy’s contribution based on real-time prediction confidence. Furthermore, it jointly leverages application-level contextual awareness and document-level semantic relationship analysis to balance accuracy and scalability. Experimental evaluation on mobile platforms demonstrates that the framework reduces average access latency by 23.6%, achieves lower memory and CPU overhead than state-of-the-art methods, and improves prefetching accuracy by 17.4%, confirming its efficiency and practicality.
To address low accuracy and susceptibility to thermal drift and dynamic noise in online monitoring of wheel flange wear depth, this paper proposes an onboard dynamic machine learning monitoring system. Methodologically, it fuses displacement and temperature sensor data, employs a real-time IIR filter to suppress nonlinear thermal drift and dynamic noise, and develops a dynamically auto-trained regression model; filter parameters are optimized via FFT, while edge acquisition, real-time transmission, and processing are implemented on an embedded IoT platform. Key contributions include a dynamic model update mechanism and an IIR-ML collaborative filtering architecture. Experimental results demonstrate a monitoring accuracy of 98.2% (post-filtering), low end-to-end latency, and effective identification of abnormal wear induced by track irregularities—significantly enhancing wheel-rail condition awareness and railway operational safety.
This work addresses core challenges in dense retrieval for Amharic—a low-resource language with 120 million speakers—including scarcity of labeled data, pretraining resources, and word embeddings. We present the first systematic feasibility study, proposing a lightweight fine-tuning and cross-lingual transfer framework built upon mBERT and XLM-R. Our approach integrates contrastive learning, pseudo-labeling, and unsupervised domain adaptation to train dense encoders. Evaluated on a newly constructed Amharic QA retrieval benchmark—the first of its kind—we achieve a 37% improvement in Recall@10 over baseline methods, substantially outperforming traditional sparse retrieval and zero-shot cross-lingual baselines. Key contributions are: (1) the first publicly available Amharic dense retrieval benchmark; (2) empirical validation of lightweight adaptation and cross-lingual transfer efficacy in low-resource settings; and (3) a reproducible methodology for information retrieval in African languages.
研究通过构建AtlasNLP解决NLP数据集中国家代表性信息缺失问题,采用人工整理与大规模自动收集相结合的方法,揭示了数据集覆盖的不均衡性和地理不对称性。
To address the insufficient accuracy of story-point-based effort estimation in agile development, this paper proposes a hybrid modeling approach integrating LASSO and Elastic Net regression, optimized via grid search and validated using 5-fold cross-validation. Empirical evaluation on 21 real-world projects demonstrates that the proposed model significantly outperforms existing methods: LASSO achieves 100.0% prediction accuracy at thresholds of 8% and 25% (PRED(8%) and PRED(25%)), with a mean magnitude of relative error (MMRE) of 0.0491 and mean squared error (MSE) of 0.0007. The primary contribution lies in the first systematic empirical validation of sparse regularized regression for story-point-driven effort estimation—establishing a new statistical modeling paradigm that is highly accurate, interpretable, and deployable in agile practice.
To address insufficient semantic information exploitation, poor generalization of historical access patterns, and resource constraints on mobile devices in Web prefetching, this paper proposes a lightweight, multi-source collaborative prefetching framework. The framework supports plug-and-play integration of diverse prefetching strategies—including semantic graph modeling and access sequence prediction—without requiring modifications to underlying algorithms. It introduces a novel adaptive weighting mechanism that dynamically adjusts each strategy’s contribution based on real-time prediction confidence. Furthermore, it jointly leverages application-level contextual awareness and document-level semantic relationship analysis to balance accuracy and scalability. Experimental evaluation on mobile platforms demonstrates that the framework reduces average access latency by 23.6%, achieves lower memory and CPU overhead than state-of-the-art methods, and improves prefetching accuracy by 17.4%, confirming its efficiency and practicality.
To address low accuracy and susceptibility to thermal drift and dynamic noise in online monitoring of wheel flange wear depth, this paper proposes an onboard dynamic machine learning monitoring system. Methodologically, it fuses displacement and temperature sensor data, employs a real-time IIR filter to suppress nonlinear thermal drift and dynamic noise, and develops a dynamically auto-trained regression model; filter parameters are optimized via FFT, while edge acquisition, real-time transmission, and processing are implemented on an embedded IoT platform. Key contributions include a dynamic model update mechanism and an IIR-ML collaborative filtering architecture. Experimental results demonstrate a monitoring accuracy of 98.2% (post-filtering), low end-to-end latency, and effective identification of abnormal wear induced by track irregularities—significantly enhancing wheel-rail condition awareness and railway operational safety.
This work addresses core challenges in dense retrieval for Amharic—a low-resource language with 120 million speakers—including scarcity of labeled data, pretraining resources, and word embeddings. We present the first systematic feasibility study, proposing a lightweight fine-tuning and cross-lingual transfer framework built upon mBERT and XLM-R. Our approach integrates contrastive learning, pseudo-labeling, and unsupervised domain adaptation to train dense encoders. Evaluated on a newly constructed Amharic QA retrieval benchmark—the first of its kind—we achieve a 37% improvement in Recall@10 over baseline methods, substantially outperforming traditional sparse retrieval and zero-shot cross-lingual baselines. Key contributions are: (1) the first publicly available Amharic dense retrieval benchmark; (2) empirical validation of lightweight adaptation and cross-lingual transfer efficacy in low-resource settings; and (3) a reproducible methodology for information retrieval in African languages.