A Differentially Private Federated Proximal Optimization Framework for Customer Churn Prediction in Heterogeneous Federated Telecom Networks
为解决电信网络中客户流失预测的隐私和异质性问题,提出了一种基于差分隐私的联邦近端优化框架DP-FedProx。
为解决电信网络中客户流失预测的隐私和异质性问题,提出了一种基于差分隐私的联邦近端优化框架DP-FedProx。
本文针对GAN生成的假脸检测难题,提出结合EfficientNet与Swin Transformer的轻量级架构方法,提高了检测准确率和效率。
This study addresses the unclear origins of domain-balancing gains in multi-domain meeting summarization by constructing budget-matched controlled corpora to decouple token distribution from data volume effects during Mistral-7B fine-tuning. Employing QLoRA, fact-level evaluation, and human verification, the research elucidates distinct mechanisms underlying token-wise versus sample-wise balancing and proposes a low-loss pruning strategy. Experimental results demonstrate that the optimized balancing approach significantly enhances minority-domain quality with minimal cost to majority domains. Furthermore, pruning 15% of ineffective tokens achieves lossless performance, establishing an efficient paradigm for data mixing and cleaning in multi-domain summarization training.
This study addresses the high sensitivity of single-run evaluations to random seeds in low-resource Garhwali speech recognition, which often obscures genuine performance gains from stochastic noise. To remedy this, the authors establish the first reproducible ASR benchmark on the official VAANI dataset using multiple random seeds and propose a new evaluation paradigm centered on multi-seed assessment and statistical significance testing. Systematic re-evaluation of various optimization objectives and transfer strategies reveals that standard CTC combined with w2v-BERT 2.0 achieves a 47.0% WER across five seeds, outperforming larger models such as MMS-1B. While speed perturbation yields consistent minor improvements, more complex approaches like Focal CTC and matra weighting fail to demonstrate statistically significant gains. The findings underscore the fragility of common enhancements in low-resource settings and highlight the superiority of thoughtful pretraining design over mere model scale expansion.
This study investigates the feasibility of leveraging a frozen DINOv2 ViT-H/16 foundation vision model for multi-label detection of tumors and cysts in 3D kidney CT scans without domain-specific pretraining. Using the KiTS23 dataset, the authors systematically evaluate three patch token aggregation strategies: CLS linear probing, gated attention-based multiple instance learning (MIL), and ProtoViT prototype heads. Results show that attention MIL achieves AUROCs of 0.74 and 0.80 for tumor and cyst detection, respectively, and demonstrates strong spatial interpretability—its attention weights concentrate 7.5–9.8× more on true lesion regions. In contrast, the prototype head fails entirely on cyst detection. The work highlights a critical trade-off between performance and interpretability when applying large-scale vision foundation models to medical multi-label classification tasks.
为解决电信网络中客户流失预测的隐私和异质性问题,提出了一种基于差分隐私的联邦近端优化框架DP-FedProx。
本文针对GAN生成的假脸检测难题,提出结合EfficientNet与Swin Transformer的轻量级架构方法,提高了检测准确率和效率。
This study addresses the unclear origins of domain-balancing gains in multi-domain meeting summarization by constructing budget-matched controlled corpora to decouple token distribution from data volume effects during Mistral-7B fine-tuning. Employing QLoRA, fact-level evaluation, and human verification, the research elucidates distinct mechanisms underlying token-wise versus sample-wise balancing and proposes a low-loss pruning strategy. Experimental results demonstrate that the optimized balancing approach significantly enhances minority-domain quality with minimal cost to majority domains. Furthermore, pruning 15% of ineffective tokens achieves lossless performance, establishing an efficient paradigm for data mixing and cleaning in multi-domain summarization training.
This study addresses the high sensitivity of single-run evaluations to random seeds in low-resource Garhwali speech recognition, which often obscures genuine performance gains from stochastic noise. To remedy this, the authors establish the first reproducible ASR benchmark on the official VAANI dataset using multiple random seeds and propose a new evaluation paradigm centered on multi-seed assessment and statistical significance testing. Systematic re-evaluation of various optimization objectives and transfer strategies reveals that standard CTC combined with w2v-BERT 2.0 achieves a 47.0% WER across five seeds, outperforming larger models such as MMS-1B. While speed perturbation yields consistent minor improvements, more complex approaches like Focal CTC and matra weighting fail to demonstrate statistically significant gains. The findings underscore the fragility of common enhancements in low-resource settings and highlight the superiority of thoughtful pretraining design over mere model scale expansion.
This study investigates the feasibility of leveraging a frozen DINOv2 ViT-H/16 foundation vision model for multi-label detection of tumors and cysts in 3D kidney CT scans without domain-specific pretraining. Using the KiTS23 dataset, the authors systematically evaluate three patch token aggregation strategies: CLS linear probing, gated attention-based multiple instance learning (MIL), and ProtoViT prototype heads. Results show that attention MIL achieves AUROCs of 0.74 and 0.80 for tumor and cyst detection, respectively, and demonstrates strong spatial interpretability—its attention weights concentrate 7.5–9.8× more on true lesion regions. In contrast, the prototype head fails entirely on cyst detection. The work highlights a critical trade-off between performance and interpretability when applying large-scale vision foundation models to medical multi-label classification tasks.