Flower Hub: A Reproducible Benchmarking Platform for Federated Learning in Simulation and Deployment
本文针对联邦学习中难以复现、比较和扩展的问题,提出Flower Hub平台,通过将基准测试打包为可执行的应用程序,支持在模拟和实际部署中统一评估。
本文针对联邦学习中难以复现、比较和扩展的问题,提出Flower Hub平台,通过将基准测试打包为可执行的应用程序,支持在模拟和实际部署中统一评估。
This study addresses the absence of a universal, domain-agnostic definition of decentralization across computer communication systems, which has led to inconsistent analyses, incommensurable comparisons, and imprecise system designs. From an ontological perspective, the work models decentralization as a relational and observer-dependent property, introducing the first domain-independent graph-based ontological framework that clearly distinguishes decentralization from distribution. The framework incorporates two novel metrics—Void Tolerance and Imperviousness—and leverages formal semantics alongside browser-based automated reasoning to enable quantitative evaluation and classification of decentralization. Validation through case studies in federated learning and blockchain demonstrates that the framework yields consistent, comparable, and logically coherent assessments.
This work addresses the high memory and communication overheads in cross-datacenter large language model (LLM) pretraining, where each site traditionally maintains a full model replica. To overcome this limitation, the authors propose an efficient MoE-based distributed pretraining approach that partitions expert layers across nodes, selectively replicates a subset of experts, incorporates a skip-token mechanism, and employs a low-communication data-parallel strategy. Notably, this is the first method to integrate partial expert replication into federated-style training, substantially reducing both communication and memory costs while ensuring routing stability. Experimental results demonstrate that the proposed method reduces communication overhead by 1.42× and 45.44× compared to strong baselines and standard DDP, respectively, achieves up to 1.4× higher throughput, and scales effectively to models with hundreds of billions of parameters.
This work addresses the challenges in federated fine-tuning caused by heterogeneous adapter ranks across clients, which lead to uncontrolled distribution of low-rank representations and suboptimal aggregation performance. To mitigate these issues, the authors propose a prefix-nested LoRA architecture integrated with a segmented aggregation rule and a multi-rank truncation training strategy. In this framework, the low-rank components prioritize learning task-critical information, while higher-rank components provide additional representational capacity. This design enables efficient and consistent information sharing and aggregation under heterogeneous rank configurations. Experimental results demonstrate that the proposed method significantly outperforms existing heterogeneous federated LoRA approaches across multiple foundation models, achieving higher accuracy and ROUGE-L scores while maintaining comparable or lower perplexity.
This work addresses the lack of a unified cross-modal and cross-domain evaluation framework for automatic poster generation, which often fails to balance informational fidelity and visual effectiveness. To bridge this gap, the authors introduce Any2Poster Bench, a comprehensive benchmark, along with the Any2Poster Agent generation system—the first to support end-to-end poster synthesis from eight input modalities across five content domains. The system integrates heterogeneous source parsing, layout planning, rendering, and an iterative refinement mechanism driven by vision-language model feedback. Experimental results demonstrate that the proposed approach achieves average cross-modal and cross-domain accuracies of 87.25% and 87.28%, respectively, on Any2Poster Bench, and attains an overall accuracy of 72.58% with a density-enhanced score of 145.16 on PaperQuiz, significantly outperforming existing methods.
本文针对联邦学习中难以复现、比较和扩展的问题,提出Flower Hub平台,通过将基准测试打包为可执行的应用程序,支持在模拟和实际部署中统一评估。
This study addresses the absence of a universal, domain-agnostic definition of decentralization across computer communication systems, which has led to inconsistent analyses, incommensurable comparisons, and imprecise system designs. From an ontological perspective, the work models decentralization as a relational and observer-dependent property, introducing the first domain-independent graph-based ontological framework that clearly distinguishes decentralization from distribution. The framework incorporates two novel metrics—Void Tolerance and Imperviousness—and leverages formal semantics alongside browser-based automated reasoning to enable quantitative evaluation and classification of decentralization. Validation through case studies in federated learning and blockchain demonstrates that the framework yields consistent, comparable, and logically coherent assessments.
This work addresses the high memory and communication overheads in cross-datacenter large language model (LLM) pretraining, where each site traditionally maintains a full model replica. To overcome this limitation, the authors propose an efficient MoE-based distributed pretraining approach that partitions expert layers across nodes, selectively replicates a subset of experts, incorporates a skip-token mechanism, and employs a low-communication data-parallel strategy. Notably, this is the first method to integrate partial expert replication into federated-style training, substantially reducing both communication and memory costs while ensuring routing stability. Experimental results demonstrate that the proposed method reduces communication overhead by 1.42× and 45.44× compared to strong baselines and standard DDP, respectively, achieves up to 1.4× higher throughput, and scales effectively to models with hundreds of billions of parameters.
This work addresses the challenges in federated fine-tuning caused by heterogeneous adapter ranks across clients, which lead to uncontrolled distribution of low-rank representations and suboptimal aggregation performance. To mitigate these issues, the authors propose a prefix-nested LoRA architecture integrated with a segmented aggregation rule and a multi-rank truncation training strategy. In this framework, the low-rank components prioritize learning task-critical information, while higher-rank components provide additional representational capacity. This design enables efficient and consistent information sharing and aggregation under heterogeneous rank configurations. Experimental results demonstrate that the proposed method significantly outperforms existing heterogeneous federated LoRA approaches across multiple foundation models, achieving higher accuracy and ROUGE-L scores while maintaining comparable or lower perplexity.
This work addresses the lack of a unified cross-modal and cross-domain evaluation framework for automatic poster generation, which often fails to balance informational fidelity and visual effectiveness. To bridge this gap, the authors introduce Any2Poster Bench, a comprehensive benchmark, along with the Any2Poster Agent generation system—the first to support end-to-end poster synthesis from eight input modalities across five content domains. The system integrates heterogeneous source parsing, layout planning, rendering, and an iterative refinement mechanism driven by vision-language model feedback. Experimental results demonstrate that the proposed approach achieves average cross-modal and cross-domain accuracies of 87.25% and 87.28%, respectively, on Any2Poster Bench, and attains an overall accuracy of 72.58% with a density-enhanced score of 145.16 on PaperQuiz, significantly outperforming existing methods.