Just Talk Once: Communication-Efficient Split Federated LLM Fine-Tuning on Edge Devices

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出L形SFT框架,通过在服务器端直接监督隐藏激活来消除双向通信瓶颈,并引入一次性SFT减少客户端持续参与,有效降低联邦大语言模型微调中的通信成本和客户端在线时间。
📝 Abstract
Large language model (LLM) fine-tuning is increasingly shifting toward data generated on edge devices, where memory, computation, bandwidth, and connectivity constraints make conventional federated learning difficult to sustain. Split federated fine-tuning (SFT) improves client-side efficiency by offloading most model parameters and computation to the server but requires step-by-step bidirectional communication loop across the split interface and forces continuous client involvement throughout training. In this paper, we present L-shaped SFT, a split fine-tuning framework that removes this bidirectional bottleneck. Our key insight is that weight tying in modern LLMs enables server-side hidden activations to be directly supervised using target embeddings, allowing the training loss to be computed on the server without returning server outputs to the client. To further eliminate the need for continuous client participation, based on L-shaped SFT, we introduce one-shot SFT, in which clients upload activations once and then go offline while the server continues optimization over cached representations. We implement our design in a real system testbed with heterogeneous edge clients, including commercial smartphones and NVIDIA developer boards. Experiments demonstrate that our schemes significantly reduce communication costs and client online time compared with existing SFT baselines.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model
Edge Devices
Federated Learning
Communication Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

L-shaped SFT
weight tying
one-shot SFT
communication-efficient
🔎 Similar Papers
No similar papers found.