Joint Partitioning and Placement of Foundation Models for Real-Time Edge AI

πŸ“… 2025-11-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
To address resource dynamics, multi-objective constraints (latency, utilization, privacy), and infrastructure instability in large language model (LLM) inference within heterogeneous edge environments, this paper proposes a runtime-reconfigurable framework for joint model partitioning and deployment optimization. It is the first to formulate layer-granular model partitioning and device placement as a dynamic constrained optimization problem, integrating model-aware capacity analysis, dynamic graph neural network–based repartitioning, and resource forecasting. Evaluated in a 6G multi-access edge computing scenario, the approach reduces end-to-end latency by 27.4% and improves average GPU utilization by 39.1% over static baselines, while enabling privacy-sensitive layers to execute locally. The core contribution lies in an online, constraint-adaptive inference scheduler that ensures theoretical rigor and practical deployability under time-varying operational conditions.

Technology Category

Application Category

πŸ“ Abstract
Inference over large-scale foundation models within heterogeneous edge environments necessitates a fundamentally reconfigurable orchestration substrate. Static partitioning of model layers presumes temporal stability across compute and network resources, which is misaligned with the volatility of real-world deployments. We introduce a framework in which both the spatial placement and internal segmentation of foundation models are elevated to runtime-resolved constructs. The orchestration problem is formalized as a constrained optimization over layer-wise assignments, subject to evolving latency, utilization, and privacy gradients. The framework implements reactive inference composition responsive to infrastructural fluctuations by integrating model-aware capacity profiling with dynamic graph re-partitioning and reallocation. We introduce architectural and algorithmic components, along with a representative use case in 6G multi-access edge computing.
Problem

Research questions and friction points this paper is trying to address.

Dynamically partitioning and placing foundation models for edge AI inference
Addressing resource volatility in heterogeneous edge environments
Optimizing real-time inference under latency, utilization, and privacy constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic model partitioning and placement at runtime
Constrained optimization for layer assignments under gradients
Reactive inference composition with dynamic graph re-partitioning
πŸ”Ž Similar Papers
No similar papers found.
A
Aladin Djuhera
Technical University of Munich, Germany
F
Fernando Koch
Florida Atlantic University, USA
A
Alecio Binotto
Carl Zeiss AG, Germany