Institution profile

Institut Polytechnique des Sciences Avancées

Academic institutioneurope · fr
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Expert Routing for Communication-Efficient MoE via Finite Expert Banks

May 06, 2026

This work investigates the information-theoretic efficiency of routing mechanisms in sparse Mixture-of-Experts (MoE) architectures, aiming to balance model accuracy with communication and computational resource utilization. The gating router is modeled as a stochastic channel, and a discrete mutual information estimator is proposed under a finite expert pool. Empirical posterior distributions \( q(W|S) \) are leveraged to compute \( I(X;T) \) and \( I(S;W) \), with the latter shown to exhibit a monotonic relationship with the generalization gap. The Blahut–Arimoto algorithm is employed to trace the accuracy–rate trade-off curve. Experiments demonstrate that the proposed mutual information estimator effectively tracks the generalization gap and significantly outperforms both the Xu–Raginsky bound and the uniform joint bound, offering a practical analytical tool for resource-aware MoE systems.

0 citationsRead paper

Wireless Broadcast Gossip for Decentralized Drone Swarms: Success Probability, Contraction, and Optimal Aloha

Mar 19, 2026

This work addresses the trade-off between communication reliability and convergence speed in decentralized drone swarms achieving consensus via gossip protocols in interference-limited wireless environments. Focusing on slotted Aloha-based wireless gossip, the authors model drone locations as a planar Poisson point process and derive a closed-form signal-to-interference ratio (SIR) success probability under Rayleigh fading channels. By decomposing the dynamics into an ideal mixing term and a wireless sparsity term, they establish a mean-square contraction bound. Building on this analysis, they propose a closed-form optimal transmission probability that jointly optimizes availability and reliability. The theoretical framework integrates stochastic geometry, mean-field approximation, and convex optimization, with the predicted optimal operating point aligning closely with the fastest convergence region observed in simulations, thereby validating the approach.

0 citationsRead paper

Mixture-of-Experts under Finite-Rate Gating: Communication--Generalization Trade-offs

Feb 16, 2026

This work investigates the trade-off between model expressivity and generalization performance in Mixture-of-Experts (MoE) architectures under communication constraints. For the first time, rate-distortion theory is introduced into MoE analysis by modeling the gating mechanism as a stochastic channel operating at a finite rate. By integrating mutual information-based generalization bounds with the rate-distortion function \(D(R_g)\), the study establishes a quantitative relationship between the gating communication rate and generalization error. A theoretical upper bound on generalization error is derived and validated through synthetic multi-expert model simulations, which demonstrate that reducing the gating rate, while limiting expressivity, can enhance generalization. Based on these insights, the paper proposes a capacity-aware design principle for MoE systems, offering theoretical guidance for efficient model construction in resource-constrained settings.

0 citationsRead paper

General Multi-User Distributed Computing

Nov 25, 2025

This paper addresses the joint optimization of computation, communication, and model accuracy in multi-user–multi-server distributed computing and inference. We propose GMUDC, a unified framework supporting personalized objective optimization over arbitrary network topologies and heterogeneous task allocations. Methodologically, we introduce spectral-covering duality—establishing, for the first time, a theoretical link between network topology and generalization capability. Leveraging reproducing kernel Hilbert spaces and translation-invariant kernels, we integrate quenched–annealed two-phase analysis to achieve topology-aware distributed optimization under resource constraints. We theoretically prove GMUDC’s verifiable efficiency under given computational and communication budgets. Our framework provides a foundational information–energy co-design principle for federated learning and edge intelligence.

0 citationsRead paper

From Initial Data to Boundary Layers: Neural Networks for Nonlinear Hyperbolic Conservation Laws

Jun 02, 2025

This work addresses the challenge of approximating entropy solutions to initial-boundary value problems for nonlinear strictly hyperbolic conservation laws. Methodologically, it introduces a physics-informed deep learning framework that—uniquely—jointly models initial data and boundary layers. The approach incorporates a boundary-layer-aware loss function, a feature-adaptive weighting scheme, and an entropy-condition regularization term, yielding a PINN variant that ensures both physical consistency and generalization capability. Evaluated on multiple one-dimensional scalar test cases, the method achieves high-fidelity entropy solution approximation, with L² errors reduced by an order of magnitude compared to state-of-the-art high-resolution numerical schemes. It also markedly accelerates training convergence and enhances prediction robustness. These results establish a novel, scalable paradigm for reliably deploying deep learning in industrial-scale, complex hyperbolic systems.

0 citationsRead paper
Recent publications

Latest Papers

Expert Routing for Communication-Efficient MoE via Finite Expert Banks

May 06, 2026

This work investigates the information-theoretic efficiency of routing mechanisms in sparse Mixture-of-Experts (MoE) architectures, aiming to balance model accuracy with communication and computational resource utilization. The gating router is modeled as a stochastic channel, and a discrete mutual information estimator is proposed under a finite expert pool. Empirical posterior distributions \( q(W|S) \) are leveraged to compute \( I(X;T) \) and \( I(S;W) \), with the latter shown to exhibit a monotonic relationship with the generalization gap. The Blahut–Arimoto algorithm is employed to trace the accuracy–rate trade-off curve. Experiments demonstrate that the proposed mutual information estimator effectively tracks the generalization gap and significantly outperforms both the Xu–Raginsky bound and the uniform joint bound, offering a practical analytical tool for resource-aware MoE systems.

0 citationsRead paper

Wireless Broadcast Gossip for Decentralized Drone Swarms: Success Probability, Contraction, and Optimal Aloha

Mar 19, 2026

This work addresses the trade-off between communication reliability and convergence speed in decentralized drone swarms achieving consensus via gossip protocols in interference-limited wireless environments. Focusing on slotted Aloha-based wireless gossip, the authors model drone locations as a planar Poisson point process and derive a closed-form signal-to-interference ratio (SIR) success probability under Rayleigh fading channels. By decomposing the dynamics into an ideal mixing term and a wireless sparsity term, they establish a mean-square contraction bound. Building on this analysis, they propose a closed-form optimal transmission probability that jointly optimizes availability and reliability. The theoretical framework integrates stochastic geometry, mean-field approximation, and convex optimization, with the predicted optimal operating point aligning closely with the fastest convergence region observed in simulations, thereby validating the approach.

0 citationsRead paper

Mixture-of-Experts under Finite-Rate Gating: Communication--Generalization Trade-offs

Feb 16, 2026

This work investigates the trade-off between model expressivity and generalization performance in Mixture-of-Experts (MoE) architectures under communication constraints. For the first time, rate-distortion theory is introduced into MoE analysis by modeling the gating mechanism as a stochastic channel operating at a finite rate. By integrating mutual information-based generalization bounds with the rate-distortion function \(D(R_g)\), the study establishes a quantitative relationship between the gating communication rate and generalization error. A theoretical upper bound on generalization error is derived and validated through synthetic multi-expert model simulations, which demonstrate that reducing the gating rate, while limiting expressivity, can enhance generalization. Based on these insights, the paper proposes a capacity-aware design principle for MoE systems, offering theoretical guidance for efficient model construction in resource-constrained settings.

0 citationsRead paper

General Multi-User Distributed Computing

Nov 25, 2025

This paper addresses the joint optimization of computation, communication, and model accuracy in multi-user–multi-server distributed computing and inference. We propose GMUDC, a unified framework supporting personalized objective optimization over arbitrary network topologies and heterogeneous task allocations. Methodologically, we introduce spectral-covering duality—establishing, for the first time, a theoretical link between network topology and generalization capability. Leveraging reproducing kernel Hilbert spaces and translation-invariant kernels, we integrate quenched–annealed two-phase analysis to achieve topology-aware distributed optimization under resource constraints. We theoretically prove GMUDC’s verifiable efficiency under given computational and communication budgets. Our framework provides a foundational information–energy co-design principle for federated learning and edge intelligence.

0 citationsRead paper

From Initial Data to Boundary Layers: Neural Networks for Nonlinear Hyperbolic Conservation Laws

Jun 02, 2025

This work addresses the challenge of approximating entropy solutions to initial-boundary value problems for nonlinear strictly hyperbolic conservation laws. Methodologically, it introduces a physics-informed deep learning framework that—uniquely—jointly models initial data and boundary layers. The approach incorporates a boundary-layer-aware loss function, a feature-adaptive weighting scheme, and an entropy-condition regularization term, yielding a PINN variant that ensures both physical consistency and generalization capability. Evaluated on multiple one-dimensional scalar test cases, the method achieves high-fidelity entropy solution approximation, with L² errors reduced by an order of magnitude compared to state-of-the-art high-resolution numerical schemes. It also markedly accelerates training convergence and enhances prediction robustness. These results establish a novel, scalable paradigm for reliably deploying deep learning in industrial-scale, complex hyperbolic systems.

0 citationsRead paper