Institution profile

Universidad Pablo de Olavide

Academic institutioneurope · es
Official website
Research library36linked papers
Opportunities0open roles
Selected work

Representative Papers

Capacity Confounds and Coverage Guarantees in Adaptive Sub-model Federated Learning

Aug 07, 2026

Existing submodel-based federated learning approaches struggle to effectively leverage client data heterogeneity when allocating model capacity based on device resources and are susceptible to confounding effects between capacity allocation and data signals. This work proposes the HAS-FL framework, which for the first time reveals the strong interference of capacity allocation on heterogeneity estimation in submodel updates, identifies a hidden failure mode in adaptive allocation strategies, and introduces a parameter coverage guarantee mechanism to prevent uncovered parameters from degrading global model performance. Through corrected update divergence estimators, reproducible data partitions, and budget-matched experiments on image and text benchmarks, we demonstrate that adaptive allocation offers no advantage over random allocation in vision tasks and performs worst—while consuming the most capacity—in language tasks; overall model performance hinges primarily on parameter coverage rather than the allocation strategy itself.

0 citationsRead paper

From Language to Navigation Goals: A Vision-Language Approach for Semantic Navigation of Mobile Robots Using RGB-D Perception

Jul 15, 2026

This work addresses the challenge of enabling non-expert users to control mobile robot navigation through natural language commands. It proposes a modular, ROS 2–based language-driven navigation framework that integrates natural language understanding, RGB-D semantic perception, and Nav2 autonomous navigation to achieve end-to-end mapping from linguistic instructions to navigational goals. The system supports context-aware instruction parsing and cross-platform deployment. Validated on both TurtleBot3 Waffle and Unitree Go2 platforms, it accurately identifies linguistically referenced targets, estimates their spatial locations, generates feasible navigation paths, and provides natural language feedback. This approach significantly enhances the intuitiveness and robustness of human–robot interaction in real-world environments.

0 citationsRead paper

Navigating the Crowd: Non-linear MPC with Social Forces Dynamics for Human-Aware Robot Navigation

Jul 11, 2026

This work addresses the challenge of enabling autonomous robots to navigate crowded environments while simultaneously avoiding collisions and adhering to human social norms. The authors propose the SFM-NMPC framework, which uniquely integrates the Social Force Model (SFM) directly into the optimization loop of Nonlinear Model Predictive Control (NMPC). This integration allows for joint prediction of human and robot trajectories and leverages a socially aware cost function to generate navigation policies consistent with human behavioral conventions. The resulting approach achieves end-to-end socially compliant trajectory planning, running in real time at 20 Hz in dense simulated environments. Experimental results demonstrate significant improvements over existing methods in terms of social compliance, trajectory smoothness, and navigation efficiency.

0 citationsRead paper

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

Jul 08, 2026

This study addresses the challenge of accurately predicting turn-taking timing in mediated human–robot interaction. To this end, the authors propose the Multimodal Voice Activity Projection (MM-VAP) framework, which extends self-supervised voice activity projection to audiovisual modalities for the first time. The approach leverages pretrained encoders enhanced with low-rank adaptation (LoRA) for efficient turn-taking prediction, incorporating a cross-speaker attention mechanism and a semantic consistency loss to model high-order conversational dynamics within a 256-dimensional output space. Evaluated on the NoXi, NoXi+J, and Haru EDR datasets, MM-VAP significantly outperforms existing baselines, demonstrating particularly strong performance in predicting critical turn-taking events and validating its effectiveness in mediated human–robot dialogue scenarios.

0 citationsRead paper
Recent publications

Latest Papers

Capacity Confounds and Coverage Guarantees in Adaptive Sub-model Federated Learning

Aug 07, 2026

Existing submodel-based federated learning approaches struggle to effectively leverage client data heterogeneity when allocating model capacity based on device resources and are susceptible to confounding effects between capacity allocation and data signals. This work proposes the HAS-FL framework, which for the first time reveals the strong interference of capacity allocation on heterogeneity estimation in submodel updates, identifies a hidden failure mode in adaptive allocation strategies, and introduces a parameter coverage guarantee mechanism to prevent uncovered parameters from degrading global model performance. Through corrected update divergence estimators, reproducible data partitions, and budget-matched experiments on image and text benchmarks, we demonstrate that adaptive allocation offers no advantage over random allocation in vision tasks and performs worst—while consuming the most capacity—in language tasks; overall model performance hinges primarily on parameter coverage rather than the allocation strategy itself.

0 citationsRead paper

From Language to Navigation Goals: A Vision-Language Approach for Semantic Navigation of Mobile Robots Using RGB-D Perception

Jul 15, 2026

This work addresses the challenge of enabling non-expert users to control mobile robot navigation through natural language commands. It proposes a modular, ROS 2–based language-driven navigation framework that integrates natural language understanding, RGB-D semantic perception, and Nav2 autonomous navigation to achieve end-to-end mapping from linguistic instructions to navigational goals. The system supports context-aware instruction parsing and cross-platform deployment. Validated on both TurtleBot3 Waffle and Unitree Go2 platforms, it accurately identifies linguistically referenced targets, estimates their spatial locations, generates feasible navigation paths, and provides natural language feedback. This approach significantly enhances the intuitiveness and robustness of human–robot interaction in real-world environments.

0 citationsRead paper

Navigating the Crowd: Non-linear MPC with Social Forces Dynamics for Human-Aware Robot Navigation

Jul 11, 2026

This work addresses the challenge of enabling autonomous robots to navigate crowded environments while simultaneously avoiding collisions and adhering to human social norms. The authors propose the SFM-NMPC framework, which uniquely integrates the Social Force Model (SFM) directly into the optimization loop of Nonlinear Model Predictive Control (NMPC). This integration allows for joint prediction of human and robot trajectories and leverages a socially aware cost function to generate navigation policies consistent with human behavioral conventions. The resulting approach achieves end-to-end socially compliant trajectory planning, running in real time at 20 Hz in dense simulated environments. Experimental results demonstrate significant improvements over existing methods in terms of social compliance, trajectory smoothness, and navigation efficiency.

0 citationsRead paper

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

Jul 08, 2026

This study addresses the challenge of accurately predicting turn-taking timing in mediated human–robot interaction. To this end, the authors propose the Multimodal Voice Activity Projection (MM-VAP) framework, which extends self-supervised voice activity projection to audiovisual modalities for the first time. The approach leverages pretrained encoders enhanced with low-rank adaptation (LoRA) for efficient turn-taking prediction, incorporating a cross-speaker attention mechanism and a semantic consistency loss to model high-order conversational dynamics within a 256-dimensional output space. Evaluated on the NoXi, NoXi+J, and Haru EDR datasets, MM-VAP significantly outperforms existing baselines, demonstrating particularly strong performance in predicting critical turn-taking events and validating its effectiveness in mediated human–robot dialogue scenarios.

0 citationsRead paper