Institution profile

University of Toulouse 2 Jean Jaurès

Academic institutioneurope · fr
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Encryption-Compatible Clustered Federated Learning via Distributed Expectation-Maximization over Metadata

Jul 30, 2026

This work addresses the trilemma in clustered federated learning—where privacy preservation, communication efficiency, and computational scalability are difficult to achieve simultaneously—by proposing an encryption-compatible, highly efficient solution. The approach reformulates metadata-driven clustering as a distributed expectation-maximization (EM) process that relies solely on additive operations at the server, thereby enabling seamless integration with practical privacy-enhancing technologies such as additive homomorphic encryption. For the first time, this method breaks the CFL trilemma without compromising efficiency, significantly improving client model performance across diverse heterogeneous datasets while simultaneously ensuring strong privacy guarantees, low computational overhead, and high communication efficiency.

0 citationsRead paper

D4R -- Exploring and Querying Relational Graphs Using Natural Language and Large Language Models -- the Case of Historical Documents

Mar 26, 2025

To address the challenge non-technical users—particularly historians—face in efficiently querying relational knowledge graphs, this paper proposes an LLM-driven end-to-end framework for natural language-to-graph query translation. Methodologically, it tightly couples large language models with domain-specific historical texts to automatically parse natural language questions into precise Cypher queries for retrieval against a Neo4j knowledge graph; it further introduces a lightweight web-based graphical interface enabling dynamic discovery and visual analytics of entities, events, and spatiotemporal relationships. The key contribution lies in the first implementation of three-layer alignment—semantic, syntactic, and graph-structural—tailored to historical scholarship, thereby balancing domain specificity with low-barrier interactivity. Experimental evaluation on real historical source materials achieves 86.3% query accuracy and demonstrates promising cross-domain transferability.

0 citationsRead paper

Dense Retrieval for Low Resource Languages -- the Case of Amharic Language

Mar 24, 2025

This work addresses core challenges in dense retrieval for Amharic—a low-resource language with 120 million speakers—including scarcity of labeled data, pretraining resources, and word embeddings. We present the first systematic feasibility study, proposing a lightweight fine-tuning and cross-lingual transfer framework built upon mBERT and XLM-R. Our approach integrates contrastive learning, pseudo-labeling, and unsupervised domain adaptation to train dense encoders. Evaluated on a newly constructed Amharic QA retrieval benchmark—the first of its kind—we achieve a 37% improvement in Recall@10 over baseline methods, substantially outperforming traditional sparse retrieval and zero-shot cross-lingual baselines. Key contributions are: (1) the first publicly available Amharic dense retrieval benchmark; (2) empirical validation of lightweight adaptation and cross-lingual transfer efficacy in low-resource settings; and (3) a reproducible methodology for information retrieval in African languages.

0 citationsRead paper
Recent publications

Latest Papers

Encryption-Compatible Clustered Federated Learning via Distributed Expectation-Maximization over Metadata

Jul 30, 2026

This work addresses the trilemma in clustered federated learning—where privacy preservation, communication efficiency, and computational scalability are difficult to achieve simultaneously—by proposing an encryption-compatible, highly efficient solution. The approach reformulates metadata-driven clustering as a distributed expectation-maximization (EM) process that relies solely on additive operations at the server, thereby enabling seamless integration with practical privacy-enhancing technologies such as additive homomorphic encryption. For the first time, this method breaks the CFL trilemma without compromising efficiency, significantly improving client model performance across diverse heterogeneous datasets while simultaneously ensuring strong privacy guarantees, low computational overhead, and high communication efficiency.

0 citationsRead paper

D4R -- Exploring and Querying Relational Graphs Using Natural Language and Large Language Models -- the Case of Historical Documents

Mar 26, 2025

To address the challenge non-technical users—particularly historians—face in efficiently querying relational knowledge graphs, this paper proposes an LLM-driven end-to-end framework for natural language-to-graph query translation. Methodologically, it tightly couples large language models with domain-specific historical texts to automatically parse natural language questions into precise Cypher queries for retrieval against a Neo4j knowledge graph; it further introduces a lightweight web-based graphical interface enabling dynamic discovery and visual analytics of entities, events, and spatiotemporal relationships. The key contribution lies in the first implementation of three-layer alignment—semantic, syntactic, and graph-structural—tailored to historical scholarship, thereby balancing domain specificity with low-barrier interactivity. Experimental evaluation on real historical source materials achieves 86.3% query accuracy and demonstrates promising cross-domain transferability.

0 citationsRead paper

Dense Retrieval for Low Resource Languages -- the Case of Amharic Language

Mar 24, 2025

This work addresses core challenges in dense retrieval for Amharic—a low-resource language with 120 million speakers—including scarcity of labeled data, pretraining resources, and word embeddings. We present the first systematic feasibility study, proposing a lightweight fine-tuning and cross-lingual transfer framework built upon mBERT and XLM-R. Our approach integrates contrastive learning, pseudo-labeling, and unsupervised domain adaptation to train dense encoders. Evaluated on a newly constructed Amharic QA retrieval benchmark—the first of its kind—we achieve a 37% improvement in Recall@10 over baseline methods, substantially outperforming traditional sparse retrieval and zero-shot cross-lingual baselines. Key contributions are: (1) the first publicly available Amharic dense retrieval benchmark; (2) empirical validation of lightweight adaptation and cross-lingual transfer efficacy in low-resource settings; and (3) a reproducible methodology for information retrieval in African languages.

0 citationsRead paper