Institution profile

Ecole Nationale Supérieure en Electrotechnique, Electronique, Informatique et Hydraulique de Toulouse

Academic institutioneurope · fr
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Automatic Extraction of Structured Information from Brain MRI Reports Using an Open-Weight Large Language Model

Jun 05, 2026

This study addresses the challenge of automatically extracting structured clinical information from Dutch-language neuroradiology reports, a key bottleneck for large-scale brain MRI research. It presents the first systematic evaluation of the open-source large language model LLaMA 3.1 for this task, introducing a structure-similarity-based strategy for few-shot example selection. By combining multilingual inputs—original Dutch reports and their English translations—with tailored prompt engineering, the approach extracts 30 key variables from 947 reports. The model achieves high accuracy on visual rating scales (e.g., 94% for Fazekas score) and detection of microbleed mentions (93%). Few-shot prompting substantially improves extraction of numeric variables such as microbleed counts (92% accuracy), though localization-related variables remain challenging.

0 citationsRead paper

KNN and K-means in Gini Prametric Spaces

Jan 29, 2025

K-means and k-nearest neighbors (k-NN) suffer from sensitivity to noise and outliers. To address this, we propose a novel clustering and classification framework grounded in Gini-based quasi-metric spaces. Our method introduces: (1) a new Gini-based distance metric that jointly captures numerical dissimilarity and ordinal structure; (2) theoretical convergence guarantees for the proposed Gini K-means algorithm; and (3) a rank–value integrated k-NN classifier leveraging the Gini distance. Extensive experiments on 14 UCI benchmark datasets demonstrate that our approach consistently outperforms standard K-means/k-NN and state-of-the-art robust alternatives—including Hassanat distance—on both clustering and classification tasks, while maintaining competitive computational efficiency.

0 citationsRead paper
Recent publications

Latest Papers

Automatic Extraction of Structured Information from Brain MRI Reports Using an Open-Weight Large Language Model

Jun 05, 2026

This study addresses the challenge of automatically extracting structured clinical information from Dutch-language neuroradiology reports, a key bottleneck for large-scale brain MRI research. It presents the first systematic evaluation of the open-source large language model LLaMA 3.1 for this task, introducing a structure-similarity-based strategy for few-shot example selection. By combining multilingual inputs—original Dutch reports and their English translations—with tailored prompt engineering, the approach extracts 30 key variables from 947 reports. The model achieves high accuracy on visual rating scales (e.g., 94% for Fazekas score) and detection of microbleed mentions (93%). Few-shot prompting substantially improves extraction of numeric variables such as microbleed counts (92% accuracy), though localization-related variables remain challenging.

0 citationsRead paper

KNN and K-means in Gini Prametric Spaces

Jan 29, 2025

K-means and k-nearest neighbors (k-NN) suffer from sensitivity to noise and outliers. To address this, we propose a novel clustering and classification framework grounded in Gini-based quasi-metric spaces. Our method introduces: (1) a new Gini-based distance metric that jointly captures numerical dissimilarity and ordinal structure; (2) theoretical convergence guarantees for the proposed Gini K-means algorithm; and (3) a rank–value integrated k-NN classifier leveraging the Gini distance. Extensive experiments on 14 UCI benchmark datasets demonstrate that our approach consistently outperforms standard K-means/k-NN and state-of-the-art robust alternatives—including Hassanat distance—on both clustering and classification tasks, while maintaining competitive computational efficiency.

0 citationsRead paper