Institution profile

Laboratoire Hubert Curien

Academic institutioneurope · fr
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

Aug 13, 2026

This study investigates how instruction tuning influences confidence expression and lexical diversity in the generated rationales of language models on question-answering tasks. By systematically comparing matched base and instruction-tuned models across multiple QA benchmarks—using confidence calibration metrics, lexical diversity measures, and controlled analyses—it reveals that instruction tuning consistently induces overconfidence and degrades likelihood calibration, even when accuracy gains are negligible. Furthermore, it significantly reduces semantic diversity in rationales across samples, while surface-level lexical diversity exhibits inconsistent changes. These findings highlight previously underappreciated, implicit effects of instruction tuning on model reasoning behavior, offering a new perspective for its refinement and evaluation.

0 citationsRead paper

Beyond Arrow's Impossibility: Fairness as an Emergent Property of Multi-Agent Collaboration

Apr 15, 2026

This study challenges the conventional view of fairness as a static, individual attribute in language models, which proves inadequate for addressing dynamic ethical conflicts in multi-agent interactions. The authors propose that fairness emerges as a procedural property through collaborative deliberation among agents. To investigate this, they design a structured three-round, two-agent debate framework within a hospital triage scenario, integrating retrieval-augmented generation (RAG) to enable ethically aligned and controlled simulation experiments. Their findings reveal that individual agents consistently violate ethical allocation principles, whereas joint decisions reached through adversarial negotiation satisfy fairness criteria unattainable by any single model. Furthermore, aligned agents partially restore equity for marginalized groups while exposing inherent framing biases. This work pioneers the application of Arrow’s Impossibility Theorem to AI fairness, underscoring negotiation—not override—as central to effective bias mitigation.

0 citationsRead paper

Speech transformer models for extracting information from baby cries

Sep 02, 2025

Infant cry analysis faces challenges including acoustic instability, data scarcity, and infant identity disambiguation—tasks poorly addressed by conventional speech models trained exclusively on linguistic signals. Method: This work systematically investigates the transferability and representational properties of Transformer-based pretrained speech models on non-speech infant cry signals. We evaluate multiple models across eight diverse datasets comprising 115 hours of audio from 960 infants, targeting cry classification, vocalization characteristic modeling, and infant identity recognition. Contribution/Results: We demonstrate that pretrained speech representations effectively encode physiological state and speaker-identity information in cries. Model architecture and pretraining strategy critically influence cross-domain generalization. Crucially, our findings reveal that speech self-supervised models implicitly learn acoustic–physiological mappings—capturing biologically grounded structure beyond phonetic content. This provides an interpretable representation foundation and a novel model design paradigm for affective computing and early-life health monitoring.

0 citationsRead paper

When Dance Video Archives Challenge Computer Vision

May 12, 2025

To address the low accuracy and poor robustness of 3D human pose estimation in dance videos—caused by high-dynamic motion, frequent occlusions, and stylized choreography—this work introduces the first end-to-end 3D pose estimation pipeline tailored for dance archival footage. Methodologically, it systematically integrates state-of-the-art monocular 3D pose estimators (e.g., VideoPose3D, PoseFormer) with a customized post-processing module incorporating motion continuity constraints and occlusion-aware reweighting, coupled with an interpretable visualization toolkit. Extensive experiments on a large-scale, multi-genre archival dataset—including ballet, modern dance, and folk dance—reveal that clothing complexity, camera viewpoint, and motion amplitude significantly impact estimation error (increasing mean per-joint position error [MPJPE] by 12.7% on average); our approach reduces MPJPE by 18.3% over baseline models. The code, annotated dataset, and evaluation benchmark are publicly released.

0 citationsRead paper
Recent publications

Latest Papers

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

Aug 13, 2026

This study investigates how instruction tuning influences confidence expression and lexical diversity in the generated rationales of language models on question-answering tasks. By systematically comparing matched base and instruction-tuned models across multiple QA benchmarks—using confidence calibration metrics, lexical diversity measures, and controlled analyses—it reveals that instruction tuning consistently induces overconfidence and degrades likelihood calibration, even when accuracy gains are negligible. Furthermore, it significantly reduces semantic diversity in rationales across samples, while surface-level lexical diversity exhibits inconsistent changes. These findings highlight previously underappreciated, implicit effects of instruction tuning on model reasoning behavior, offering a new perspective for its refinement and evaluation.

0 citationsRead paper

Beyond Arrow's Impossibility: Fairness as an Emergent Property of Multi-Agent Collaboration

Apr 15, 2026

This study challenges the conventional view of fairness as a static, individual attribute in language models, which proves inadequate for addressing dynamic ethical conflicts in multi-agent interactions. The authors propose that fairness emerges as a procedural property through collaborative deliberation among agents. To investigate this, they design a structured three-round, two-agent debate framework within a hospital triage scenario, integrating retrieval-augmented generation (RAG) to enable ethically aligned and controlled simulation experiments. Their findings reveal that individual agents consistently violate ethical allocation principles, whereas joint decisions reached through adversarial negotiation satisfy fairness criteria unattainable by any single model. Furthermore, aligned agents partially restore equity for marginalized groups while exposing inherent framing biases. This work pioneers the application of Arrow’s Impossibility Theorem to AI fairness, underscoring negotiation—not override—as central to effective bias mitigation.

0 citationsRead paper

Speech transformer models for extracting information from baby cries

Sep 02, 2025

Infant cry analysis faces challenges including acoustic instability, data scarcity, and infant identity disambiguation—tasks poorly addressed by conventional speech models trained exclusively on linguistic signals. Method: This work systematically investigates the transferability and representational properties of Transformer-based pretrained speech models on non-speech infant cry signals. We evaluate multiple models across eight diverse datasets comprising 115 hours of audio from 960 infants, targeting cry classification, vocalization characteristic modeling, and infant identity recognition. Contribution/Results: We demonstrate that pretrained speech representations effectively encode physiological state and speaker-identity information in cries. Model architecture and pretraining strategy critically influence cross-domain generalization. Crucially, our findings reveal that speech self-supervised models implicitly learn acoustic–physiological mappings—capturing biologically grounded structure beyond phonetic content. This provides an interpretable representation foundation and a novel model design paradigm for affective computing and early-life health monitoring.

0 citationsRead paper

When Dance Video Archives Challenge Computer Vision

May 12, 2025

To address the low accuracy and poor robustness of 3D human pose estimation in dance videos—caused by high-dynamic motion, frequent occlusions, and stylized choreography—this work introduces the first end-to-end 3D pose estimation pipeline tailored for dance archival footage. Methodologically, it systematically integrates state-of-the-art monocular 3D pose estimators (e.g., VideoPose3D, PoseFormer) with a customized post-processing module incorporating motion continuity constraints and occlusion-aware reweighting, coupled with an interpretable visualization toolkit. Extensive experiments on a large-scale, multi-genre archival dataset—including ballet, modern dance, and folk dance—reveal that clothing complexity, camera viewpoint, and motion amplitude significantly impact estimation error (increasing mean per-joint position error [MPJPE] by 12.7% on average); our approach reduces MPJPE by 18.3% over baseline models. The code, annotated dataset, and evaluation benchmark are publicly released.

0 citationsRead paper