Institution profile

Laboratoire Jean Kuntzmann

Academic institutioneurope · fr
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?

Aug 13, 2026

This work addresses critical limitations in existing audio-visual social understanding benchmarks, which suffer from high noise levels and poorly designed questions, while complex reasoning approaches often incur substantial costs with marginal gains. The authors systematically evaluate the reasoning capabilities of multimodal large language models on social audio-visual question answering, introducing IntentBench-Prime—a high-quality benchmark constructed through rigorous data cleaning—and comparing diverse training strategies. Their findings reveal that a simple vanilla supervised fine-tuning (SFT) baseline matches or surpasses state-of-the-art complex methods across three benchmarks. Notably, using only textual captions achieves performance comparable to full video inputs, suggesting that linguistic modalities encode strong social priors. The study further proposes a cost-effective evaluation paradigm and publicly releases the denoised IntentBench-Prime benchmark to support future research.

0 citationsRead paper

When Imbalance Comes Twice: Active Learning under Simulated Class Imbalance and Label Shift in Binary Semantic Segmentation

Jan 08, 2026arXiv.org

This study addresses the degradation of active learning performance in binary semantic segmentation caused by the coexistence of class imbalance and label shift. For the first time, it systematically simulates both challenges jointly on open-source datasets to evaluate the effectiveness of three active learning strategies: random sampling, entropy maximization, and core-set selection. Experimental results demonstrate that entropy-based and core-set methods remain robust under severe class imbalance; however, strong label shift significantly impairs their performance. By revealing distinct behavioral patterns of these strategies under compound distribution shifts, this work provides critical insights for deploying active learning in real-world scenarios where multiple data biases may co-occur.

0 citationsRead paper
Recent publications

Latest Papers

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?

Aug 13, 2026

This work addresses critical limitations in existing audio-visual social understanding benchmarks, which suffer from high noise levels and poorly designed questions, while complex reasoning approaches often incur substantial costs with marginal gains. The authors systematically evaluate the reasoning capabilities of multimodal large language models on social audio-visual question answering, introducing IntentBench-Prime—a high-quality benchmark constructed through rigorous data cleaning—and comparing diverse training strategies. Their findings reveal that a simple vanilla supervised fine-tuning (SFT) baseline matches or surpasses state-of-the-art complex methods across three benchmarks. Notably, using only textual captions achieves performance comparable to full video inputs, suggesting that linguistic modalities encode strong social priors. The study further proposes a cost-effective evaluation paradigm and publicly releases the denoised IntentBench-Prime benchmark to support future research.

0 citationsRead paper

When Imbalance Comes Twice: Active Learning under Simulated Class Imbalance and Label Shift in Binary Semantic Segmentation

Jan 08, 2026arXiv.org

This study addresses the degradation of active learning performance in binary semantic segmentation caused by the coexistence of class imbalance and label shift. For the first time, it systematically simulates both challenges jointly on open-source datasets to evaluate the effectiveness of three active learning strategies: random sampling, entropy maximization, and core-set selection. Experimental results demonstrate that entropy-based and core-set methods remain robust under severe class imbalance; however, strong label shift significantly impairs their performance. By revealing distinct behavioral patterns of these strategies under compound distribution shifts, this work provides critical insights for deploying active learning in real-world scenarios where multiple data biases may co-occur.

0 citationsRead paper