Institution profile

École de technologie supérieure

Academic institutionnorthamerica · ca
Official website
Research library285linked papers
Opportunities0open roles
Selected work

Representative Papers

Weakly Supervised Learning for Facial Behavior Analysis : A Review

Jan 25, 2021arXiv.org

Facial behavior analysis faces significant weakly supervised learning challenges, including high annotation costs, reliance on domain experts, ambiguous intensity labeling, and expert bias. To address these issues, this paper presents a systematic survey of weakly supervised learning methods for facial expression recognition and action unit detection in real-world scenarios. We propose the first unified weak supervision taxonomy encompassing both categorical and dimensional labels, explicitly characterizing core challenges such as label ambiguity and intensity bias. Methodologically, we introduce an integrated framework combining label-noise-robust training, multiple-instance learning, bag-level ranking supervision, self-training, and consistency regularization. Extensive evaluation—standardized across 12+ benchmark datasets and 30+ baseline methods—demonstrates the effectiveness and generalizability of our approach. Our work advances the practical deployment of facial behavior models under limited or imperfect supervision.

6 citationsRead paper

Generalizable and Explainable Deep Learning for Medical Image Computing: An Overview

Nov 01, 2024Current Opinion in Biomedical Engineering

Clinical deployment of AI in medical imaging is hindered by poor generalizability across devices, institutions, and diseases, alongside insufficient interpretability and decision transparency. Method: We propose the first unified framework that jointly integrates domain generalization and self-supervised pretraining—enhancing cross-domain robustness—with concept bottleneck models, disentangled attention, counterfactual reasoning, and uncertainty quantification—to improve decision interpretability and trustworthiness. Contribution/Results: We systematically survey over 100 state-of-the-art works to clarify technical evolution and clinical translation bottlenecks. We introduce a novel, clinically grounded evaluation paradigm for trustworthy AI, spanning four orthogonal dimensions: performance, robustness, interpretability, and uncertainty. Our framework provides both a methodological foundation and a reproducible implementation roadmap for deploying reliable, clinically viable AI systems in medical imaging.

3 citationsRead paper

Words Matter: Leveraging Individual Text Embeddings for Code Generation in CLIP Test-Time Adaptation

Nov 26, 2024arXiv.org

To address the severe performance degradation of vision-language models (e.g., CLIP) under distribution shift during testing, this paper proposes a fine-tuning-free test-time adaptation method. The approach treats predefined class text embeddings as fixed semantic centroids and formulates label assignment as an optimal transport problem to generate high-quality pseudo-labels. Additionally, it introduces a multi-template knowledge distillation mechanism that effectively emulates multi-view contrastive learning at zero additional computational cost. This work is the first to explicitly model text embeddings as fixed semantic centroids for test-time adaptation. Evaluated on multiple standard benchmarks, the method achieves an average accuracy improvement of 7%, significantly outperforming existing state-of-the-art methods, while maintaining minimal computational and memory overhead.

1 citations1 influentialRead paper

Are foundation models for computer vision good conformal predictors?

Dec 08, 2024arXiv.org

This study systematically evaluates the uncertainty calibration capabilities of vision and multimodal foundation models within the conformal prediction (CP) framework, with emphasis on risk-sensitive applications. We examine prevalent Vision Transformer (ViT) architectures, three canonical CP methods—Adaptive Prediction Sets (APS), Split CP, and Adaptive CP—and multiple image classification benchmarks. Our key contributions are: (i) ViT-based models exhibit inherent compatibility with CP, achieving well-calibrated uncertainty estimates without retraining; (ii) adapter-based fine-tuning substantially outperforms prompt learning for CP adaptation; (iii) APS achieves the optimal trade-off between theoretical guarantees and empirical performance—strictly maintaining marginal coverage while yielding more compact prediction sets; and (iv) post-hoc confidence calibration degrades Adaptive CP’s efficacy. Collectively, these findings demonstrate that modern foundation models possess strong inherent conformalizability, offering a robust pathway for uncertainty quantification in high-stakes visual recognition tasks.

1 citationsRead paper
Recent publications

Latest Papers