Towards Unified Dynamic Face Landmark Detection

πŸ“… 2026-08-10
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limitations of conventional facial landmark detection methods, which require separate models for different β€œN-point” datasets and can only predict a fixed number of landmarks, lacking both universality and flexibility. To overcome this, the authors propose a unified dynamic framework that introduces Face Part-Anchored Landmark Positions (FPALP), a novel representation that models arbitrary landmarks as normalized positions along facial contours. By integrating query-based dynamic prediction with a cross-modal decoder, the framework enables end-to-end regression of any number of landmarks. It supports joint training across multiple datasets and allows on-demand inference of user-specified landmarks. Extensive experiments demonstrate that the method achieves state-of-the-art or comparable performance across multiple benchmarks, significantly enhancing model generalizability and deployment flexibility.
πŸ“ Abstract
Although advancements in face landmark detection (FLD) methods continue to push performance boundaries, they overlook two major functional limitations: (1) different network parameters need to be trained independently for each ``$N$-point'' benchmark dataset, and (2) a model trained on an ``$N$-point'' dataset reliably outputs only the $N$ landmarks. In our work, we first conceptualize Face Part-Anchored Landmark Positions (FPALPs), wherein each landmark is treated as a progression value between zero (start) and one (end) along a face part's contour. Every landmark can be expressed in the FPALP format, irrespective of its source dataset, hence unlocking the ability to unify all ``$N$-point'' datasets into a single dataset. Secondly, we represent each landmark with an FPALP-based query, refine it progressively with a cross-modality decoder, and predict its coordinates based on the final representation. Our approach, called Unified Dynamic FLD, embodies these two design choices and streamlines the landmark detection pipeline by enabling (1) a single model to learn on any number of ``$N$-point'' datasets, and (2) yield any number of specific landmark predictions by loading the designated landmark queries at runtime. Extensive experiments on multiple benchmark datasets show that our method delivers these benefits while remaining competitive with, and in several cases outperforming existing state-of-the-art methods.
Problem

Research questions and friction points this paper is trying to address.

face landmark detection
N-point datasets
unified model
dynamic prediction
landmark generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Face Landmark Detection
Unified Representation
Dynamic Prediction
FPALP
Cross-modality Decoder
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
S
Sebastian Regalado
University of Toronto
V
Varshanth R. Rao
ModiFace
R
Ruowei Jiang
ModiFace
P
Parham Aarabi
University of Toronto
Igor Gilitschenski
Igor Gilitschenski
Assistant Professor, University of Toronto
RoboticsMachine LearningComputer Vision