POLARIS: Training-Free Audio Fingerprinting with Saliency-Based Landmarks and Delaunay Grouping

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
POLARIS通过从局部归一化的显著性场选择地标并使用Delaunay三角剖分进行分组,实现无训练音频指纹识别,有效处理查询失真问题。
📝 Abstract
This work presents POLARIS, a training-free audio fingerprinting system that selects landmarks from a locally normalized saliency field and groups them into sparse fingerprints using Delaunay triangulation. To deal with query distortion, POLARIS adds fingerprints from two-hop Delaunay neighborhoods only at query time, without enlarging the reference index. An adaptive configuration applies this expansion only when the original fingerprints do not produce a confident match. We evaluate POLARIS on synthetic distortions from the public PEX Hard Medium benchmark, excluding queries with pitch or tempo shifts, and on a new benchmark of real re-recorded music. POLARIS achieves the best performance among the evaluated training-free methods on both benchmarks. On the real recordings, its adaptive configuration also outperforms the neural NMFP baseline with a comparable measured query time and a smaller logical reference payload. Code, dataset, and instructions for reproducing all experiments are available at https://github.com/JihengLi/POLARIS.git.
Problem

Research questions and friction points this paper is trying to address.

Audio Fingerprinting
Training-Free
Query Distortion
Innovation

Methods, ideas, or system contributions that make the work stand out.

training-free audio fingerprinting
saliency-based landmarks
Delaunay triangulation
adaptive configuration
query distortion handling
🔎 Similar Papers
No similar papers found.
J
Jiheng Li
Department of Computer Science, Vanderbilt University, Nashville, TN, 37235