🤖 AI Summary
This study addresses the challenges of score interpretability and strong energy dependence in collider anomaly detection by proposing the ORCA framework. Integrating supervised contrastive learning with autoencoders, this method constructs a geometric embedding space to generate anomaly scores while enabling template-fitting attribution and uncertainty quantification. Experimental results demonstrate that ORCA significantly enhances sensitivity to new physics searches, accurately recovers signal yields, and effectively characterizes unknown signal features. Consequently, this work establishes a novel paradigm for high-energy physics anomaly detection that successfully combines high performance with robust interpretability, overcoming limitations inherent in previous approaches.
📝 Abstract
Generic event-level anomaly detection for collider physics has two recurring problems: anomaly scores are hard to interpret, and they correlate strongly with energy scale and object multiplicity. We present Organized Representation via Contrastive learning for Anomaly detection (ORCA), a two-stage framework that first learns an embedding space via supervised contrastive learning across a diverse set of physics processes, then runs a standard autoencoder in that space to generate event-level anomaly scores. On a simulated dataset consistent with conditions at the High-Luminosity Large Hadron Collider, ORCA delivers significant gains in both breadth and depth of sensitivity to new physics signals with respect to a baseline autoencoder architecture. Beyond improved sensitivity, the contrastive embedding makes the anomalous sample interpretable: because known processes occupy distinct regions of the space, a maximum-likelihood template fit to the embedding distributions can attribute events in an anomalous sample to template physics processes with quantified uncertainties. We demonstrate that the fit accurately recovers injected signal yields, including for signals excluded from the training of the embedding, and characterizes signals absent from the template library through the known processes they most resemble. These results establish ORCA as a route to interpretable anomaly detection-based searches at colliders, where the embedding geometry carries higher dimensional physics information compared to standard one-dimensional output fits, enhancing downstream statistical analysis.