Panda: Self-distillation of Reusable Sensor-level Representations for High Energy Physics

πŸ“… 2025-12-01
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Liquid argon time projection chambers (LArTPCs) rely heavily on large volumes of labeled simulation data and precise detector calibration for particle reconstruction, limiting scalability and realism. Method: We propose an unsupervised sensor-level representation learning framework for LArTPCs, employing a hierarchical sparse 3D encoder coupled with multi-view prototype self-distillation and an ensemble prediction headβ€”trained end-to-end without physics priors. Contribution/Results: Our self-distillation-driven prototype clustering learns transferable, sparse 3D feature representations across downstream tasks. Experiments show that with only 0.1% labeled data, our method surpasses state-of-the-art semantic segmentation models. With the backbone frozen, a lightweight prediction head achieves particle identification accuracy comparable to mainstream reconstruction tools. Full fine-tuning further yields significant gains in both reconstruction and classification performance, substantially reducing dependence on simulated labels and fine-grained calibration.

Technology Category

Application Category

πŸ“ Abstract
Liquid argon time projection chambers (LArTPCs) provide dense, high-fidelity 3D measurements of particle interactions and underpin current and future neutrino and rare-event experiments. Physics reconstruction typically relies on complex detector-specific pipelines that use tens of hand-engineered pattern recognition algorithms or cascades of task-specific neural networks that require extensive, labeled simulation that requires a careful, time-consuming calibration process. We introduce extbf{Panda}, a model that learns reusable sensor-level representations directly from raw unlabeled LArTPC data. Panda couples a hierarchical sparse 3D encoder with a multi-view, prototype-based self-distillation objective. On a simulated dataset, Panda substantially improves label efficiency and reconstruction quality, beating the previous state-of-the-art semantic segmentation model with 1,000$ imes$ fewer labels. We also show that a single set-prediction head 1/20th the size of the backbone with no physical priors trained on frozen outputs from Panda can result in particle identification that is comparable with state-of-the-art (SOTA) reconstruction tools. Full fine-tuning further improves performance across all tasks.
Problem

Research questions and friction points this paper is trying to address.

Learns reusable sensor-level representations from raw LArTPC data
Improves label efficiency and reconstruction quality in particle detection
Reduces reliance on complex, hand-engineered pipelines and labeled simulations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-distillation from raw unlabeled LArTPC data
Hierarchical sparse 3D encoder with multi-view prototypes
Small set-prediction head on frozen representations for particle ID
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.