Efficient Human-Contact Representation for Human-Scene Interaction
This work addresses the challenges of high-dimensional input redundancy and low computational efficiency in existing methods for modeling human-scene interaction. To overcome these limitations, the authors propose an efficient representation based on sparse contact masks that retain only the critical contact regions, along with dedicated sparse neural network operators designed to replace conventional dense operations. This approach substantially reduces data redundancy while achieving state-of-the-art reconstruction accuracy across three public benchmarks. Moreover, it delivers a computational speedup of at least 12× compared to baseline methods, effectively balancing reconstruction fidelity with computational efficiency.