🤖 AI Summary
This work addresses the challenge that independently trained visual watermarks often cause uncontrolled interference when coexisting, degrading both decoding robustness and visual quality. For the first time, watermark coexistence is explicitly formulated as an optimization objective rather than treated as an incidental phenomenon. The authors propose a decoder-aware joint training framework that enables controlled superposition of multi-layer watermarks in images and videos through end-to-end differentiable embedding and extraction networks, complemented by tailored loss functions. A novel “landmark watermark” is introduced to signal the presence of a multi-source provenance system. Experimental results demonstrate that the proposed method significantly improves decoding accuracy and visual fidelity under multi-watermark coexistence, thereby validating the feasibility of hierarchical content provenance for both images and video.
📝 Abstract
We present a method for training imperceptible visual watermarks to coexist with other such watermarks. Recent work has shown that independently trained image watermarking models can coexist with surprisingly limited interference, enabling watermark ensembling. However, this coexistence is a serendipitous property rather than an explicit optimization objective, leaving interference uncontrolled and potentially reducing decoding robustness or visual quality. We first show empirically that the same coexistence property extends to video watermarking. We then show that both image and video watermarks can be trained with a decoder-aware objective to improve coexistence. Our results suggest a practical path to signpost watermarks that indicate the presence of independently deployed provenance watermarking systems, supporting layered provenance signaling for content authenticity and rights.