MarkNull: Model-Agnostic Watermark Removal in AI-Generated Images via On-Manifold Latent Manipulation

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited robustness of existing AI image watermarking methods under model-agnostic attacks, which often compromise visual quality or are restricted to specific generative models. We propose a universal watermark removal approach based on latent space manipulation within the data manifold. By analyzing the statistical dependency between the latent representation of watermarked images and the initial noise, we formulate an optimization objective that disentangles watermark information while preserving semantic fidelity. To quantify this dependency, we introduce the Noise-Latent Alignment Score and develop an efficient framework, MarkNull, along with a fast, optimization-free variant, MarkNull-A, complemented by a dedicated detection mechanism. Our method reduces average bit accuracy to 53.14%—near random chance—across diverse watermarking schemes, successfully breaks Google’s SynthID-Image, achieves a processing speed of 0.50 seconds per image without perceptible distortion, and extends naturally to video domains.
📝 Abstract
Digital watermarking has emerged as a critical technique for provenance and copyright attribution in AI-generated imagery, yet its robustness against realistic, model-agnostic removal attacks remains poorly explored. Existing attacks either succeed only against specific generative models or achieve removal at the cost of severe visual degradation. In this paper, we propose MarkNull, a model-agnostic watermark removal attack via on-manifold latent manipulation. MarkNull is grounded in a key observation: watermarked images exhibit a strong statistical dependency between the generated latent representation and the embedded initial noise. To quantify this dependency, we introduce the Noise-Latent Alignment Score (NLAS) and formulate an optimization objective that selectively decorrelates the latent representation from the embedded watermark while preserving semantic fidelity. Extensive evaluations across different categories of watermarking paradigms, including post-hoc, fine-tuning-based, and initial-noise-based schemes, demonstrate that MarkNull reduces average bit accuracy to 53.14%, approaching random-guessing (50%), without perceptible image degradation. To further improve scalability, we propose MarkNull-A, an amortized, optimization-free variant that distills the attack into a single forward pass, achieving 0.50 s/image with modest computational overhead. Notably, our attacks successfully compromise Google's SynthID-Image system while preserving high visual quality and transfer effectively to video watermarking. Finally, we present an attack detection mechanism as a defensive counterpart to MarkNull and MarkNull-A, highlighting the necessity of developing watermark designs resilient to model-agnostic latent-space attacks.
Problem

Research questions and friction points this paper is trying to address.

watermark removal
AI-generated images
model-agnostic
latent manipulation
digital watermarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

model-agnostic
on-manifold latent manipulation
watermark removal
Noise-Latent Alignment Score
amortized attack
💼 Related Jobs
No related jobs found.