🤖 AI Summary
Existing text-to-image diffusion models often disregard physical constraints in projection mapping, leading to generated content that misaligns with real 3D scenes. To address this, this work proposes ConPhyG, a framework that formalizes two complementary generation paradigms: a cooperative mode that guides the diffusion process using pixel-level physical priors—such as depth, edges, and color gamut—and an adversarial mode that employs bounded numerical optimization to achieve radiometric compensation across multiple projectors, with dynamic switching between modes to balance artistic freedom and physical plausibility. ConPhyG further introduces a sequential generation strategy to ensure 360-degree multi-view consistent, physics-aware image synthesis. Experiments on a real four-projector system demonstrate that ConPhyG significantly outperforms existing methods, achieving state-of-the-art performance in geometric alignment, color gamut utilization, and semantic fidelity.
📝 Abstract
Projection Mapping (PM) enables seamless superimposition of digital content onto real-world 3D objects, serving as a fundamental technique for immersive visualization, digital twins, and interactive art. Although text-to-image diffusion models have greatly facilitated customized content creation, directly integrating them into practical PM pipelines remains challenging due to the mismatch between idealized 2D generation and physical constraints. To bridge this gap, this paper formalizes two application-level generative paradigms: the cooperative paradigm (harmonizing generated semantics with physical attributes) and the adversarial paradigm (eliminating surface interference via radiometric compensation). Based on this, we propose ConPhyG, a unified controllable physically-guided generative multi-projection mapping framework that enables creators to interactively adjust physical constraints and flexibly switch generative paradigms. In cooperative mode, multi-dimensional physical priors (per-pixel gamut, depth, and edges) are injected into the diffusion process. In adversarial mode, the framework releases the generative potential and applies bounded numerical optimization for multi-projector radiometric compensation. It allows users to dynamically switch constraints to balance artistic freedom with physical feasibility. Furthermore, we extend ConPhyG to 360-degree multi-view consistent PM using a sequential generation strategy. Quantitative and qualitative evaluations on a real-world four-projector setup demonstrate that ConPhyG significantly outperforms state-of-the-art methods in geometric alignment, gamut utilization, and semantic fidelity.