PE-Field 4D: Video Generation Models as Canvas
This work addresses the challenge of spatial control in video generation arising from viewpoint changes and camera motion by proposing a geometry-aware diffusion Transformer architecture. By integrating projected positional encoding and a depth-aware disambiguation mechanism, the method effectively fuses 3D depth information with 2D reprojection. It further introduces structured context tokens and geometry-guided cross-attention to enable precise spatial manipulation directly within the native latent space. The proposed approach significantly enhances controllability for viewpoint-dependent editing tasks, supporting camera trajectory redirection, novel view synthesis, and geometry-consistent video editing, all while preserving the strong generative priors of the underlying foundation model.