Arbor: Explicit Geometric Conditioning for Controllable 3D Asset Generation
Existing text- or image-conditioned 3D generation models struggle to explicitly express user intent regarding spatial regions an object should occupy or avoid. To address this limitation, this work proposes Arbor—a trainable plug-in module that, for the first time, introduces locally typed geometric constraint grids (specifying presence, avoidance, and contact zones) as non-target evidence into latent 3D diffusion models. Arbor learns to inject these geometric constraints positionally within a frozen denoiser via a constraint-to-token transformation and a spatial routing attention mechanism, enabling fine-grained control over the generated object’s spatial layout. Experiments demonstrate that Arbor significantly improves adherence to spatial constraints in both automatic and human evaluations, without requiring specialized compliance losses, while preserving generation quality and diversity under fixed constraints.