SOV-CAD: Stepwise Orthographic Views Guided CAD Modeling Sequence Reconstruction
Reconstructing CAD modeling sequences from images often fails to capture the iterative, feedback-driven nature of human design. This work formulates the task as a sequential decision-making problem and introduces a novel mechanism that combines stepwise orthographic view supervision with geometric alignment rewards. At each step, continuous visual feedback—comprising orthographic views, incremental model projections, and the current sketch—guides action selection. Built upon offline reinforcement learning and the Decision Transformer architecture, the proposed method significantly outperforms existing approaches in both reconstruction accuracy and data efficiency, achieving state-of-the-art performance while better reflecting authentic design workflows.