TailorCoPilot: Enabling Agentic Pattern Making with Version-Controlled State Tracking
为解决服装制版中因隐性知识导致的技术断层问题,TailorCoPilot通过版本控制状态跟踪技术记录专家的操作过程,并辅助新手学习。
为解决服装制版中因隐性知识导致的技术断层问题,TailorCoPilot通过版本控制状态跟踪技术记录专家的操作过程,并辅助新手学习。
This work addresses the significant challenge of simulating realistic curly hair, where macroscopic deformations are tightly coupled with high-frequency geometric details such as waves and spirals. The authors propose a coiled finite element representation that models each hair strand as a rod-like base augmented with analytically defined high-frequency wrinkles, whose deformation is governed primarily by bending or twisting. To enhance numerical stability, they introduce a curvature energy splitting strategy that decouples stretching, buckling, and bending energies. Combined with a hybrid collision handling scheme and guide-hair-based interpolation, the method efficiently generates dense, structurally intricate curls. The approach achieves stable, visually plausible simulations across diverse scenarios while balancing computational efficiency and geometric fidelity.
This work addresses the challenge that digital sewing patterns typically lack explicit stitching annotations, necessitating expert intervention to define seam relationships for 3D garment modeling. To overcome this limitation, the authors propose a graph learning–based framework that automatically reconstructs two-level stitching information—coarse-grained inter-pattern connectivity and fine-grained seam correspondence—solely from the 2D pattern geometry. The method uniquely integrates pattern semantics, human body structural constraints, and garment design conventions within a unified model, enabling robust handling of complex topologies such as many-to-one, intra-pattern, and curved seams. By leveraging graph neural networks to fuse local geometric features with global contextual cues, the framework generates high-fidelity edge embeddings and decodes precise seam correspondences. Experiments demonstrate that the approach achieves high reconstruction accuracy across diverse garment styles, exhibits strong generalization, and effectively manages intricate sewing configurations.
This work addresses the limitations of existing Real2Sim approaches, which rely heavily on manual intervention and struggle to efficiently construct high-fidelity physics-based simulation environments. The authors propose the first unified framework that leverages a vision–language agent to automatically translate real-world videos of robot–object interactions into fully simulatable scenes, supporting diverse interaction types including rigid bodies, deformable objects, and human-like motions. By integrating geometric reconstruction, physical parameter inference, state reasoning, and an open-source vision–language model, the system achieves end-to-end automation. Experiments demonstrate that the method successfully reproduces a wide range of complex interactive scenarios with high fidelity, significantly reducing reliance on large-scale models and providing a high-quality, low-cost simulation foundation for robot policy learning and evaluation.
Existing virtual try-on methods rely on segmentation masks, which struggle to preserve fine-grained textures and cannot support arbitrary combinations of multiple garments, limiting their applicability in e-commerce scenarios. This work proposes a unified generative framework that operates without human parsing or mask supervision. By leveraging frequency-domain consistency constraints, lightweight Mixture-of-Experts (MoE) fine-tuning, and adaptively constructed inpainting data, the method achieves high-fidelity texture preservation while maintaining structural coherence with the human body. It flexibly accommodates any number and category of garments, outperforming state-of-the-art approaches in both quantitative metrics and perceptual quality. Moreover, with INT4 quantization, the model enables inference within 15 seconds per image on an RTX 4090 GPU, demonstrating strong practical deployability.
为解决服装制版中因隐性知识导致的技术断层问题,TailorCoPilot通过版本控制状态跟踪技术记录专家的操作过程,并辅助新手学习。
This work addresses the significant challenge of simulating realistic curly hair, where macroscopic deformations are tightly coupled with high-frequency geometric details such as waves and spirals. The authors propose a coiled finite element representation that models each hair strand as a rod-like base augmented with analytically defined high-frequency wrinkles, whose deformation is governed primarily by bending or twisting. To enhance numerical stability, they introduce a curvature energy splitting strategy that decouples stretching, buckling, and bending energies. Combined with a hybrid collision handling scheme and guide-hair-based interpolation, the method efficiently generates dense, structurally intricate curls. The approach achieves stable, visually plausible simulations across diverse scenarios while balancing computational efficiency and geometric fidelity.
This work addresses the challenge that digital sewing patterns typically lack explicit stitching annotations, necessitating expert intervention to define seam relationships for 3D garment modeling. To overcome this limitation, the authors propose a graph learning–based framework that automatically reconstructs two-level stitching information—coarse-grained inter-pattern connectivity and fine-grained seam correspondence—solely from the 2D pattern geometry. The method uniquely integrates pattern semantics, human body structural constraints, and garment design conventions within a unified model, enabling robust handling of complex topologies such as many-to-one, intra-pattern, and curved seams. By leveraging graph neural networks to fuse local geometric features with global contextual cues, the framework generates high-fidelity edge embeddings and decodes precise seam correspondences. Experiments demonstrate that the approach achieves high reconstruction accuracy across diverse garment styles, exhibits strong generalization, and effectively manages intricate sewing configurations.
This work addresses the limitations of existing Real2Sim approaches, which rely heavily on manual intervention and struggle to efficiently construct high-fidelity physics-based simulation environments. The authors propose the first unified framework that leverages a vision–language agent to automatically translate real-world videos of robot–object interactions into fully simulatable scenes, supporting diverse interaction types including rigid bodies, deformable objects, and human-like motions. By integrating geometric reconstruction, physical parameter inference, state reasoning, and an open-source vision–language model, the system achieves end-to-end automation. Experiments demonstrate that the method successfully reproduces a wide range of complex interactive scenarios with high fidelity, significantly reducing reliance on large-scale models and providing a high-quality, low-cost simulation foundation for robot policy learning and evaluation.
Existing virtual try-on methods rely on segmentation masks, which struggle to preserve fine-grained textures and cannot support arbitrary combinations of multiple garments, limiting their applicability in e-commerce scenarios. This work proposes a unified generative framework that operates without human parsing or mask supervision. By leveraging frequency-domain consistency constraints, lightweight Mixture-of-Experts (MoE) fine-tuning, and adaptively constructed inpainting data, the method achieves high-fidelity texture preservation while maintaining structural coherence with the human body. It flexibly accommodates any number and category of garments, outperforming state-of-the-art approaches in both quantitative metrics and perceptual quality. Moreover, with INT4 quantization, the model enables inference within 15 seconds per image on an RTX 4090 GPU, demonstrating strong practical deployability.