Flow3D-OPD: Multi-Teacher On-Policy Distillation for 3D Geometry Generation with Flow-Matching Diffusion Transformer

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过引入多教师在线策略蒸馏方法Flow3D-OPD,解决了3D几何生成中定义奖励困难和梯度干扰问题,提升了3D模型的质量。
📝 Abstract
Recent image-to-3D generation models built on flow-matching diffusion Transformers (DiT) can produce high-fidelity meshes, yet their post-training strategy remains largely unexplored. There exist several critical bottlenecks in reinforcement learning: the inherent difficulty of defining comprehensive rewards for 3D geometric quality, and the gradient interference that arises when jointly optimizing heterogeneous objectives. Inspired by the practicability of on-policy distillation (OPD) in large language models and image generation, we propose \textbf{Flow3D-OPD}, a two-stage post-training framework that introduces multi-teacher distillation into 3D geometry generation. In the first stage, we utilize the semi-policy to enhance the foundational capability of the pretrained model and then design an agentic verifier for 3D geometric quality evaluation. Based on the verifier, we could cultivate domain-specialized teacher models via direct preference optimization (DPO). In the second stage, we consolidate heterogeneous expertise into a unified student model through on-policy distillation with hard task-routing sampling and gradient accumulation, which could mitigate the gradient interference in joint optimization. Without relying on elaborate modifications, our straightforward yet effective design achieves consistent improvements across all geometric quality dimensions and surpasses all teacher models in the average metric. Extensive experiments demonstrate that our approach provides an effective paradigm for reinforcement learning in 3D generation.
Problem

Research questions and friction points this paper is trying to address.

3D Geometry Generation
Reinforcement Learning
Gradient Interference
Flow-Matching Diffusion Transformer
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-teacher distillation
on-policy distillation (OPD)
direct preference optimization (DPO)
gradient interference mitigation
🔎 Similar Papers
2024-03-18European Conference on Computer VisionCitations: 70
Z
Zhiwei Ning
Shanghai Jiao Tong University
Z
Zhen Zhou
Tencent Hunyuan3D
Puhua Jiang
Puhua Jiang
Tencent
Computer VisionGraphicsGenerative model
Xintong Han
Xintong Han
Huya Inc
Computer VisionMachine LearningImage Processing
G
Gengming Zhang
Shanghai Jiao Tong University
Jie Yang
Jie Yang
Shanghai Jiao Tong University
Image ProcessingMedical Image ProcessingPattern Recognition
Z
Zhonglong Zheng
Zhejiang Normal University
Y
Yuanjie Zheng
Shandong Normal University
W
Wei Liu
Shanghai Jiao Tong University
C
Chunchao Guo
Tencent Hunyuan3D