GOLF: Global Observation with Local Focus for Calibration-Aware Stereo Interaction Field Estimation
本文提出GOLF方法,通过结合全局上下文和局部特征,利用改进的DINOv3 ViT-H+/16模型预测手部关节与操作物体间的3D向量,解决立体交互场估计问题。
本文提出GOLF方法,通过结合全局上下文和局部特征,利用改进的DINOv3 ViT-H+/16模型预测手部关节与操作物体间的3D向量,解决立体交互场估计问题。
本文提出EpaCache方法,通过自适应分配重用预算来减少扩散模型生成图像和视频时的推理延迟,同时提高生成质量。
This work addresses the slow convergence of denoising diffusion Transformers by proposing a parameter-free structural regularization method that circumvents reliance on external encoders and better preserves intrinsic data structure. The approach uniquely leverages pairwise similarity relationships among tokens in clean latent representations as a direct supervision signal and innovatively incorporates cross-image token pairs to capture spatial structure. Built upon the flow matching framework, it introduces a parameter-free affinity regularization term that jointly calibrates intra-layer and inter-sample relationships. Evaluated on ImageNet 256×256 with a SiT backbone, the method incurs only 0.08 GB additional GPU memory without introducing new parameters, achieving superior FID compared to all existing parameter-free approaches at 400K training iterations and reaching an FID of 1.90 at 1M iterations when combined with REPA.
This work addresses a critical issue in high-dimensional flow matching: the systematic underestimation of velocity magnitude at trajectory initialization induces integration lag, preventing generated samples from accurately reaching the data manifold. The study uncovers an asymmetric mechanism wherein velocity contraction is detrimental at the start but beneficial near the end of trajectories. To mitigate this without retraining, the authors propose a joint strategy combining a Scale Scheduling Corrector (SSC) with Magnitude-Aware Flow Matching (MAFM), implementable with just a single line of code. The approach substantially improves performance—on ImageNet-1k, it reduces FID from 13.68 to 7.58 (a 44.6% improvement), achieves a 5× speedup, and yields 50-step generation quality surpassing the original 250-step baseline; on MS-COCO text-to-image generation, FID also improves by approximately 22%.
This work addresses the inefficiency and poor scalability of existing multimodal visual object tracking methods, which typically require separate modeling due to modality heterogeneity. To overcome these limitations, the authors propose OneTrackerV2, the first unified tracking framework capable of handling arbitrary visual modalities as input. Its core innovations include a Meta Merger for multimodal fusion under a unified representation and a dual-path mixture-of-experts mechanism—comprising T-MoE and M-MoE—to model spatiotemporal dynamics and cross-modal knowledge, respectively. Trained end-to-end, OneTrackerV2 achieves state-of-the-art performance across five RGB and RGB+X tracking tasks on twelve benchmarks, while maintaining high inference efficiency. Notably, the model remains competitive even after compression and demonstrates exceptional robustness and generalization under missing-modality scenarios.
本文提出GOLF方法,通过结合全局上下文和局部特征,利用改进的DINOv3 ViT-H+/16模型预测手部关节与操作物体间的3D向量,解决立体交互场估计问题。
本文提出EpaCache方法,通过自适应分配重用预算来减少扩散模型生成图像和视频时的推理延迟,同时提高生成质量。
This work addresses the slow convergence of denoising diffusion Transformers by proposing a parameter-free structural regularization method that circumvents reliance on external encoders and better preserves intrinsic data structure. The approach uniquely leverages pairwise similarity relationships among tokens in clean latent representations as a direct supervision signal and innovatively incorporates cross-image token pairs to capture spatial structure. Built upon the flow matching framework, it introduces a parameter-free affinity regularization term that jointly calibrates intra-layer and inter-sample relationships. Evaluated on ImageNet 256×256 with a SiT backbone, the method incurs only 0.08 GB additional GPU memory without introducing new parameters, achieving superior FID compared to all existing parameter-free approaches at 400K training iterations and reaching an FID of 1.90 at 1M iterations when combined with REPA.
This work addresses a critical issue in high-dimensional flow matching: the systematic underestimation of velocity magnitude at trajectory initialization induces integration lag, preventing generated samples from accurately reaching the data manifold. The study uncovers an asymmetric mechanism wherein velocity contraction is detrimental at the start but beneficial near the end of trajectories. To mitigate this without retraining, the authors propose a joint strategy combining a Scale Scheduling Corrector (SSC) with Magnitude-Aware Flow Matching (MAFM), implementable with just a single line of code. The approach substantially improves performance—on ImageNet-1k, it reduces FID from 13.68 to 7.58 (a 44.6% improvement), achieves a 5× speedup, and yields 50-step generation quality surpassing the original 250-step baseline; on MS-COCO text-to-image generation, FID also improves by approximately 22%.
This work addresses the inefficiency and poor scalability of existing multimodal visual object tracking methods, which typically require separate modeling due to modality heterogeneity. To overcome these limitations, the authors propose OneTrackerV2, the first unified tracking framework capable of handling arbitrary visual modalities as input. Its core innovations include a Meta Merger for multimodal fusion under a unified representation and a dual-path mixture-of-experts mechanism—comprising T-MoE and M-MoE—to model spatiotemporal dynamics and cross-modal knowledge, respectively. Trained end-to-end, OneTrackerV2 achieves state-of-the-art performance across five RGB and RGB+X tracking tasks on twelve benchmarks, while maintaining high inference efficiency. Notably, the model remains competitive even after compression and demonstrates exceptional robustness and generalization under missing-modality scenarios.