ZipMVS: Multi-View Stereo with Compressed Cost Volumes
ZipMVS通过创新的深度假设策略压缩成本体积,解决了多视图立体视觉方法内存需求大的问题,实现了高效高质量的3D重建。
ZipMVS通过创新的深度假设策略压缩成本体积,解决了多视图立体视觉方法内存需求大的问题,实现了高效高质量的3D重建。
This work addresses the computational redundancy in existing single-step diffusion models for multitask dense prediction, which typically rely on parameter-heavy adapters or learnable task tokens. The study is the first to reveal and exploit the fixed sinusoidal timestep embeddings inherent in diffusion models as endogenous task-conditioning signals, proposing a unified multitask learning paradigm that requires no additional parameters. Built upon pretrained diffusion models, the method leverages timestep embeddings for task guidance and incorporates manifold disentanglement to enable task-specific generation, compatible with both U-Net and DiT architectures. Experiments across ten datasets demonstrate that the approach achieves performance on par with state-of-the-art methods in monocular depth and surface normal estimation, confirming its effectiveness and broad applicability.
Modeling human hands in digital twins requires balancing anatomical fidelity with real-time physical simulation—a longstanding challenge. Method: This paper proposes a personalized multi-rigid-body hand modeling framework: (1) fitting subject-specific MANO models to optical motion capture data and mapping them to anatomically consistent URDF representations; (2) introducing a novel iterative projection method that combines closed-form SO(3) rotation solutions with Baker–Campbell–Hausdorff (BCH) formula-based corrections to efficiently project rotations onto physiologically valid joint constraint spaces using Lie group algebra. Results: Experiments show hand reconstruction error < 1 cm; the resulting models enable high-fidelity, real-time physics simulation. Reinforcement learning policies trained on these models successfully reproduce diverse human grasping behaviors, demonstrating effectiveness and generalization capability in digital twin interaction tasks.
Existing continual knowledge editing methods for large language models suffer from error accumulation due to parameter interference, degrading both editing accuracy and generalization. This paper proposes a fine-grained neuron localization framework coupled with an entropy-guided dynamic sparse masking mechanism. First, neurons are functionally attributed and categorized into *knowledge-general* and *knowledge-specific* types. Then, only critical knowledge-specific neurons undergo sparse, adaptive weight updates—minimizing parameter perturbation while preserving model integrity. The method requires no full model retraining and maintains high editing success rates (+12.7%) and strong generalization stability (38.5% reduction in forgetting) over thousands of sequential edits—outperforming state-of-the-art approaches. Its core innovations lie in (i) interpretable, function-based neuron partitioning and (ii) an information-theoretic, sparsity-aware editing strategy that balances fidelity and plasticity.
Existing cosine-similarity- and Softmax-based face recognition methods exhibit insufficient discriminative power on challenging high-quality samples. To address this, we propose LH2Face, a novel loss function. Our approach models face features on the hypersphere using the von Mises–Fisher (vMF) distribution—replacing conventional Euclidean or cosine metrics—and introduces an uncertainty-aware adaptive margin function that jointly characterizes sample difficulty and quality. Furthermore, LH2Face integrates proxy-based classification loss with a face reconstruction task within a multi-task optimization framework. Evaluated on the IJB-B benchmark, LH2Face achieves 49.39% true positive rate at a false acceptance rate of 1e−4, outperforming the second-best method by 2.37%. This demonstrates substantial improvement in recognizing high-quality yet difficult samples.
ZipMVS通过创新的深度假设策略压缩成本体积,解决了多视图立体视觉方法内存需求大的问题,实现了高效高质量的3D重建。
This work addresses the computational redundancy in existing single-step diffusion models for multitask dense prediction, which typically rely on parameter-heavy adapters or learnable task tokens. The study is the first to reveal and exploit the fixed sinusoidal timestep embeddings inherent in diffusion models as endogenous task-conditioning signals, proposing a unified multitask learning paradigm that requires no additional parameters. Built upon pretrained diffusion models, the method leverages timestep embeddings for task guidance and incorporates manifold disentanglement to enable task-specific generation, compatible with both U-Net and DiT architectures. Experiments across ten datasets demonstrate that the approach achieves performance on par with state-of-the-art methods in monocular depth and surface normal estimation, confirming its effectiveness and broad applicability.
Modeling human hands in digital twins requires balancing anatomical fidelity with real-time physical simulation—a longstanding challenge. Method: This paper proposes a personalized multi-rigid-body hand modeling framework: (1) fitting subject-specific MANO models to optical motion capture data and mapping them to anatomically consistent URDF representations; (2) introducing a novel iterative projection method that combines closed-form SO(3) rotation solutions with Baker–Campbell–Hausdorff (BCH) formula-based corrections to efficiently project rotations onto physiologically valid joint constraint spaces using Lie group algebra. Results: Experiments show hand reconstruction error < 1 cm; the resulting models enable high-fidelity, real-time physics simulation. Reinforcement learning policies trained on these models successfully reproduce diverse human grasping behaviors, demonstrating effectiveness and generalization capability in digital twin interaction tasks.
Existing continual knowledge editing methods for large language models suffer from error accumulation due to parameter interference, degrading both editing accuracy and generalization. This paper proposes a fine-grained neuron localization framework coupled with an entropy-guided dynamic sparse masking mechanism. First, neurons are functionally attributed and categorized into *knowledge-general* and *knowledge-specific* types. Then, only critical knowledge-specific neurons undergo sparse, adaptive weight updates—minimizing parameter perturbation while preserving model integrity. The method requires no full model retraining and maintains high editing success rates (+12.7%) and strong generalization stability (38.5% reduction in forgetting) over thousands of sequential edits—outperforming state-of-the-art approaches. Its core innovations lie in (i) interpretable, function-based neuron partitioning and (ii) an information-theoretic, sparsity-aware editing strategy that balances fidelity and plasticity.
Existing cosine-similarity- and Softmax-based face recognition methods exhibit insufficient discriminative power on challenging high-quality samples. To address this, we propose LH2Face, a novel loss function. Our approach models face features on the hypersphere using the von Mises–Fisher (vMF) distribution—replacing conventional Euclidean or cosine metrics—and introduces an uncertainty-aware adaptive margin function that jointly characterizes sample difficulty and quality. Furthermore, LH2Face integrates proxy-based classification loss with a face reconstruction task within a multi-task optimization framework. Evaluated on the IJB-B benchmark, LH2Face achieves 49.39% true positive rate at a false acceptance rate of 1e−4, outperforming the second-best method by 2.37%. This demonstrates substantial improvement in recognizing high-quality yet difficult samples.