Conditioned Initialization for Attention
本文提出了一种条件初始化方法,通过优化注意力层的谱特性来改进Transformer模型的训练稳定性和收敛速度。
本文提出了一种条件初始化方法,通过优化注意力层的谱特性来改进Transformer模型的训练稳定性和收敛速度。
为解决X射线到CT重建中深度信息缺失问题,提出LiftXR框架,通过恢复3D解剖布局指导CT强度重建,提高重建精度。
Existing human pose estimation methods struggle with atypical limb structures such as amputated or prosthetic limbs, primarily due to the absence of a unified topological representation and insufficient training data. This work introduces ProPose, a novel benchmark that establishes the first unified keypoint topology protocol accommodating biological limbs, prostheses, and missing limbs. To address anatomical and mechanical constraints inherent in such diverse structures, we propose ProLoss, a structure-aware loss function that explicitly models dependencies among keypoints, thereby preventing implausible pose predictions on non-biological configurations. Combined with a Real-to-Synthetic data augmentation strategy, our approach maintains high spatial localization accuracy while improving classification accuracy for prosthetic joints under long-tailed data distributions by 2%–6%.
This work addresses the challenge of reliably distinguishing correct from incorrect reasoning in large language models without relying on superficial shortcuts. It introduces the first approach that models reasoning errors as region- and direction-specific signals within residual streams and proposes a three-stream detector that integrates residual trajectory dynamics, vector-quantized coarse-grained regional information, and fine-grained directional cues from normalized multi-layer states to reconstruct rich contextual representations. By moving beyond methods limited to token-level shifts or single-layer probing, the proposed framework achieves up to a 12% improvement in selection accuracy over existing shift-based methods and a 21% gain over single-layer baselines on unseen reasoning benchmarks, while consistently outperforming competing probes in factual completion and verification tasks.
3D Gaussian Splatting suffers from geometric inconsistency and lack of geometric interpretability under pose-free supervision. Method: This work introduces the first self-supervised framework that explicitly models geometric consistency by decoupling view synthesis from geometry reconstruction. It incorporates triple geometric priors—depth, surface normals, and relative camera poses—optimized jointly via a multi-view structure prediction network (pretrained on RE10K) embedded within a differentiable Gaussian rendering pipeline. Contribution/Results: The approach enables end-to-end learning of geometrically faithful Gaussian scene representations with zero-shot cross-dataset generalization. Experiments establish new state-of-the-art performance on RE10K for novel-view synthesis, geometric reconstruction, and relative pose estimation. When transferred to ScanNet, it reduces geometric error by 37% and improves pose estimation accuracy by 2.1× over prior methods.
本文提出了一种条件初始化方法,通过优化注意力层的谱特性来改进Transformer模型的训练稳定性和收敛速度。
为解决X射线到CT重建中深度信息缺失问题,提出LiftXR框架,通过恢复3D解剖布局指导CT强度重建,提高重建精度。
Existing human pose estimation methods struggle with atypical limb structures such as amputated or prosthetic limbs, primarily due to the absence of a unified topological representation and insufficient training data. This work introduces ProPose, a novel benchmark that establishes the first unified keypoint topology protocol accommodating biological limbs, prostheses, and missing limbs. To address anatomical and mechanical constraints inherent in such diverse structures, we propose ProLoss, a structure-aware loss function that explicitly models dependencies among keypoints, thereby preventing implausible pose predictions on non-biological configurations. Combined with a Real-to-Synthetic data augmentation strategy, our approach maintains high spatial localization accuracy while improving classification accuracy for prosthetic joints under long-tailed data distributions by 2%–6%.
This work addresses the challenge of reliably distinguishing correct from incorrect reasoning in large language models without relying on superficial shortcuts. It introduces the first approach that models reasoning errors as region- and direction-specific signals within residual streams and proposes a three-stream detector that integrates residual trajectory dynamics, vector-quantized coarse-grained regional information, and fine-grained directional cues from normalized multi-layer states to reconstruct rich contextual representations. By moving beyond methods limited to token-level shifts or single-layer probing, the proposed framework achieves up to a 12% improvement in selection accuracy over existing shift-based methods and a 21% gain over single-layer baselines on unseen reasoning benchmarks, while consistently outperforming competing probes in factual completion and verification tasks.
3D Gaussian Splatting suffers from geometric inconsistency and lack of geometric interpretability under pose-free supervision. Method: This work introduces the first self-supervised framework that explicitly models geometric consistency by decoupling view synthesis from geometry reconstruction. It incorporates triple geometric priors—depth, surface normals, and relative camera poses—optimized jointly via a multi-view structure prediction network (pretrained on RE10K) embedded within a differentiable Gaussian rendering pipeline. Contribution/Results: The approach enables end-to-end learning of geometrically faithful Gaussian scene representations with zero-shot cross-dataset generalization. Experiments establish new state-of-the-art performance on RE10K for novel-view synthesis, geometric reconstruction, and relative pose estimation. When transferred to ScanNet, it reduces geometric error by 37% and improves pose estimation accuracy by 2.1× over prior methods.