CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes via Geometric Pseudo Labels
为解决3D场景中根据文本指令自动定位相机视角的问题,提出CapFrame框架,通过将文本转换为几何伪标签来优化相机姿态。
为解决3D场景中根据文本指令自动定位相机视角的问题,提出CapFrame框架,通过将文本转换为几何伪标签来优化相机姿态。
This work addresses the limitation of existing vision backbones, which prioritize semantic content (“what”) over spatial location (“where”) in image classification, thereby hindering performance on localization tasks. To remedy this, the authors propose a novel Vision Transformer backbone that incorporates a “what–where” disentangled inductive bias: tokens encode semantic representations while attention maps serve as spatial representations. A multi-stream slot architecture processes these two streams in parallel. Remarkably, with only single-label ImageNet supervision, the model directly yields localization-aware features at its final layer. It achieves substantial improvements over current ViT-based methods in zero-shot object discovery and weakly supervised semantic segmentation, and demonstrates strong transferability across diverse localization scenarios.
This work addresses the instability and optimization challenges commonly encountered in training sparse vision Mixture-of-Experts (MoE) models, which often stem from gradient blocking and insufficient feedback during routing. To mitigate these issues, the authors propose a Teacher-Guided Routing mechanism (TGR-MoE), which, for the first time, leverages intermediate representations from a pretrained dense teacher model to provide knowledge-driven pseudo-supervisory signals for the student router, thereby enabling stable routing decisions early in training. By integrating knowledge distillation, sparse MoE architecture, and a representation-similarity-based routing strategy, TGR-MoE achieves substantial improvements in both accuracy and routing consistency on ImageNet-1K and CIFAR-100, maintaining robust training stability even under highly sparse configurations.
This work addresses the challenge of characterizing the generalization capability of deep neural networks (DNNs), which, as singular statistical models, defy accurate description by conventional metrics. The study establishes, for the first time, a theoretical foundation for the validity of the Takeuchi Information Criterion (TIC) in estimating the generalization gap of DNNs within the neural tangent kernel (NTK) regime, while also delineating the boundary conditions under which TIC fails outside this regime. Furthermore, the authors propose a computationally efficient approximation of TIC to enable trial-and-error pruning in hyperparameter optimization. Extensive experiments across 12 architectures and over 5,000 DNN models demonstrate that TIC exhibits strong correlation with the true generalization gap within the NTK regime and significantly outperforms existing pruning methods in hyperparameter tuning.
This work addresses the limited applicability of large language models in safety-critical automotive systems engineering due to concerns regarding trustworthiness, traceability, and compatibility with established verification workflows. The authors propose workflow-level design principles for trustworthy generative AI and implement them within an end-to-end automotive engineering pipeline encompassing requirement change identification, SysML v2 architecture updates, and regression testing. To enhance completeness in change detection, they employ segmented prompt decomposition, diversity sampling, and lightweight NLP-based validation. Traceable test generation is achieved through explicit variable-to-port mappings. Experimental results demonstrate that the approach significantly improves the detection rate of critical changes in large-scale specifications, ensures correctness of architectural updates, and enables automated, traceable regression testing—providing a practical foundation for deploying generative AI in safety-critical contexts.
为解决3D场景中根据文本指令自动定位相机视角的问题,提出CapFrame框架,通过将文本转换为几何伪标签来优化相机姿态。
This work addresses the limitation of existing vision backbones, which prioritize semantic content (“what”) over spatial location (“where”) in image classification, thereby hindering performance on localization tasks. To remedy this, the authors propose a novel Vision Transformer backbone that incorporates a “what–where” disentangled inductive bias: tokens encode semantic representations while attention maps serve as spatial representations. A multi-stream slot architecture processes these two streams in parallel. Remarkably, with only single-label ImageNet supervision, the model directly yields localization-aware features at its final layer. It achieves substantial improvements over current ViT-based methods in zero-shot object discovery and weakly supervised semantic segmentation, and demonstrates strong transferability across diverse localization scenarios.
This work addresses the instability and optimization challenges commonly encountered in training sparse vision Mixture-of-Experts (MoE) models, which often stem from gradient blocking and insufficient feedback during routing. To mitigate these issues, the authors propose a Teacher-Guided Routing mechanism (TGR-MoE), which, for the first time, leverages intermediate representations from a pretrained dense teacher model to provide knowledge-driven pseudo-supervisory signals for the student router, thereby enabling stable routing decisions early in training. By integrating knowledge distillation, sparse MoE architecture, and a representation-similarity-based routing strategy, TGR-MoE achieves substantial improvements in both accuracy and routing consistency on ImageNet-1K and CIFAR-100, maintaining robust training stability even under highly sparse configurations.
This work addresses the challenge of characterizing the generalization capability of deep neural networks (DNNs), which, as singular statistical models, defy accurate description by conventional metrics. The study establishes, for the first time, a theoretical foundation for the validity of the Takeuchi Information Criterion (TIC) in estimating the generalization gap of DNNs within the neural tangent kernel (NTK) regime, while also delineating the boundary conditions under which TIC fails outside this regime. Furthermore, the authors propose a computationally efficient approximation of TIC to enable trial-and-error pruning in hyperparameter optimization. Extensive experiments across 12 architectures and over 5,000 DNN models demonstrate that TIC exhibits strong correlation with the true generalization gap within the NTK regime and significantly outperforms existing pruning methods in hyperparameter tuning.
This work addresses the limited applicability of large language models in safety-critical automotive systems engineering due to concerns regarding trustworthiness, traceability, and compatibility with established verification workflows. The authors propose workflow-level design principles for trustworthy generative AI and implement them within an end-to-end automotive engineering pipeline encompassing requirement change identification, SysML v2 architecture updates, and regression testing. To enhance completeness in change detection, they employ segmented prompt decomposition, diversity sampling, and lightweight NLP-based validation. Traceable test generation is achieved through explicit variable-to-port mappings. Experimental results demonstrate that the approach significantly improves the detection rate of critical changes in large-scale specifications, ensures correctness of architectural updates, and enables automated, traceable regression testing—providing a practical foundation for deploying generative AI in safety-critical contexts.