Hub-Spectral Activation of Latent Multimodal Knowledge
本文提出Hub-Spectral Activation方法,通过恢复和激活冻结表示中的多模态潜在知识,提高跨模态检索和原型分类的性能,无需成对监督或梯度优化。
本文提出Hub-Spectral Activation方法,通过恢复和激活冻结表示中的多模态潜在知识,提高跨模态检索和原型分类的性能,无需成对监督或梯度优化。
研究提出C2Nav框架,通过比较控制器构建的选项而非直接生成几何输出来解决零样本视觉-语言导航问题,提高导航准确性。
本文提出CoAL-RAG方法,通过量化问题复杂度和选择合适的检索策略来解决法律咨询中简单问题过度推理和复杂问题解释性差的问题。
Existing methods struggle to accurately generate and edit fine-grained geometric details of 3D faces—such as eyebrow tension or cheek contraction—from long textual descriptions. To address this challenge, this work introduces FaME-G2E, a large-scale multimodal dataset, and proposes RAGMesh, a retrieval-augmented framework that integrates text-guided global and regional geometric priors in blendshape space. The framework innovatively combines a multi-scale retrieval fusion (MSRF) module with an adaptive RAG-guided supervision (AdaRAGS) mechanism to achieve precise semantic alignment and localized deformation control. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches in terms of local geometric accuracy, text controllability, regional editing precision, and inference efficiency.
This work addresses the limitations of existing vision token pruning methods, which rely on handcrafted heuristics and struggle to generalize across diverse model architectures and pruning objectives. To overcome this, the authors propose AutoPrune, a training-free AI4AI framework that leverages large language models (LLMs) to automatically design effective pruning strategies. Key innovations include TPDSL, a domain-specific language tailored for token pruning; a residual representation of search states to emphasize critical strategy components; and structured guidance for LLMs to generate algorithms within a constrained space. Extensive experiments demonstrate AutoPrune’s superior performance across 14 benchmarks and three multimodal large language models: it retains over 99% of original accuracy even after removing 94.4% of visual tokens, achieving a 9.9× reduction in FLOPs and a 6.4× speedup in prefill latency.
本文提出Hub-Spectral Activation方法,通过恢复和激活冻结表示中的多模态潜在知识,提高跨模态检索和原型分类的性能,无需成对监督或梯度优化。
研究提出C2Nav框架,通过比较控制器构建的选项而非直接生成几何输出来解决零样本视觉-语言导航问题,提高导航准确性。
本文提出CoAL-RAG方法,通过量化问题复杂度和选择合适的检索策略来解决法律咨询中简单问题过度推理和复杂问题解释性差的问题。
Existing methods struggle to accurately generate and edit fine-grained geometric details of 3D faces—such as eyebrow tension or cheek contraction—from long textual descriptions. To address this challenge, this work introduces FaME-G2E, a large-scale multimodal dataset, and proposes RAGMesh, a retrieval-augmented framework that integrates text-guided global and regional geometric priors in blendshape space. The framework innovatively combines a multi-scale retrieval fusion (MSRF) module with an adaptive RAG-guided supervision (AdaRAGS) mechanism to achieve precise semantic alignment and localized deformation control. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches in terms of local geometric accuracy, text controllability, regional editing precision, and inference efficiency.
This work addresses the limitations of existing vision token pruning methods, which rely on handcrafted heuristics and struggle to generalize across diverse model architectures and pruning objectives. To overcome this, the authors propose AutoPrune, a training-free AI4AI framework that leverages large language models (LLMs) to automatically design effective pruning strategies. Key innovations include TPDSL, a domain-specific language tailored for token pruning; a residual representation of search states to emphasize critical strategy components; and structured guidance for LLMs to generate algorithms within a constrained space. Extensive experiments demonstrate AutoPrune’s superior performance across 14 benchmarks and three multimodal large language models: it retains over 99% of original accuracy even after removing 94.4% of visual tokens, achieving a 9.9× reduction in FLOPs and a 6.4× speedup in prefill latency.