E2S-Pruner: Progressive Two-Stage Evidence Fusion for Visual Token Pruning in Vision-Language Models
为解决视觉-语言模型中视觉令牌修剪问题,提出E2S-Pruner方法,通过两阶段证据融合策略,在无需额外训练的情况下有效减少计算延迟和GPU内存消耗。
为解决视觉-语言模型中视觉令牌修剪问题,提出E2S-Pruner方法,通过两阶段证据融合策略,在无需额外训练的情况下有效减少计算延迟和GPU内存消耗。
本文探讨了通过整合大型语言模型、知识库和推理能力来构建下一代AI代理,以实现通用具身智能,并分析了面临的五大挑战。
This study investigates whether the effectiveness of frozen vision-language model prompts—such as those from CLIP—remains consistent after visual adaptation to a target domain in unsupervised cross-domain few-shot learning. By evaluating performance shifts of simple class-name templates versus rich semantic descriptions before and after adaptation across multiple medical and remote sensing datasets, and employing LoRA fine-tuning, paired comparisons, semantic shuffling controls, and multi-seed validation, the work uncovers two novel phenomena: “semantic saturation,” where adaptation yields diminishing gains (e.g., EuroSAT, CropDisease), and “semantic emergence,” where only detailed descriptions become effective post-adaptation (e.g., ISIC, ChestX). These findings challenge the prevailing assumption that zero-shot prompt quality reliably predicts its efficacy as a semantic anchor after adaptation, revealing instead a dynamic evolution of semantic utility.
This work addresses key challenges in multispectral object detection—inefficient cross-modal feature alignment, high computational cost of global attention, and the inability of fixed receptive fields to model nonlinear spatial relationships—by proposing the PNAFusion framework. The method introduces Pixel Neighborhood Cross-Attention (PNCA) to eliminate redundant global matching, designs an Adaptive Deformable Alignment (ADA) module to capture nonlinear spatial mappings, and incorporates a progressive feedback mechanism to iteratively refine fusion between visible and thermal infrared features. Integrated into YOLOv5 and Co-DETR, PNAFusion achieves 84.2, 90.5, and 85.5 mAP@0.5 on the FLIR, M3FD, and DroneVehicle datasets, respectively, while reducing GPU memory consumption by 33.0% compared to ICAFusion and lowering theoretical FLOPs from 194.8G to 156.4G.
Quantum machine learning faces two major bottlenecks: barren plateaus and noise sensitivity, compounded by the absence of a unified theoretical framework. This work proposes a novel paradigm based on Lie-algebraic generator dynamics, modeling parameterized quantum circuits as Lie subalgebras of $\mathfrak{u}(2^n)$ and characterizing trainability and expressivity through the induced Riemannian manifold geometry. The key innovation is structured Lie algebra truncation (LieTrunc), which contracts the manifold to circumvent concentration of measure while preserving non-vanishing gradients. We establish the first “geometric capacity–plateau” principle, proving that expressivity is governed by the span of generators rather than parameter count, and identify trainable regions where gradient variance decays polynomially. Experiments on 2–6 qubit systems demonstrate that LieTrunc-QNN maintains stable gradients and high effective dimensionality, fully preserving the Fubini–Study metric rank (rank=16 at $n=6$) and significantly outperforming random truncation, thereby validating the scaling law between gradient variance and effective dimension.
为解决视觉-语言模型中视觉令牌修剪问题,提出E2S-Pruner方法,通过两阶段证据融合策略,在无需额外训练的情况下有效减少计算延迟和GPU内存消耗。
本文探讨了通过整合大型语言模型、知识库和推理能力来构建下一代AI代理,以实现通用具身智能,并分析了面临的五大挑战。
This study investigates whether the effectiveness of frozen vision-language model prompts—such as those from CLIP—remains consistent after visual adaptation to a target domain in unsupervised cross-domain few-shot learning. By evaluating performance shifts of simple class-name templates versus rich semantic descriptions before and after adaptation across multiple medical and remote sensing datasets, and employing LoRA fine-tuning, paired comparisons, semantic shuffling controls, and multi-seed validation, the work uncovers two novel phenomena: “semantic saturation,” where adaptation yields diminishing gains (e.g., EuroSAT, CropDisease), and “semantic emergence,” where only detailed descriptions become effective post-adaptation (e.g., ISIC, ChestX). These findings challenge the prevailing assumption that zero-shot prompt quality reliably predicts its efficacy as a semantic anchor after adaptation, revealing instead a dynamic evolution of semantic utility.
This work addresses key challenges in multispectral object detection—inefficient cross-modal feature alignment, high computational cost of global attention, and the inability of fixed receptive fields to model nonlinear spatial relationships—by proposing the PNAFusion framework. The method introduces Pixel Neighborhood Cross-Attention (PNCA) to eliminate redundant global matching, designs an Adaptive Deformable Alignment (ADA) module to capture nonlinear spatial mappings, and incorporates a progressive feedback mechanism to iteratively refine fusion between visible and thermal infrared features. Integrated into YOLOv5 and Co-DETR, PNAFusion achieves 84.2, 90.5, and 85.5 mAP@0.5 on the FLIR, M3FD, and DroneVehicle datasets, respectively, while reducing GPU memory consumption by 33.0% compared to ICAFusion and lowering theoretical FLOPs from 194.8G to 156.4G.
Quantum machine learning faces two major bottlenecks: barren plateaus and noise sensitivity, compounded by the absence of a unified theoretical framework. This work proposes a novel paradigm based on Lie-algebraic generator dynamics, modeling parameterized quantum circuits as Lie subalgebras of $\mathfrak{u}(2^n)$ and characterizing trainability and expressivity through the induced Riemannian manifold geometry. The key innovation is structured Lie algebra truncation (LieTrunc), which contracts the manifold to circumvent concentration of measure while preserving non-vanishing gradients. We establish the first “geometric capacity–plateau” principle, proving that expressivity is governed by the span of generators rather than parameter count, and identify trainable regions where gradient variance decays polynomially. Experiments on 2–6 qubit systems demonstrate that LieTrunc-QNN maintains stable gradients and high effective dimensionality, fully preserving the Fubini–Study metric rank (rank=16 at $n=6$) and significantly outperforming random truncation, thereby validating the scaling law between gradient variance and effective dimension.