Generative Late-Interaction Embeddings For Visual Document Retrieval
为解决视觉文档检索中存储效率问题,提出生成式后期交互嵌入(GLIE)方法,通过少量向量重建全页嵌入集,提高存储效率同时保持检索精度。
为解决视觉文档检索中存储效率问题,提出生成式后期交互嵌入(GLIE)方法,通过少量向量重建全页嵌入集,提高存储效率同时保持检索精度。
研究提出了一种无乘法器的Shift-Accumulate Attention方法,通过将键缓存量化为有符号的2的幂固定点编码,使QK^T中的每个标量乘法变为符号翻转、位移和整数累加,从而提高Transformer解码效率。
This work addresses the challenge of fine-grained human action recognition, where visual similarity among actions impedes discriminability. To this end, we propose a structured multimodal fusion framework that jointly models three complementary representations: RGB appearance, pose heatmap geometry, and skeletal graph topology. Our approach employs pairwise cross-attention to enable symmetric inter-stream interaction and introduces a stream-wise latent sparse Mixture-of-Experts (MoE) mechanism that dynamically routes inputs to a shared subset of experts based on content, augmented with load-balancing regularization. Notably, our method achieves state-of-the-art performance on Gym99, Gym288, and Diving48 without relying on textual supervision or large-scale vision-language pretraining. On the long-tailed Gym288 benchmark, it improves mean class accuracy by 7.6 percentage points, from 68.6% to 76.2%.
This study addresses the challenge of learning threat-aware and adaptive driving policies in safety-critical scenarios using only observable kinematic information and natural language descriptions. The authors propose a language-structured relational Q-learning approach, introducing language-guided scene modeling into relational reinforcement learning for the first time. They design an ego-centric relational Q-network (ERQ-Net) to jointly learn inter-vehicle dynamic dependencies and action values. The work formally identifies and articulates the “perception-control gap” problem, demonstrating improved performance in 2,500 CARLA simulation scenarios—raising success rates from 49–52% to 55–58% and enhancing adversarial target attention by 1.2–2.1×. Nevertheless, experiments reveal that 76% of scenarios remain solvable by simple policy compositions, highlighting a fundamental performance bottleneck in current methods.
This work addresses the challenges posed by high structural heterogeneity, large intra-class variation, and subtle visual differences between benign and malignant lesions in dermoscopic images. To this end, the authors propose a superpixel-based multimodal fusion approach that models lesions as graphs whose nodes correspond to superpixels. Node features are extracted using a frozen CNN, while geometric relationships are incorporated as edge attributes. A novel metadata context node is introduced to enable native graph-level fusion of clinical information with visual features. Discriminative classification embeddings are generated through an edge-aware Graph Transformer coupled with an attention propagation mechanism. This method, which uniquely integrates superpixel graph structure, geometric edge attributes, and metadata context, achieves state-of-the-art performance across four public datasets, significantly improving both accuracy and robustness in benign–malignant skin lesion classification.
为解决视觉文档检索中存储效率问题,提出生成式后期交互嵌入(GLIE)方法,通过少量向量重建全页嵌入集,提高存储效率同时保持检索精度。
研究提出了一种无乘法器的Shift-Accumulate Attention方法,通过将键缓存量化为有符号的2的幂固定点编码,使QK^T中的每个标量乘法变为符号翻转、位移和整数累加,从而提高Transformer解码效率。
This work addresses the challenge of fine-grained human action recognition, where visual similarity among actions impedes discriminability. To this end, we propose a structured multimodal fusion framework that jointly models three complementary representations: RGB appearance, pose heatmap geometry, and skeletal graph topology. Our approach employs pairwise cross-attention to enable symmetric inter-stream interaction and introduces a stream-wise latent sparse Mixture-of-Experts (MoE) mechanism that dynamically routes inputs to a shared subset of experts based on content, augmented with load-balancing regularization. Notably, our method achieves state-of-the-art performance on Gym99, Gym288, and Diving48 without relying on textual supervision or large-scale vision-language pretraining. On the long-tailed Gym288 benchmark, it improves mean class accuracy by 7.6 percentage points, from 68.6% to 76.2%.
This study addresses the challenge of learning threat-aware and adaptive driving policies in safety-critical scenarios using only observable kinematic information and natural language descriptions. The authors propose a language-structured relational Q-learning approach, introducing language-guided scene modeling into relational reinforcement learning for the first time. They design an ego-centric relational Q-network (ERQ-Net) to jointly learn inter-vehicle dynamic dependencies and action values. The work formally identifies and articulates the “perception-control gap” problem, demonstrating improved performance in 2,500 CARLA simulation scenarios—raising success rates from 49–52% to 55–58% and enhancing adversarial target attention by 1.2–2.1×. Nevertheless, experiments reveal that 76% of scenarios remain solvable by simple policy compositions, highlighting a fundamental performance bottleneck in current methods.
This work addresses the challenges posed by high structural heterogeneity, large intra-class variation, and subtle visual differences between benign and malignant lesions in dermoscopic images. To this end, the authors propose a superpixel-based multimodal fusion approach that models lesions as graphs whose nodes correspond to superpixels. Node features are extracted using a frozen CNN, while geometric relationships are incorporated as edge attributes. A novel metadata context node is introduced to enable native graph-level fusion of clinical information with visual features. Discriminative classification embeddings are generated through an edge-aware Graph Transformer coupled with an attention propagation mechanism. This method, which uniquely integrates superpixel graph structure, geometric edge attributes, and metadata context, achieves state-of-the-art performance across four public datasets, significantly improving both accuracy and robustness in benign–malignant skin lesion classification.