Quality Recovery for Quantized KV Caches via Low-Rank Attention Adaptation
研究通过低秩注意力适应方法解决量化KV缓存导致的质量损失问题,恢复了多个模型的困惑度差距。
研究通过低秩注意力适应方法解决量化KV缓存导致的质量损失问题,恢复了多个模型的困惑度差距。
This work addresses the challenge that, in pruned Vision Transformers (ViTs), attention latency fails to scale proportionally with reduced computation due to scheduling overhead dominating on short sequences. To overcome this, the authors propose the first low-overhead attention implementation tailored for pruned ViTs, featuring a lightweight bidirectional Triton attention kernel and a pack-attend-unpack pipeline compatible with diverse token pruning strategies (e.g., Threshold-L2, DynamicViT). The approach reduces scheduling latency to approximately 40 microseconds and achieves up to 2.24× end-to-end throughput improvement on DeiT-T/S/B models while preserving bit-level prediction consistency—demonstrated by a maximum logit discrepancy below 0.007—thereby closely approaching the theoretical acceleration limit.
This study investigates how cognitive task types influence token acceptance probability in tree-based speculative decoding. Leveraging TinyLlama-1.1B as the draft model and Llama-2-7B-Chat-GPTQ as the target model with tree attention, the authors analyze acceptance dynamics across 99,768 speculative nodes spanning four task categories: code generation, mathematical reasoning, logical reasoning, and open-ended dialogue. The work reveals, for the first time, that task type is a stronger predictor of acceptance probability than tree depth. Notably, despite exhibiting the highest entropy, chat tasks achieve the highest acceptance rate (mean accepted length > 1.0), attributed to the stylistic predictability induced by RLHF alignment. A weak negative correlation between domain entropy and acceptance rate (ρ ∈ [−0.20, −0.15]) further provides empirical grounding for domain-adaptive speculative budget allocation.
This study addresses the scarcity of hope speech detection research for low-resource languages—particularly Urdu—and the limited cross-lingual generalization of existing models. We propose the first lightweight, multilingual adaptation framework for hope speech recognition. Methodologically, we systematically evaluate the cross-lingual transfer performance of pretrained multilingual models—including XLM-RoBERTa, mBERT, EuroBERT, and UrduBERT—on hope speech detection, employing simple text preprocessing and supervised fine-tuning for efficient binary/multiclass classification. On the PolyHope-M 2025 benchmark, our approach achieves 95.2% F1 for Urdu binary classification and 65.2% F1 for multiclass classification, with robust performance also observed for Spanish, German, and English. Our key contributions are threefold: (1) filling critical resource and methodological gaps in hope speech detection for low-resource languages; (2) providing the first empirical validation of multilingual Transformers’ generalization capability for positive discourse detection; and (3) introducing an extensible, lightweight adaptation paradigm.
研究通过低秩注意力适应方法解决量化KV缓存导致的质量损失问题,恢复了多个模型的困惑度差距。
This work addresses the challenge that, in pruned Vision Transformers (ViTs), attention latency fails to scale proportionally with reduced computation due to scheduling overhead dominating on short sequences. To overcome this, the authors propose the first low-overhead attention implementation tailored for pruned ViTs, featuring a lightweight bidirectional Triton attention kernel and a pack-attend-unpack pipeline compatible with diverse token pruning strategies (e.g., Threshold-L2, DynamicViT). The approach reduces scheduling latency to approximately 40 microseconds and achieves up to 2.24× end-to-end throughput improvement on DeiT-T/S/B models while preserving bit-level prediction consistency—demonstrated by a maximum logit discrepancy below 0.007—thereby closely approaching the theoretical acceleration limit.
This study investigates how cognitive task types influence token acceptance probability in tree-based speculative decoding. Leveraging TinyLlama-1.1B as the draft model and Llama-2-7B-Chat-GPTQ as the target model with tree attention, the authors analyze acceptance dynamics across 99,768 speculative nodes spanning four task categories: code generation, mathematical reasoning, logical reasoning, and open-ended dialogue. The work reveals, for the first time, that task type is a stronger predictor of acceptance probability than tree depth. Notably, despite exhibiting the highest entropy, chat tasks achieve the highest acceptance rate (mean accepted length > 1.0), attributed to the stylistic predictability induced by RLHF alignment. A weak negative correlation between domain entropy and acceptance rate (ρ ∈ [−0.20, −0.15]) further provides empirical grounding for domain-adaptive speculative budget allocation.
This study addresses the scarcity of hope speech detection research for low-resource languages—particularly Urdu—and the limited cross-lingual generalization of existing models. We propose the first lightweight, multilingual adaptation framework for hope speech recognition. Methodologically, we systematically evaluate the cross-lingual transfer performance of pretrained multilingual models—including XLM-RoBERTa, mBERT, EuroBERT, and UrduBERT—on hope speech detection, employing simple text preprocessing and supervised fine-tuning for efficient binary/multiclass classification. On the PolyHope-M 2025 benchmark, our approach achieves 95.2% F1 for Urdu binary classification and 65.2% F1 for multiclass classification, with robust performance also observed for Spanish, German, and English. Our key contributions are threefold: (1) filling critical resource and methodological gaps in hope speech detection for low-resource languages; (2) providing the first empirical validation of multilingual Transformers’ generalization capability for positive discourse detection; and (3) introducing an extensible, lightweight adaptation paradigm.