ABLE: Representing and Mapping LLMs via Attribution-Based Large-model Embedding

📅 2026-04-20
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决大语言模型比较难题,提出ABLE方法,通过基于梯度的特征归因构建模型表示,实现模型特性的有效捕捉。
📝 Abstract
The explosive growth of large language models (LLMs) has created a heterogeneous and poorly documented ecosystem, making systematic model comparison increasingly important for provenance auditing, security analysis, and model selection. Existing representation methods struggle to address this setting efficiently. Approaches analyzing internal parameters are powerful when architectures are compatible, but face scalability barriers under structural heterogeneity, while methods relying on external outputs may conflate models with similar behaviors and are difficult to align in richer output spaces across different tokenizers. To bridge this gap, we propose ABLE (Attribution-Based Large-model Embedding), a framework that leverages the interpretability space to construct model representations. By aggregating gradient-based feature attributions via a tokenizer-agnostic word-level alignment, ABLE captures model-specific input-sensitivity patterns rather than only surface-level outputs. Beyond empirical utility, we provide a stability analysis showing that, under standard regularity assumptions for differentiable Transformer-style models, ABLE induces a Lipschitz-continuous parameter-to-embedding map with finite-sample convergence guarantees. Extensive experiments on 239 open-source LLMs demonstrate that our training-free approach achieves competitive or superior performance in relation prediction, model routing, and benchmark score prediction.
Problem

Research questions and friction points this paper is trying to address.

large language models
systematic model comparison
structural heterogeneity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attribution-Based Large-model Embedding
Tokenizer-agnostic Word-level Alignment
Gradient-based Feature Attributions
Lipschitz-Continuous Parameter-to-Embedding Map
Zirui Wang
Zirui Wang
City University of Hong Kong
infrared image enhancementadversarial attack
Y
Yusen Hou
The Hong Kong University of Science and Technology (Guangzhou)
S
Shaofeng Liang
The Hong Kong University of Science and Technology (Guangzhou); Deep Interdisciplinary Intelligence Lab (DI 2Lab)
Bowen Tian
Bowen Tian
The Hong Kong University of Science and Technology (Guangzhou)
Model FusionNeural Network FunctionalsSemi-Supervised Learning
Y
Yanlin Zhang
The Hong Kong University of Science and Technology (Guangzhou)
Wenshuo Chen
Wenshuo Chen
Shandong University undergraduate student
Generative ModelsXAI
Y
Yutao Yue
The Hong Kong University of Science and Technology (Guangzhou); Deep Interdisciplinary Intelligence Lab (DI 2Lab)