Beyond QA Matching: Perturbation-Response Fingerprinting via Probability Distributions for Large Language Models

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出BReF方法,通过比较语言模型在文本扰动下的概率分布变化来识别模型的来源,解决了细粒度溯源难题。
📝 Abstract
Large language models are often instruction-tuned, specialized, quantized, or otherwise transformed, making fine-grained provenance difficult. In this paper, we introduce BReF, a training-free fingerprint that compares how probability distributions over four answer-option labels A/B/C/D move under controlled textual perturbations. For each pair of models, BReF selects 25 jointly responsive probes and compares their perturbation log-ratio (PLR) response directions by global cosine similarity. On a unified benchmark with 34 checkpoints, 22 documented direct-parent relations, and 411 suspect-candidate pairs, BReF retrieves the documented parent in 22/22 cases (MRR=1.0000), with DP-DF AUC 1.0000. Same-family discrimination is harder (DP-SF AUC 0.8969), and paired tests show a significant exact-retrieval gain over a magnitude-only Top-25 control. Together with static, random-probe, permutation, calibration, and transformation-level controls, the results show that strong pooled separation does not guarantee correct parent ranking among closely related checkpoints, verifying the superiority of our work.
Problem

Research questions and friction points this paper is trying to address.

large language models
instruction-tuned
provenance
Innovation

Methods, ideas, or system contributions that make the work stand out.

BReF
Perturbation-Response Fingerprinting
Probability Distributions
Perturbation Log-Ratio (PLR)
Cosine Similarity
💼 Related Jobs
No related jobs found.