Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses key challenges in verifying ownership of black-box large language models (LLMs)—including open-ended generation, response instability, and the fragility of existing fingerprinting methods—by proposing the Targeted Counterfactual Fingerprinting (TCF) framework. TCF transforms open-ended verification into a targeted counterfactual transfer task within a constrained answer space, leveraging prompt perturbations to steer model outputs toward predefined target answers and enabling verification through simple comparison of parsed final responses. The method innovatively introduces the Source-model Counterfactual Margin (SCM) as a reference-free metric to guide target selection, perturbation termination, and fingerprint filtering. Under a local behavioral similarity assumption, the authors theoretically demonstrate a significant gap in target accuracy between derived and independently trained models. Experiments across four LLM families show that TCF achieves an average AUC of 0.9861, substantially outperforming TRAP, ProFLingo, and ZeroPrint by margins of 0.07–0.19.
📝 Abstract
Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment. Because deployed LLMs are commonly exposed only through query APIs, ownership verification must often rely on black-box text responses. This setting is difficult: generations are open-ended and can vary across repeated queries, while existing black-box fingerprints rely on signals that are fragile under a final-response interface, including full-text matching, soft behavioral features, or model-specific prompts designed not to transfer. We propose TCF (Targeted Counterfactual Fingerprinting), a black-box LLM fingerprinting framework that converts open-ended generation comparison into constrained-answer targeted counterfactual transfer. TCF restricts each verification query to a finite answer space, reducing the surface-form ambiguity that enters the verification score, and optimizes a prompt perturbation toward a counterfactual target different from the protected model's clean answer on the original prompt. Verification reduces to checking whether the suspect model's parsed final answer matches the recorded target. We introduce the source-model counterfactual margin (SCM), a protected-model-only quantity that certifies the target is unlikely before the perturbation and likely after it; SCM controls target selection, perturbation stopping, and fingerprint filtering. Under explicit derived-preservation and independent-transfer budgets motivated by local behavioral closeness, we derive a target-accuracy gap between derived and independent models. Across four LLM families, TCF achieves an average AUC of 0.9861, improving over TRAP, ProFLingo, and ZeroPrint by 0.07 to 0.19.
Problem

Research questions and friction points this paper is trying to address.

black-box LLM
ownership verification
fingerprinting
counterfactual
model derivation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Targeted Counterfactual Fingerprinting
Black-Box Verification
Source-Model Counterfactual Margin
Prompt Perturbation
LLM Ownership
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yutong Wu
Huazhong University of Science and Technology
X
Xiaofan Bai
Alibaba Group
S
Shixin Li
Huazhong University of Science and Technology
P
Pingyi Hu
Huazhong University of Science and Technology
Ziqi Zhou
Ziqi Zhou
Huazhong University of Science and Technology (HUST)
Trustworthy AI
Z
Zilong Wang
Huazhong University of Science and Technology
Xiaojing Ma
Xiaojing Ma
Associate Professor, School of Computer, Huazhong University of Sci&Tech
multimedia security
Songfeng Lu
Songfeng Lu
HuaZhong University of Science and Technology
Y
Yuhong Li
Alibaba Group
J
Jin Xuan
Alibaba Group
Y
Yi Wang
Microsoft Corporation
Dongmei Zhang
Dongmei Zhang
Microsoft Research
Software EngineeringMachine LearningInformation Visualization
B
Bin Benjamin Zhu
Microsoft Corporation