Token-Level Likelihood-Array Regression for Membership Inference and AI-Generated Text Detection

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过提出一种基于似然数组回归的方法(LAR),利用不同上下文窗口下的标记级别概率特征,提高了成员推理和AI生成文本检测的准确性。
📝 Abstract
Membership inference asks whether a text was used to train a language model, whereas AI-generated text detection asks whether it was generated by a language model rather than written by a human. Existing likelihood-based methods typically compress token-level probabilities into a few prespecified scores, most often using only probabilities conditioned on the full preceding context. We propose likelihood-array regression (LAR), which evaluates each target token under nested left-context windows and organizes the resulting likelihood-derived features into a structured array. After aligning arrays across texts of different lengths, LAR learns how detection information varies with context scale, token position, and likelihood features. LAR-1 aggregates learned contributions from individual aligned cells, while LAR-2 adds second-order features formed from pairs of evaluations of the same target token across context lengths. For within-path quadratic model, we establish matching minimax lower and upper bounds, characterize errors from finite-dimensional approximation and random squared projections, and derive conditions under which an oracle spectral sieve attains the minimax rate. Across multiple scoring language models, LAR substantially improves membership inference and AI-generated text detection over likelihood-based baselines. The analyses further show that shorter-context likelihoods contain information beyond conventional full-context probabilities, while second-order features provide additional gains for membership inference.
Problem

Research questions and friction points this paper is trying to address.

Membership Inference
AI-Generated Text Detection
Likelihood-based Methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

likelihood-array regression
membership inference
AI-generated text detection
nested left-context windows
second-order features
🔎 Similar Papers
2024-06-21Journal of Artificial Intelligence ResearchCitations: 6
💼 Related Jobs
No related jobs found.
J
Jiajun Sun
Department of Statistics and Data Science, School of Economics, Xiamen University
Zhanrui Cai
Zhanrui Cai
The University of Hong Kong
Statistical inference with privacy