In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过设计满足特权访问要求的神经反馈范式,探讨大型语言模型是否能控制其内部表征,结果表明之前报道的控制可能依赖于表面机制。
📝 Abstract
Whether large language models (LLMs) can control their own internal representations matters for both machine metacognition and AI safety. A recent study applied neurofeedback to LLMs and claimed that they can control their internal representations. However, the reported control may rely on superficial mechanisms rather than genuine internal access because the control targets in that study are not privileged, meaning that a third party can infer them from the prompt. We redesign the neurofeedback paradigm for LLMs so that the control target satisfies the privileged access requirement, which is closer to neurofeedback experiments in human cognitive neuroscience. Under this stricter setting, the models do not demonstrate reliable control over privileged internal representations, suggesting that previously reported control cannot exclude the possibility that it relies on superficial mechanisms. Our results indicate that rigorous assessments of metacognition in LLMs require evaluation methods that demand privileged access.
Problem

Research questions and friction points this paper is trying to address.

large language models
internal representations
privileged access
neurofeedback
metacognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Neurofeedback
Privileged Access
Internal Representations
🔎 Similar Papers
No similar papers found.