The Calls are Coming from Inside the Model: Investigating Probe-based Detection of Tool-Calling Errors in LLMs

πŸ“… 2026-08-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
η ”η©Άδ½Ώη”¨ηΊΏζ€§ζŽ’ι’ˆζ£€ζ΅‹ε€§εž‹θ―­θ¨€ζ¨‘εž‹δΈ­ε·₯ε…·θ°ƒη”¨ι”™θ――ηš„ζœ‰ζ•ˆζ€§οΌŒι€šθΏ‡18δΈͺζ¨‘εž‹ζ΅‹θ―•οΌŒε‘ηŽ°θ―₯ζ–Ήζ³•θƒ½ζœ‰ζ•ˆθ―†εˆ«ε€šη§ι”™θ――γ€‚
πŸ“ Abstract
The hidden states of large language models (LLMs) are known to capture rich information relating to model knowledge and behavior that can be hard to extract from examination of input and output alone. As LLM-based systems increasingly interface with the external world, one area of concern is detecting incorrect or improper use of tools. Motivated by this, we study the effectiveness of using linear probes to detect incorrect tool-calls, measuring probe efficacy across 18 tool-calling LLMs evaluated on the Berkeley Function Calling Leaderboard. Overall, we find that probing is an effective means to catch a range of different tool-calling errors, including errors arising from using an argument that has the wrong value but the correct type, which might not be recorded by standard logging frameworks. Important factors in success include model size, probing layer, and model post-training type. We also show that probes are capable of generalizing to novel types of errors, which is critical in real world deployments.
Problem

Research questions and friction points this paper is trying to address.

tool-calling errors
large language models
probing
Innovation

Methods, ideas, or system contributions that make the work stand out.

linear probes
tool-calling errors
generalization
πŸ”Ž Similar Papers
E
Eric Yeats
Pacific Northwest National Laboratory
Brendan Kennedy
Brendan Kennedy
Professor of Chemistry, The University of Sydney
CrystallographyInorganic ChemistryStructural Phase Tranitions
L
Loc Truong
Pacific Northwest National Laboratory
J
John Buckheit
Pacific Northwest National Laboratory
J
Jung Lee
Pacific Northwest National Laboratory
J
Jesse Friedbaum
Pacific Northwest National Laboratory
J
John Emanuello
National Security Agency
Henry Kvinge
Henry Kvinge
Pacific Northwest National Lab/University of Washington
representation learningadversarial machine learninggeometric deep learningrepresentation theorycombinatorics