AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents
This work addresses the limitation of existing embodied language model agents, which respond only to literal queries and fail to capture users’ implicit intentions—such as inferring availability or emotional state from a question like “Where is Lin Wei?” To bridge this gap, the authors propose IntentFrame, a framework that introduces structured intent reasoning between scene perception and tool invocation to explicitly model latent user needs. IntentFrame incorporates a gap-calibration mechanism that dynamically allocates exploration budgets and selects appropriate tools. This approach enables the first targeted probing for implicit intents without relying on memorized answers. Evaluated on a benchmark of 100 queries across four scenarios, IntentFrame improves implicit intent coverage by 0.07 over ReAct (p < 10⁻⁶), reduces probing actions by 82%, and entirely eliminates privacy-violating tool calls.