TPTU: Large Language Model-based AI Agents for Task Planning and Tool Usage

📅 2023-08-07
📈 Citations: 49
Influential: 1
📄 PDF
🤖 AI Summary
To address LLMs’ weak task planning capability and brittle tool invocation in complex scenarios, this paper proposes TPTU—the first structured LLM agent framework that explicitly decouples and formalizes task planning and tool utilization as dual core competencies. It introduces a collaborative reasoning mechanism between single-step and sequential agents, supports extensible agent-type specialization, and integrates prompt engineering, dynamic tool selection, multi-step reasoning scheduling, and structured output parsing—ensuring compatibility with diverse mainstream LLMs. A systematic evaluation of 12 LLMs across representative tasks reveals three fundamental bottlenecks: insufficient planning depth, poor tool generalization, and weak error recovery. Based on these findings, we establish the first benchmark suite tailored for practical AI agents, providing both theoretical foundations and empirical pathways for advancing agent architecture design and capability enhancement. (149 words)
📝 Abstract
With recent advancements in natural language processing, Large Language Models (LLMs) have emerged as powerful tools for various real-world applications. Despite their prowess, the intrinsic generative abilities of LLMs may prove insufficient for handling complex tasks which necessitate a combination of task planning and the usage of external tools. In this paper, we first propose a structured framework tailored for LLM-based AI Agents and discuss the crucial capabilities necessary for tackling intricate problems. Within this framework, we design two distinct types of agents (i.e., one-step agent and sequential agent) to execute the inference process. Subsequently, we instantiate the framework using various LLMs and evaluate their Task Planning and Tool Usage (TPTU) abilities on typical tasks. By highlighting key findings and challenges, our goal is to provide a helpful resource for researchers and practitioners to leverage the power of LLMs in their AI applications. Our study emphasizes the substantial potential of these models, while also identifying areas that need more investigation and improvement.
Problem

Research questions and friction points this paper is trying to address.

Develops a structured framework for LLM-based AI agents
Designs agents for task planning and tool usage
Evaluates LLMs on complex tasks requiring external tools
Innovation

Methods, ideas, or system contributions that make the work stand out.

Framework for LLM-based AI agents with task planning
Two agent types: one-step and sequential for inference
Evaluation of Task Planning and Tool Usage (TPTU) abilities
J
Jingqing Ruan
University of Chinese Academy of Sciences, China; SenseTime Research, HKUST, Hong Kong
Y
Yihong Chen
University of Chinese Academy of Sciences, China; SenseTime Research, HKUST, Hong Kong
B
Bin Zhang
University of Chinese Academy of Sciences, China; SenseTime Research, HKUST, Hong Kong
Z
Zhiwei Xu
University of Chinese Academy of Sciences, China; SenseTime Research, HKUST, Hong Kong
T
Tianpeng Bao
SenseTime Research, HKUST, Hong Kong
G
Guoqing Du
SenseTime Research, HKUST, Hong Kong
S
Shiwei Shi
SenseTime Research, HKUST, Hong Kong
H
Hangyu Mao
SenseTime Research, HKUST, Hong Kong
Xingyu Zeng
Xingyu Zeng
Shenzhen University of Advanced Technology
Computer VisionDeep Learning
R
Rui Zhao
SenseTime Research, HKUST, Hong Kong