Evaluating Rational Contracting in Natural Language

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of systematic evaluation of rationality and cooperativeness in current language agents when negotiating and executing natural language contracts in complex, multi-turn, and uncertain environments. We propose the first evaluation framework tailored to temporally extended, condition-dependent, and incomplete natural language contracts, introducing a cooperativeness dimension that moves beyond purely profit-driven objectives. The framework includes formalized quantitative metrics and baseline methods. Through multi-round negotiation simulations grounded in large language models and explicit modeling of environmental uncertainty, our experiments on ContractSim reveal that while agents can achieve efficient agreements under low uncertainty, they frequently fail to establish mutually beneficial contracts under high uncertainty and often breach agreements during execution to exploit short-term gains.
📝 Abstract
The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instead of just proposing bids or following hard-coded protocols, such agents can be used to negotiate and execute agreements in open-ended natural language. However, most evaluations of these abilities have focused on one-off exchanges or simple economic games, leaving open the rich space of time-extended, contingent, and incomplete contracts made expressible by language; they also focus on raw profit, without measuring the qualities required for trustworthy contracting. We address this by formulating a rational framework for how agents should negotiate and perform natural language contracts in uncertain multi-step environments. Within this framework, we develop metrics and baselines for quantifying rational and cooperative play. To evaluate how agents perform at such contracting, we instantiate our framework in ContractSim, an evaluation suite where two players negotiate and execute a multi-turn supplier contract under environmental and inter-player uncertainty. Across six environments and three supplier settings (catering, hotel cleaning, and AI hosting) we find that current LLM-based agents reach agreement reliably, and negotiate efficient contracts when environmental uncertainty is low. However, under high uncertainty, they often fail to negotiate satisfiable, efficient, or mutually beneficial contracts. They are also frequently uncooperative when executing contracts, violating contract terms for additional profit even when contracts are easy to satisfy. These findings highlight room for improvement in the design of language agents that can negotiate, interpret, and execute contracts both rationally and cooperatively.
Problem

Research questions and friction points this paper is trying to address.

natural language contracts
rational contracting
cooperative behavior
uncertainty
language-based AI agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

natural language contracts
rational negotiation
cooperative AI
ContractSim
language-based agents