Cattle Trade: A Multi-Agent Benchmark for LLM Bluffing, Bidding, and Bargaining

📅 2026-05-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing benchmarks struggle to evaluate large language models’ (LLMs’) integrated strategic reasoning under conditions of incomplete information, adversarial interaction, and resource constraints. This work proposes the first multi-agent, long-horizon economic game environment that combines auction mechanisms, hidden-bid trading, bargaining, bluffing, and opponent modeling. The framework evaluates LLMs alongside rule-based agents over 50–60 dynamic, competitive rounds. Leveraging fine-grained behavioral logs, the study enables performance analysis beyond win rates, revealing that strategic consistency—such as expenditure efficiency and phase-adaptive bidding—is a stronger predictor of ranking than isolated skills or total spending. The analysis also uncovers characteristic failure modes of LLMs and demonstrates that heuristic code-based agents outperform LLMs in most scenarios.
📝 Abstract
We introduce \textsc{Cattle Trade, a multi-agent benchmark for evaluating large language models (LLMs) as agents in strategic reasoning under imperfect information, adversarial interaction, and resource constraints. The benchmark combines auctions, hidden-offer trade challenges (TCs), bargaining, bluffing, opponent modeling, and resource allocation within a single long-horizon game lasting 50--60 turns. Unlike prior agent benchmarks that test these abilities in isolation, \textsc{Cattle Trade} evaluates whether agents integrate them across a competitive, multi-agent economic game with conflicting incentives. The benchmark logs every bid, TC offer, counteroffer, and card selection, enabling behavioural analysis beyond final scores or win rates. We evaluate seven cost-efficient language models and three deterministic code agents across 242 games. Strategic coherence, in particular spending efficiency, resource discipline, and phase-adaptive bidding, is associated with rank more strongly than spending volume or any single subskill. Two heuristic code agents outperform most tested LLMs, and behavioural traces surface recurring LLM failure modes including overbidding, self-bidding, bankrupt TC initiation, and weak opponent-state adaptation. Evaluating agentic competence requires benchmarks that test the joint deployment of multiple capabilities in multi-agent environments with conflicting incentives, uncertainty, and economic dynamics.
Problem

Research questions and friction points this paper is trying to address.

multi-agent benchmark
strategic reasoning
imperfect information
adversarial interaction
resource constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-agent benchmark
strategic reasoning
imperfect information
LLM bluffing
behavioral analysis
🔎 Similar Papers
No similar papers found.