World Models Unlock Optimal Foraging Strategies in Reinforcement Learning Agents

📅 2025-12-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the computational mechanisms underlying the “when-to-leave” decision in biological patch foraging—a core ecological decision—and leverages these insights to design more interpretable, biologically grounded AI decision models. We propose a model-driven reinforcement learning framework integrating sparse world modeling, predictive representation learning, and forward simulation. We provide the first theoretical proof and empirical validation that agents equipped with a learned world model autonomously converge to the biologically optimal patch-leaving strategy predicted by the Marginal Value Theorem (MVT), without explicit programming. Critically, their decisions are driven by minimization of prediction error—not merely reward maximization. Compared to conventional model-free RL, our approach significantly improves fidelity to observed foraging behavior, ecological plausibility, and decision interpretability. This work establishes a novel paradigm bridging computational neuroscience and trustworthy AI.

Technology Category

Application Category

📝 Abstract
Patch foraging involves the deliberate and planned process of determining the optimal time to depart from a resource-rich region and investigate potentially more beneficial alternatives. The Marginal Value Theorem (MVT) is frequently used to characterize this process, offering an optimality model for such foraging behaviors. Although this model has been widely used to make predictions in behavioral ecology, discovering the computational mechanisms that facilitate the emergence of optimal patch-foraging decisions in biological foragers remains under investigation. Here, we show that artificial foragers equipped with learned world models naturally converge to MVT-aligned strategies. Using a model-based reinforcement learning agent that acquires a parsimonious predictive representation of its environment, we demonstrate that anticipatory capabilities, rather than reward maximization alone, drive efficient patch-leaving behavior. Compared with standard model-free RL agents, these model-based agents exhibit decision patterns similar to many of their biological counterparts, suggesting that predictive world models can serve as a foundation for more explainable and biologically grounded decision-making in AI systems. Overall, our findings highlight the value of ecological optimality principles for advancing interpretable and adaptive AI.
Problem

Research questions and friction points this paper is trying to address.

Discover computational mechanisms for optimal patch-foraging decisions in biology
Show model-based RL agents converge to Marginal Value Theorem strategies
Demonstrate predictive world models enable explainable, biologically grounded AI decisions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model-based RL agents learn predictive world models
Anticipatory capabilities drive efficient patch-leaving behavior
World models align with Marginal Value Theorem strategies
💼 Related Jobs
No related jobs found.
Y
Yesid Fonseca
Department of Biomedical Engineering, Universidad de los Andes, Bogotá, Colombia
M
Manuel S. Ríos
Center of Excellence in Analytics, Artificial Intelligence, and Information Governance, Bancolombia, Colombia
N
Nicanor Quijano
Department of Electrical and Electronics Engineering, Universidad de los Andes, Bogotá, Colombia
L
Luis F. Giraldo
Department of Biomedical Engineering, Universidad de los Andes, Bogotá, Colombia