Institution profile

Grupo Bancolombia

Industry researchsouthamerica · co
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

World Models Unlock Optimal Foraging Strategies in Reinforcement Learning Agents

Dec 13, 2025

This study investigates the computational mechanisms underlying the “when-to-leave” decision in biological patch foraging—a core ecological decision—and leverages these insights to design more interpretable, biologically grounded AI decision models. We propose a model-driven reinforcement learning framework integrating sparse world modeling, predictive representation learning, and forward simulation. We provide the first theoretical proof and empirical validation that agents equipped with a learned world model autonomously converge to the biologically optimal patch-leaving strategy predicted by the Marginal Value Theorem (MVT), without explicit programming. Critically, their decisions are driven by minimization of prediction error—not merely reward maximization. Compared to conventional model-free RL, our approach significantly improves fidelity to observed foraging behavior, ecological plausibility, and decision interpretability. This work establishes a novel paradigm bridging computational neuroscience and trustworthy AI.

0 citationsRead paper
Recent publications

Latest Papers

World Models Unlock Optimal Foraging Strategies in Reinforcement Learning Agents

Dec 13, 2025

This study investigates the computational mechanisms underlying the “when-to-leave” decision in biological patch foraging—a core ecological decision—and leverages these insights to design more interpretable, biologically grounded AI decision models. We propose a model-driven reinforcement learning framework integrating sparse world modeling, predictive representation learning, and forward simulation. We provide the first theoretical proof and empirical validation that agents equipped with a learned world model autonomously converge to the biologically optimal patch-leaving strategy predicted by the Marginal Value Theorem (MVT), without explicit programming. Critically, their decisions are driven by minimization of prediction error—not merely reward maximization. Compared to conventional model-free RL, our approach significantly improves fidelity to observed foraging behavior, ecological plausibility, and decision interpretability. This work establishes a novel paradigm bridging computational neuroscience and trustworthy AI.

0 citationsRead paper