CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

๐Ÿ“… 2026-08-07
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the inefficiency of fixed offline datasets in reinforcement learning, where redundant transitions hinder training while naive subsampling risks discarding rare but critical transitions essential for long-horizon credit assignment. To tackle this, the authors propose CODS, a method that iteratively constructs a static yet high-quality data subset by alternately updating a critic network and selecting transitions with high Bellman residuals. By dynamically refining residual-based scores while maintaining a fixed dataset, CODS effectively preserves sparse rewards and ensures computational efficiency. Empirical results demonstrate that, under a stringent 10% data budget, CODS retains 96.6% of full-dataset performance on average across 20 D4RL taskโ€“algorithm combinations, substantially outperforming existing baselines, and further exhibits strong generalization on ALFWorld and GSM8K benchmarks.
๐Ÿ“ Abstract
Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can remove rare transitions needed for long-horizon credit assignment. We introduce CODS, a critic-guided selector that alternates between fitting an algorithm-matched critic and acquiring high-residual transitions before freezing a reusable subset. Unlike prioritized replay, CODS produces a static artifact; unlike one-shot residual selection, it refreshes scores as the critic changes. At a 10\% budget, CODS retains 96.6\% of eligible-pool performance across 20 valid D4RL task--algorithm cells. It exceeds ReDOR and OPER on 19/20 cells and every other subset baseline on 20/20; all six subset advantages remain significant under predeclared hierarchical inference with Holm correction. Holding total selector updates fixed, five acquisition rounds improve four representative cells by 11.23 points over one round and saturate thereafter. Equal-pass and equal-hour evaluations clarify that reuse, rather than a single-run speedup, creates the compute advantage. Mechanism and corruption interventions expose both useful sparse-reward enrichment and sensitivity to outliers. Finally, a whole-trace extension retains 95.4\% of pooled ALFWorld success and 96.5\% of pooled GSM8K exact match. CODS is therefore a reusable selection procedure, not a formal coreset guarantee.
Problem

Research questions and friction points this paper is trying to address.

offline reinforcement learning
data selection
reusable subset
long-horizon credit assignment
transition pool
Innovation

Methods, ideas, or system contributions that make the work stand out.

offline reinforcement learning
data selection
Bellman residual
reusable dataset
critic-guided sampling
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Ibne Farabi Shihab
Ibne Farabi Shihab
Iowa State University
Deep LearningroboticsLarge Language Model
Sanjeda Akter
Sanjeda Akter
Iowa State University
Quantum ComputingRLDLLLM
A
Abu Sa-Adat Mohamed Moon-Im Al Ahsan
Department of Computer Science & Engineering, BRAC University
M
Md Najmus Swaqeeb
Department of Computer Science & Engineering, BRAC University
Anuj Sharma
Anuj Sharma
Iowa State University
Transportation