Abstract
Model-free offline reinforcement learning learns policies from static, pre-collected datasets but often suffers from overestimating the values of out-of-distribution (OOD) data, which leads to biased value predictions and degraded performance. Existing methods mainly address this issue by penalizing OOD actions, but their extrapolation capabilities are limited by two key factors: (i) action-only penalties do not handle state-space extrapolation, causing cascading errors in unseen states; and (ii) uniform penalties suppress potentially beneficial extrapolation, biasing learning toward suboptimal dataset behavior. To overcome these limitations, we introduce ExtrapoBoost Offline Reinforcement Learning (EBORL), a method that improves extrapolation by jointly addressing both OOD states and actions. EBORL leverages an ensemble of Q-networks, each independently minimizing Temporal Difference error on in-distribution data, while maximizing the diversity of input gradients across the Q-networks for OOD state-action pairs. Theoretically, we provide a decomposition of the extrapolation error and prove that EBORL achieves a smaller upper bound on extrapolation error compared to baseline methods. Empirically, EBORL demonstrates strong performance on D4RL tasks with fewer ensemble heads and lower computational cost, and real-world evaluations further highlight its practical applicability.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Cognitive and Developmental Systems |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Extrapolation
- Model-free offline reinforcement learning
- Out-of-distribution
Fingerprint
Dive into the research topics of 'Boosting Extrapolation Capabilities in Model-free Offline Reinforcement Learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver