Orienteering Problem with Uncertain Time-Varying Rewards: Framework and Benchmark for Everyday Service Robotics

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文解决了具有不确定时变奖励的定向问题,通过三种不同规划视野和在线适应性的规划器来应对奖励的随机变化,并在服务机器人上进行了实验验证。
📝 Abstract
We present the orienteering problem with uncertain time-varying rewards (OP-UTVR), a novel variant of the orienteering problem (OP). While most existing OP formulations assume rewards to be known in advance, practical applications involve uncertain and time-varying rewards, as with shifting customer demand for delivery agents. OP-UTVR relaxes this assumption by allowing agents to estimate reward dynamics from observations and forecast future rewards. This enables informed routing decisions despite stochastic reward changes and inevitable prediction errors. We address this problem using three planners that differ in planning horizon and online adaptivity, and derive theoretical bounds on their performance under reward stochasticity. We further introduce a mobile service robot benchmark for OP-UTVR, where a robot navigates among pedestrians in indoor environments. Experiments reveal trade-offs between planning horizon and adaptivity, and demonstrate the effectiveness of long-horizon planning with online adaptation.
Problem

Research questions and friction points this paper is trying to address.

orienteering problem
uncertain time-varying rewards
reward dynamics
stochastic reward changes
planning horizon
Innovation

Methods, ideas, or system contributions that make the work stand out.

Orienteering Problem
Uncertain Time-Varying Rewards
Reward Dynamics Estimation
Online Adaptation