Institution profile

Indian Institute of Technology Bombay

Academic institutionasia · in
Official website
Research library336linked papers
Opportunities0open roles
Selected work

Representative Papers

Semi-Bandit Learning for Monotone Stochastic Optimization*

Dec 24, 2023IEEE Annual Symposium on Foundations of Computer Science

This paper addresses monotone stochastic optimization problems with unknown distributions—such as prophet inequalities and Pandora’s box—where classical approaches rely on full distributional knowledge. Method: We propose the first unified semi-bandit online learning framework that learns and approximates the optimal policy solely from observed samples of probed random variables, without any prior distributional information. Our approach integrates semi-bandit feedback modeling, a monotonicity-aware stochastic probing strategy, and refined regret analysis. Contribution/Results: We establish a near-optimal regret bound of $O(sqrt{T log T})$, which strictly improves upon fundamental lower bounds under both full-information and pure bandit settings for multiple canonical problems. This is the first result to break the long-standing dependence on distributional priors in stochastic optimization, providing both a general theoretical foundation and an efficient algorithmic pathway for online approximation in distribution-agnostic environments.

2 citations1 influentialRead paper

Label-Efficient Hyperspectral Image Classification via Spectral FiLM Modulation of Low-Level Pretrained Diffusion Features

Dec 03, 2025

Hyperspectral image (HSI) classification suffers from low spatial resolution and severe label scarcity. To address these challenges, we propose a label-efficient framework that freezes a diffusion model pre-trained on natural images to extract robust low-level spatial features; introduces a spectral-aware FiLM modulation module to dynamically condition these spatial features with spectral information, enabling cross-modal and cross-domain feature fusion; and employs a lightweight classification head optimized end-to-end. To our knowledge, this is the first work to effectively transfer low-level spatial representations from pre-trained diffusion models to HSI classification. Our method achieves state-of-the-art performance on two recent benchmarks using only extremely sparse annotations. Ablation studies confirm the critical roles of both the diffusion-based feature transfer mechanism and the FiLM-based fusion strategy, significantly improving fine-grained land-cover classification under few-shot settings.

1 citationsRead paper

Motion Planning of Nonholonomic Cooperative Mobile Manipulators

Feb 08, 2025

This work addresses the cooperative object transport task for nonholonomic mobile manipulator robots (MMRs) in environments with static and dynamic obstacles. Methodologically, it proposes an integrated framework combining offline global planning with online coordinated control. It introduces, for the first time, a convex polygonal constraint-space modeling technique based on visibility vertices; unifies nonlinear model predictive control (NMPC) for simultaneous base and manipulator trajectory planning, explicitly coupling kinematics, dynamics, and nonholonomic constraints; and integrates real-time trajectory optimization with torque limiting for safety and feasibility. Evaluated in simulation and on multiple hardware platforms, the system achieves multi-arm cooperative manipulation, dynamic obstacle avoidance, and high-precision trajectory tracking at a planning frequency ≥10 Hz. All solutions strictly satisfy kinematic feasibility, dynamic feasibility, and safety requirements throughout execution.

1 citationsRead paper

Lagrangian Index Policy for Restless Bandits with Average Reward

Dec 17, 2024arXiv.org

This paper investigates the non-stationary Restless Multi-Armed Bandit (RMAB) problem under the average-reward criterion. To address the failure of Whittle’s Index Policy (WIP) in degenerate scenarios and its high memory overhead, we propose the Lagrangian Index Policy (LIP)—the first index policy grounded in Lagrangian duality and exchangeability analysis. Leveraging the de Finetti theorem, we establish its asymptotic optimality in the homogeneous-arm limit. LIP supports online learning under model uncertainty and admits unified implementation via tabular Q-learning or neural-network-based RL, reducing memory consumption by an order of magnitude. For restart-type models—including web crawling and weighted Age-of-Information minimization—we derive closed-form LIP indices analytically. Experiments demonstrate that LIP maintains high robustness and near-optimality even when WIP collapses, significantly outperforming existing approaches in challenging non-stationary regimes.

1 citationsRead paper
Recent publications

Latest Papers