Convergence rates for generative drifting flows: fixed-scale obstructions and multihead acceleration
研究探讨了生成漂移流的收敛速度问题,提出了一种多头方法来处理不同尺度的信息,从而恢复指数级收敛速度,解决了固定分辨率导致的慢收敛问题。
研究探讨了生成漂移流的收敛速度问题,提出了一种多头方法来处理不同尺度的信息,从而恢复指数级收敛速度,解决了固定分辨率导致的慢收敛问题。
本文提出DynEoMT方法,通过在线查询增强视频分割模型以预测区域动态性,解决了无法仅从语义推断物体是否独立于相机移动的问题。
研究了带有多个步骤转换预览的强化学习问题,证明了其在固定折扣因子下的NP难性,并提出了一种近似方案以实现高效近优规划。
This work addresses the performance degradation of the classical single-threshold strategy in single-choice prophet inequalities under ill-behaved distributions by introducing a relative variance constraint on the maximum as a nonparametric complexity measure. It pioneers the application of kernel methods to this domain, constructing a linear functional optimization framework over quantile function spaces. By establishing a strong minimax duality and leveraging infinite-dimensional convex programming alongside variational analysis, the paper precisely characterizes the optimal threshold under bounded variance conditions. Key contributions include an exact performance curve under the i.i.d. setting, an asymptotically optimal threshold for finite horizons, closed-form solutions for non-i.i.d. cases, and a rigorous separation of the performance bounds between the prophet-secretary model and the i.i.d. benchmark.
This work addresses the challenges of policy learning in contextual bandits with extremely large action spaces, where inefficient exploration, high variance of importance weights, and optimization difficulties commonly arise. To improve exploration efficiency in online settings, the authors propose two approaches—mixed-effects Thompson Sampling (meTS) and diffusion Thompson Sampling (dTS)—that explicitly model dependencies among actions. For offline settings, they introduce a latent-variable-based method, sDM, which integrates a differentiable pessimism mechanism with a concave policy-weighted log-likelihood objective to mitigate extrapolation bias and variance issues. Theoretical analysis yields regret bounds that scale with the effective number of actions, and empirical results demonstrate that the proposed methods significantly enhance both stability and performance of policy learning in large action spaces.
研究探讨了生成漂移流的收敛速度问题,提出了一种多头方法来处理不同尺度的信息,从而恢复指数级收敛速度,解决了固定分辨率导致的慢收敛问题。
本文提出DynEoMT方法,通过在线查询增强视频分割模型以预测区域动态性,解决了无法仅从语义推断物体是否独立于相机移动的问题。
研究了带有多个步骤转换预览的强化学习问题,证明了其在固定折扣因子下的NP难性,并提出了一种近似方案以实现高效近优规划。
This work addresses the performance degradation of the classical single-threshold strategy in single-choice prophet inequalities under ill-behaved distributions by introducing a relative variance constraint on the maximum as a nonparametric complexity measure. It pioneers the application of kernel methods to this domain, constructing a linear functional optimization framework over quantile function spaces. By establishing a strong minimax duality and leveraging infinite-dimensional convex programming alongside variational analysis, the paper precisely characterizes the optimal threshold under bounded variance conditions. Key contributions include an exact performance curve under the i.i.d. setting, an asymptotically optimal threshold for finite horizons, closed-form solutions for non-i.i.d. cases, and a rigorous separation of the performance bounds between the prophet-secretary model and the i.i.d. benchmark.
This work addresses the challenges of policy learning in contextual bandits with extremely large action spaces, where inefficient exploration, high variance of importance weights, and optimization difficulties commonly arise. To improve exploration efficiency in online settings, the authors propose two approaches—mixed-effects Thompson Sampling (meTS) and diffusion Thompson Sampling (dTS)—that explicitly model dependencies among actions. For offline settings, they introduce a latent-variable-based method, sDM, which integrates a differentiable pessimism mechanism with a concave policy-weighted log-likelihood objective to mitigate extrapolation bias and variance issues. Theoretical analysis yields regret bounds that scale with the effective number of actions, and empirical results demonstrate that the proposed methods significantly enhance both stability and performance of policy learning in large action spaces.