Institution profile

Profluent Bio

Industry researchnorthamerica · us
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebysheff Scalarization

Apr 14, 2026

This work addresses the limitation of linear scalarization in offline multi-objective reinforcement learning, which fails to recover non-convex Pareto fronts. To overcome this, the authors propose STOMP, the first algorithm to integrate smooth Chebyshev scalarization into offline multi-objective preference optimization. STOMP combines reward distribution normalization with Direct Preference Optimization (DPO) to establish a general framework for multi-objective alignment, effectively circumventing the theoretical constraints of linear scalarization and enabling full recovery of non-convex Pareto fronts. Evaluated on three protein language models and real-world experimental datasets, STOMP achieves state-of-the-art performance, attaining the highest hypervolume metric in eight out of nine settings and demonstrating superior capability in multi-attribute joint optimization.

0 citationsRead paper
Recent publications

Latest Papers

Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebysheff Scalarization

Apr 14, 2026

This work addresses the limitation of linear scalarization in offline multi-objective reinforcement learning, which fails to recover non-convex Pareto fronts. To overcome this, the authors propose STOMP, the first algorithm to integrate smooth Chebyshev scalarization into offline multi-objective preference optimization. STOMP combines reward distribution normalization with Direct Preference Optimization (DPO) to establish a general framework for multi-objective alignment, effectively circumventing the theoretical constraints of linear scalarization and enabling full recovery of non-convex Pareto fronts. Evaluated on three protein language models and real-world experimental datasets, STOMP achieves state-of-the-art performance, attaining the highest hypervolume metric in eight out of nine settings and demonstrating superior capability in multi-attribute joint optimization.

0 citationsRead paper