Risk-Conditioned Fine-Tuning of Large Language Models

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出风险条件RLHF框架,训练单一策略提供连续风险控制接口,解决大型语言模型在部署中遇到的罕见但严重的有害生成问题。
📝 Abstract
Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences. Existing Risk-Averse RLHF addresses this issue by optimizing Conditional Value-at-Risk (CVaR), but it trains policies for fixed risk levels and therefore cannot adjust the desired degree of risk aversion at inference time. In this paper, we propose risk-conditioned RLHF, a framework that trains a single policy that provides a continuous risk-control interface, enabling users to select different degrees of risk aversion without retraining or deploying multiple risk-specific models. Experiments across multiple benchmarks demonstrate that a single risk-conditioned policy can adapt to different risk levels at inference time, enabling more flexible and risk-aware LLM deployment.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Risk-Averse RLHF
Conditional Value-at-Risk
risk aversion
Innovation

Methods, ideas, or system contributions that make the work stand out.

risk-conditioned RLHF
continuous risk-control interface
single policy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zixuan Liu
Department of Computer Science, Tulane University, New Orleans, LA, 70118, USA
F
Fangzheng Wu
Department of Computer Science, Tulane University, New Orleans, LA, 70118, USA
Brian Summa
Brian Summa
Associate Professor, Tulane University
Zizhan Zheng
Zizhan Zheng
Tulane University
AI Security and SafetyReinforcement LearningGenerative AINetworks