Cross-Relational Preference Learning for Better LLM Instruction Following

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为提高大型语言模型遵循复杂指令的能力,提出跨关系偏好学习框架,通过建模指令间关系生成多样化偏好数据,并引入基于原子约束的验证机制。
📝 Abstract
Large Language Models (LLMs) still exhibit limited capability in following complex instructions. While existing approaches often rely on preference learning to enhance this ability, they typically overlook the relationships between the permissible response spaces of different instructions, which restricts a model to align with subtle and diverse constraint variations. To address this, we propose Cross-Relational Preference Learning (CRPL), a novel framework for constructing preference data that explicitly models inter-instruction relationships through two key techniques: Cross-Relationship Perturbation and Cross-Region Pair Sampling. This enables the generation of more diverse preference data that captures a wide spectrum of constraint variations. Additionally, we introduce an atomic constraint-based verification mechanism to rigorously assess response satisfaction, ensuring high-quality preference pair construction. Extensive experiments across multiple preference learning methods (e.g., DPO, KTO), LLM backbones and four instruction-following benchmarks demonstrate that our approach achieves substantial improvements over prior baselines and exhibits strong generalization.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Complex Instructions
Preference Learning
Inter-Instruction Relationships
Constraint Variations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Relational Preference Learning (CRPL)
Cross-Relationship Perturbation
Cross-Region Pair Sampling
Atomic Constraint-Based Verification
🔎 Similar Papers
No similar papers found.