From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers

๐Ÿ“… 2026-08-24
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡้€š่ฟ‡ๅŸบไบŽ่ฏ„ๅˆ†ๆ ‡ๅ‡†็š„ๅฅ–ๅŠฑๆก†ๆžถ๏ผŒๅˆฉ็”จๆฃ€็ดขๅˆฐ็š„่ฏๆฎ็”Ÿๆˆๅคš็ปดๅบฆ็š„่ดจ้‡่ฏ„ไผฐ๏ผŒไปฅๆ”น่ฟ›ๅผ€ๆ”พ้ข†ๅŸŸ้—ฎ้ข˜ๅ›ž็ญ”็š„่ดจ้‡ใ€‚
๐Ÿ“ Abstract
Designing effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multiple aspects of answer quality that are difficult to capture with a holistic scalar objective. We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during post-training. Averaged across three evaluation axes (composition, grounding, and instruction-following), our approach improves over the instruction-tuned baseline by 6.5% and over flat rubric variants by 4%, with consistent gains across all evaluation datasets. Conditioning rubrics on retrieved evidence improves factual support, while decomposing rubrics into quality-specific dimensions further improves coherence, organization, and adherence to query requirements. Our results show that grounded, multi-dimensional rubrics provide more effective reward supervision for complex open-domain question answering.
Problem

Research questions and friction points this paper is trying to address.

reward signals
open-domain question answering
answer quality
rubric-based reward
Innovation

Methods, ideas, or system contributions that make the work stand out.

rubric-based reward
query-specific rubrics
multi-dimensional quality
retrieved evidence
fine-grained supervision