Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过约束训练目标与基础模型的知识对齐来减少事实性幻觉,提出两种新方法:证据重写和回忆重写,实验表明这些方法有效。
📝 Abstract
Supervised fine-tuning (SFT) trains a base language model to imitate target responses, and these targets may require knowledge the base model has not robustly internalized. We study this as a source of hallucinations and frame a group of mitigation methods as \emph{knowledge-aligned SFT}: constraining SFT training targets to the base model's parametric knowledge. Under a unified setup, we compare existing generation-based and estimation-based knowledge-alignment methods and introduce two new variants: Evidence Rewrite, which verifies base-model generations using external evidence, and Recall Rewrite, which retains claims only when they can be consistently recalled by the base model. Experiments with Qwen 3 4B and OLMo 3 7B show that knowledge-aligned SFT can reduce factual hallucinations on WildHalu and Biography while largely preserving general capabilities. Recall Rewrite yields the strongest factuality gains and improves refusal behavior on UnknownBench. It thereby confirms that SFT targets beyond the base model's knowledge drive hallucination behavior.
Problem

Research questions and friction points this paper is trying to address.

Supervised Fine-Tuning
Hallucinations
Knowledge-Aligned
Base Model
Factual
Innovation

Methods, ideas, or system contributions that make the work stand out.

knowledge-aligned SFT
Evidence Rewrite
Recall Rewrite
💼 Related Jobs
No related jobs found.
A
Arthur Becker
AppTek GmbH, Aachen, Germany
J
Jakob Kemmler
AppTek GmbH, Aachen, Germany
David Thulke
David Thulke
RWTH Aachen University | AppTek
large language modelsretrieval augmented generation
C
Christine Schäfer
AppTek GmbH, Aachen, Germany
C
Christian Dugast
AppTek GmbH, Aachen, Germany
Hermann Ney
Hermann Ney
RWTH Aachen University
Machine LearningSpeech RecognitionMachine TranslationComputer Vision