Prompt-Robust Language Models: Which Training Strategies Work?

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究对比了不同训练策略对大型语言模型提示敏感性的改善效果,发现现有方法虽优于标准微调,但最佳与最差提示间性能差距仍大。
📝 Abstract
Despite their strong performance, large language models remain highly sensitive to prompt formulation. Prior work addresses this through refined data construction or through dedicated robustness objectives. We reproduce and compare these strategies under controlled conditions, and measure how effective they are in addressing models' prompt sensitivity. We find the current robustness fine-tuning methods improve over standard fine-tuning and in-context learning, but the best-to-worst prompt gap remains as high as 40-57% of performance. Moreover, the recent robustness-enhancing methods we test - CoIN for contrastive alignment and PPCL for consistency regularization - often fail to outperform the simplest data construction strategy: training on one template per batch. Our diagnostics explain these results. The auxiliary objectives move the quantity they penalize, but do not generalize beyond it. Additionally, data construction strategies differ due to the conflicting signs of per-template gradients on 57-64% of parameters. Thus, batches that mix formulations force the optimizer to reconcile competing updates instead of finding a shared, prompt-agnostic one.
Problem

Research questions and friction points this paper is trying to address.

large language models
prompt sensitivity
robustness fine-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prompt Sensitivity
Robustness Fine-tuning
Data Construction Strategy
Contrastive Alignment
Consistency Regularization
🔎 Similar Papers
2023-11-15Conference on Empirical Methods in Natural Language ProcessingCitations: 2