From Errors to Rules: Iterative Prompt Optimization for Text Classification

📅 2026-06-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited generalization of existing prompt optimization methods and their lack of systematic mechanisms for correcting error patterns. The authors propose ERGO, a novel approach that iteratively traverses the training set, diagnoses classification errors, and establishes a “diagnosis–prescription–rewrite” feedback loop to automatically generate interpretable decision rules for refining prompt templates. ERGO introduces the first complementary framework linking task characteristics with prompt optimization paradigms, revealing that tasks exhibiting concentrated errors among label pairs are particularly amenable to error-driven strategies. Experimental results demonstrate that ERGO achieves state-of-the-art accuracy on learnable-boundary tasks such as TREC (90.0%) and CLINC150 (94.4%), converges within 3–5 iterations, and produces highly interpretable rules.
📝 Abstract
Prompt optimization for text classification spans diverse approaches, from demonstration selection to exploration-based search to error-driven diagnosis, each with known but incompletely characterized strengths and limitations. We conduct a comprehensive empirical study across diverse classification benchmarks (2 to 150 classes) comparing these paradigms through both quantitative evaluation and qualitative analysis of optimization traces, revealing that each paradigm excels on structurally different task types and that no single method dominates. Guided by these insights, we propose Error-Guided Optimization (ERGO), an error-driven method that iterates over the full training set in non-overlapping batches, diagnoses classification failures, and generates targeted decision rules through a diagnose-prescribe-rewrite feedback loop. ERGO achieves the best accuracy on tasks where errors concentrate in specific confused label pairs (which we term boundary-learnable tasks): TREC: 90.0%, CLINC150: 94.4%, converges in 3-5 iterations, and produces interpretable decision rules. While ERGO does not achieve the highest overall average, it fills a complementary role: demonstration-based ICL wins on coverage-dependent tasks, exploration-based search wins on many-class intent, and ERGO wins where decision boundaries are learnable from error patterns. We provide a complementarity framework linking task characteristics to optimal paradigm selection, offering practical guidance for practitioners.
Problem

Research questions and friction points this paper is trying to address.

prompt optimization
text classification
error-driven diagnosis
task characteristics
decision boundaries
Innovation

Methods, ideas, or system contributions that make the work stand out.

Error-Guided Optimization
Prompt Optimization
Text Classification
Decision Rules
Interpretable AI
💼 Related Jobs
No related jobs found.
Y
Yueying Cui
Amazon Web Services
R
Renhao Xue
Amazon Web Services
Yi Zhang
Yi Zhang
Principal Applied Scientist, AWS Agentic AI Labs
Natural Language ProcessingComputational LinguisticsSyntaxParsingArtificial Intelligence
M
Mukul Prasad
Amazon Web Services