LLM-Driven Data Generation and a Novel Soft Metric for Evaluating Text-to-SQL in Aviation MRO
In aviation MRO text-to-SQL, coarse-grained evaluation metrics (e.g., binary execution accuracy) and scarcity of high-quality annotated data jointly hinder progress. To address this, we propose: (1) an F1-based soft evaluation metric quantifying SQL semantic correctness via information overlap—enabling fine-grained, attributable assessment; and (2) a schema-driven LLM synthetic framework that leverages database structure-aware prompting and execution-result semantic alignment to generate high-fidelity question-SQL pairs. Evaluated on a real-world aviation MRO database, our soft metric significantly improves error localization capability. The synthesized data constitutes the first domain-specific text-to-SQL benchmark for aviation MRO, demonstrating superior reliability and validity over conventional evaluation paradigms.