A Cost-Aware Agentic Architecture for NL-to-SQL over Nested Enterprise Schemas, with a New Benchmark
本文针对企业级复杂模式下的自然语言转SQL问题,提出了一种成本感知的代理架构及新的基准测试DevRev NL2SQL。
本文针对企业级复杂模式下的自然语言转SQL问题,提出了一种成本感知的代理架构及新的基准测试DevRev NL2SQL。
In multi-turn reasoning fine-tuning, reasoning tokens from prior turns are excluded from subsequent inputs, preventing single-pass forward propagation and inducing computational redundancy. Method: We propose a response token replication mechanism coupled with a customized attention masking scheme. During training, historical response tokens are explicitly reused as input, while structured attention masks restrict attention computation to valid contextual spans. Contribution/Results: This enables end-to-end, single-pass forward propagation over multi-turn reasoning data for the first time. The approach is fully compatible with standard Transformer architectures—requiring no architectural modifications—while preserving inference capability. It significantly reduces training latency, improves fine-tuning efficiency and scalability, and establishes a new paradigm for efficiently aligning large language models with multi-turn reasoning capabilities.
本文针对企业级复杂模式下的自然语言转SQL问题,提出了一种成本感知的代理架构及新的基准测试DevRev NL2SQL。
In multi-turn reasoning fine-tuning, reasoning tokens from prior turns are excluded from subsequent inputs, preventing single-pass forward propagation and inducing computational redundancy. Method: We propose a response token replication mechanism coupled with a customized attention masking scheme. During training, historical response tokens are explicitly reused as input, while structured attention masks restrict attention computation to valid contextual spans. Contribution/Results: This enables end-to-end, single-pass forward propagation over multi-turn reasoning data for the first time. The approach is fully compatible with standard Transformer architectures—requiring no architectural modifications—while preserving inference capability. It significantly reduces training latency, improves fine-tuning efficiency and scalability, and establishes a new paradigm for efficiently aligning large language models with multi-turn reasoning capabilities.