Institution profile

DevRev

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Efficient Single-Pass Training for Multi-Turn Reasoning

Apr 25, 2025

In multi-turn reasoning fine-tuning, reasoning tokens from prior turns are excluded from subsequent inputs, preventing single-pass forward propagation and inducing computational redundancy. Method: We propose a response token replication mechanism coupled with a customized attention masking scheme. During training, historical response tokens are explicitly reused as input, while structured attention masks restrict attention computation to valid contextual spans. Contribution/Results: This enables end-to-end, single-pass forward propagation over multi-turn reasoning data for the first time. The approach is fully compatible with standard Transformer architectures—requiring no architectural modifications—while preserving inference capability. It significantly reduces training latency, improves fine-tuning efficiency and scalability, and establishes a new paradigm for efficiently aligning large language models with multi-turn reasoning capabilities.

0 citationsRead paper
Recent publications

Latest Papers

Efficient Single-Pass Training for Multi-Turn Reasoning

Apr 25, 2025

In multi-turn reasoning fine-tuning, reasoning tokens from prior turns are excluded from subsequent inputs, preventing single-pass forward propagation and inducing computational redundancy. Method: We propose a response token replication mechanism coupled with a customized attention masking scheme. During training, historical response tokens are explicitly reused as input, while structured attention masks restrict attention computation to valid contextual spans. Contribution/Results: This enables end-to-end, single-pass forward propagation over multi-turn reasoning data for the first time. The approach is fully compatible with standard Transformer architectures—requiring no architectural modifications—while preserving inference capability. It significantly reduces training latency, improves fine-tuning efficiency and scalability, and establishes a new paradigm for efficiently aligning large language models with multi-turn reasoning capabilities.

0 citationsRead paper