ForgeTrain: Forging Production-Grade Training Frameworks via Harness-Driven AI Development
本文提出ForgeTrain,通过AI驱动的专用实现和迭代优化方法,解决了通用框架在特定场景下的性能优化问题,提升了训练效率。
本文提出ForgeTrain,通过AI驱动的专用实现和迭代优化方法,解决了通用框架在特定场景下的性能优化问题,提升了训练效率。
本文提出ForgeMegakernel,通过编码代理生成高性能解码巨核,解决自回归模型解码带宽受限问题,确保跨模型通用性和正确性。
本文通过揭示现有LLM的问题并引入递归自我改进(RSI)的概念及发展路线图,探讨了在不同场景下实现真正RSI的关键挑战与方法。
研究通过构建LexIssue基准和法律知识库,利用检索增强生成方法解决中文民事诉讼中法律争议点识别问题。
This study addresses the challenges of parameter memorization dependency, insufficient library knowledge, and inadequate feedback mechanisms in automatic mathematical formalization. We propose an iterative optimization framework integrating Mathlib retrieval with compiler-driven feedback. Leveraging the newly constructed FormalVerse dataset comprising 367K samples, we train an 8B model via supervised fine-tuning and reinforcement learning, introducing a novel retrieval planner and verification-guided refinement mechanism to overcome single-pass generation limitations. The resulting model achieves Pass@8 scores of 88.06% (SC) and 72.37% (CC) across six benchmarks, significantly outperforming multiple specialized 32B models. These results demonstrate substantial improvements in both syntactic correctness and semantic consistency for mathematical formalization tasks.
本文提出ForgeTrain,通过AI驱动的专用实现和迭代优化方法,解决了通用框架在特定场景下的性能优化问题,提升了训练效率。
本文提出ForgeMegakernel,通过编码代理生成高性能解码巨核,解决自回归模型解码带宽受限问题,确保跨模型通用性和正确性。
本文通过揭示现有LLM的问题并引入递归自我改进(RSI)的概念及发展路线图,探讨了在不同场景下实现真正RSI的关键挑战与方法。
研究通过构建LexIssue基准和法律知识库,利用检索增强生成方法解决中文民事诉讼中法律争议点识别问题。
This study addresses the challenges of parameter memorization dependency, insufficient library knowledge, and inadequate feedback mechanisms in automatic mathematical formalization. We propose an iterative optimization framework integrating Mathlib retrieval with compiler-driven feedback. Leveraging the newly constructed FormalVerse dataset comprising 367K samples, we train an 8B model via supervised fine-tuning and reinforcement learning, introducing a novel retrieval planner and verification-guided refinement mechanism to overcome single-pass generation limitations. The resulting model achieves Pass@8 scores of 88.06% (SC) and 72.37% (CC) across six benchmarks, significantly outperforming multiple specialized 32B models. These results demonstrate substantial improvements in both syntactic correctness and semantic consistency for mathematical formalization tasks.