From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
为解决企业自托管LLM导致的GPU资源碎片化问题,通过分析生产错误并训练特定领域的专家模型,使用两阶段SLERP方法合并,提高了模型性能和效率。
为解决企业自托管LLM导致的GPU资源碎片化问题,通过分析生产错误并训练特定领域的专家模型,使用两阶段SLERP方法合并,提高了模型性能和效率。
本文提出一种基于混合专家视觉语言模型的文档理解系统,通过难度感知的数据整理和质量调整的成本分析方法,有效降低了受监管行业文档处理成本。
为解决语言模型推理时计算成本高和离散化问题,提出使用轻量级投影器替代传统头部的方法,在连续空间中进行推理,实验显示该方法有效提高了性能。
Existing JEPA-based world models are constrained by a single inference paradigm, limiting their flexibility at deployment. This work proposes Qantara, the first single JEPA model capable of unifying latent-space planning, behavioral cloning, and inverse dynamics inference without retraining. Central to Qantara is a bridging flow training framework that jointly optimizes Brownian bridge interpolation along the state axis and noise-to-data flow matching along the action axis, augmented with a video inverse synthesis query mechanism that focuses on squared-noise boundary regions to enhance training efficiency. Experiments demonstrate that Qantara achieves an average success rate of 91.2% on the LeWM control suite, sets a new state of the art on OGBench-Cube with a 7.7% absolute improvement in success rate, and simultaneously attains 82–83% behavioral cloning accuracy and 71–73% video inverse trajectory success on Push-T and Cube tasks.
This work addresses the challenge of learning control policies from expert video demonstrations in the absence of environmental reward signals. To this end, the authors propose the Rank-Then-Act framework, which leverages a pretrained vision-language model as an ordinal progress scorer. A reinforcement learning reward signal is constructed by measuring the Spearman rank correlation between the model’s predicted ordering and the true temporal sequence of demonstration frames. This approach decouples reward learning from absolute calibration, enabling stable transfer across tasks and environments. Experimental results on benchmarks including PyBoy, PointMaze, and MetaWorld demonstrate that the method significantly outperforms existing video-based reward learning and ranking baselines. Furthermore, the study validates that a single pretrained scorer can be effectively reused across multiple tasks without retraining.
为解决企业自托管LLM导致的GPU资源碎片化问题,通过分析生产错误并训练特定领域的专家模型,使用两阶段SLERP方法合并,提高了模型性能和效率。
本文提出一种基于混合专家视觉语言模型的文档理解系统,通过难度感知的数据整理和质量调整的成本分析方法,有效降低了受监管行业文档处理成本。
为解决语言模型推理时计算成本高和离散化问题,提出使用轻量级投影器替代传统头部的方法,在连续空间中进行推理,实验显示该方法有效提高了性能。
Existing JEPA-based world models are constrained by a single inference paradigm, limiting their flexibility at deployment. This work proposes Qantara, the first single JEPA model capable of unifying latent-space planning, behavioral cloning, and inverse dynamics inference without retraining. Central to Qantara is a bridging flow training framework that jointly optimizes Brownian bridge interpolation along the state axis and noise-to-data flow matching along the action axis, augmented with a video inverse synthesis query mechanism that focuses on squared-noise boundary regions to enhance training efficiency. Experiments demonstrate that Qantara achieves an average success rate of 91.2% on the LeWM control suite, sets a new state of the art on OGBench-Cube with a 7.7% absolute improvement in success rate, and simultaneously attains 82–83% behavioral cloning accuracy and 71–73% video inverse trajectory success on Push-T and Cube tasks.
This work addresses the challenge of learning control policies from expert video demonstrations in the absence of environmental reward signals. To this end, the authors propose the Rank-Then-Act framework, which leverages a pretrained vision-language model as an ordinal progress scorer. A reinforcement learning reward signal is constructed by measuring the Spearman rank correlation between the model’s predicted ordering and the true temporal sequence of demonstration frames. This approach decouples reward learning from absolute calibration, enabling stable transfer across tasks and environments. Experimental results on benchmarks including PyBoy, PointMaze, and MetaWorld demonstrate that the method significantly outperforms existing video-based reward learning and ranking baselines. Furthermore, the study validates that a single pretrained scorer can be effectively reused across multiple tasks without retraining.