Agentic Search Spaces for Tabular Machine Learning
研究利用先进的AI系统扩展表格模型的超参数搜索空间,通过代理提出模块化管道的候选实现,并使用经典HPO算法优化,从而在多个数据集上提高了模型性能。
研究利用先进的AI系统扩展表格模型的超参数搜索空间,通过代理提出模块化管道的候选实现,并使用经典HPO算法优化,从而在多个数据集上提高了模型性能。
研究了在给定Hamming距离r内字符串的最大Kolmogorov复杂度g_x(r)的可能取值及函数形式,通过界定g_x(r)的上下限,并探讨了其性质。
为解决语言模型推理时计算成本高和离散化问题,提出使用轻量级投影器替代传统头部的方法,在连续空间中进行推理,实验显示该方法有效提高了性能。
为解决多模态大模型在处理专业文档时的推理能力评估不足问题,通过构建包含1000个基于英俄双语文本丰富的企业和学术文档的问题基准BEAR-Bench进行评测。
This study addresses ongoing debates regarding the efficacy of training-free inference enhancement methods in modern Transformers by systematically evaluating eight Classifier-Free Guidance (CFG) derivatives on open-weight Rectified-Flow Transformers and compositional alignment benchmarks. The findings reveal that no evaluated method consistently outperforms standard CFG; improvements from Adaptive Projected Guidance largely fall within error margins, while attention perturbation techniques demonstrate unstable performance. By delineating the effectiveness boundaries of training-free guidance approaches, this work establishes that standard CFG remains the most robust and cost-effective baseline for current applications. These results provide critical empirical evidence to inform future research directions in inference-time scaling and model alignment, clarifying misconceptions about newer alternatives' superiority over established guidance mechanisms.
研究利用先进的AI系统扩展表格模型的超参数搜索空间,通过代理提出模块化管道的候选实现,并使用经典HPO算法优化,从而在多个数据集上提高了模型性能。
研究了在给定Hamming距离r内字符串的最大Kolmogorov复杂度g_x(r)的可能取值及函数形式,通过界定g_x(r)的上下限,并探讨了其性质。
为解决语言模型推理时计算成本高和离散化问题,提出使用轻量级投影器替代传统头部的方法,在连续空间中进行推理,实验显示该方法有效提高了性能。
为解决多模态大模型在处理专业文档时的推理能力评估不足问题,通过构建包含1000个基于英俄双语文本丰富的企业和学术文档的问题基准BEAR-Bench进行评测。
This study addresses ongoing debates regarding the efficacy of training-free inference enhancement methods in modern Transformers by systematically evaluating eight Classifier-Free Guidance (CFG) derivatives on open-weight Rectified-Flow Transformers and compositional alignment benchmarks. The findings reveal that no evaluated method consistently outperforms standard CFG; improvements from Adaptive Projected Guidance largely fall within error margins, while attention perturbation techniques demonstrate unstable performance. By delineating the effectiveness boundaries of training-free guidance approaches, this work establishes that standard CFG remains the most robust and cost-effective baseline for current applications. These results provide critical empirical evidence to inform future research directions in inference-time scaling and model alignment, clarifying misconceptions about newer alternatives' superiority over established guidance mechanisms.