A.X K2 Technical Report
本文介绍了A.X K2,一个688B参数的语言模型,通过Sparse Gated Attention和Gated Norm技术提高长文本处理效率与质量,超越前代模型。
本文介绍了A.X K2,一个688B参数的语言模型,通过Sparse Gated Attention和Gated Norm技术提高长文本处理效率与质量,超越前代模型。
本文提出DuELRec模型,通过领域门控双专家框架和双重采样对比学习方法,解决跨域序列推荐中的负迁移问题。
This work addresses the issue of policy update deviation in generative large language model–based recommender systems under continual learning, where only biased contextual bandit feedback—characterized by reliable positive signals and ambiguous non-responses due to exposure bias—is available. To mitigate this, the paper proposes the Anchored Bandit Policy Optimization (ABPO) framework, which innovatively treats historically exposed items as fixed anchors within a group-wise relative policy optimization scheme. ABPO jointly alleviates exposure bias and negative signal ambiguity by integrating inverse propensity score weighting with an asymmetric feedback reliability model grounded in the recommender’s output confidence. Experiments across five domains from Amazon Reviews and MovieLens demonstrate that ABPO significantly improves recommendation accuracy while effectively reducing bias induced by historical deployment policies.
This work addresses the inefficiencies in existing distributed training frameworks for multimodal large language models, which overlook the heterogeneity of input data modalities, leading to imbalanced computational loads and suboptimal GPU utilization. To tackle this issue, the study introduces data-characteristic awareness into training scheduling—a novel approach that leverages runtime performance profiling and data feature modeling to construct a predictive scheduling policy. This policy enables dynamic load balancing across pipeline stages and micro-batches, effectively mitigating computation skew caused by modality disparities. Evaluated on large-scale multimodal benchmarks, the proposed method achieves up to a 3.6× speedup in training throughput compared to state-of-the-art distributed training frameworks.
This work proposes A.X K1, a 519-billion-parameter mixture-of-experts (MoE) language model trained from scratch under constrained computational budgets, designed to simultaneously enhance multilingual—particularly Korean—reasoning capabilities and inference efficiency. Leveraging a 10-trillion-token corpus, multi-stage data curation, scaling-law-informed training configurations, and an innovative Think-Fusion training strategy, the model enables users to explicitly control reasoning mode switching. Experimental results demonstrate that A.X K1 achieves state-of-the-art performance among open-source models across multiple benchmarks, significantly outperforming existing approaches—especially on Korean-language tasks—while maintaining high inference efficiency and deployment flexibility.
本文介绍了A.X K2,一个688B参数的语言模型,通过Sparse Gated Attention和Gated Norm技术提高长文本处理效率与质量,超越前代模型。
本文提出DuELRec模型,通过领域门控双专家框架和双重采样对比学习方法,解决跨域序列推荐中的负迁移问题。
This work addresses the issue of policy update deviation in generative large language model–based recommender systems under continual learning, where only biased contextual bandit feedback—characterized by reliable positive signals and ambiguous non-responses due to exposure bias—is available. To mitigate this, the paper proposes the Anchored Bandit Policy Optimization (ABPO) framework, which innovatively treats historically exposed items as fixed anchors within a group-wise relative policy optimization scheme. ABPO jointly alleviates exposure bias and negative signal ambiguity by integrating inverse propensity score weighting with an asymmetric feedback reliability model grounded in the recommender’s output confidence. Experiments across five domains from Amazon Reviews and MovieLens demonstrate that ABPO significantly improves recommendation accuracy while effectively reducing bias induced by historical deployment policies.
This work addresses the inefficiencies in existing distributed training frameworks for multimodal large language models, which overlook the heterogeneity of input data modalities, leading to imbalanced computational loads and suboptimal GPU utilization. To tackle this issue, the study introduces data-characteristic awareness into training scheduling—a novel approach that leverages runtime performance profiling and data feature modeling to construct a predictive scheduling policy. This policy enables dynamic load balancing across pipeline stages and micro-batches, effectively mitigating computation skew caused by modality disparities. Evaluated on large-scale multimodal benchmarks, the proposed method achieves up to a 3.6× speedup in training throughput compared to state-of-the-art distributed training frameworks.
This work proposes A.X K1, a 519-billion-parameter mixture-of-experts (MoE) language model trained from scratch under constrained computational budgets, designed to simultaneously enhance multilingual—particularly Korean—reasoning capabilities and inference efficiency. Leveraging a 10-trillion-token corpus, multi-stage data curation, scaling-law-informed training configurations, and an innovative Think-Fusion training strategy, the model enables users to explicitly control reasoning mode switching. Experimental results demonstrate that A.X K1 achieves state-of-the-art performance among open-source models across multiple benchmarks, significantly outperforming existing approaches—especially on Korean-language tasks—while maintaining high inference efficiency and deployment flexibility.