Institution profile

SK Telecom

Industry researchasia · kr
Official website
Research library20linked papers
Opportunities0open roles
Selected work

Representative Papers

A.X K2 Technical Report

Aug 30, 2026

本文介绍了A.X K2,一个688B参数的语言模型,通过Sparse Gated Attention和Gated Norm技术提高长文本处理效率与质量,超越前代模型。

0 citationsRead paper

Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target

May 17, 2026

This work addresses the issue of policy update deviation in generative large language model–based recommender systems under continual learning, where only biased contextual bandit feedback—characterized by reliable positive signals and ambiguous non-responses due to exposure bias—is available. To mitigate this, the paper proposes the Anchored Bandit Policy Optimization (ABPO) framework, which innovatively treats historically exposed items as fixed anchors within a group-wise relative policy optimization scheme. ABPO jointly alleviates exposure bias and negative signal ambiguity by integrating inverse propensity score weighting with an asymmetric feedback reliability model grounded in the recommender’s output confidence. Experiments across five domains from Amazon Reviews and MovieLens demonstrate that ABPO significantly improves recommendation accuracy while effectively reducing bias induced by historical deployment policies.

0 citationsRead paper

DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization

Mar 26, 2026

This work addresses the inefficiencies in existing distributed training frameworks for multimodal large language models, which overlook the heterogeneity of input data modalities, leading to imbalanced computational loads and suboptimal GPU utilization. To tackle this issue, the study introduces data-characteristic awareness into training scheduling—a novel approach that leverages runtime performance profiling and data feature modeling to construct a predictive scheduling policy. This policy enables dynamic load balancing across pipeline stages and micro-batches, effectively mitigating computation skew caused by modality disparities. Evaluated on large-scale multimodal benchmarks, the proposed method achieves up to a 3.6× speedup in training throughput compared to state-of-the-art distributed training frameworks.

0 citationsRead paper

A.X K1 Technical Report

Jan 14, 2026

This work proposes A.X K1, a 519-billion-parameter mixture-of-experts (MoE) language model trained from scratch under constrained computational budgets, designed to simultaneously enhance multilingual—particularly Korean—reasoning capabilities and inference efficiency. Leveraging a 10-trillion-token corpus, multi-stage data curation, scaling-law-informed training configurations, and an innovative Think-Fusion training strategy, the model enables users to explicitly control reasoning mode switching. Experimental results demonstrate that A.X K1 achieves state-of-the-art performance among open-source models across multiple benchmarks, significantly outperforming existing approaches—especially on Korean-language tasks—while maintaining high inference efficiency and deployment flexibility.

0 citationsRead paper
Recent publications

Latest Papers

A.X K2 Technical Report

Aug 30, 2026

本文介绍了A.X K2,一个688B参数的语言模型,通过Sparse Gated Attention和Gated Norm技术提高长文本处理效率与质量,超越前代模型。

0 citationsRead paper

Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target

May 17, 2026

This work addresses the issue of policy update deviation in generative large language model–based recommender systems under continual learning, where only biased contextual bandit feedback—characterized by reliable positive signals and ambiguous non-responses due to exposure bias—is available. To mitigate this, the paper proposes the Anchored Bandit Policy Optimization (ABPO) framework, which innovatively treats historically exposed items as fixed anchors within a group-wise relative policy optimization scheme. ABPO jointly alleviates exposure bias and negative signal ambiguity by integrating inverse propensity score weighting with an asymmetric feedback reliability model grounded in the recommender’s output confidence. Experiments across five domains from Amazon Reviews and MovieLens demonstrate that ABPO significantly improves recommendation accuracy while effectively reducing bias induced by historical deployment policies.

0 citationsRead paper

DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization

Mar 26, 2026

This work addresses the inefficiencies in existing distributed training frameworks for multimodal large language models, which overlook the heterogeneity of input data modalities, leading to imbalanced computational loads and suboptimal GPU utilization. To tackle this issue, the study introduces data-characteristic awareness into training scheduling—a novel approach that leverages runtime performance profiling and data feature modeling to construct a predictive scheduling policy. This policy enables dynamic load balancing across pipeline stages and micro-batches, effectively mitigating computation skew caused by modality disparities. Evaluated on large-scale multimodal benchmarks, the proposed method achieves up to a 3.6× speedup in training throughput compared to state-of-the-art distributed training frameworks.

0 citationsRead paper

A.X K1 Technical Report

Jan 14, 2026

This work proposes A.X K1, a 519-billion-parameter mixture-of-experts (MoE) language model trained from scratch under constrained computational budgets, designed to simultaneously enhance multilingual—particularly Korean—reasoning capabilities and inference efficiency. Leveraging a 10-trillion-token corpus, multi-stage data curation, scaling-law-informed training configurations, and an innovative Think-Fusion training strategy, the model enables users to explicitly control reasoning mode switching. Experimental results demonstrate that A.X K1 achieves state-of-the-art performance among open-source models across multiple benchmarks, significantly outperforming existing approaches—especially on Korean-language tasks—while maintaining high inference efficiency and deployment flexibility.

0 citationsRead paper