Institution profile

NCSOFT

Industry researchasia · kr
Official website
Research library10linked papers
Opportunities0open roles
Selected work

Representative Papers

Dream to Generalize: Zero-Shot Model-Based Reinforcement Learning for Unseen Visual Distractions

Jun 26, 2023AAAI Conference on Artificial Intelligence

To address the degradation of model-based reinforcement learning (MBRL) generalization under high-dimensional visual observations corrupted by clouds, shadows, and illumination variations, this paper proposes Dr. G—a zero-shot model-based RL framework. Our approach tackles this challenge through three key contributions: (1) a novel dual-contrastive self-supervised learning mechanism that disentangles and encodes task-relevant features from multi-view augmented data; (2) recurrent state-wise inverse dynamics modeling to enhance the world model’s temporal causal understanding; and (3) zero-shot cross-background transfer without fine-tuning. Evaluated on DeepMind Control (with complex video backgrounds) and Robosuite (with randomized environments), Dr. G achieves performance gains of 117% and 14%, respectively, over state-of-the-art methods. The implementation is publicly available.

5 citationsRead paper

Investor risk profiles of large language models

Mar 10, 2026

This study systematically investigates how large language models (LLMs) form and express investor risk preferences in the context of retail investment advice. Using standardized financial risk questionnaires combined with prompt engineering and persona-based role assignments—such as age, wealth, and investment experience—the authors conduct multi-round evaluations of GPT, Gemini, and Llama. The work reveals, for the first time, that LLMs exhibit identifiable and adjustable default risk preferences, and quantifies their heterogeneous responsiveness to persona-induced interventions. Results show that Gemini displays a neutral and highly consistent risk profile, Llama leans conservative, and GPT tends toward aggressiveness but with notable volatility. All three models adapt their risk profiles according to assigned personas, yet the magnitude of adjustment varies significantly across models.

0 citationsRead paper

Human-AI Collaborative Bot Detection in MMORPGs

Aug 28, 2025

To address the challenge of detecting human-mimicking, fairness-compromising, and inherently uninterpretable leveling bots in MMORPGs, this paper proposes a human-in-the-loop unsupervised detection framework. Methodologically, it integrates contrastive representation learning with clustering to identify anomalous leveling patterns and—novelty—employs large language models (LLMs) as auxiliary reviewers, generating interpretable decision criteria via growth-curve visualization. The framework establishes an end-to-end detection–explanation–validation闭环 without requiring labeled data and supports scalable deployment. Experiments demonstrate significant improvements over baselines in both detection accuracy and traceability: manual review efficiency increases by 42%. This work introduces a new paradigm for game anti-cheating systems that is robust, transparent, and operationally viable.

0 citationsRead paper

Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models

Jun 25, 2025

Existing 3D texture synthesis methods suffer from single-view input constraints and inadequate geometric understanding, leading to inconsistent inter-patch seams, semantic distortion in occluded regions, and temporal flickering under dynamic scenes. To address these limitations, we propose the first video-generation-based framework for 3D texture synthesis, innovatively integrating geometry-aware diffusion with structural UV-space modeling. Our method employs mesh-structured conditioning to enforce geometric awareness during generation and introduces a spatiotemporal diffusion strategy explicitly defined over the UV parameterization—thereby modeling both inter-patch topological relationships and inter-frame temporal dependencies. Experiments demonstrate significant improvements over state-of-the-art approaches in three key metrics: texture fidelity, seam coherence, and temporal stability. The framework enables high-fidelity, dynamically consistent, and real-time 3D content generation.

0 citationsRead paper

Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues

Jun 01, 2025

Current large language models struggle to model nonverbal cues—such as gestures, facial expressions, and body language—hindering the development of immersive conversational AI. To address this, we introduce VENUS, the first large-scale multimodal dataset for video-driven dialogue, featuring time-aligned dialogue transcripts, fine-grained facial action units, and full-body pose annotations. Methodologically, we propose a novel nonverbal vector quantization (VQ) representation built upon a VQ-VAE, and design MARS: a text-video-action tri-modal joint modeling framework trained end-to-end via multimodal next-token prediction. Experimental results demonstrate that MARS significantly outperforms baselines in nonverbal cue generation and contextual consistency. Quantitative and qualitative analyses confirm VENUS’s broad coverage and high-quality annotations. Collectively, this work establishes a foundational resource and methodology for nonverbal understanding and generation in conversational AI.

0 citationsRead paper
Recent publications

Latest Papers

Investor risk profiles of large language models

Mar 10, 2026

This study systematically investigates how large language models (LLMs) form and express investor risk preferences in the context of retail investment advice. Using standardized financial risk questionnaires combined with prompt engineering and persona-based role assignments—such as age, wealth, and investment experience—the authors conduct multi-round evaluations of GPT, Gemini, and Llama. The work reveals, for the first time, that LLMs exhibit identifiable and adjustable default risk preferences, and quantifies their heterogeneous responsiveness to persona-induced interventions. Results show that Gemini displays a neutral and highly consistent risk profile, Llama leans conservative, and GPT tends toward aggressiveness but with notable volatility. All three models adapt their risk profiles according to assigned personas, yet the magnitude of adjustment varies significantly across models.

0 citationsRead paper

Human-AI Collaborative Bot Detection in MMORPGs

Aug 28, 2025

To address the challenge of detecting human-mimicking, fairness-compromising, and inherently uninterpretable leveling bots in MMORPGs, this paper proposes a human-in-the-loop unsupervised detection framework. Methodologically, it integrates contrastive representation learning with clustering to identify anomalous leveling patterns and—novelty—employs large language models (LLMs) as auxiliary reviewers, generating interpretable decision criteria via growth-curve visualization. The framework establishes an end-to-end detection–explanation–validation闭环 without requiring labeled data and supports scalable deployment. Experiments demonstrate significant improvements over baselines in both detection accuracy and traceability: manual review efficiency increases by 42%. This work introduces a new paradigm for game anti-cheating systems that is robust, transparent, and operationally viable.

0 citationsRead paper

Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models

Jun 25, 2025

Existing 3D texture synthesis methods suffer from single-view input constraints and inadequate geometric understanding, leading to inconsistent inter-patch seams, semantic distortion in occluded regions, and temporal flickering under dynamic scenes. To address these limitations, we propose the first video-generation-based framework for 3D texture synthesis, innovatively integrating geometry-aware diffusion with structural UV-space modeling. Our method employs mesh-structured conditioning to enforce geometric awareness during generation and introduces a spatiotemporal diffusion strategy explicitly defined over the UV parameterization—thereby modeling both inter-patch topological relationships and inter-frame temporal dependencies. Experiments demonstrate significant improvements over state-of-the-art approaches in three key metrics: texture fidelity, seam coherence, and temporal stability. The framework enables high-fidelity, dynamically consistent, and real-time 3D content generation.

0 citationsRead paper

Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues

Jun 01, 2025

Current large language models struggle to model nonverbal cues—such as gestures, facial expressions, and body language—hindering the development of immersive conversational AI. To address this, we introduce VENUS, the first large-scale multimodal dataset for video-driven dialogue, featuring time-aligned dialogue transcripts, fine-grained facial action units, and full-body pose annotations. Methodologically, we propose a novel nonverbal vector quantization (VQ) representation built upon a VQ-VAE, and design MARS: a text-video-action tri-modal joint modeling framework trained end-to-end via multimodal next-token prediction. Experimental results demonstrate that MARS significantly outperforms baselines in nonverbal cue generation and contextual consistency. Quantitative and qualitative analyses confirm VENUS’s broad coverage and high-quality annotations. Collectively, this work establishes a foundational resource and methodology for nonverbal understanding and generation in conversational AI.

0 citationsRead paper

EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild

Feb 17, 2025

This work addresses the critical challenge of “when to initiate speech” in real-world, first-person streaming video. We propose the first end-to-end framework for egocentric speech initiation timing prediction. Methodologically, we design a multimodal temporal online learning architecture that integrates real-time RGB visual feature extraction, context-aware attention, and sliding-window streaming inference; it is pre-trained on our large-scale in-the-wild dialogue dataset, YT-Conversation, enabling uncropped, full-temporal video understanding. Our key contribution is the first unified modeling of egocentric visual perception and dynamic speech initiation decision-making. Experiments on the EasyCom and Ego4D benchmarks demonstrate significant improvements over random and silence baselines. Ablation studies confirm the essential roles of multimodal input and context length in prediction accuracy.

0 citationsRead paper