Institution profile

KRAFTON Inc.

Industry researchasia · kr
Official website
Research library53linked papers
Opportunities0open roles
Selected work

Representative Papers

Imperceptible Protection against Style Imitation from Diffusion Models

Mar 28, 2024arXiv.org

To address copyright infringement and artistic style appropriation risks posed by diffusion models, this paper proposes a visually lossless copyright protection method. The approach comprises three key contributions: (1) perception-sensitive map-guided instance-aware fine-tuning, enabling fine-grained stylistic perturbation; (2) difficulty-aware dynamic intensity modulation, which adaptively adjusts perturbation magnitude based on the sample’s stylistic mimicability; and (3) a multi-scale perceptual constraint library, jointly optimizing defense robustness and image fidelity. Without introducing perceptible visual artifacts, the method achieves over 92% style imitation suppression, reduces LPIPS by 41%, and improves FID by 27%, significantly outperforming existing state-of-the-art methods.

7 citationsRead paper

Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers

Jun 30, 2026

Existing latent chain-of-thought (Latent CoT) methods significantly underperform explicit CoT when model size exceeds 1 billion parameters, with the performance gap widening as scale increases. This work proposes a recurrent deep Transformer architecture that enhances computational depth through weight sharing and applies explicit CoT supervision via cross-entropy loss in parallel across multiple latent state positions. For the first time at the 3B parameter scale, this approach closes the performance gap between latent and explicit CoT, achieving comparable accuracy while reducing inference latency by 2.5–6.9×. Experiments demonstrate the critical roles of recurrence and parallel supervision in latent reasoning, revealing that latent states are interpretable and well-aligned with CoT: a base language model head can recover both correct and alternative intermediate reasoning steps from these representations.

0 citationsRead paper

AsyncOPD: How Stale Can On-Policy Distillation Be?

Jun 23, 2026

This work addresses the instability and performance degradation in asynchronous online policy distillation (OPD) caused by training on stale policy data. It is the first to demonstrate that the reverse KL divergence is highly sensitive to outdated data, whereas the forward KL divergence exhibits greater robustness. Building on this insight, the authors propose an efficient alternative that recomputes the KL signal using the current student model and introduces a multi-sample Monte Carlo estimator to balance bias and variance. The resulting open-source AsyncOPD framework achieves comparable accuracy to synchronous training while improving throughput by 1.6–3.8×.

0 citationsRead paper

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views

Jun 22, 2026

This work addresses the challenges of inconsistent cross-view reasoning and fragile view selection in multi-view 3D visual question answering, which arise when relying solely on sparse answer-level supervision. To overcome these issues, the authors decouple the reasoning process into three stages: global map construction, question-guided view trajectory planning, and egocentric answer prediction. They introduce a dense, annotation-free reward mechanism that provides process-level supervision by combining global geometric consistency with local sequential view-selection rewards. Built upon frozen 3D vision foundation models (e.g., VGGT+SAM3), their approach incorporates pseudo-target generation, trajectory-level policy optimization (GRPO), and map-anchored learning. Experiments on MindCube, VSI-Bench, and BLINK (MV) demonstrate substantial improvements over strong multi-image baselines, validating the efficacy of dense process supervision.

0 citationsRead paper
Recent publications

Latest Papers

Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers

Jun 30, 2026

Existing latent chain-of-thought (Latent CoT) methods significantly underperform explicit CoT when model size exceeds 1 billion parameters, with the performance gap widening as scale increases. This work proposes a recurrent deep Transformer architecture that enhances computational depth through weight sharing and applies explicit CoT supervision via cross-entropy loss in parallel across multiple latent state positions. For the first time at the 3B parameter scale, this approach closes the performance gap between latent and explicit CoT, achieving comparable accuracy while reducing inference latency by 2.5–6.9×. Experiments demonstrate the critical roles of recurrence and parallel supervision in latent reasoning, revealing that latent states are interpretable and well-aligned with CoT: a base language model head can recover both correct and alternative intermediate reasoning steps from these representations.

0 citationsRead paper

AsyncOPD: How Stale Can On-Policy Distillation Be?

Jun 23, 2026

This work addresses the instability and performance degradation in asynchronous online policy distillation (OPD) caused by training on stale policy data. It is the first to demonstrate that the reverse KL divergence is highly sensitive to outdated data, whereas the forward KL divergence exhibits greater robustness. Building on this insight, the authors propose an efficient alternative that recomputes the KL signal using the current student model and introduces a multi-sample Monte Carlo estimator to balance bias and variance. The resulting open-source AsyncOPD framework achieves comparable accuracy to synchronous training while improving throughput by 1.6–3.8×.

0 citationsRead paper

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views

Jun 22, 2026

This work addresses the challenges of inconsistent cross-view reasoning and fragile view selection in multi-view 3D visual question answering, which arise when relying solely on sparse answer-level supervision. To overcome these issues, the authors decouple the reasoning process into three stages: global map construction, question-guided view trajectory planning, and egocentric answer prediction. They introduce a dense, annotation-free reward mechanism that provides process-level supervision by combining global geometric consistency with local sequential view-selection rewards. Built upon frozen 3D vision foundation models (e.g., VGGT+SAM3), their approach incorporates pseudo-target generation, trajectory-level policy optimization (GRPO), and map-anchored learning. Experiments on MindCube, VSI-Bench, and BLINK (MV) demonstrate substantial improvements over strong multi-image baselines, validating the efficacy of dense process supervision.

0 citationsRead paper

Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents

Jun 06, 2026

This work addresses the limitations of existing evaluation methods that rely on passive observation and struggle to capture critical behaviors—such as conflict resolution—exhibited by large language model–driven social agents under specific triggering conditions. To overcome this, the authors propose an active evaluation framework that deploys embedded assessment agents within the environment to proactively induce scenarios aligned with social norms, thereby eliciting evaluable interaction trajectories from target agents. The framework integrates a context-generation mechanism, a dialogue-action protocol, and an evidence-driven scoring strategy, shifting evaluation from passive to active. Evaluated in a life-simulation environment encompassing 32 human-designed social norms, the approach significantly improves both coverage and alignment with human annotations, effectively uncovering latent social competencies invisible to passive assessment methods.

0 citationsRead paper