Institution profile

Kookmin University

Academic institutionasia · kr
Official website
Research library58linked papers
Opportunities0open roles
Selected work

Representative Papers

Design of Seamless Multi-modal Interaction Framework for Intelligent Virtual Agents in Wearable Mixed Reality Environment

Jul 01, 2019International Conference on Computer Animation and Social Agents

To address high interaction latency, modality fragmentation, and weak immersion in wearable mixed reality (MR) environments, this paper proposes a lightweight multimodal intelligent agent framework. Methodologically, it integrates spatial mapping, automatic speech recognition (ASR), gaze estimation, object detection, and knowledge-graph-driven dialogue, underpinned by a cloud-edge collaborative computing architecture for efficient computational offloading; it further introduces novel mechanisms for automatic speech–animation synchronization and human-like gaze modeling. The key contributions are: (1) the first realization of low-latency (2–4 seconds), high-naturalness virtual–physical interaction on resource-constrained edge devices; (2) a modular, cross-device-compatible framework supporting all SLAM-capable MR headsets. Evaluation in real-world museum and botanical garden deployments demonstrates significant improvements in user engagement and content retention rates.

22 citationsRead paper

FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge

Aug 15, 2026

This study addresses the lack of flood rescue-specific benchmarks and edge resource constraints in existing VLM reasoning segmentation research by constructing the first dedicated dataset and benchmark for this domain. Leveraging lightweight visual encoding, hierarchical reasoning, and intermediate representation compression, we evaluate system performance on real-world edge devices. Our analysis reveals partition-dependent accuracy variations and elucidates multidimensional trade-offs among accuracy, latency, energy consumption, and communication overhead. By achieving joint task- and system-level characterization, this work effectively facilitates the deployment of embodied intelligence in resource-constrained environments, bridging the gap between high-level semantic reasoning and practical edge implementation for emergency response scenarios.

0 citationsRead paper

SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features

Aug 11, 2026

This work addresses a fundamental limitation in combining quantization-aware training (QAT) with knowledge distillation (KD) in label-free scenarios: the mismatch in feature ranges between teacher and student models introduces an irreducible lower bound on distillation loss. To resolve this, the authors propose SQuaT, a framework that applies the student’s quantization parameters to the teacher’s features, thereby performing student-aware quantization of teacher features and eliminating the residual error caused by range misalignment. SQuaT is the first method to theoretically remove this distillation loss lower bound, enabling significant performance gains in extreme low-bit settings (e.g., 1–2 bits) without requiring labels. Extensive experiments demonstrate that SQuaT consistently outperforms strong baselines across diverse architectures and quantization configurations, particularly excelling at ultra-low bitwidths, while maintaining broad applicability due to its architecture-agnostic design.

0 citationsRead paper
Recent publications

Latest Papers

FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge

Aug 15, 2026

This study addresses the lack of flood rescue-specific benchmarks and edge resource constraints in existing VLM reasoning segmentation research by constructing the first dedicated dataset and benchmark for this domain. Leveraging lightweight visual encoding, hierarchical reasoning, and intermediate representation compression, we evaluate system performance on real-world edge devices. Our analysis reveals partition-dependent accuracy variations and elucidates multidimensional trade-offs among accuracy, latency, energy consumption, and communication overhead. By achieving joint task- and system-level characterization, this work effectively facilitates the deployment of embodied intelligence in resource-constrained environments, bridging the gap between high-level semantic reasoning and practical edge implementation for emergency response scenarios.

0 citationsRead paper

SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features

Aug 11, 2026

This work addresses a fundamental limitation in combining quantization-aware training (QAT) with knowledge distillation (KD) in label-free scenarios: the mismatch in feature ranges between teacher and student models introduces an irreducible lower bound on distillation loss. To resolve this, the authors propose SQuaT, a framework that applies the student’s quantization parameters to the teacher’s features, thereby performing student-aware quantization of teacher features and eliminating the residual error caused by range misalignment. SQuaT is the first method to theoretically remove this distillation loss lower bound, enabling significant performance gains in extreme low-bit settings (e.g., 1–2 bits) without requiring labels. Extensive experiments demonstrate that SQuaT consistently outperforms strong baselines across diverse architectures and quantization configurations, particularly excelling at ultra-low bitwidths, while maintaining broad applicability due to its architecture-agnostic design.

0 citationsRead paper

Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models

Aug 07, 2026

This work addresses the challenge of achieving general-purpose, efficient compression of vision-language models without relying on task-specific data or costly retraining. The authors propose PORTA, a novel framework that enables universal pruning by estimating cross-modal activation variability using a generic calibration set, thereby constructing a task- and modality-agnostic importance metric. PORTA further incorporates an adaptive layer-wise sparsity allocation mechanism to optimize compression efficiency. Evaluated on prominent vision-language models—including CLIP, BLIP, and Qwen2-VL—PORTA consistently outperforms existing retraining-free pruning methods under high compression ratios while preserving strong downstream task performance.

0 citationsRead paper