Institution profile

Ewha Womans University

Academic institutionasia · kr
Official website
Research library124linked papers
Opportunities0open roles
Selected work

Representative Papers

Seeing What You Say: Expressive Image Generation from Speech

Nov 05, 2025

This work addresses end-to-end speech-to-image generation—producing semantically accurate and emotionally consistent images directly from raw speech, bypassing automatic speech recognition (ASR) as an intermediate step. To this end, we propose VoxStudio, a unified framework featuring: (i) a speech information bottleneck module that compresses raw speech into compact tokens encoding semantic, prosodic, and affective information; (ii) VoxEmoset, the first large-scale emotional speech–image paired dataset; and (iii) a cross-modal alignment mechanism enabling joint optimization of speech representations and the image generator. Experiments on SpokenCOCO, Flickr8kAudio, and VoxEmoset demonstrate substantial improvements in emotional consistency and visual fidelity of generated images. Results validate the critical role of paralinguistic modeling—particularly prosody and emotion—in speech-driven multimodal generation, establishing a novel paradigm for direct speech-conditioned visual synthesis.

1 citationsRead paper

What You Ask is What You Ground: Bridging Question Intent to Temporal Evidence for Grounded VideoQA

Aug 16, 2026

This study addresses temporal localization failures in video question answering caused by modality isolation and insufficient question injection. To overcome these limitations, we propose GroundFormer, which injects question intent prior to localization via learnable communication tokens. By integrating decomposed MIL cross-attention with Gaussian smoothing, the method achieves precise alignment of QA-aware temporal evidence, further optimized through a hierarchical multimodal contrastive loss. Experimental results demonstrate that GroundFormer attains state-of-the-art performance on both NExT-GQA and STAR datasets. The proposed approach significantly enhances discriminative question-aware temporal localization capabilities, effectively overcoming the constraints inherent in traditional methods.

0 citationsRead paper

Deciding When to Switch: E-Processes for Adaptive Minimax Training for Generative Adversarial Nets

Aug 10, 2026

This work addresses the limitation in generative adversarial network (GAN) training where discriminator and generator updates rely on fixed ratios or heuristic criteria. It introduces e-processes into GAN training for the first time, proposing an adaptive switching mechanism grounded in sequential hypothesis testing. By constructing e-processes via conditional e-values, the method provides rigorous Type I error control at any stopping time and enables data-driven dynamic updates through monitoring the degree of distributional separation. Integrated with stochastic min-max optimization and latent variable resampling, the proposed approach matches or surpasses the performance of the best fixed-ratio baselines across multimodal synthetic and standard image benchmarks, while remaining compatible with various mainstream GAN objective functions.

0 citationsRead paper
Recent publications

Latest Papers

What You Ask is What You Ground: Bridging Question Intent to Temporal Evidence for Grounded VideoQA

Aug 16, 2026

This study addresses temporal localization failures in video question answering caused by modality isolation and insufficient question injection. To overcome these limitations, we propose GroundFormer, which injects question intent prior to localization via learnable communication tokens. By integrating decomposed MIL cross-attention with Gaussian smoothing, the method achieves precise alignment of QA-aware temporal evidence, further optimized through a hierarchical multimodal contrastive loss. Experimental results demonstrate that GroundFormer attains state-of-the-art performance on both NExT-GQA and STAR datasets. The proposed approach significantly enhances discriminative question-aware temporal localization capabilities, effectively overcoming the constraints inherent in traditional methods.

0 citationsRead paper

Deciding When to Switch: E-Processes for Adaptive Minimax Training for Generative Adversarial Nets

Aug 10, 2026

This work addresses the limitation in generative adversarial network (GAN) training where discriminator and generator updates rely on fixed ratios or heuristic criteria. It introduces e-processes into GAN training for the first time, proposing an adaptive switching mechanism grounded in sequential hypothesis testing. By constructing e-processes via conditional e-values, the method provides rigorous Type I error control at any stopping time and enables data-driven dynamic updates through monitoring the degree of distributional separation. Integrated with stochastic min-max optimization and latent variable resampling, the proposed approach matches or surpasses the performance of the best fixed-ratio baselines across multimodal synthetic and standard image benchmarks, while remaining compatible with various mainstream GAN objective functions.

0 citationsRead paper

LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs

Aug 09, 2026

This study addresses the tendency of large language models (LLMs) to overlook reference data embedded in server instructions within Model Context Protocol (MCP) environments, instead inefficiently invoking search tools and wasting computational resources. Through 54,000 controlled trials, the authors systematically evaluate the behavior of 24 prominent LLMs on legal information retrieval tasks under MCP, employing tool ablation, a 2³ factorial design, and cross-model-family analysis. The work reveals— for the first time—that this behavior stems from preference rather than capability deficits. Notably, removing the search tool yields over 98% accuracy in 23 out of 24 models, and combining three targeted prompting interventions restores accuracy above 86% for 20 out of 24 models even when the tool is present. The findings advocate for MCP hosts to explicitly prioritize server-provided instructions.

0 citationsRead paper