Institution profile

Maum.AI

Industry researchnorthamerica · us
Official website
Research library12linked papers
Opportunities0open roles
Selected work

Representative Papers

Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles

Aug 05, 2026

This work addresses the computational bottleneck in vision-language models during inference caused by processing a large number of visual tokens, noting that existing pruning methods fail to account for the varying functional roles of redundant tokens. Building on EmbedLens-based token role identification, the study reveals that prevailing pruning strategies implicitly favor certain token roles, yet this preference shows no direct correlation with downstream performance. The authors propose a role-aware pruning strategy that deliberately preserves specific non-“alive” tokens—such as those with weaker semantic alignment—and demonstrate that doing so can maintain or even enhance model performance. These findings underscore the critical influence of functional token roles on importance assessment and offer a novel perspective for designing efficient vision-language models.

0 citationsRead paper

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades

May 17, 2026

This study investigates the propagation of automatic speech recognition (ASR) errors in Korean spoken question-answering systems that employ an ASR–large language model (LLM) cascade, and the resulting semantic failures. Through ASR error analysis, semantic failure evaluation, and comparative experiments with end-to-end audio-language models, the authors demonstrate that even single-character ASR errors can cause complete downstream QA failure. They find that information loss during ASR is the primary driver of performance degradation, and LLMs of varying capabilities exhibit similar sensitivity to such errors. The results indicate that end-to-end models directly processing audio inputs significantly outperform conventional cascaded architectures in noisy conditions, effectively mitigating semantic information loss caused by transcription errors.

0 citationsRead paper

EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model

Apr 11, 2026

Existing diffusion models struggle to edit high-resolution images with arbitrary aspect ratios or resolutions significantly exceeding their training scale (e.g., 512×512), as naive tiling often introduces structural distortions and content duplication. This work proposes a fine-tuning-free editing framework that integrates tiled latent-space inversion with an enhanced noise-damped classifier-free guidance strategy (NDCFG++). By effectively leveraging the generative priors of pre-trained text-to-image diffusion models, the method achieves coherent and photorealistic high-resolution edits while preserving image identity. It supports inputs of arbitrary dimensions and consistently produces structurally consistent and detail-rich results across diverse resolutions, without requiring model fine-tuning or additional optimization.

0 citationsRead paper
Recent publications

Latest Papers

Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles

Aug 05, 2026

This work addresses the computational bottleneck in vision-language models during inference caused by processing a large number of visual tokens, noting that existing pruning methods fail to account for the varying functional roles of redundant tokens. Building on EmbedLens-based token role identification, the study reveals that prevailing pruning strategies implicitly favor certain token roles, yet this preference shows no direct correlation with downstream performance. The authors propose a role-aware pruning strategy that deliberately preserves specific non-“alive” tokens—such as those with weaker semantic alignment—and demonstrate that doing so can maintain or even enhance model performance. These findings underscore the critical influence of functional token roles on importance assessment and offer a novel perspective for designing efficient vision-language models.

0 citationsRead paper

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades

May 17, 2026

This study investigates the propagation of automatic speech recognition (ASR) errors in Korean spoken question-answering systems that employ an ASR–large language model (LLM) cascade, and the resulting semantic failures. Through ASR error analysis, semantic failure evaluation, and comparative experiments with end-to-end audio-language models, the authors demonstrate that even single-character ASR errors can cause complete downstream QA failure. They find that information loss during ASR is the primary driver of performance degradation, and LLMs of varying capabilities exhibit similar sensitivity to such errors. The results indicate that end-to-end models directly processing audio inputs significantly outperform conventional cascaded architectures in noisy conditions, effectively mitigating semantic information loss caused by transcription errors.

0 citationsRead paper

EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model

Apr 11, 2026

Existing diffusion models struggle to edit high-resolution images with arbitrary aspect ratios or resolutions significantly exceeding their training scale (e.g., 512×512), as naive tiling often introduces structural distortions and content duplication. This work proposes a fine-tuning-free editing framework that integrates tiled latent-space inversion with an enhanced noise-damped classifier-free guidance strategy (NDCFG++). By effectively leveraging the generative priors of pre-trained text-to-image diffusion models, the method achieves coherent and photorealistic high-resolution edits while preserving image identity. It supports inputs of arbitrary dimensions and consistently produces structurally consistent and detail-rich results across diverse resolutions, without requiring model fine-tuning or additional optimization.

0 citationsRead paper