PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control
本文提出PonderPounce方法,利用预训练的多模态大语言模型作为机器人控制的情境记忆引擎,通过双系统协作实现高效任务执行。
本文提出PonderPounce方法,利用预训练的多模态大语言模型作为机器人控制的情境记忆引擎,通过双系统协作实现高效任务执行。
研究针对语音和文本间结构差异导致的语义对齐问题,提出一种新框架以增强语音语言模型性能。
This work addresses the computational bottleneck in vision-language models during inference caused by processing a large number of visual tokens, noting that existing pruning methods fail to account for the varying functional roles of redundant tokens. Building on EmbedLens-based token role identification, the study reveals that prevailing pruning strategies implicitly favor certain token roles, yet this preference shows no direct correlation with downstream performance. The authors propose a role-aware pruning strategy that deliberately preserves specific non-“alive” tokens—such as those with weaker semantic alignment—and demonstrate that doing so can maintain or even enhance model performance. These findings underscore the critical influence of functional token roles on importance assessment and offer a novel perspective for designing efficient vision-language models.
This study investigates the propagation of automatic speech recognition (ASR) errors in Korean spoken question-answering systems that employ an ASR–large language model (LLM) cascade, and the resulting semantic failures. Through ASR error analysis, semantic failure evaluation, and comparative experiments with end-to-end audio-language models, the authors demonstrate that even single-character ASR errors can cause complete downstream QA failure. They find that information loss during ASR is the primary driver of performance degradation, and LLMs of varying capabilities exhibit similar sensitivity to such errors. The results indicate that end-to-end models directly processing audio inputs significantly outperform conventional cascaded architectures in noisy conditions, effectively mitigating semantic information loss caused by transcription errors.
Existing diffusion models struggle to edit high-resolution images with arbitrary aspect ratios or resolutions significantly exceeding their training scale (e.g., 512×512), as naive tiling often introduces structural distortions and content duplication. This work proposes a fine-tuning-free editing framework that integrates tiled latent-space inversion with an enhanced noise-damped classifier-free guidance strategy (NDCFG++). By effectively leveraging the generative priors of pre-trained text-to-image diffusion models, the method achieves coherent and photorealistic high-resolution edits while preserving image identity. It supports inputs of arbitrary dimensions and consistently produces structurally consistent and detail-rich results across diverse resolutions, without requiring model fine-tuning or additional optimization.
本文提出PonderPounce方法,利用预训练的多模态大语言模型作为机器人控制的情境记忆引擎,通过双系统协作实现高效任务执行。
研究针对语音和文本间结构差异导致的语义对齐问题,提出一种新框架以增强语音语言模型性能。
This work addresses the computational bottleneck in vision-language models during inference caused by processing a large number of visual tokens, noting that existing pruning methods fail to account for the varying functional roles of redundant tokens. Building on EmbedLens-based token role identification, the study reveals that prevailing pruning strategies implicitly favor certain token roles, yet this preference shows no direct correlation with downstream performance. The authors propose a role-aware pruning strategy that deliberately preserves specific non-“alive” tokens—such as those with weaker semantic alignment—and demonstrate that doing so can maintain or even enhance model performance. These findings underscore the critical influence of functional token roles on importance assessment and offer a novel perspective for designing efficient vision-language models.
This study investigates the propagation of automatic speech recognition (ASR) errors in Korean spoken question-answering systems that employ an ASR–large language model (LLM) cascade, and the resulting semantic failures. Through ASR error analysis, semantic failure evaluation, and comparative experiments with end-to-end audio-language models, the authors demonstrate that even single-character ASR errors can cause complete downstream QA failure. They find that information loss during ASR is the primary driver of performance degradation, and LLMs of varying capabilities exhibit similar sensitivity to such errors. The results indicate that end-to-end models directly processing audio inputs significantly outperform conventional cascaded architectures in noisy conditions, effectively mitigating semantic information loss caused by transcription errors.
Existing diffusion models struggle to edit high-resolution images with arbitrary aspect ratios or resolutions significantly exceeding their training scale (e.g., 512×512), as naive tiling often introduces structural distortions and content duplication. This work proposes a fine-tuning-free editing framework that integrates tiled latent-space inversion with an enhanced noise-damped classifier-free guidance strategy (NDCFG++). By effectively leveraging the generative priors of pre-trained text-to-image diffusion models, the method achieves coherent and photorealistic high-resolution edits while preserving image identity. It supports inputs of arbitrary dimensions and consistently produces structurally consistent and detail-rich results across diverse resolutions, without requiring model fine-tuning or additional optimization.