SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem
本文通过构建包含15000个积木堆叠问题的合成数据集SpatialBlock-15k,以提升大视觉语言模型的空间智能能力。
本文通过构建包含15000个积木堆叠问题的合成数据集SpatialBlock-15k,以提升大视觉语言模型的空间智能能力。
本文通过分析配体与目标口袋间的几何不匹配,提出SurfSpec框架,在无需了解脱靶结构的情况下优化先导化合物的特异性。
This study addresses the underestimation of automatic speech recognition (ASR) performance in non-English clinical settings, where valid orthographic variants of the same term—written in multiple scripts—are erroneously penalized by conventional string-matching evaluation metrics. To remedy this, the authors propose MultiClin, the first multi-script evaluation framework tailored for clinical ASR, which employs multi-reference assessment to more fairly measure model robustness to orthographic variation. The work systematically investigates the impact of script consistency in training data, revealing that script unification substantially improves recognition accuracy, while a 50% script-mixing ratio induces maximal model uncertainty. Experiments demonstrate that the proposed multi-script-aware evaluation significantly enhances the fidelity of performance measurement, and that script normalization consistently yields optimal results across diverse ASR architectures. The dataset and code are publicly released.
This work addresses the inefficiency of knowledge fusion in ensembles of multi-masked diffusion language models (MDLMs) by proposing the TIE framework. It reveals, for the first time, the dynamic confidence patterns at answer-relevant positions during MDLM decoding and leverages this insight to design an iterative ensemble mechanism based on confidence trajectory tracking. By integrating partial denoising sequence relaying with multi-model collaborative decoding, TIE dynamically selects the optimal model at each denoising stage to achieve complementary strengths across phases. Experimental results demonstrate that TIE significantly enhances generation quality across diverse reasoning tasks, validating its effectiveness and practicality in MDLM ensembling.
This work addresses the high computational cost and cross-modal interference inherent in existing vision-language reinforcement learning approaches based on autoregressive models, which require full image regeneration and employ shared reward mechanisms. To overcome these limitations, we introduce, for the first time, a multimodal discrete diffusion model into this task, enabling efficient inference through localized visual editing. We further propose a decoupled reward allocation strategy that assigns separate rewards to textual and visual segments. Integrated with the GRPO algorithm, our method achieves substantial computational savings—reducing inference FLOPs by 26.9% compared to autoregressive baselines—while significantly improving performance: the decoupled reward scheme yields an 11.2% gain over joint reward assignment and a 38.04% improvement over the base model, all without compromising task effectiveness.
本文通过构建包含15000个积木堆叠问题的合成数据集SpatialBlock-15k,以提升大视觉语言模型的空间智能能力。
本文通过分析配体与目标口袋间的几何不匹配,提出SurfSpec框架,在无需了解脱靶结构的情况下优化先导化合物的特异性。
This study addresses the underestimation of automatic speech recognition (ASR) performance in non-English clinical settings, where valid orthographic variants of the same term—written in multiple scripts—are erroneously penalized by conventional string-matching evaluation metrics. To remedy this, the authors propose MultiClin, the first multi-script evaluation framework tailored for clinical ASR, which employs multi-reference assessment to more fairly measure model robustness to orthographic variation. The work systematically investigates the impact of script consistency in training data, revealing that script unification substantially improves recognition accuracy, while a 50% script-mixing ratio induces maximal model uncertainty. Experiments demonstrate that the proposed multi-script-aware evaluation significantly enhances the fidelity of performance measurement, and that script normalization consistently yields optimal results across diverse ASR architectures. The dataset and code are publicly released.
This work addresses the inefficiency of knowledge fusion in ensembles of multi-masked diffusion language models (MDLMs) by proposing the TIE framework. It reveals, for the first time, the dynamic confidence patterns at answer-relevant positions during MDLM decoding and leverages this insight to design an iterative ensemble mechanism based on confidence trajectory tracking. By integrating partial denoising sequence relaying with multi-model collaborative decoding, TIE dynamically selects the optimal model at each denoising stage to achieve complementary strengths across phases. Experimental results demonstrate that TIE significantly enhances generation quality across diverse reasoning tasks, validating its effectiveness and practicality in MDLM ensembling.
This work addresses the high computational cost and cross-modal interference inherent in existing vision-language reinforcement learning approaches based on autoregressive models, which require full image regeneration and employ shared reward mechanisms. To overcome these limitations, we introduce, for the first time, a multimodal discrete diffusion model into this task, enabling efficient inference through localized visual editing. We further propose a decoupled reward allocation strategy that assigns separate rewards to textual and visual segments. Integrated with the GRPO algorithm, our method achieves substantial computational savings—reducing inference FLOPs by 26.9% compared to autoregressive baselines—while significantly improving performance: the decoupled reward scheme yields an 11.2% gain over joint reward assignment and a 38.04% improvement over the base model, all without compromising task effectiveness.