OPUS: A Simple yet Effective Unified Framework for Open-Vocabulary Detection
本文提出OPUS框架,通过语义丰富的视觉表示和可扩展的定位监督简化开放词汇检测,支持多种提示方式并在多项基准测试中取得领先性能。
本文提出OPUS框架,通过语义丰富的视觉表示和可扩展的定位监督简化开放词汇检测,支持多种提示方式并在多项基准测试中取得领先性能。
研究提出SOLO框架,通过Query Reconstructor和TA-MSE Distillation方法解决长距离复杂地形中人形机器人行走不稳定的问题。
This work addresses the inefficiency of existing high-resolution visual question answering (HR-VQA) methods, which suffer from redundant image cropping or re-encoding and neglect the dilution of fine-grained intermediate evidence in subsequent processing. The authors propose Thinking-Once, a training-free, single-pass framework that preserves critical entities and compact background context through question-conditioned attention reshaping and token selection within a single visual forward pass. Crucially, it routes intermediate-layer evidence directly to higher layers without additional training or repeated visual processing. This approach reveals, for the first time, that the performance bottleneck in HR-VQA stems from evidence dilution rather than insufficient input resolution. Evaluated across five multimodal large language models, Thinking-Once improves average scores by 3.1, 3.0, and 2.7 points on V*Bench, HRBench-4K, and HRBench-8K, respectively, reduces peak memory usage by approximately 4 GB, accelerates inference by 97.2% over DeepScan, and achieves an average cross-benchmark score of 82.7.
This work addresses the challenge of achieving both high perceptual quality and temporal consistency in video compression at extremely low bitrates (<0.005 bpp), where existing methods struggle. The authors propose a unified generative framework that introduces a causal tokenizer to decompose latent representations into I-latents and P-latents, and employs a Group-of-Latents strategy for structured modeling. Key latents are efficiently encoded via an I-frame Deep Compression Module (I-DCM), while a unified latent denoising module (U-LDM), built upon a pre-trained Diffusion Transformer (DiT), reconstructs high-fidelity intra-frame textures and coherent temporal dynamics directly from noise. Notably, this approach incurs no additional bitrate overhead and significantly outperforms current state-of-the-art methods under extreme low-bitrate constraints, delivering spatially detailed and temporally stable visual quality.
This work addresses the challenge of simulating human-like multi-step logical reasoning with auxiliary constructions in geometric problem solving by proposing a novel framework that integrates mathematical reasoning with procedural representations. The approach employs program code as an intermediate visual representation, decoupling discovery reasoning from code generation in a latent space and structuring the reasoning manifold through supervised fine-tuning. The study demonstrates that hierarchical syntactic code structures effectively encode rich mathematical semantics, offering greater expressiveness than purely visual representations. Experimental results show that the proposed method significantly enhances geometric reasoning performance while yielding clearer and more interpretable multi-step derivations.
本文提出OPUS框架,通过语义丰富的视觉表示和可扩展的定位监督简化开放词汇检测,支持多种提示方式并在多项基准测试中取得领先性能。
研究提出SOLO框架,通过Query Reconstructor和TA-MSE Distillation方法解决长距离复杂地形中人形机器人行走不稳定的问题。
This work addresses the inefficiency of existing high-resolution visual question answering (HR-VQA) methods, which suffer from redundant image cropping or re-encoding and neglect the dilution of fine-grained intermediate evidence in subsequent processing. The authors propose Thinking-Once, a training-free, single-pass framework that preserves critical entities and compact background context through question-conditioned attention reshaping and token selection within a single visual forward pass. Crucially, it routes intermediate-layer evidence directly to higher layers without additional training or repeated visual processing. This approach reveals, for the first time, that the performance bottleneck in HR-VQA stems from evidence dilution rather than insufficient input resolution. Evaluated across five multimodal large language models, Thinking-Once improves average scores by 3.1, 3.0, and 2.7 points on V*Bench, HRBench-4K, and HRBench-8K, respectively, reduces peak memory usage by approximately 4 GB, accelerates inference by 97.2% over DeepScan, and achieves an average cross-benchmark score of 82.7.
This work addresses the challenge of achieving both high perceptual quality and temporal consistency in video compression at extremely low bitrates (<0.005 bpp), where existing methods struggle. The authors propose a unified generative framework that introduces a causal tokenizer to decompose latent representations into I-latents and P-latents, and employs a Group-of-Latents strategy for structured modeling. Key latents are efficiently encoded via an I-frame Deep Compression Module (I-DCM), while a unified latent denoising module (U-LDM), built upon a pre-trained Diffusion Transformer (DiT), reconstructs high-fidelity intra-frame textures and coherent temporal dynamics directly from noise. Notably, this approach incurs no additional bitrate overhead and significantly outperforms current state-of-the-art methods under extreme low-bitrate constraints, delivering spatially detailed and temporally stable visual quality.
This work addresses the challenge of simulating human-like multi-step logical reasoning with auxiliary constructions in geometric problem solving by proposing a novel framework that integrates mathematical reasoning with procedural representations. The approach employs program code as an intermediate visual representation, decoupling discovery reasoning from code generation in a latent space and structuring the reasoning manifold through supervised fine-tuning. The study demonstrates that hierarchical syntactic code structures effectively encode rich mathematical semantics, offering greater expressiveness than purely visual representations. Experimental results show that the proposed method significantly enhances geometric reasoning performance while yielding clearer and more interpretable multi-step derivations.