OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization
针对NVFP4量化中激活异常值导致的精度下降问题,提出OCGQuant方法,通过异常-伴生组策略优化量化块组成,提高量化精度。
针对NVFP4量化中激活异常值导致的精度下降问题,提出OCGQuant方法,通过异常-伴生组策略优化量化块组成,提高量化精度。
This study addresses the limited generalization and high data dependency of Vision-Language-Action (VLA) models by proposing ART, a framework that integrates tool use into VLAs to reduce action space complexity. By combining few-shot dataset construction with long-horizon reasoning training, ART achieves modular model enhancement. Experiments demonstrate that ART outperforms mainstream baselines by 20% in success rate across both simulated and real-world tasks while significantly reducing data requirements. The approach markedly improves robotic generalization in complex environments and enables lightweight, efficient deployment. Ultimately, this work establishes a scalable paradigm for embodied intelligence, effectively mitigating key bottlenecks in current VLA architectures through strategic tool integration and optimized training methodologies.
This work addresses the challenge that existing open-vocabulary object recognition methods struggle to simultaneously achieve coarse-grained generality and fine-grained discriminative capability. To this end, we propose HyMOR, a hybrid framework that synergistically integrates multimodal large language models (MLLMs) and CLIP for the first time: the MLLM handles open-domain coarse-grained recognition, while CLIP specializes in fine-grained discrimination within domains such as flora and fauna. A Sentence-BERT-guided hybrid reasoning mechanism enables multi-granularity semantic alignment between the two streams. Built upon this architecture, we introduce a unified perception system tailored for educational games and release TBO, a textbook-derived dataset. Our approach achieves a 23.2% improvement in average SBert score, narrows the fine-grained recognition gap to 0.2%, and enhances general recognition performance by 2.5%, substantially strengthening the perceptual foundation for multimodal interactive learning.
This work addresses the open-set challenge in generalized category discovery (GCD), where only a subset of known classes is labeled and unknown categories must be identified. To tackle this problem, we propose the SSR²-GCD framework, which introduces semi-supervised rate reduction into GCD for the first time. By optimizing multimodal representation learning, our method enhances intra-modal alignment to construct structured feature distributions and leverages the prompt candidate mechanism of vision-language models (VLMs) to strengthen cross-modal knowledge transfer. Extensive experiments on both generic and fine-grained benchmark datasets demonstrate that SSR²-GCD significantly outperforms existing approaches, achieving state-of-the-art performance.
Existing methods for image and video object manipulation struggle to simultaneously preserve background content, maintain geometric consistency across viewpoints, and offer fine-grained user control. This work proposes Ctrl&Shift, an end-to-end diffusion framework that decomposes manipulation into two stages—object removal followed by camera-pose-guided reference-based inpainting—enabling geometrically consistent editing within a unified diffusion process without explicit 3D modeling. By integrating explicit camera pose control, reference-guided inpainting, and a multi-task, multi-stage training strategy, the method effectively disentangles background, identity, and pose signals. This design preserves generalization to real-world scenes while supporting precise geometric manipulation. Experiments demonstrate that Ctrl&Shift significantly outperforms existing geometry-based and diffusion-based approaches in terms of generation fidelity, viewpoint consistency, and user controllability.
针对NVFP4量化中激活异常值导致的精度下降问题,提出OCGQuant方法,通过异常-伴生组策略优化量化块组成,提高量化精度。
This study addresses the limited generalization and high data dependency of Vision-Language-Action (VLA) models by proposing ART, a framework that integrates tool use into VLAs to reduce action space complexity. By combining few-shot dataset construction with long-horizon reasoning training, ART achieves modular model enhancement. Experiments demonstrate that ART outperforms mainstream baselines by 20% in success rate across both simulated and real-world tasks while significantly reducing data requirements. The approach markedly improves robotic generalization in complex environments and enables lightweight, efficient deployment. Ultimately, this work establishes a scalable paradigm for embodied intelligence, effectively mitigating key bottlenecks in current VLA architectures through strategic tool integration and optimized training methodologies.
This work addresses the challenge that existing open-vocabulary object recognition methods struggle to simultaneously achieve coarse-grained generality and fine-grained discriminative capability. To this end, we propose HyMOR, a hybrid framework that synergistically integrates multimodal large language models (MLLMs) and CLIP for the first time: the MLLM handles open-domain coarse-grained recognition, while CLIP specializes in fine-grained discrimination within domains such as flora and fauna. A Sentence-BERT-guided hybrid reasoning mechanism enables multi-granularity semantic alignment between the two streams. Built upon this architecture, we introduce a unified perception system tailored for educational games and release TBO, a textbook-derived dataset. Our approach achieves a 23.2% improvement in average SBert score, narrows the fine-grained recognition gap to 0.2%, and enhances general recognition performance by 2.5%, substantially strengthening the perceptual foundation for multimodal interactive learning.
This work addresses the open-set challenge in generalized category discovery (GCD), where only a subset of known classes is labeled and unknown categories must be identified. To tackle this problem, we propose the SSR²-GCD framework, which introduces semi-supervised rate reduction into GCD for the first time. By optimizing multimodal representation learning, our method enhances intra-modal alignment to construct structured feature distributions and leverages the prompt candidate mechanism of vision-language models (VLMs) to strengthen cross-modal knowledge transfer. Extensive experiments on both generic and fine-grained benchmark datasets demonstrate that SSR²-GCD significantly outperforms existing approaches, achieving state-of-the-art performance.
Existing methods for image and video object manipulation struggle to simultaneously preserve background content, maintain geometric consistency across viewpoints, and offer fine-grained user control. This work proposes Ctrl&Shift, an end-to-end diffusion framework that decomposes manipulation into two stages—object removal followed by camera-pose-guided reference-based inpainting—enabling geometrically consistent editing within a unified diffusion process without explicit 3D modeling. By integrating explicit camera pose control, reference-guided inpainting, and a multi-task, multi-stage training strategy, the method effectively disentangles background, identity, and pose signals. This design preserves generalization to real-world scenes while supporting precise geometric manipulation. Experiments demonstrate that Ctrl&Shift significantly outperforms existing geometry-based and diffusion-based approaches in terms of generation fidelity, viewpoint consistency, and user controllability.