Institution profile

Intellifusion

Industry researchasia · cn
Official website
Research library25linked papers
Opportunities0open roles
Selected work

Representative Papers

Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use

Aug 14, 2026

This study addresses the limited generalization and high data dependency of Vision-Language-Action (VLA) models by proposing ART, a framework that integrates tool use into VLAs to reduce action space complexity. By combining few-shot dataset construction with long-horizon reasoning training, ART achieves modular model enhancement. Experiments demonstrate that ART outperforms mainstream baselines by 20% in success rate across both simulated and real-world tasks while significantly reducing data requirements. The approach markedly improves robotic generalization in complex environments and enables lightweight, efficient deployment. Ultimately, this work establishes a scalable paradigm for embodied intelligence, effectively mitigating key bottlenecks in current VLA architectures through strategic tool integration and optimized training methodologies.

0 citationsRead paper

Bridging Coarse and Fine Recognition: A Hybrid Approach for Open-Ended Multi-Granularity Object Recognition in Interactive Educational Games

Apr 17, 2026

This work addresses the challenge that existing open-vocabulary object recognition methods struggle to simultaneously achieve coarse-grained generality and fine-grained discriminative capability. To this end, we propose HyMOR, a hybrid framework that synergistically integrates multimodal large language models (MLLMs) and CLIP for the first time: the MLLM handles open-domain coarse-grained recognition, while CLIP specializes in fine-grained discrimination within domains such as flora and fauna. A Sentence-BERT-guided hybrid reasoning mechanism enables multi-granularity semantic alignment between the two streams. Built upon this architecture, we introduce a unified perception system tailored for educational games and release TBO, a textbook-derived dataset. Our approach achieves a 23.2% improvement in average SBert score, narrows the fine-grained recognition gap to 0.2%, and enhances general recognition performance by 2.5%, substantially strengthening the perceptual foundation for multimodal interactive learning.

0 citationsRead paper

Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery

Feb 23, 2026

This work addresses the open-set challenge in generalized category discovery (GCD), where only a subset of known classes is labeled and unknown categories must be identified. To tackle this problem, we propose the SSR²-GCD framework, which introduces semi-supervised rate reduction into GCD for the first time. By optimizing multimodal representation learning, our method enhances intra-modal alignment to construct structured feature distributions and leverages the prompt candidate mechanism of vision-language models (VLMs) to strengthen cross-modal knowledge transfer. Extensive experiments on both generic and fine-grained benchmark datasets demonstrate that SSR²-GCD significantly outperforms existing approaches, achieving state-of-the-art performance.

0 citationsRead paper

Ctrl&Shift: High-Quality Geometry-Aware Object Manipulation in Visual Generation

Feb 11, 2026

Existing methods for image and video object manipulation struggle to simultaneously preserve background content, maintain geometric consistency across viewpoints, and offer fine-grained user control. This work proposes Ctrl&Shift, an end-to-end diffusion framework that decomposes manipulation into two stages—object removal followed by camera-pose-guided reference-based inpainting—enabling geometrically consistent editing within a unified diffusion process without explicit 3D modeling. By integrating explicit camera pose control, reference-guided inpainting, and a multi-task, multi-stage training strategy, the method effectively disentangles background, identity, and pose signals. This design preserves generalization to real-world scenes while supporting precise geometric manipulation. Experiments demonstrate that Ctrl&Shift significantly outperforms existing geometry-based and diffusion-based approaches in terms of generation fidelity, viewpoint consistency, and user controllability.

0 citationsRead paper
Recent publications

Latest Papers

Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use

Aug 14, 2026

This study addresses the limited generalization and high data dependency of Vision-Language-Action (VLA) models by proposing ART, a framework that integrates tool use into VLAs to reduce action space complexity. By combining few-shot dataset construction with long-horizon reasoning training, ART achieves modular model enhancement. Experiments demonstrate that ART outperforms mainstream baselines by 20% in success rate across both simulated and real-world tasks while significantly reducing data requirements. The approach markedly improves robotic generalization in complex environments and enables lightweight, efficient deployment. Ultimately, this work establishes a scalable paradigm for embodied intelligence, effectively mitigating key bottlenecks in current VLA architectures through strategic tool integration and optimized training methodologies.

0 citationsRead paper

Bridging Coarse and Fine Recognition: A Hybrid Approach for Open-Ended Multi-Granularity Object Recognition in Interactive Educational Games

Apr 17, 2026

This work addresses the challenge that existing open-vocabulary object recognition methods struggle to simultaneously achieve coarse-grained generality and fine-grained discriminative capability. To this end, we propose HyMOR, a hybrid framework that synergistically integrates multimodal large language models (MLLMs) and CLIP for the first time: the MLLM handles open-domain coarse-grained recognition, while CLIP specializes in fine-grained discrimination within domains such as flora and fauna. A Sentence-BERT-guided hybrid reasoning mechanism enables multi-granularity semantic alignment between the two streams. Built upon this architecture, we introduce a unified perception system tailored for educational games and release TBO, a textbook-derived dataset. Our approach achieves a 23.2% improvement in average SBert score, narrows the fine-grained recognition gap to 0.2%, and enhances general recognition performance by 2.5%, substantially strengthening the perceptual foundation for multimodal interactive learning.

0 citationsRead paper

Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery

Feb 23, 2026

This work addresses the open-set challenge in generalized category discovery (GCD), where only a subset of known classes is labeled and unknown categories must be identified. To tackle this problem, we propose the SSR²-GCD framework, which introduces semi-supervised rate reduction into GCD for the first time. By optimizing multimodal representation learning, our method enhances intra-modal alignment to construct structured feature distributions and leverages the prompt candidate mechanism of vision-language models (VLMs) to strengthen cross-modal knowledge transfer. Extensive experiments on both generic and fine-grained benchmark datasets demonstrate that SSR²-GCD significantly outperforms existing approaches, achieving state-of-the-art performance.

0 citationsRead paper

Ctrl&Shift: High-Quality Geometry-Aware Object Manipulation in Visual Generation

Feb 11, 2026

Existing methods for image and video object manipulation struggle to simultaneously preserve background content, maintain geometric consistency across viewpoints, and offer fine-grained user control. This work proposes Ctrl&Shift, an end-to-end diffusion framework that decomposes manipulation into two stages—object removal followed by camera-pose-guided reference-based inpainting—enabling geometrically consistent editing within a unified diffusion process without explicit 3D modeling. By integrating explicit camera pose control, reference-guided inpainting, and a multi-task, multi-stage training strategy, the method effectively disentangles background, identity, and pose signals. This design preserves generalization to real-world scenes while supporting precise geometric manipulation. Experiments demonstrate that Ctrl&Shift significantly outperforms existing geometry-based and diffusion-based approaches in terms of generation fidelity, viewpoint consistency, and user controllability.

0 citationsRead paper