Institution profile

Chinese Academy of Sciences Institute of Automation

Academic institutionasia · cn
Official website
Research library273linked papers
Opportunities0open roles
Selected work

Representative Papers

TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers

Jan 20, 2026

This work addresses the challenge that fine-tuning general-purpose vision-language models (VLMs) for robotic control often leads to degradation of semantic understanding and interference with fine-grained motor learning. To mitigate this, the authors propose TwinBrainVLA, an architecture that freezes a pre-trained VLM as the “left brain” to preserve open-world semantic comprehension, while introducing a trainable, embodied perception-specific model as the “right brain” to learn high-precision continuous actions. The two components are integrated via an Asymmetric Mixture-of-Transformers (AsyMoT) mechanism and further enhanced by a Flow-Matching action expert module, enabling effective fusion of high-level semantics and low-level control. Experiments demonstrate that the approach outperforms existing methods on SimplerEnv and RoboCasa benchmarks, significantly alleviates catastrophic forgetting, and maintains the pretrained VLM’s general visual understanding capabilities.

2 citationsRead paper

From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning

Jan 19, 2026

Existing code completion methods, such as Fill-in-the-Middle (FIM), struggle to correct contextual errors and rely on potentially unsafe base models. Meanwhile, chat-based large language models suffer from performance degradation, and agent-based workflows incur high latency. To address these limitations, this work proposes the Search-and-Replace Infilling (SRI) framework, which extends code completion from static infilling to context-aware dynamic editing. SRI internalizes the agent-like verify-and-edit mechanism into a single inference pass, preserving low latency and general programming proficiency while maintaining instruction-following capabilities. Leveraging a synthetically constructed SRI-200K dataset and structured search-replace instructions, a model fine-tuned with only 20,000 samples—SRI-Coder—outperforms base models in completion accuracy while matching the inference speed of standard FIM.

2 citationsRead paper

One Size, Many Fits: Aligning Diverse Group-Wise Click Preferences in Large-Scale Advertising Image Generation

Feb 02, 2026

This work addresses the limitations of existing advertising image generation methods, which adopt a one-size-fits-all strategy and neglect inter-group differences in click preferences, leading to suboptimal performance for certain user segments. To overcome this, the authors propose OSMF, a unified framework that enables personalized ad content generation through product-aware adaptive grouping and preference-conditioned image synthesis. The key contributions include the first introduction of a group-level click preference alignment mechanism, the construction of GAIP—the first large-scale dataset capturing group-specific advertising image preferences—and the development of Group-DPO optimization integrated with a group-aware multimodal large language model (G-MLLM). Both offline evaluations and online experiments demonstrate that the proposed approach significantly improves click-through rates across diverse user groups, achieving state-of-the-art performance.

1 citationsRead paper
Recent publications

Latest Papers