Institution profile

Hangzhou Dianzi University

Academic institutionasia · cn
Official website
Research library407linked papers
Opportunities0open roles
Selected work

Representative Papers

How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing

Feb 02, 2026

This work addresses the limitation of existing image editing models and evaluation benchmarks, which predominantly rely on textual instructions and struggle to support visual directives—such as sketches—that are integral to human multimodal interaction. To bridge this gap, we introduce VIBE, the first systematic benchmark for vision-instructed image editing, which defines a three-tiered hierarchy of task complexity ranging from referential localization and shape manipulation to causal reasoning. We further develop a fine-grained automatic evaluation framework based on multimodal large language models (LMMs) and use it to assess 17 open- and closed-source models across diverse visual instructions. Our evaluation reveals that closed-source models exhibit初步 stronger instruction-following capabilities, yet all models suffer significant performance degradation on higher-order tasks, highlighting critical limitations and pointing toward promising directions for future research.

2 citationsRead paper

ToolACE-MCP: Generalizing History-Aware Routing from MCP Tools to the Agent Web

Jan 13, 2026

This work addresses the scalability and generalization limitations of existing agent architectures in the face of exponentially growing tool ecosystems. The authors propose ToolACE-MCP, a lightweight routing agent endowed with historical awareness, which models multi-turn interaction trajectories through graph-based representations and learns a context-aware routing policy within the Model Context Protocol (MCP) framework. The approach enables plug-and-play deployment and zero-shot transfer to multi-agent collaborative settings, achieving significant advances in noise robustness and scalability over large candidate tool spaces. Experimental evaluation on the MCP-Universe and MCP-Mark real-world benchmarks demonstrates its superior capability for general-purpose orchestration in open-ended tool environments.

1 citationsRead paper

Beyond Static Tools: Test-Time Tool Evolution for Scientific Reasoning

Jan 12, 2026

This work addresses the limitations of existing large language model (LLM) agents, which rely on static tool libraries and struggle with the sparsity, heterogeneity, and incompleteness of tools in scientific reasoning. To overcome this, we propose a novel Test-Time Tool Evolution (TTE) paradigm that dynamically synthesizes, verifies, and optimizes executable tools during inference, transforming tools from predefined resources into problem-driven artifacts. Our approach integrates LLMs with program synthesis and automated verification to establish an end-to-end pipeline for tool generation and evolution. Evaluated on SciEvo—a newly curated benchmark comprising 1,590 tasks and 925 evolved tools—our method achieves state-of-the-art performance, significantly improving accuracy, tool efficiency, and cross-domain transferability.

1 citationsRead paper

DATransNet: Dynamic Attention Transformer Network for Infrared Small Target Detection

Sep 29, 2024

To address the challenge of detecting weak and small infrared targets that are easily obscured by complex backgrounds, this paper proposes the Dynamic Attention Transformer (DATrans). DATrans first employs center-difference convolution to model edge-gradient features, then adaptively fuses multi-scale gradient representations with deep semantic features via a dynamic attention mechanism. Furthermore, a Global Background-Aware Module (GBAM) is introduced to explicitly model target-background contextual relationships, thereby enhancing discriminability. The architecture achieves a synergistic optimization of fine-grained detail sensitivity and global semantic understanding while maintaining computational efficiency. Extensive experiments on multiple benchmark infrared small-target datasets demonstrate significant improvements over state-of-the-art methods in both detection accuracy and robustness. The source code is publicly available.

1 citationsRead paper
Recent publications

Latest Papers