Institution profile

Shanghai University of Engineering Science

Academic institutionasia · cn
Official website
Research library16linked papers
Opportunities0open roles
Selected work

Representative Papers

Multimodal-Aware Fusion Network for Referring Remote Sensing Image Segmentation

Mar 14, 2025IEEE Geoscience and Remote Sensing Letters

To address coarse-grained multimodal alignment and insufficient feature fusion in referring segmentation of remote sensing images, this paper proposes a fine-grained cross-modal collaborative segmentation framework. The method introduces two key components: (1) a Correlation Fusion Module (CFM) that enables pixel-wise semantic alignment between textual and visual features via cross-modal correlation modeling; and (2) a Multi-Scale Refinement Convolution (MSRC) integrated with an adaptive noise-augmented Transformer-based visual encoder, which jointly captures multi-directional, multi-scale object structures and orientation invariance. Evaluated on the RRSIS-D benchmark, the proposed approach achieves significant improvements over existing state-of-the-art methods, attaining a 3.2% gain in mean Intersection-over-Union (mIoU). The source code is publicly available.

1 citationsRead paper

Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training

May 11, 2026

This work addresses the lack of effective, interpretable, and low-cost guidance for determining layer-wise freezing or training strategies during continual pre-training of large language models. The authors propose LayerTracer, a framework that identifies task-executing layers, quantifies layer sensitivity, and analyzes the stability of representational evolution across layers. Their analysis reveals that deeper layers are primarily responsible for task execution and exhibit greater robustness to perturbations. Based on these insights, they introduce an interpretable and cost-efficient continual pre-training paradigm—freezing deeper layers while fine-tuning shallower ones—and extend it to hybrid model construction. The approach is architecture-agnostic and, through representation analysis and controlled experiments, consistently outperforms full-parameter fine-tuning and alternative strategies on C-Eval and CMMLU benchmarks, while demonstrating that embedding high-quality deep modules effectively preserves original knowledge.

0 citationsRead paper

From Answers to Arguments: Toward Trustworthy Clinical Diagnostic Reasoning with Toulmin-Guided Curriculum Goal-Conditioned Learning

Apr 13, 2026

This study addresses the challenge that large language models often produce correct diagnostic conclusions through opaque or flawed reasoning, failing to meet the high standards of interpretability and reliability required in clinical decision-making. To tackle this issue, the authors introduce the Toulmin model of argumentation into clinical diagnosis for the first time and propose a Curriculum Goal-Conditioned Learning (CGCL) framework. This approach employs a three-stage progressive training strategy to guide models in constructing structured, verifiable diagnostic arguments. Integrated with the T-Eval evaluation framework, the method achieves diagnostic accuracy and reasoning quality comparable to reinforcement learning baselines while significantly enhancing reasoning transparency, reliability, and training stability.

0 citationsRead paper

Expert Pyramid Tuning: Efficient Parameter Fine-Tuning for Expertise-Driven Task Allocation

Mar 12, 2026

This work addresses the limitation of existing Mixture-of-Experts (MoE)-based parameter-efficient fine-tuning (PEFT) methods, which struggle to simultaneously capture high-level semantics and fine-grained syntactic requirements due to their neglect of the hierarchical complexity inherent in tasks. To overcome this, we propose Expert Pyramid Tuning (EPT), a novel architecture that, for the first time, integrates a multi-scale feature pyramid mechanism into the PEFT framework. EPT generates multi-scale features through a shared meta-knowledge subspace and pyramid projections, dynamically composing them via a task-aware router. By synergistically combining LoRA, MoE, and learnable up-projection operators, EPT establishes a two-stage tuning pipeline and supports post-training reparameterization for parameter compression. Extensive experiments demonstrate that EPT significantly outperforms current MoE-LoRA approaches across multiple multitask benchmarks while reducing the number of trainable parameters.

0 citationsRead paper
Recent publications

Latest Papers

Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training

May 11, 2026

This work addresses the lack of effective, interpretable, and low-cost guidance for determining layer-wise freezing or training strategies during continual pre-training of large language models. The authors propose LayerTracer, a framework that identifies task-executing layers, quantifies layer sensitivity, and analyzes the stability of representational evolution across layers. Their analysis reveals that deeper layers are primarily responsible for task execution and exhibit greater robustness to perturbations. Based on these insights, they introduce an interpretable and cost-efficient continual pre-training paradigm—freezing deeper layers while fine-tuning shallower ones—and extend it to hybrid model construction. The approach is architecture-agnostic and, through representation analysis and controlled experiments, consistently outperforms full-parameter fine-tuning and alternative strategies on C-Eval and CMMLU benchmarks, while demonstrating that embedding high-quality deep modules effectively preserves original knowledge.

0 citationsRead paper

From Answers to Arguments: Toward Trustworthy Clinical Diagnostic Reasoning with Toulmin-Guided Curriculum Goal-Conditioned Learning

Apr 13, 2026

This study addresses the challenge that large language models often produce correct diagnostic conclusions through opaque or flawed reasoning, failing to meet the high standards of interpretability and reliability required in clinical decision-making. To tackle this issue, the authors introduce the Toulmin model of argumentation into clinical diagnosis for the first time and propose a Curriculum Goal-Conditioned Learning (CGCL) framework. This approach employs a three-stage progressive training strategy to guide models in constructing structured, verifiable diagnostic arguments. Integrated with the T-Eval evaluation framework, the method achieves diagnostic accuracy and reasoning quality comparable to reinforcement learning baselines while significantly enhancing reasoning transparency, reliability, and training stability.

0 citationsRead paper

Expert Pyramid Tuning: Efficient Parameter Fine-Tuning for Expertise-Driven Task Allocation

Mar 12, 2026

This work addresses the limitation of existing Mixture-of-Experts (MoE)-based parameter-efficient fine-tuning (PEFT) methods, which struggle to simultaneously capture high-level semantics and fine-grained syntactic requirements due to their neglect of the hierarchical complexity inherent in tasks. To overcome this, we propose Expert Pyramid Tuning (EPT), a novel architecture that, for the first time, integrates a multi-scale feature pyramid mechanism into the PEFT framework. EPT generates multi-scale features through a shared meta-knowledge subspace and pyramid projections, dynamically composing them via a task-aware router. By synergistically combining LoRA, MoE, and learnable up-projection operators, EPT establishes a two-stage tuning pipeline and supports post-training reparameterization for parameter compression. Extensive experiments demonstrate that EPT significantly outperforms current MoE-LoRA approaches across multiple multitask benchmarks while reducing the number of trainable parameters.

0 citationsRead paper

Three-dimensional visualization of X-ray micro-CT with large-scale datasets: Efficiency and accuracy for real-time interaction

Jan 21, 2026Computer Science Review

This study addresses the critical trade-off between accuracy and efficiency in processing large-scale 3D data for ultra-precision industrial micro-CT inspection. By systematically reviewing and integrating advances in reconstruction and volume rendering techniques—from medical imaging to industrial non-destructive testing—the work proposes a high-fidelity, efficient 3D visualization framework tailored for digital twin–enabled structural health monitoring. The framework synergistically combines analytical and deep learning–based reconstruction methods, accelerated volume rendering, dimensionality reduction, and physically accurate lighting models. Beyond offering researchers a practical guide for method selection, this approach facilitates real-time interactive analysis and online monitoring of internal material defects. The paper further outlines a promising direction toward co-optimizing deep learning architectures with advanced illumination models to enhance both visual realism and diagnostic precision.

0 citationsRead paper