Institution profile

Shanghai Institute of Technology

Academic institutionasia · cn
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

HVPNet: A Bio-Inspired Network for General Salient and Camouflaged Object Detection

Jun 30, 2026

Existing multimodal salient and camouflaged object detection methods suffer from structural complexity and large parameter counts, making it challenging to balance accuracy and efficiency. Inspired by the human visual system, this work proposes a lightweight, unified architecture that integrates a Retinal Integration Module (RIM) for hierarchical, multi-stage cross-modal feature fusion and a Cortical Decoder (CD) that mimics visual cortical mechanisms for layered decoding. This approach establishes a biologically inspired, simplified modeling paradigm capable of supporting diverse modalities and tasks within a single framework. Evaluated across four modalities, seven tasks, and 22 datasets, the model achieves an excellent trade-off between accuracy and efficiency with a compact structure, demonstrating strong generalization capability.

0 citationsRead paper

Adaptive Output Steps: FlexiSteps Network for Dynamic Trajectory Prediction

Aug 25, 2025

Traditional trajectory prediction models are constrained by fixed output horizons, limiting their adaptability to dynamic real-world scenarios. To address this, we propose FlexiSteps—a novel framework introducing the first adaptive timestep selection module for trajectory forecasting. Our method jointly optimizes geometric trajectory similarity (measured via Fréchet distance) and timestep plausibility (via a weighted scoring mechanism) to enable context-aware, variable-length prediction horizons. FlexiSteps comprises a pre-trained adaptive prediction head, a dynamic decoder, and an end-to-end differentiable timestep selection mechanism. Extensive experiments on Argoverse and INTERACTION demonstrate that FlexiSteps significantly outperforms baseline models in both average displacement error (ADE) and final displacement error (FDE), while improving computational efficiency and scene adaptability. By enabling flexible, input-dependent horizon lengths, FlexiSteps establishes a scalable, multimodal paradigm for long-horizon trajectory prediction.

0 citationsRead paper

Machine Learning-Assisted Surrogate Modeling with Multi-Objective Optimization and Decision-Making of a Steam Methane Reforming Reactor

Jul 10, 2025

Steam methane reforming (SMR) reactor multi-objective optimization faces high computational cost and inherent trade-offs among conflicting objectives—e.g., methane conversion, hydrogen production rate, and CO₂ emissions. Method: This paper proposes an integrated framework comprising mechanistic modeling, data-driven surrogate modeling, multi-objective optimization, and decision-making support. An artificial neural network (ANN)-hybrid surrogate model replaces computationally expensive high-fidelity 1D fixed-bed simulations; non-dominated sorting genetic algorithm II (NSGA-II) efficiently computes the Pareto-optimal front; and a dual-criteria decision strategy—combining Technique for Order Preference by Similarity to Ideal Solution (TOPSIS) and stochastic PROBID (sPROBID)—supports optimal operating condition selection. Contribution/Results: The surrogate model reduces average simulation time by 93.8%. The identified Pareto-optimal solution achieves methane conversion of 0.988, H₂ production of 3.335 mol/s, and CO₂ emissions of 0.781 mol/s—demonstrating substantial gains in optimization efficiency and engineering applicability.

0 citationsRead paper

GAMDTP: Dynamic Trajectory Prediction with Graph Attention Mamba Network

Apr 07, 2025

To address the insufficient accuracy and efficiency of dynamic trajectory prediction for traffic participants in autonomous driving, this paper proposes an end-to-end multimodal trajectory prediction framework integrating graph attention mechanisms with the Mamba state space model (SSM). The framework jointly encodes high-definition maps and historical trajectories, and introduces—novelly within graph convolutional layers—a gated fusion mechanism combining self-attention and Mamba-SSM. It further incorporates a two-stage proposal-refinement architecture and a prediction-quality rescorer to enhance robustness and balance diversity with accuracy. Evaluated on the Argoverse 2 benchmark, our method achieves state-of-the-art performance with significantly fewer parameters: it reduces average displacement error (ADE) and final displacement error (FDE) by 12.3% and 9.7%, respectively, demonstrating both effectiveness and strong generalization capability.

0 citationsRead paper

DyTTP: Trajectory Prediction with Normalization-Free Transformers

Apr 07, 2025

Transformer-based approaches for multi-agent trajectory prediction in autonomous driving exhibit strong modeling capacity but suffer from training instability and high computational overhead due to layer normalization. Method: This paper proposes a normalization-free Transformer architecture: (i) replacing layer normalization with a dynamic Tanh activation—introduced to trajectory prediction for the first time—and (ii) designing a lightweight snapshot ensemble framework that synergistically integrates cyclical learning rate scheduling with model weight averaging. Contribution/Results: Evaluated on the Argoverse dataset, our method achieves significant improvements: 8.2% reduction in average displacement error (ADE) and 6.7% reduction in final displacement error (FDE), 23% faster inference speed, and enhanced robustness in complex traffic scenarios. The proposed approach establishes a new paradigm for efficient, stable, and deployable trajectory prediction—eliminating normalization-induced bottlenecks while maintaining high accuracy and generalization.

0 citationsRead paper
Recent publications

Latest Papers

HVPNet: A Bio-Inspired Network for General Salient and Camouflaged Object Detection

Jun 30, 2026

Existing multimodal salient and camouflaged object detection methods suffer from structural complexity and large parameter counts, making it challenging to balance accuracy and efficiency. Inspired by the human visual system, this work proposes a lightweight, unified architecture that integrates a Retinal Integration Module (RIM) for hierarchical, multi-stage cross-modal feature fusion and a Cortical Decoder (CD) that mimics visual cortical mechanisms for layered decoding. This approach establishes a biologically inspired, simplified modeling paradigm capable of supporting diverse modalities and tasks within a single framework. Evaluated across four modalities, seven tasks, and 22 datasets, the model achieves an excellent trade-off between accuracy and efficiency with a compact structure, demonstrating strong generalization capability.

0 citationsRead paper

Adaptive Output Steps: FlexiSteps Network for Dynamic Trajectory Prediction

Aug 25, 2025

Traditional trajectory prediction models are constrained by fixed output horizons, limiting their adaptability to dynamic real-world scenarios. To address this, we propose FlexiSteps—a novel framework introducing the first adaptive timestep selection module for trajectory forecasting. Our method jointly optimizes geometric trajectory similarity (measured via Fréchet distance) and timestep plausibility (via a weighted scoring mechanism) to enable context-aware, variable-length prediction horizons. FlexiSteps comprises a pre-trained adaptive prediction head, a dynamic decoder, and an end-to-end differentiable timestep selection mechanism. Extensive experiments on Argoverse and INTERACTION demonstrate that FlexiSteps significantly outperforms baseline models in both average displacement error (ADE) and final displacement error (FDE), while improving computational efficiency and scene adaptability. By enabling flexible, input-dependent horizon lengths, FlexiSteps establishes a scalable, multimodal paradigm for long-horizon trajectory prediction.

0 citationsRead paper

Machine Learning-Assisted Surrogate Modeling with Multi-Objective Optimization and Decision-Making of a Steam Methane Reforming Reactor

Jul 10, 2025

Steam methane reforming (SMR) reactor multi-objective optimization faces high computational cost and inherent trade-offs among conflicting objectives—e.g., methane conversion, hydrogen production rate, and CO₂ emissions. Method: This paper proposes an integrated framework comprising mechanistic modeling, data-driven surrogate modeling, multi-objective optimization, and decision-making support. An artificial neural network (ANN)-hybrid surrogate model replaces computationally expensive high-fidelity 1D fixed-bed simulations; non-dominated sorting genetic algorithm II (NSGA-II) efficiently computes the Pareto-optimal front; and a dual-criteria decision strategy—combining Technique for Order Preference by Similarity to Ideal Solution (TOPSIS) and stochastic PROBID (sPROBID)—supports optimal operating condition selection. Contribution/Results: The surrogate model reduces average simulation time by 93.8%. The identified Pareto-optimal solution achieves methane conversion of 0.988, H₂ production of 3.335 mol/s, and CO₂ emissions of 0.781 mol/s—demonstrating substantial gains in optimization efficiency and engineering applicability.

0 citationsRead paper

GAMDTP: Dynamic Trajectory Prediction with Graph Attention Mamba Network

Apr 07, 2025

To address the insufficient accuracy and efficiency of dynamic trajectory prediction for traffic participants in autonomous driving, this paper proposes an end-to-end multimodal trajectory prediction framework integrating graph attention mechanisms with the Mamba state space model (SSM). The framework jointly encodes high-definition maps and historical trajectories, and introduces—novelly within graph convolutional layers—a gated fusion mechanism combining self-attention and Mamba-SSM. It further incorporates a two-stage proposal-refinement architecture and a prediction-quality rescorer to enhance robustness and balance diversity with accuracy. Evaluated on the Argoverse 2 benchmark, our method achieves state-of-the-art performance with significantly fewer parameters: it reduces average displacement error (ADE) and final displacement error (FDE) by 12.3% and 9.7%, respectively, demonstrating both effectiveness and strong generalization capability.

0 citationsRead paper

DyTTP: Trajectory Prediction with Normalization-Free Transformers

Apr 07, 2025

Transformer-based approaches for multi-agent trajectory prediction in autonomous driving exhibit strong modeling capacity but suffer from training instability and high computational overhead due to layer normalization. Method: This paper proposes a normalization-free Transformer architecture: (i) replacing layer normalization with a dynamic Tanh activation—introduced to trajectory prediction for the first time—and (ii) designing a lightweight snapshot ensemble framework that synergistically integrates cyclical learning rate scheduling with model weight averaging. Contribution/Results: Evaluated on the Argoverse dataset, our method achieves significant improvements: 8.2% reduction in average displacement error (ADE) and 6.7% reduction in final displacement error (FDE), 23% faster inference speed, and enhanced robustness in complex traffic scenarios. The proposed approach establishes a new paradigm for efficient, stable, and deployable trajectory prediction—eliminating normalization-induced bottlenecks while maintaining high accuracy and generalization.

0 citationsRead paper