Institution profile

Dali University

Academic institutionasia · cn
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Semantic-Enhanced Cross-Modal Place Recognition for Robust Robot Localization

Sep 16, 2025

To address the insufficient robustness of vision–LiDAR cross-modal place recognition under illumination, weather, and viewpoint variations in GPS-denied environments, this paper proposes SCM-PR, a semantic-enhanced cross-modal localization framework. SCM-PR innovatively adopts VMamba as the visual backbone, designs a semantic-aware feature fusion module and a semantic-guided LiDAR descriptor, and introduces a cross-modal semantic attention mechanism alongside a multi-view semantic–geometric matching strategy. Furthermore, a semantic consistency loss is proposed to enforce cross-modal semantic alignment. Evaluated on KITTI and KITTI-360, SCM-PR achieves state-of-the-art performance—significantly improving matching accuracy and robustness under complex scenes, high-resolution inputs, and large viewpoint changes.

0 citationsRead paper

LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation

Aug 05, 2025

Existing text-to-image (T2I) models exhibit significant limitations in fine-grained control—such as text rendering, human pose alignment, and complex compositional reasoning—as well as deep semantic consistency. To address these challenges, we propose LumiGen, an iterative T2I generation framework grounded in vision-language models (VLMs). Its core innovation is a closed-loop feedback mechanism endowed with “visual critique” capability: leveraging VLM-driven intelligent prompt parsing and enhancement, coupled with multi-round visual feedback refinement, to enable precise, stepwise control over the generation process. LumiGen end-to-end integrates diffusion models, VLMs, and iterative optimization modules. Evaluated on the LongBench-T2I benchmark, it achieves a mean score of 3.08—substantially outperforming prior methods—particularly excelling in textual accuracy and pose fidelity, two critical dimensions of fine-grained semantic alignment.

0 citationsRead paper

Time-EAPCR-T: A Universal Deep Learning Approach for Anomaly Detection in Industrial Equipment

Mar 16, 2025

Addressing the challenge of anomaly detection in multi-source, heterogeneous, strongly coupled, and noisy time-series data prevalent in Industry 4.0, conventional methods suffer from information loss due to reliance on dimensionality reduction and feature selection, and fail to capture high-order temporal dependencies. This paper proposes Time-EAPCR-T, an end-to-end deep learning framework. It innovatively replaces LSTM with Transformer in the temporal modeling module to enable lossless fusion of multi-source features and effective capture of high-order dynamic interactions. Furthermore, it introduces the EAPCR (Enhanced Adaptive Principal Component Reconstruction) feature disentanglement mechanism and an improved temporal encoding structure to enhance cross-device and cross-operating-condition generalization. Evaluated on four real-world industrial datasets, Time-EAPCR-T achieves an average 9.2% improvement in F1-score and a 37% reduction in false positive rate over state-of-the-art methods, demonstrating strong practicality and engineering deployability.

0 citationsRead paper

Time-EAPCR: A Deep Learning-Based Novel Approach for Anomaly Detection Applied to the Environmental Field

Mar 12, 2025

Traditional environmental monitoring methods suffer from delayed response, poor generalizability, and difficulty modeling long-term temporal dependencies and high variability. To address these limitations, this paper proposes Time-EAPCR—the first end-to-end anomaly detection framework integrating time embedding, self-attention, permutation convolution, and residual learning—explicitly capturing spatiotemporal evolution patterns and multidimensional feature correlations in environmental data. The architecture significantly enhances robustness and interpretability of anomaly discrimination. Evaluated on four public environmental datasets, it achieves an average 12.6% improvement in F1-score. Deployment feasibility is further validated in a real-world river monitoring system. Experimental results demonstrate strong cross-scenario generalization capability, establishing a novel paradigm for intelligent anomaly monitoring in aquatic ecosystems and water treatment infrastructure.

0 citationsRead paper

Inorganic Catalyst Efficiency Prediction Based on EAPCR Model: A Deep Learning Solution for Multi-Source Heterogeneous Data

Mar 10, 2025

Predicting inorganic catalyst efficiency faces challenges in integrating heterogeneous multi-source data and insufficient mechanistic modeling. To address these, we propose EAPCR, a novel deep learning model featuring an Embedding–Attention–Permutation-Convolution–Residual (EAPCR) architecture. It employs embedding layers to unify heterogeneous inputs, self-attention to enhance cross-condition feature interaction, permutation-invariant convolutional kernels to accommodate catalytic site arrangements, and residual connections to improve training stability. Additionally, EAPCR constructs a feature correlation matrix to explicitly capture multi-scale structure–property relationships. Evaluated on TiO₂ photocatalysis, thermocatalysis, and electrocatalysis tasks, EAPCR consistently outperforms linear regression, random forest, and artificial neural networks—reducing MAE, MSE, and RMSE by 18.7% on average and increasing R² by 0.23. The model demonstrates strong generalizability and enhanced physical interpretability.

0 citationsRead paper
Recent publications

Latest Papers

Semantic-Enhanced Cross-Modal Place Recognition for Robust Robot Localization

Sep 16, 2025

To address the insufficient robustness of vision–LiDAR cross-modal place recognition under illumination, weather, and viewpoint variations in GPS-denied environments, this paper proposes SCM-PR, a semantic-enhanced cross-modal localization framework. SCM-PR innovatively adopts VMamba as the visual backbone, designs a semantic-aware feature fusion module and a semantic-guided LiDAR descriptor, and introduces a cross-modal semantic attention mechanism alongside a multi-view semantic–geometric matching strategy. Furthermore, a semantic consistency loss is proposed to enforce cross-modal semantic alignment. Evaluated on KITTI and KITTI-360, SCM-PR achieves state-of-the-art performance—significantly improving matching accuracy and robustness under complex scenes, high-resolution inputs, and large viewpoint changes.

0 citationsRead paper

LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation

Aug 05, 2025

Existing text-to-image (T2I) models exhibit significant limitations in fine-grained control—such as text rendering, human pose alignment, and complex compositional reasoning—as well as deep semantic consistency. To address these challenges, we propose LumiGen, an iterative T2I generation framework grounded in vision-language models (VLMs). Its core innovation is a closed-loop feedback mechanism endowed with “visual critique” capability: leveraging VLM-driven intelligent prompt parsing and enhancement, coupled with multi-round visual feedback refinement, to enable precise, stepwise control over the generation process. LumiGen end-to-end integrates diffusion models, VLMs, and iterative optimization modules. Evaluated on the LongBench-T2I benchmark, it achieves a mean score of 3.08—substantially outperforming prior methods—particularly excelling in textual accuracy and pose fidelity, two critical dimensions of fine-grained semantic alignment.

0 citationsRead paper

Time-EAPCR-T: A Universal Deep Learning Approach for Anomaly Detection in Industrial Equipment

Mar 16, 2025

Addressing the challenge of anomaly detection in multi-source, heterogeneous, strongly coupled, and noisy time-series data prevalent in Industry 4.0, conventional methods suffer from information loss due to reliance on dimensionality reduction and feature selection, and fail to capture high-order temporal dependencies. This paper proposes Time-EAPCR-T, an end-to-end deep learning framework. It innovatively replaces LSTM with Transformer in the temporal modeling module to enable lossless fusion of multi-source features and effective capture of high-order dynamic interactions. Furthermore, it introduces the EAPCR (Enhanced Adaptive Principal Component Reconstruction) feature disentanglement mechanism and an improved temporal encoding structure to enhance cross-device and cross-operating-condition generalization. Evaluated on four real-world industrial datasets, Time-EAPCR-T achieves an average 9.2% improvement in F1-score and a 37% reduction in false positive rate over state-of-the-art methods, demonstrating strong practicality and engineering deployability.

0 citationsRead paper

Time-EAPCR: A Deep Learning-Based Novel Approach for Anomaly Detection Applied to the Environmental Field

Mar 12, 2025

Traditional environmental monitoring methods suffer from delayed response, poor generalizability, and difficulty modeling long-term temporal dependencies and high variability. To address these limitations, this paper proposes Time-EAPCR—the first end-to-end anomaly detection framework integrating time embedding, self-attention, permutation convolution, and residual learning—explicitly capturing spatiotemporal evolution patterns and multidimensional feature correlations in environmental data. The architecture significantly enhances robustness and interpretability of anomaly discrimination. Evaluated on four public environmental datasets, it achieves an average 12.6% improvement in F1-score. Deployment feasibility is further validated in a real-world river monitoring system. Experimental results demonstrate strong cross-scenario generalization capability, establishing a novel paradigm for intelligent anomaly monitoring in aquatic ecosystems and water treatment infrastructure.

0 citationsRead paper

Inorganic Catalyst Efficiency Prediction Based on EAPCR Model: A Deep Learning Solution for Multi-Source Heterogeneous Data

Mar 10, 2025

Predicting inorganic catalyst efficiency faces challenges in integrating heterogeneous multi-source data and insufficient mechanistic modeling. To address these, we propose EAPCR, a novel deep learning model featuring an Embedding–Attention–Permutation-Convolution–Residual (EAPCR) architecture. It employs embedding layers to unify heterogeneous inputs, self-attention to enhance cross-condition feature interaction, permutation-invariant convolutional kernels to accommodate catalytic site arrangements, and residual connections to improve training stability. Additionally, EAPCR constructs a feature correlation matrix to explicitly capture multi-scale structure–property relationships. Evaluated on TiO₂ photocatalysis, thermocatalysis, and electrocatalysis tasks, EAPCR consistently outperforms linear regression, random forest, and artificial neural networks—reducing MAE, MSE, and RMSE by 18.7% on average and increasing R² by 0.23. The model demonstrates strong generalizability and enhanced physical interpretability.

0 citationsRead paper