Institution profile

ENN Group

Industry researchasia · cn
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models

May 02, 2026

This work addresses the susceptibility of current vision-language models to visual perception errors and hallucinations during complex reasoning, a challenge exacerbated by the low sampling efficiency and sparse rewards of conventional reinforcement learning approaches that struggle to disentangle error sources. To overcome these limitations, the authors propose MIRL, a novel framework that introduces mutual information into vision-language reinforcement learning for the first time. Specifically, MIRL leverages the mutual information between generated captions and visual inputs as a pre-screening signal, employs a trajectory-forking mechanism to intelligently allocate sampling budgets, and adopts a decoupled training strategy that provides dedicated rewards for the visual perception stage. Evaluated across six benchmarks, MIRL achieves an average accuracy of 70.22%, surpassing the performance of methods using 16 full trajectories with only 10 pre-samples followed by top-6 selection—reducing sampling cost by 25% while significantly improving both training efficiency and accuracy.

0 citationsRead paper

Fusing Large Language Models with Temporal Transformers for Time Series Forecasting

Jul 14, 2025

This work addresses two key limitations in time-series forecasting: the insufficient temporal modeling capability of large language models (LLMs) and the lack of high-level semantic understanding in conventional Transformers. To bridge this gap, we propose a novel collaborative fusion architecture integrating LLMs and time-series Transformers. Methodologically, we design a dual-stream encoder: one stream leverages prompt learning to guide the LLM in extracting sequence-level semantic representations, while the other employs a time-series Transformer to capture dynamic temporal dependencies. These streams undergo learnable cross-modal representation fusion at intermediate layers and are jointly optimized end-to-end. Our key contribution is the first explicit, differentiable, and structurally grounded integration of LLM-derived semantic priors with time-series dynamics. Extensive experiments on multiple benchmark datasets demonstrate that our method significantly outperforms both pure-LLM baselines and state-of-the-art time-series models (e.g., Informer, Autoformer), achieving average MAE reductions of 12.3%–18.7%.

0 citationsRead paper

Semi-Supervised Federated Learning via Dual Contrastive Learning and Soft Labeling for Intelligent Fault Diagnosis

Jul 12, 2025

To address the challenges of scarce labeled data, heterogeneous client data distributions, and privacy constraints in industrial intelligent fault diagnosis, this paper proposes SSFL-DCSL, a semi-supervised federated learning framework. Methodologically, SSFL-DCSL innovatively integrates dual contrastive learning (local and global) with a soft-labeling mechanism; introduces a Laplacian-weighted function to mitigate pseudo-label bias; and employs momentum-updated prototype aggregation to enable cross-client knowledge transfer. By jointly optimizing semi-supervised and federated learning objectives, the framework significantly enhances feature representation consistency and generalization under low-resource conditions. Extensive experiments on three public benchmark datasets demonstrate that, using only 10% labeled data, SSFL-DCSL achieves accuracy improvements of 1.15–7.85% over state-of-the-art methods. The framework thus effectively reconciles the tripartite challenges of few-shot learning, non-IID data, and privacy preservation in industrial fault diagnosis.

0 citationsRead paper
Recent publications

Latest Papers

MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models

May 02, 2026

This work addresses the susceptibility of current vision-language models to visual perception errors and hallucinations during complex reasoning, a challenge exacerbated by the low sampling efficiency and sparse rewards of conventional reinforcement learning approaches that struggle to disentangle error sources. To overcome these limitations, the authors propose MIRL, a novel framework that introduces mutual information into vision-language reinforcement learning for the first time. Specifically, MIRL leverages the mutual information between generated captions and visual inputs as a pre-screening signal, employs a trajectory-forking mechanism to intelligently allocate sampling budgets, and adopts a decoupled training strategy that provides dedicated rewards for the visual perception stage. Evaluated across six benchmarks, MIRL achieves an average accuracy of 70.22%, surpassing the performance of methods using 16 full trajectories with only 10 pre-samples followed by top-6 selection—reducing sampling cost by 25% while significantly improving both training efficiency and accuracy.

0 citationsRead paper

Fusing Large Language Models with Temporal Transformers for Time Series Forecasting

Jul 14, 2025

This work addresses two key limitations in time-series forecasting: the insufficient temporal modeling capability of large language models (LLMs) and the lack of high-level semantic understanding in conventional Transformers. To bridge this gap, we propose a novel collaborative fusion architecture integrating LLMs and time-series Transformers. Methodologically, we design a dual-stream encoder: one stream leverages prompt learning to guide the LLM in extracting sequence-level semantic representations, while the other employs a time-series Transformer to capture dynamic temporal dependencies. These streams undergo learnable cross-modal representation fusion at intermediate layers and are jointly optimized end-to-end. Our key contribution is the first explicit, differentiable, and structurally grounded integration of LLM-derived semantic priors with time-series dynamics. Extensive experiments on multiple benchmark datasets demonstrate that our method significantly outperforms both pure-LLM baselines and state-of-the-art time-series models (e.g., Informer, Autoformer), achieving average MAE reductions of 12.3%–18.7%.

0 citationsRead paper

Semi-Supervised Federated Learning via Dual Contrastive Learning and Soft Labeling for Intelligent Fault Diagnosis

Jul 12, 2025

To address the challenges of scarce labeled data, heterogeneous client data distributions, and privacy constraints in industrial intelligent fault diagnosis, this paper proposes SSFL-DCSL, a semi-supervised federated learning framework. Methodologically, SSFL-DCSL innovatively integrates dual contrastive learning (local and global) with a soft-labeling mechanism; introduces a Laplacian-weighted function to mitigate pseudo-label bias; and employs momentum-updated prototype aggregation to enable cross-client knowledge transfer. By jointly optimizing semi-supervised and federated learning objectives, the framework significantly enhances feature representation consistency and generalization under low-resource conditions. Extensive experiments on three public benchmark datasets demonstrate that, using only 10% labeled data, SSFL-DCSL achieves accuracy improvements of 1.15–7.85% over state-of-the-art methods. The framework thus effectively reconciles the tripartite challenges of few-shot learning, non-IID data, and privacy preservation in industrial fault diagnosis.

0 citationsRead paper