Institution profile

Henan Polytechnic University

Academic institutionasia · cn
Official website
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

Multi-Granularity Reasoning for Image Quality Assessment via Attribute-Aware Reinforcement Learning to Rank

Apr 07, 2026

This work addresses the limitation of existing image quality assessment (IQA) methods, which typically predict only a single scalar score while neglecting the multidimensional attributes—such as sharpness, color fidelity, noise, and composition—that underpin human perception. To overcome this, the authors propose MG-IQA, a novel framework that jointly predicts overall quality and fine-grained perceptual attributes in a single inference pass. MG-IQA integrates vision-language models with reinforcement learning–based ranking, leveraging attribute-aware prompts, a multidimensional Thurstone reward model, and a cross-domain alignment mechanism to enable interpretable, multi-granularity evaluation without requiring perceptual scale realignment. Experiments demonstrate that MG-IQA consistently outperforms state-of-the-art methods across eight IQA benchmarks, achieving an average 2.1% improvement in SRCC for overall quality prediction and generating explanations highly aligned with human judgments.

0 citationsRead paper

HSI Image Enhancement Classification Based on Knowledge Distillation: A Study on Forgetting

Mar 18, 2026

This work addresses catastrophic forgetting in incremental hyperspectral image classification by proposing a knowledge retention method that operates without storing samples from previously seen classes. The approach leverages a teacher model and employs a masking mechanism to perform partial-class knowledge distillation using only the new-class data available in each incremental phase. This design effectively decouples the distillation process and filters out misleading information, thereby preserving relevant knowledge from earlier tasks. Without relying on rehearsal or replay of old samples, the proposed method significantly enhances both classification accuracy and model robustness. Its effectiveness is consistently validated through comprehensive comparative and ablation experiments across multiple benchmarks.

0 citationsRead paper

DeCoT: Decomposing Complex Instructions for Enhanced Text-to-Image Generation with Large Language Models

Aug 17, 2025

Current text-to-image (T2I) models struggle to accurately interpret complex, lengthy textual prompts, exhibiting significant semantic gaps in fine-grained detail reconstruction, spatial relationship modeling, and multi-constraint coordination. To address this, we propose DeCoT—a two-stage framework. In Stage I, a large language model (LLM) performs structured decomposition and semantic clarification of the original prompt. In Stage II, hierarchical prompt construction and chain-of-thought–driven semantic enhancement jointly model text embeddings, compositional logic, and constraint conditions. DeCoT is architecture-agnostic and seamlessly integrates with mainstream T2I models without architectural modification. Evaluated on the LongBench-T2I benchmark, DeCoT integrated with Infinity-8B achieves a score of 3.52—substantially outperforming baseline methods. Both multimodal LLM (MLLM)-based and human evaluations confirm consistent improvements across textual fidelity, spatial consistency, and compositional plausibility.

0 citationsRead paper
Recent publications

Latest Papers

Multi-Granularity Reasoning for Image Quality Assessment via Attribute-Aware Reinforcement Learning to Rank

Apr 07, 2026

This work addresses the limitation of existing image quality assessment (IQA) methods, which typically predict only a single scalar score while neglecting the multidimensional attributes—such as sharpness, color fidelity, noise, and composition—that underpin human perception. To overcome this, the authors propose MG-IQA, a novel framework that jointly predicts overall quality and fine-grained perceptual attributes in a single inference pass. MG-IQA integrates vision-language models with reinforcement learning–based ranking, leveraging attribute-aware prompts, a multidimensional Thurstone reward model, and a cross-domain alignment mechanism to enable interpretable, multi-granularity evaluation without requiring perceptual scale realignment. Experiments demonstrate that MG-IQA consistently outperforms state-of-the-art methods across eight IQA benchmarks, achieving an average 2.1% improvement in SRCC for overall quality prediction and generating explanations highly aligned with human judgments.

0 citationsRead paper

HSI Image Enhancement Classification Based on Knowledge Distillation: A Study on Forgetting

Mar 18, 2026

This work addresses catastrophic forgetting in incremental hyperspectral image classification by proposing a knowledge retention method that operates without storing samples from previously seen classes. The approach leverages a teacher model and employs a masking mechanism to perform partial-class knowledge distillation using only the new-class data available in each incremental phase. This design effectively decouples the distillation process and filters out misleading information, thereby preserving relevant knowledge from earlier tasks. Without relying on rehearsal or replay of old samples, the proposed method significantly enhances both classification accuracy and model robustness. Its effectiveness is consistently validated through comprehensive comparative and ablation experiments across multiple benchmarks.

0 citationsRead paper

DeCoT: Decomposing Complex Instructions for Enhanced Text-to-Image Generation with Large Language Models

Aug 17, 2025

Current text-to-image (T2I) models struggle to accurately interpret complex, lengthy textual prompts, exhibiting significant semantic gaps in fine-grained detail reconstruction, spatial relationship modeling, and multi-constraint coordination. To address this, we propose DeCoT—a two-stage framework. In Stage I, a large language model (LLM) performs structured decomposition and semantic clarification of the original prompt. In Stage II, hierarchical prompt construction and chain-of-thought–driven semantic enhancement jointly model text embeddings, compositional logic, and constraint conditions. DeCoT is architecture-agnostic and seamlessly integrates with mainstream T2I models without architectural modification. Evaluated on the LongBench-T2I benchmark, DeCoT integrated with Infinity-8B achieves a score of 3.52—substantially outperforming baseline methods. Both multimodal LLM (MLLM)-based and human evaluations confirm consistent improvements across textual fidelity, spatial consistency, and compositional plausibility.

0 citationsRead paper