Institution profile

Huazhong Agricultural University

Academic institutionasia · cn
Official website
Research library61linked papers
Opportunities0open roles
Selected work

Representative Papers

Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing

Aug 03, 2026

This work addresses the scarcity of training-free, open-vocabulary methods for semantic segmentation in remote sensing imagery, a task typically hindered by the high cost of pixel-level annotations. The authors propose DinoSplat-OV, a novel framework that, for the first time, directly leverages DINOv3 for open-vocabulary segmentation of remote sensing images without requiring fine-tuning or additional pretraining. To tackle the challenges posed by the dense, multi-scale, and large-size nature of such imagery, the method integrates text-guided denoising Laplacian propagation, RGB-guided anisotropic feature aggregation, Gaussian lattice upsampling, and a global anchor sliding-window mechanism. Evaluated on UDD5, DOTA, and LoveDA benchmarks, DinoSplat-OV achieves performance on par with or superior to existing zero-shot approaches, thereby filling a critical gap in the application of DINO-based models to this domain.

0 citationsRead paper

From Understanding to Action: Feedback-Grounded Policy Discovery for Generative Recommendation

Jul 30, 2026

This work addresses the "understanding-action gap" in large language models (LLMs) when applied to recommender systems by proposing a feedback-driven agent framework. The approach first infers task-oriented user intent and then discovers effective recommendation strategies based on incremental utility and outcome feedback—rather than linguistic plausibility. It innovatively decouples the modeling of intent and policy knowledge, and compresses both into a lightweight semantic ID generator via dual-space relational distillation, enabling efficient LLM-free online inference. Evaluated on public benchmarks, the method significantly outperforms existing baselines, and large-scale online A/B tests demonstrate a 4.506% increase in revenue and a 4.621% improvement in ADVV (Average Daily Value per Visitor).

0 citationsRead paper

Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis

Jul 30, 2026

This work addresses the limitation of current large language models in multimodal sentiment analysis, which struggle to capture deep emotional semantics arising from structural dynamics and contextual interactions, often relying only on superficial features. To overcome this, the authors propose SentiLLM, a novel framework that first converts continuous non-verbal signals into compact, semantically grounded text-like tokens through structure-aware semantic abstraction. It then introduces a dual-stream salience–context calibration mechanism that decouples focal and ambient streams, leveraging textual priors to guide the detection of sentiment shifts and achieve cross-modal semantic alignment. Finally, a lightweight, plug-and-play adapter module is integrated for efficient adaptation. Evaluated on MOSI, MOSEI, CH-SIMS, and CH-SIMS v2 benchmarks, SentiLLM achieves state-of-the-art performance with significantly improved discriminative capability while requiring only a minimal number of trainable parameters.

0 citationsRead paper
Recent publications

Latest Papers

Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing

Aug 03, 2026

This work addresses the scarcity of training-free, open-vocabulary methods for semantic segmentation in remote sensing imagery, a task typically hindered by the high cost of pixel-level annotations. The authors propose DinoSplat-OV, a novel framework that, for the first time, directly leverages DINOv3 for open-vocabulary segmentation of remote sensing images without requiring fine-tuning or additional pretraining. To tackle the challenges posed by the dense, multi-scale, and large-size nature of such imagery, the method integrates text-guided denoising Laplacian propagation, RGB-guided anisotropic feature aggregation, Gaussian lattice upsampling, and a global anchor sliding-window mechanism. Evaluated on UDD5, DOTA, and LoveDA benchmarks, DinoSplat-OV achieves performance on par with or superior to existing zero-shot approaches, thereby filling a critical gap in the application of DINO-based models to this domain.

0 citationsRead paper

From Understanding to Action: Feedback-Grounded Policy Discovery for Generative Recommendation

Jul 30, 2026

This work addresses the "understanding-action gap" in large language models (LLMs) when applied to recommender systems by proposing a feedback-driven agent framework. The approach first infers task-oriented user intent and then discovers effective recommendation strategies based on incremental utility and outcome feedback—rather than linguistic plausibility. It innovatively decouples the modeling of intent and policy knowledge, and compresses both into a lightweight semantic ID generator via dual-space relational distillation, enabling efficient LLM-free online inference. Evaluated on public benchmarks, the method significantly outperforms existing baselines, and large-scale online A/B tests demonstrate a 4.506% increase in revenue and a 4.621% improvement in ADVV (Average Daily Value per Visitor).

0 citationsRead paper

Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis

Jul 30, 2026

This work addresses the limitation of current large language models in multimodal sentiment analysis, which struggle to capture deep emotional semantics arising from structural dynamics and contextual interactions, often relying only on superficial features. To overcome this, the authors propose SentiLLM, a novel framework that first converts continuous non-verbal signals into compact, semantically grounded text-like tokens through structure-aware semantic abstraction. It then introduces a dual-stream salience–context calibration mechanism that decouples focal and ambient streams, leveraging textual priors to guide the detection of sentiment shifts and achieve cross-modal semantic alignment. Finally, a lightweight, plug-and-play adapter module is integrated for efficient adaptation. Evaluated on MOSI, MOSEI, CH-SIMS, and CH-SIMS v2 benchmarks, SentiLLM achieves state-of-the-art performance with significantly improved discriminative capability while requiring only a minimal number of trainable parameters.

0 citationsRead paper