Institution profile

Guangzhou University

Academic institutionasia · cn
Official website
Research library217linked papers
Opportunities0open roles
Selected work

Representative Papers

ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation

Jan 20, 2026

This work addresses the challenge of "dialogue drift" in large language models during multi-turn conversations, where ambiguous initial instructions lead to persistent erroneous assumptions that are difficult to correct. To mitigate this, the paper proposes a novel reinforcement learning framework that, for the first time, incorporates pragmatic awareness into the reward mechanism. By detecting instruction ambiguity, the framework dynamically modulates reward signals to encourage the model to express epistemic humility or proactively seek clarification under uncertainty. This approach combines verifiable rewards with data augmentation using ambiguous prompts to enable fine-grained control over response style. Experimental results demonstrate that the proposed method improves average performance by 75% on multi-turn dialogue tasks while maintaining robustness on single-turn benchmarks, significantly enhancing both the cooperativeness and robustness of conversational agents.

1 citationsRead paper

Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer Network

Dec 20, 2024arXiv.org

Existing temporal sentence grounding (TSG) methods train on untrimmed video–sentence query pairs independently, neglecting inter-pair correlations—leading to knowledge redundancy, inefficient training, and limited generalization. This paper proposes a novel multi-pair joint TSG paradigm, enabling a single model to collaboratively optimize multiple video–query pairs simultaneously. To this end, we design a multi-threaded knowledge transfer network featuring: (i) cross-modal contrastive learning to strengthen fine-grained alignment; (ii) a dual-granularity prototype matching mechanism—operating at both object/phrase level (spatial) and action/sentence level (temporal); and (iii) adaptive threshold-based hard negative mining coupled with self-supervised representation learning. Extensive experiments on multiple benchmarks demonstrate substantial improvements in both grounding accuracy and inference efficiency, achieving new state-of-the-art performance. Ablation studies confirm the effectiveness of inter-pair knowledge transfer and the model’s strong generalization capability across diverse queries and videos.

1 citationsRead paper
Recent publications

Latest Papers