MOCC-R1: Reinforcing Reasoning-Response Consistency for Multimodal Counselor Response Generation

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态心理咨询响应生成的一致性问题,提出基于新数据集MOCC的两阶段框架MOCC-R1,通过冷启动监督微调和强化学习优化推理-响应一致性。
📝 Abstract
Multimodal counselor response generation (MCRG) aims to generate an appropriate counselor response from multimodal dialogue histories. Progress is limited by two gaps: first, existing datasets rarely capture sustained, human-recorded counseling interactions conducted by qualified counselors; Second, existing methods do not explicitly optimize consistency between counseling reasoning and the generated response, potentially undermining the reliability of MCRG systems. Thus, we introduce MOCC, a multimodal counseling conversation corpus containing over 200 hours of interactions involving 154 credential-verified counselors. Based on MOCC, we propose MOCC-R1, a two-stage framework for optimizing reasoning-response consistency. Cold-start supervised fine-tuning trains the model to generate a structured trajectory consisting of client-state understanding, a response intent that links a counseling principle to a planned action, and the final response. Reinforcement learning (RL) then rewards grounded plan coherence and plan execution, encouraging the inferred state and plan to be supported by the dialogue context and the response to realize that plan. Experiments demonstrate the effectiveness of the proposed MOCC-R1.
Problem

Research questions and friction points this paper is trying to address.

multimodal counselor response generation
reasoning-response consistency
counseling interactions
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal counseling
reasoning-response consistency
reinforcement learning
supervised fine-tuning
🔎 Similar Papers
No similar papers found.