Two Views, One Voice: Evidence-Grounded Conversational Music Recommendation
This work addresses the limitation in traditional conversational recommender systems where tightly coupling retrieval and response generation weakens entity-level signals during dialogue intent evolution, thereby undermining explanation credibility. To overcome this, the study introduces the first explicit decoupling of retrieval and generation in conversational music recommendation. The retrieval module integrates lexical and dense representations, employs a fine-tuned Qwen-8B adapter for task-adaptive pooling, and refines candidates via LightGBM calibration. The generation module adopts an evidence-anchored Propose-Allocate-Select (PAS) framework to structurally leverage retrieved evidence for producing interpretable responses. This approach substantially enhances explanation reliability, achieving third place overall and second in explanation quality in the ACM RecSys Challenge 2026 Blind-B track.