MORTAR: Metamorphic Multi-turn Testing for LLM-based Dialogue Systems

📅 2024-12-20
🏛️ arXiv.org
📈 Citations: 2
Influential: 0
📄 PDF
🤖 AI Summary
To address the long-standing test oracle problem in multi-turn LLM-based dialogue systems, this paper proposes the first knowledge graph–driven metamorphic testing framework for multi-turn dialogues. Methodologically: (1) it constructs dialogue-level knowledge graphs to model semantic structures across turns; and (2) it designs dialogue-level perturbations and metamorphic relations to enable reference-free, LLM-free automated test case generation. The key contribution is the first systematic adaptation of metamorphic testing to multi-turn dialogue scenarios—circumventing evaluation bias introduced by LLM-based oracles. Experiments on multiple mainstream LLM dialogue systems demonstrate that our approach detects significantly more unique defects than existing methods; notably, its detection rate for severe defects is four times higher than that of the best-performing single-turn metamorphic testing baseline, while maintaining low computational cost and high reliability.

Technology Category

Application Category

📝 Abstract
With the widespread application of LLM-based dialogue systems in daily life, quality assurance has become more important than ever. Recent research has successfully introduced methods to identify unexpected behaviour in single-turn scenarios. However, multi-turn dialogue testing remains underexplored, with the Oracle problem in multi-turn testing posing a persistent challenge for dialogue system developers and researchers. In this paper, we propose MORTAR, a MetamORphic multi-TuRn diAlogue testing appRoach, which mitigates the test oracle problem in the assessment of LLM-based dialogue systems. MORTAR automates the generation of follow-up question-answer (QA) dialogue test cases with multiple dialogue-level perturbations and metamorphic relations. MORTAR employs a novel knowledge graph-based dialogue information model which effectively generates perturbed dialogue test datasets and detects bugs of multi-turn dialogue systems in a low-cost manner. The proposed approach does not require an LLM as a judge, eliminating potential of any biases in the evaluation step. According to the experiment results on multiple LLM-based dialogue systems and comparisons with single-turn metamorphic testing approaches, MORTAR explores more unique bugs in LLM-based dialogue systems, especially for severe bugs that MORTAR detects up to four times more unique bugs than the most effective existing metamorphic testing approach.
Problem

Research questions and friction points this paper is trying to address.

Addresses multi-turn testing challenges in LLM-based dialogue systems
Automates test case generation with perturbations and metamorphic relations
Improves bug detection effectiveness and quality without LLM judges
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automated multi-turn dialogue test generation
Perturbation-MR matching for flexible testing
No reliance on biased LLM test oracles
💼 Related Jobs
No related jobs found.
G
Guoxiang Guo
Faculty of Information Technology, Monash University, Australia
A
A. Aleti
Faculty of Information Technology, Monash University, Australia
N
Neelofar Neelofar
School of Computing Technologies, RMIT University, Australia
C
C. Tantithamthavorn
Faculty of Information Technology, Monash University, Australia