🤖 AI Summary
This study investigates which types of code review comments developers are more likely to adopt, aiming to guide LLMs in generating high-adoption-rate review suggestions. We propose a five-dimensional taxonomy—readability, defects, maintainability, design, and style—and develop an LLM-as-a-Judge framework to automatically classify both human-written and LLM-generated review comments from real-world projects. Empirical analysis reveals that comments targeting readability, defects, and maintainability achieve significantly higher resolution rates; LLM-generated comments exhibit strong actionability overall and demonstrate complementary strengths relative to human reviewers across diverse project contexts. Our key contributions are: (1) the first systematic empirical characterization of the relationship between comment type and developer adoption behavior, and (2) validation of the LLM-as-a-Judge paradigm for code review quality assessment, demonstrating its effectiveness and generalizability across projects and comment categories.
📝 Abstract
Large language model (LLM)-powered code review automation tools have been introduced to generate code review comments. However, not all generated comments will drive code changes. Understanding what types of generated review comments are likely to trigger code changes is crucial for identifying those that are actionable. In this paper, we set out to investigate (1) the types of review comments written by humans and LLMs, and (2) the types of generated comments that are most frequently resolved by developers. To do so, we developed an LLM-as-a-Judge to automatically classify review comments based on our own taxonomy of five categories. Our empirical study confirms that (1) the LLM reviewer and human reviewers exhibit distinct strengths and weaknesses depending on the project context, and (2) readability, bugs, and maintainability-related comments had higher resolution rates than those focused on code design. These results suggest that a substantial proportion of LLM-generated comments are actionable and can be resolved by developers. Our work highlights the complementarity between LLM and human reviewers and offers suggestions to improve the practical effectiveness of LLM-powered code review tools.