🤖 AI Summary
This study addresses the high labor costs of semi-structured interviews, the opaque behavioral mechanisms of existing multimodal large language models (MLLMs), and the challenges in establishing trust during human–AI interaction. To this end, the authors developed InterviewBot—a system that integrates researcher-designed interview protocols—and conducted the first empirical analysis of MLLM turn-by-turn behaviors in a real-world deployment setting. Leveraging voice-driven real-time MLLM interaction, a semi-structured interview framework, and inductive qualitative methods, the study identifies four data collection failure modes, including information loss and premature termination, and reveals that only 4.9% of AI-generated questions constituted probing follow-ups while 28.7% violated single-question instructions. The findings uncover three key socio-dynamic mechanisms—disclosure calibration, institutional legitimacy, and conversational anchoring—and distill design principles centered on enhanced depth control and non-scripted listening.
📝 Abstract
Semi-structured interviews are a cornerstone of qualitative research but remain labor-intensive. We report an empirical study of what actually happens when the interviewer is an off-the-shelf real-time multimodal LLM (MLLM). We built InterviewBot, a voice-based interviewing system that wraps a real-time MLLM with a researcher-authored outline, and deployed it not as a novel architecture but as a research instrument for observing default MLLM interviewing behavior. In a practice study (N=15), participants completed a bot-led semi-structured interview and then a human-led reflection session about that experience. We contribute (i) a turn-level behavioral analysis of an MLLM interviewer (N_turns=428) showing that it is acknowledgment-heavy but probe-light (deepening probes account for 4.9% of all turns), and that 28.7% of question-bearing turns pack multiple questions into one turn despite an explicit one-question-at-a-time instruction; (ii) an inductive catalogue of four data-collection breakdowns (information loss, premature termination, latency, and interruption) observed in a deployed rather than simulated system; and (iii) three social dynamics from participants' reflections: disclosure calibration, where reduced social pressure coincided with shallower elaboration; institutional legitimacy, where trust tracked perceived stakes and what delegation to AI signaled about the organizer rather than conversational competence; and conversational grounding, where content-grounded paraphrase, not generic social filler, was what participants read as listening. We conclude with design implications for depth control, transparent handoffs, and non-templated listening mechanisms in human-centered interview automation.