🤖 AI Summary
This study pioneers the application of large language models (LLMs) to phenomenological analysis of borderline personality disorder (BPD), specifically examining their capacity to interpret first-person narratives concerning patients’ lived experiences of temporality and selfhood.
Method: We employed GPT-4o, Gemini 2.5 Pro, and Claude Opus 4 to analyze life-story interview transcripts, evaluating outputs quantitatively via semantic consistency, Jaccard similarity, and qualitatively through multidimensional validity criteria—credibility, coherence, substantive adequacy, and data grounding.
Contribution/Results: Gemini 2.5 Pro achieved the highest thematic overlap (58%, *p* < 0.0001 vs. others), optimal validity scores, and was indistinguishable from human analysts in blind expert evaluation. Thematic extraction quality correlated strongly with input text length (*R* > 0.78). Critically, LLMs successfully recovered themes missed by human analysts, demonstrating potential to mitigate interpretive bias and enhance rigor in phenomenological research.
📝 Abstract
This study examines the capacity of large language models (LLMs) to support phenomenological qualitative analysis of first-person experience in Borderline Personality Disorder (BPD), understood as a disorder of temporality and selfhood. Building on a prior human-led thematic analysis of 24 inpatients' life-story interviews, we compared three LLMs (OpenAI GPT-4o, Google Gemini 2.5 Pro, Anthropic Claude Opus 4) prompted to mimic the interpretative style of the original investigators. The models were evaluated with blinded and non-blinded expert judges in phenomenology and clinical psychology. Assessments included semantic congruence, Jaccard coefficients, and multidimensional validity ratings (credibility, coherence, substantiveness, and groundness in data). Results showed variable overlap with the human analysis, from 0 percent in GPT to 42 percent in Claude and 58 percent in Gemini, and a low Jaccard coefficient (0.21-0.28). However, the models recovered themes omitted by humans. Gemini's output most closely resembled the human analysis, with validity scores significantly higher than GPT and Claude (p < 0.0001), and was judged as human by blinded experts. All scores strongly correlated (R > 0.78) with the quantity of text and words per theme, highlighting both the variability and potential of AI-augmented thematic analysis to mitigate human interpretative bias.