Measuring the practice of shared-decision making (OPTION12): An Investigation into Open-sourced Smaller LLMs (OS-sLLMs) for Better Privacy and Sustainability
This study explores the feasibility of using locally deployed, small open-source language models to automatically evaluate shared decision-making (SDM) in clinical settings while preserving privacy and sustainability. Grounded in the Observer OPTION12 framework, we assess multiple general-purpose and medical-domain small language models on Dutch melanoma consultation transcripts and introduce, for the first time, a Judge-LLM multi-model consensus mechanism to resolve scoring discrepancies. Experimental results indicate that general-purpose models outperform medical-specific ones, with Gemma3:12b achieving the highest correlation with human ratings (Pearson r = 0.51, Spearman ρ = 0.59). This work presents a novel, privacy-preserving, on-premises approach to automated SDM assessment under the OPTION12 framework and reveals systematic model limitations in temporal reasoning, role attribution, and evidence anchoring.