🤖 AI Summary
This study addresses the limitation of current legal AI evaluations in capturing authentic witness behavioral dynamics for cross-examination training. We propose WitnessSim, a controllable simulator driven by distinct legal personas, and introduce a novel evaluation framework that decouples behavioral authenticity from pedagogical utility. Through adversarial testing, blind attorney assessments, and longitudinal trajectory analysis, we demonstrate that generated testimonies exhibit no significant preference difference compared to real ones, while maintaining reasonable behavioral responses and stable personalities. This work achieves high-fidelity simulation of legal behaviors alongside systematic performance evaluation, effectively bridging the gap in independently assessing dynamic behavioral authenticity within legal AI systems.
📝 Abstract
Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simulator driven by controllable legal personas. We use an evaluation framework separating behavioral realism from pedagogical usefulness. We assess realism through adversarial testing, blinded attorney comparison, and analysis of longitudinal behavioral trajectories. WitnessSim generally maintained plausible behavioral boundaries, and attorneys did not systematically prefer either original testimony or WitnessSim generated testimony. Pedagogical tests showed that witness behavior changed meaningfully in response to question form and attorney intervention without uniformly collapsing the assigned persona. Together, these results showcase a model of behavioral fidelity in legal simulations, and provide a framework for evaluating its performance.