When Robots Say No: The Empathic Ethical Disobedience Benchmark
Robots must balance instruction-following with adherence to safety and social norms, yet existing safety reinforcement learning benchmarks emphasize physical risks, while human-robot trust studies suffer from limited scale and poor reproducibility. Method: We propose the Empathic Ethical Disobedience (EED) benchmark and introduce EED Gym—a standardized, multi-role, multi-scenario testbed enabling systematic evaluation of compliance, refusal, clarification, and alternative-action decisions. Contribution/Results: We jointly quantify refusal behavior along three dimensions: safety, user trust, and empathy. We integrate empirically grounded blame/trust models and a personified role framework, and define verifiable credibility tiers for constructive, empathic, and other refusal styles. Experiments show that action masking eliminates unsafe compliance; explanatory refusals preserve trust; constructive refusals achieve highest credibility scores, while empathic refusals yield highest empathy scores; safety-aware RL improves robustness but often induces excessive conservatism.