Differentially Private Paired Table-Image Multimodal Synthesis
本文提出DP-TabImage框架,通过专用模型分别处理表格和图像数据并保持其依赖性,以解决在差分隐私下合成配对表-图数据的难题。
本文提出DP-TabImage框架,通过专用模型分别处理表格和图像数据并保持其依赖性,以解决在差分隐私下合成配对表-图数据的难题。
This work addresses the critical challenge that safe reinforcement learning (RL) algorithms, while satisfying safety constraints during training, often fail to maintain safety under distributional shifts—such as those arising from different patient populations—during deployment. Using diabetes management as a safety-critical testbed, the study systematically evaluates the cross-population safety generalization of eight prominent safe RL algorithms, including PPO-Lag and CPO, and reveals a significant performance gap for the first time. To mitigate this issue, the authors propose a test-time shielding mechanism based on learned dynamics models. Evaluated within a unified clinical simulator, the approach consistently improves Time-in-Range by 13–14% across three diabetes types and three age groups, while substantially reducing clinical risk indices and glycemic variability, thereby effectively restoring safety across diverse populations.
Existing health AI benchmarks lack personalized evaluation frameworks tailored to daily decision-making for diabetic patients, failing to reflect large language models’ (LLMs) real-world supportive capabilities. To address this gap, we introduce DexBench—the first LLM benchmark dedicated to diabetes self-management—constructed from 15,000 patients’ real-world longitudinal physiological and behavioral data, comprising 360,000 personalized question-answer instances across seven task categories: glycemic interpretation, behavior–glucose association, long-term planning, and more. We propose a novel multidimensional evaluation framework assessing accuracy, safety, actionability, credibility, and clarity, and systematically benchmark eight state-of-the-art LLMs. Results reveal pronounced performance imbalances across models, with no single model dominating across all dimensions. DexBench fills a critical void in patient-facing health AI evaluation, providing both a rigorous assessment tool and actionable insights to enhance LLM reliability and practical utility in metabolic health management.
本文提出DP-TabImage框架,通过专用模型分别处理表格和图像数据并保持其依赖性,以解决在差分隐私下合成配对表-图数据的难题。
This work addresses the critical challenge that safe reinforcement learning (RL) algorithms, while satisfying safety constraints during training, often fail to maintain safety under distributional shifts—such as those arising from different patient populations—during deployment. Using diabetes management as a safety-critical testbed, the study systematically evaluates the cross-population safety generalization of eight prominent safe RL algorithms, including PPO-Lag and CPO, and reveals a significant performance gap for the first time. To mitigate this issue, the authors propose a test-time shielding mechanism based on learned dynamics models. Evaluated within a unified clinical simulator, the approach consistently improves Time-in-Range by 13–14% across three diabetes types and three age groups, while substantially reducing clinical risk indices and glycemic variability, thereby effectively restoring safety across diverse populations.
Existing health AI benchmarks lack personalized evaluation frameworks tailored to daily decision-making for diabetic patients, failing to reflect large language models’ (LLMs) real-world supportive capabilities. To address this gap, we introduce DexBench—the first LLM benchmark dedicated to diabetes self-management—constructed from 15,000 patients’ real-world longitudinal physiological and behavioral data, comprising 360,000 personalized question-answer instances across seven task categories: glycemic interpretation, behavior–glucose association, long-term planning, and more. We propose a novel multidimensional evaluation framework assessing accuracy, safety, actionability, credibility, and clarity, and systematically benchmark eight state-of-the-art LLMs. Results reveal pronounced performance imbalances across models, with no single model dominating across all dimensions. DexBench fills a critical void in patient-facing health AI evaluation, providing both a rigorous assessment tool and actionable insights to enhance LLM reliability and practical utility in metabolic health management.