Institution profile

Dynamo AI

Industry researchnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs

Oct 14, 2025

Existing medical large language models (LLMs) lack systematic robustness evaluation in multi-turn clinical dialogues; conventional single-turn benchmarks fail to capture real-world challenges such as contradictory inputs, misleading contextual cues, and authoritative bias. Method: We propose MedQA-Followup—a novel framework that formally defines and distinguishes shallow versus deep robustness in multi-turn medical question answering, introducing the “indirect–direct intervention” analytical axis. Leveraging MedQA, we construct a controllable multi-turn test suite simulating realistic clinical consultation disruptions. Contribution/Results: Experiments on five state-of-the-art medical LLMs reveal a dramatic accuracy drop—from 91.2% in single-turn settings to as low as 13.5% in multi-turn scenarios—with indirect contextual interference proving more detrimental than direct prompt manipulation. These findings expose structural fragility in sequential clinical interaction, offering critical risk awareness for clinical deployment and establishing a new evaluation paradigm for dialogue robustness in medical AI.

0 citationsRead paper

Unified Multi-Task Learning&Model Fusion for Efficient Language Model Guardrailing

Apr 27, 2025

To address high latency, substantial memory overhead, prohibitive deployment costs, and unstructured outputs in large language model (LLM) input moderation, this paper proposes UniGuard—a lightweight, efficient safety guard system. Methodologically, it introduces a novel task-customized synthetic data generation mechanism, constructs the multitask pre-trained model MultiTaskGuard, and designs a search-based parameter-space fusion framework to jointly optimize diverse safety policies within a single model. Evaluated on seven public datasets and four internally curated guard benchmarks, UniGuard achieves F1 scores 29.92 points higher than Aegis-LlamaGuard and 21.62 points higher than GPT-4o, significantly outperforming existing LLMs and third-party API-based solutions. Key contributions include: (1) a scalable multitask modeling paradigm; (2) a human-annotation-free synthetic data generation strategy; and (3) an end-to-end guard architecture delivering low computational overhead and high output consistency.

0 citationsRead paper

Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges

Mar 06, 2025

This work exposes severe robustness deficiencies in current LLM safety evaluator models under realistic deployment conditions. Addressing three key challenges—prompt sensitivity, distributional shift, and generation-layer adversarial attacks—the authors conduct the first systematic quantification of failure modes across mainstream evaluators (e.g., Self-Check, SAFETY-LLM): minor stylistic perturbations in model outputs increase false-negative rates by 0.24; targeted adversarial attacks cause 100% misclassification of harmful content as safe. Through prompt sensitivity analysis, stylistic perturbation experiments, and generation-directed adversarial attacks, the study empirically demonstrates the unreliability of offline benchmark evaluations and reveals fundamental flaws in existing meta-evaluation frameworks. The core contribution is the introduction of the first empirical evaluation framework specifically designed to assess the robustness of safety evaluators. The findings critically challenge the validity of prevailing safety assessment paradigms and underscore an urgent need for their foundational rethinking and reconstruction.

0 citationsRead paper
Recent publications

Latest Papers

Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs

Oct 14, 2025

Existing medical large language models (LLMs) lack systematic robustness evaluation in multi-turn clinical dialogues; conventional single-turn benchmarks fail to capture real-world challenges such as contradictory inputs, misleading contextual cues, and authoritative bias. Method: We propose MedQA-Followup—a novel framework that formally defines and distinguishes shallow versus deep robustness in multi-turn medical question answering, introducing the “indirect–direct intervention” analytical axis. Leveraging MedQA, we construct a controllable multi-turn test suite simulating realistic clinical consultation disruptions. Contribution/Results: Experiments on five state-of-the-art medical LLMs reveal a dramatic accuracy drop—from 91.2% in single-turn settings to as low as 13.5% in multi-turn scenarios—with indirect contextual interference proving more detrimental than direct prompt manipulation. These findings expose structural fragility in sequential clinical interaction, offering critical risk awareness for clinical deployment and establishing a new evaluation paradigm for dialogue robustness in medical AI.

0 citationsRead paper

Unified Multi-Task Learning&Model Fusion for Efficient Language Model Guardrailing

Apr 27, 2025

To address high latency, substantial memory overhead, prohibitive deployment costs, and unstructured outputs in large language model (LLM) input moderation, this paper proposes UniGuard—a lightweight, efficient safety guard system. Methodologically, it introduces a novel task-customized synthetic data generation mechanism, constructs the multitask pre-trained model MultiTaskGuard, and designs a search-based parameter-space fusion framework to jointly optimize diverse safety policies within a single model. Evaluated on seven public datasets and four internally curated guard benchmarks, UniGuard achieves F1 scores 29.92 points higher than Aegis-LlamaGuard and 21.62 points higher than GPT-4o, significantly outperforming existing LLMs and third-party API-based solutions. Key contributions include: (1) a scalable multitask modeling paradigm; (2) a human-annotation-free synthetic data generation strategy; and (3) an end-to-end guard architecture delivering low computational overhead and high output consistency.

0 citationsRead paper

Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges

Mar 06, 2025

This work exposes severe robustness deficiencies in current LLM safety evaluator models under realistic deployment conditions. Addressing three key challenges—prompt sensitivity, distributional shift, and generation-layer adversarial attacks—the authors conduct the first systematic quantification of failure modes across mainstream evaluators (e.g., Self-Check, SAFETY-LLM): minor stylistic perturbations in model outputs increase false-negative rates by 0.24; targeted adversarial attacks cause 100% misclassification of harmful content as safe. Through prompt sensitivity analysis, stylistic perturbation experiments, and generation-directed adversarial attacks, the study empirically demonstrates the unreliability of offline benchmark evaluations and reveals fundamental flaws in existing meta-evaluation frameworks. The core contribution is the introduction of the first empirical evaluation framework specifically designed to assess the robustness of safety evaluators. The findings critically challenge the validity of prevailing safety assessment paradigms and underscore an urgent need for their foundational rethinking and reconstruction.

0 citationsRead paper