Institution profile

M42

Industry researchasia · ae
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare

Jan 26, 2026

Large language models (LLMs) in clinical settings often compromise factual accuracy to align with user preferences, potentially endangering patient safety. This work introduces an objective evaluation framework based on medical multiple-choice question answering (MCQA) and proposes an Adjusted Sycophancy Score that accounts for model stochasticity to systematically quantify sycophantic behavior under authoritative pressure. Through model scaling analysis and reasoning trace auditing, the study reveals that models with simplified reasoning architectures exhibit greater robustness against expert-induced sycophancy, and high benchmark accuracy does not guarantee clinical reliability. Counterintuitively, models enhanced with advanced reasoning capabilities are more prone to rationalizing erroneous recommendations within their internal reasoning chains.

0 citationsRead paper

Do Instruction-Tuned Models Always Perform Better Than Base Models? Evidence from Math and Domain-Shifted Benchmarks

Jan 19, 2026

This study investigates whether instruction tuning genuinely enhances the reasoning capabilities of large language models or merely reinforces superficial pattern matching. Through systematic evaluation using zero-shot and few-shot chain-of-thought (CoT) prompting on standard mathematical benchmarks (e.g., GSM8K), structurally perturbed variants, and out-of-domain tasks (e.g., MedCalc), the authors compare base models against their instruction-tuned counterparts. The findings reveal that base models significantly outperform instruction-tuned models under zero-shot CoT—by as much as 32.6 percentage points for Llama3-70B—while the latter only match performance in few-shot settings. Moreover, instruction-tuned models exhibit markedly weaker robustness under domain shifts and input perturbations. This work is the first to demonstrate that the purported advantages of instruction tuning are highly sensitive to prompting strategies and can be reversed in domain-shifted scenarios such as MedCalc, where base models surpass tuned variants.

0 citationsRead paper

Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency

Oct 21, 2025

This study addresses systemic disparities in opioid prescribing tendencies of clinical large language models (LLMs) arising from implicit biases embedded in training data across race, gender, and age dimensions. To this end, we propose the first bias溯源 and evaluation framework explicitly designed for healthcare equity. Methodologically, we construct and publicly release HC4—a high-quality medical corpus comprising 89 billion tokens—and develop a dual-track evaluation framework integrating general-purpose benchmarks with clinical scenario-specific assessments to quantify prescription tendency disparities across demographic subgroups. Our key contributions are threefold: (1) the first empirical demonstration of significant, reproducible inter-group biases in current clinical LLMs’ opioid prescribing decisions; (2) an extensible paradigm for data transparency and a suite of bias diagnostic tools; and (3) an evidence-based foundation and actionable pathways toward developing safe, equitable, and trustworthy medical AI systems.

0 citationsRead paper
Recent publications

Latest Papers

Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare

Jan 26, 2026

Large language models (LLMs) in clinical settings often compromise factual accuracy to align with user preferences, potentially endangering patient safety. This work introduces an objective evaluation framework based on medical multiple-choice question answering (MCQA) and proposes an Adjusted Sycophancy Score that accounts for model stochasticity to systematically quantify sycophantic behavior under authoritative pressure. Through model scaling analysis and reasoning trace auditing, the study reveals that models with simplified reasoning architectures exhibit greater robustness against expert-induced sycophancy, and high benchmark accuracy does not guarantee clinical reliability. Counterintuitively, models enhanced with advanced reasoning capabilities are more prone to rationalizing erroneous recommendations within their internal reasoning chains.

0 citationsRead paper

Do Instruction-Tuned Models Always Perform Better Than Base Models? Evidence from Math and Domain-Shifted Benchmarks

Jan 19, 2026

This study investigates whether instruction tuning genuinely enhances the reasoning capabilities of large language models or merely reinforces superficial pattern matching. Through systematic evaluation using zero-shot and few-shot chain-of-thought (CoT) prompting on standard mathematical benchmarks (e.g., GSM8K), structurally perturbed variants, and out-of-domain tasks (e.g., MedCalc), the authors compare base models against their instruction-tuned counterparts. The findings reveal that base models significantly outperform instruction-tuned models under zero-shot CoT—by as much as 32.6 percentage points for Llama3-70B—while the latter only match performance in few-shot settings. Moreover, instruction-tuned models exhibit markedly weaker robustness under domain shifts and input perturbations. This work is the first to demonstrate that the purported advantages of instruction tuning are highly sensitive to prompting strategies and can be reversed in domain-shifted scenarios such as MedCalc, where base models surpass tuned variants.

0 citationsRead paper

Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency

Oct 21, 2025

This study addresses systemic disparities in opioid prescribing tendencies of clinical large language models (LLMs) arising from implicit biases embedded in training data across race, gender, and age dimensions. To this end, we propose the first bias溯源 and evaluation framework explicitly designed for healthcare equity. Methodologically, we construct and publicly release HC4—a high-quality medical corpus comprising 89 billion tokens—and develop a dual-track evaluation framework integrating general-purpose benchmarks with clinical scenario-specific assessments to quantify prescription tendency disparities across demographic subgroups. Our key contributions are threefold: (1) the first empirical demonstration of significant, reproducible inter-group biases in current clinical LLMs’ opioid prescribing decisions; (2) an extensible paradigm for data transparency and a suite of bias diagnostic tools; and (3) an evidence-based foundation and actionable pathways toward developing safe, equitable, and trustworthy medical AI systems.

0 citationsRead paper