Institution profile

Northwestern University

Academic institutionnorthamerica · us
Official website
Research library1,167linked papers
Opportunities0open roles
Selected work

Representative Papers

Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

Jan 14, 2026Research Square

Current safety evaluations of mental health–focused AI systems predominantly rely on small-scale simulated benchmarks, which inadequately capture the linguistic and contextual diversity of real-world scenarios. This study presents the first systematic safety assessment combining replication across four established benchmarks with an ecological audit of 20,000 real user conversations, comparing specialized mental health AI against six state-of-the-art general-purpose large language models on high-risk topics. Employing clinical expert blind review, LLM-based adjudicators, automated crisis resource triggering, and statistical confidence interval analysis, the findings reveal that the specialized system exhibits significantly lower rates of harmful content in response to prompts involving self-harm, eating disorders, and substance abuse. In live deployment, it achieved zero end-to-end missed detections, with the LLM adjudicator demonstrating 100% sensitivity and 99.2% specificity. The work advocates for ecological auditing as a critical complement to pre-deployment safety testing.

7 citationsRead paper

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation

May 29, 2025

To address the challenges of cross-domain transfer in multi-agent systems—namely, the need for redesign and retraining—we propose Workforce, a hierarchical multi-agent framework that decouples domain-agnostic planning (Planner-Coordinator) from domain-specific execution (modular, plug-and-play Workers), enabling zero-shot cross-domain adaptation. We introduce OWL (Online Weighted Learning), a novel reinforcement learning method that enables the planner to achieve domain-invariant generalization through optimization guided by real-world feedback. Workforce integrates tool invocation, modular Worker architecture, and hierarchical coordination mechanisms. On the GAIA benchmark, Workforce achieves state-of-the-art open-source performance (69.70%). A 32B model trained with OWL attains 52.73% accuracy—16.37 percentage points higher than the baseline—and approaches the performance of GPT-4o.

4 citationsRead paper

Demand Estimation with Text and Image Data

Mar 26, 2025Social Science Research Network

Traditional demand estimation struggles to quantify implicit product attributes—such as visual design—leading to biased modeling of substitution relationships. Method: We propose a novel method that automatically infers consumer substitution preferences from multimodal unstructured data (images and text) by jointly embedding visual and textual features via CLIP and BERT, and directly integrating these embeddings into a random-coefficients logit model—eliminating the need for manual attribute specification and enabling end-to-end learning of latent preferences and substitution structures. We combine choice experiment data with a counterfactual prediction framework. Results: Our approach significantly improves prediction accuracy for second-best alternatives in experiments. Empirically applied across 40 Amazon product categories, it consistently identifies substitution sets better aligned with market intuition. The method establishes a scalable, interpretable paradigm for demand modeling in high-dimensional, unstructured environments.

4 citationsRead paper

Persuasion with Ambiguous Communication

Jul 08, 2024ACM Conference on Economics and Computation

This paper investigates whether a sender can enhance persuasion effectiveness through ambiguous communication under ambiguity aversion. Method: Extending the Bayesian persuasion framework, we incorporate the smooth ambiguity model of Klibanoff et al. (2005) and systematically characterize—using concavification techniques and incentive-compatible signal design—the necessary and sufficient conditions for effective ambiguous communication. Contribution/Results: We establish that ambiguous communication never improves the sender’s payoff in binary-action settings; however, in settings with three or more actions, Pareto-ranked experimental split structures enable substantial sender gains. Crucially, these gains are robust to perturbations in the receiver’s degree of ambiguity aversion. Our findings provide novel theoretical foundations and design principles for real-world ambiguous decision-making contexts—such as bank stress testing—where ambiguity is pervasive and agents exhibit systematic ambiguity aversion.

3 citations1 influentialRead paper

Addressing Overthinking in Large Vision-Language Models via Gated Perception-Reasoning Optimization

Jan 07, 2026arXiv.org

This work addresses the inefficiency and accuracy degradation of large vision-language models (VLMs) on simple tasks, where over-reasoning often leads to unnecessarily verbose responses. While prior approaches overlook visual perception failure as a fundamental bottleneck, this paper proposes GPRO, a novel framework that decouples perception failures from reasoning errors for the first time. GPRO constructs supervision signals based on failure attribution and introduces a meta-reasoning controller that dynamically selects among a lightweight fast path, a slow perception path, or a slow reasoning path. Leveraging a teacher model to generate approximately 790,000 failure-attribution labels, the path selection strategy is optimized via multi-objective reinforcement learning. Experiments demonstrate that GPRO significantly improves both accuracy and inference efficiency across five benchmarks, outperforming existing "slow thinking" methods while producing more concise responses.

3 citationsRead paper
Recent publications

Latest Papers