Institution profile

University of California, San Francisco

Academic institutionnorthamerica · us
Official website
Research library230linked papers
Opportunities0open roles
Selected work

Representative Papers

Human-AI Co-design for Clinical Prediction Models

Jan 14, 2026

This work addresses the limitations of traditional clinical prediction models, which rely heavily on expert involvement and struggle to effectively leverage unstructured clinical text, hindering real-world deployment. To overcome this, the authors propose HACHI, a novel framework that deeply integrates human feedback with AI agent exploration. Through iterative human–AI collaboration, HACHI extracts interpretable clinical concepts from free-text notes using simple yes/no queries and leverages expert feedback to construct transparent, verifiable linear prediction models. Evaluated on acute kidney injury and traumatic brain injury prediction tasks, HACHI not only outperforms existing methods but also discovers novel clinically relevant concepts, substantially improving generalizability across institutions and time. Furthermore, the framework effectively identifies data biases and leakage issues, enhancing model reliability and trustworthiness.

1 citationsRead paper

ScaleMAI: Accelerating the Development of Trusted Datasets and AI Models

Jan 06, 2025

Medical AI development is hindered by lengthy dataset curation cycles and the decoupling of annotation from model training. To address this, we propose an AI-driven collaborative co-evolution framework for medical data, instantiated on pancreatic tumor CT analysis. Our approach introduces a novel human-in-the-loop, progressive “data flywheel” mechanism that jointly enhances annotation quality and model performance. Methodologically, it integrates multi-round human-in-the-loop iteration, 3D voxel-level semi-automatic annotation, domain-adaptive few-shot learning, and cross-task joint modeling (detection, segmentation, classification). We construct a high-quality, multi-task dataset comprising 25,362 CT scans. Our flagship model achieves annotation accuracy comparable to that of experts with 30 years of experience, delivering performance gains of 14%, 5%, and 72% over prior state-of-the-art on detection, segmentation, and classification benchmarks, respectively. This work transcends static dataset paradigms, enabling dynamic, scalable, and trustworthy medical AI infrastructure.

1 citationsRead paper

Generative causal testing to bridge data-driven models and scientific theories in language neuroscience

Oct 01, 2024

This study investigates the stimulus features driving language selectivity across brain regions in LLM-predicted BOLD fMRI responses, aiming to generate testable neuroscientific explanations. To this end, we propose the Generative Causal Testing (GCT) framework—first repurposing LLMs from predictive tools to hypothesis generators—by leveraging controllable text generation to formulate formal, falsifiable hypotheses about neural language selectivity, followed by closed-loop fMRI validation. Our approach integrates fMRI encoding modeling, causal intervention design, controllable LLM generation, and interpretable neural representational analysis. Key contributions include: (1) high-accuracy, empirically verifiable causal explanations at both voxel- and ROI-levels; (2) discovery of fine-grained functional subdivisions within prefrontal cortex and their precise linguistic selectivities; and (3) empirical confirmation that explanation accuracy strongly correlates with model performance and stability, thereby establishing a bidirectional bridge between data-driven modeling and formal neurocognitive theory.

1 citationsRead paper

LTSM-Bundle: A Toolbox and Benchmark on Large Language Models for Time Series Forecasting

Jun 20, 2024

To address modeling and evaluation challenges posed by heterogeneous data—characterized by multi-frequency sampling, high dimensionality, and multimodality—in large time-series models (LTSMs), this paper introduces the first unified toolbox and benchmark platform for time-series forecasting. Methodologically, it achieves full-stack decoupling and co-evaluation across preprocessing, tokenization, prompt learning, training paradigms, and data diversity; it further proposes a Transformer-based autoregressive architecture, multi-granularity tokenization, instruction-style prompting, and cross-frequency/dimension adaptation techniques. The core contribution lies in systematically uncovering strong synergistic effects among design choices and identifying an optimal configuration, which significantly improves zero-shot and few-shot generalization performance across multiple standard benchmarks—outperforming both state-of-the-art LTSMs and conventional time-series models.

1 citationsRead paper
Recent publications

Latest Papers