Institution profile

Indiana University

Academic institutionnorthamerica · us
Official website
Research library385linked papers
Opportunities0open roles
Selected work

Representative Papers

Analysis of Distributional Dynamics for Repeated Cross-Sectional and Intra-Period Observations

May 21, 2025

This paper addresses the challenge of modeling the dynamic evolution of state density functions in two distinct data structures: repeated cross-sections (e.g., monthly stock return distributions) and high-frequency intra-period time series (e.g., intraday GBP/USD return distributions). We propose the first unified functional dynamic framework compatible with both. Methodologically, we embed density functions into a Hilbert space and formulate a functional autoregressive (FAR) model, integrating kernel density estimation with asymptotic statistical inference. Theoretically, we establish asymptotic theory for density forecasting and distributional moment dynamics, and prove strong consistency of the estimators. Empirically, our approach significantly outperforms conventional methods on GBP/USD and NYSE datasets, delivering high-accuracy density forecasts. The core innovation lies in overcoming data-type barriers—enabling, for the first time, unified modeling and theoretical analysis of distributional evolution across both inter-period cross-sectional and intra-period temporal dimensions.

4 citationsRead paper

Learning in Random Utility Models Via Online Decision Problems

Dec 21, 2021Social Science Research Network

This paper studies how decision-makers learn online in repeated stochastic choice settings without knowledge of true option utilities, under the Random Utility Model (RUM). We embed RUM into an online decision-making framework and design gradient-based learning algorithms. We establish, for the first time, that multi-class RUM satisfies Hannan consistency. Moreover, we prove a rigorous equivalence between RUM-based learning and Follow-the-Regularized-Leader (FTRL), thereby providing a microeconomic foundation for FTRL. The framework is extended to model recency bias, no-regret learning in games, and prediction market mechanism design. Theoretically, it guarantees convergence of long-run average payoff to that of the optimal fixed strategy and satisfies no-regret properties. Empirically, the approach significantly improves behavioral prediction accuracy and mechanistic interpretability across three canonical economic domains.

2 citations1 influentialRead paper

Emergence of Structural Disparities in the Web of Scientific Citations

Jan 19, 2026

This study addresses the persistent structural inequities in the allocation of scholarly attention within scientific citation networks, which are shaped by gender and institutional prestige. The authors propose a generative model of citation network growth that integrates homophily-driven preferences, preferential attachment, and group size effects to systematically uncover the mechanisms underlying these disparities. Their analysis reveals that merely increasing the representation of underrepresented groups is insufficient to redress imbalance; effective mitigation requires simultaneously reducing homophilic citation biases and enhancing the visibility of their scholarly contributions. By quantifying the distinct impacts of gender and institutional prestige on citation distributions, the work advances a multidimensional intervention framework that offers both theoretical grounding and actionable pathways toward fostering a more equitable, transparent, and inclusive scientific communication ecosystem.

1 citationsRead paper

Generative diffusion model surrogates for mechanistic agent-based biological models

May 01, 2025

Cellular-Potts model (CPM)-based in vitro angiogenesis simulation suffers from inherent stochasticity and massive spatiotemporal computational demands, hindering effective surrogate modeling. To address this, we propose the first generative surrogate model for CPM grounded in denoising diffusion probabilistic models (DDPMs). Our method treats CPM’s stochastic outputs as image distributions, enabling DDPMs to learn high-dimensional configuration dynamics; a 2D image classifier further guides parameter-space partitioning and surrogate validation. Experiments demonstrate that the surrogate generates 20,000-step, single-cell-resolution spatial configurations autoregressively—achieving ~22× speedup in inference time versus full CPM simulation. This work extends the applicability of diffusion models to stochastic multicellular systems and establishes a verifiable, high-fidelity, and computationally efficient simulation paradigm for digital twins of biological systems.

1 citationsRead paper

Uncovering the universal dynamics of citation systems: From science of science to law of law and patterns of patents

Jan 26, 2025

Existing citation dynamics models are domain-specific and mechanistically fragmented, failing to explain cross-domain phenomena such as “delayed recognition” and the ubiquity of “sleeping beauties.” Method: We propose the first unified model integrating three fundamental mechanisms—cumulative advantage, temporal decay, and structural embedding—and validate it via multi-source citation network mining, temporal modeling, and cross-domain comparative analysis across scientific publications, legal cases, and patents. Contribution/Results: We empirically establish the cross-domain universality of skewed citation distributions—including sleeping beauties—across all three knowledge systems. Our model significantly outperforms state-of-the-art baselines in reproducing empirical citation evolution and achieves 12–28% higher accuracy in predicting high-impact nodes. This work reveals universal principles governing citation dynamics beyond disciplinary boundaries and provides a generalizable framework for modeling knowledge diffusion and impact accumulation.

1 citationsRead paper
Recent publications

Latest Papers

Locating and Controlling Implicit Personalization in Large Language Models

Aug 12, 2026

This study addresses the phenomenon wherein large language models exhibit output biases in response to implicit demographic cues when users do not explicitly declare their identities—a mechanism that remains poorly understood. The work establishes, for the first time, a strong association between such implicit personalization behaviors and localized internal activation signals within the models. Through a systematic empirical investigation across five mainstream large language models—employing controlled experiments, activation analysis, signal ablation, and multi-cue combination tests—the authors demonstrate that this activation signal is highly correlated with biased recommendations (r = 0.87). Ablating the signal effectively mitigates cue-induced bias while preserving general task performance, achieving causal control that surpasses conventional prompt-based interventions.

0 citationsRead paper

GenFAR: A generalized representation of brain structure, derived from 49,246 multi-cohort MRIs via deep learning

Aug 12, 2026

This study addresses the limitation of existing deep learning models in neuroimaging, which are typically confined to single tasks and struggle with cross-task knowledge transfer. To overcome this, the authors propose GenFAR, a modular deep learning framework trained jointly on 17 cognitive, clinical, and diagnostic tasks using brain MRI data from 49,246 individuals across 11 cohorts. The framework leverages an innovative task-sequencing mechanism and a novel Donor Score metric to identify five pivotal source tasks that substantially enhance sample efficiency and performance on downstream tasks. The resulting general-purpose brain representations demonstrate strong clinical relevance, significantly improving model accuracy on unseen tasks while markedly reducing the required training sample size.

0 citationsRead paper

Algorithm Transparency and Search Manipulation: Steering vs. Persuasion

Aug 12, 2026

This study addresses how digital platforms simultaneously steer user attention and convey match quality through recommendation algorithms. By constructing a Bayesian persuasion model, the paper disentangles—for the first time—the algorithm’s dual roles of “steering” and “informing,” integrating insights from information design and consumer search theory. It examines how profit-maximizing platforms choose ranking strategies and algorithmic transparency levels, and how these choices affect consumers’ understanding, search behavior, and welfare. The analysis reveals a dual effect of transparency: moderate transparency enhances welfare by improving information transmission, yet excessive or mandated transparency disrupts the synergy between steering and signaling, ultimately reducing welfare below that under full opacity.

0 citationsRead paper

Empirical Simulation of Survival and Mixed-Type Data for Clinical Trial Design

Aug 12, 2026

Traditional parametric approaches often fail to accurately model complex time-to-event data in clinical trials due to their reliance on prespecified hazard function forms. This work proposes a joint simulation framework based on empirical copulas that integrates nonparametric reconstruction with parametric tail modeling. The method handles right censoring via conditional Kaplan–Meier imputation, models marginal distributions through log-scale location–scale transformations and power distortion of quantile functions, and preserves multivariate rank correlation structures using a Gaussian copula. Remarkably, the approach can reproduce treatment-group survival curves using only a few target quantiles. Applied to a non-small cell lung cancer trial, it successfully replicated both overall survival and progression-free survival curves, yielding a simulated censored Kendall’s tau of 0.522—closely approximating the observed value of 0.549.

0 citationsRead paper

An Agentic Generative Large Language Model for Treatment Planning of Colorectal Cancer

Aug 10, 2026

This study addresses the challenges of heterogeneous data integration, dynamic adherence to clinical guidelines, and safe decision-making in precision treatment for colorectal cancer by proposing GatorOnco—the first large language model that integrates agent-based reasoning with large-scale domain adaptation. The approach leverages domain-adaptive pretraining, model fusion, two-stage post-training, agent-based reinforcement learning, and retrieval-augmented generation (RAG) to enable dynamic incorporation of clinical guidelines and generate safe, controllable treatment plans. In blinded evaluations, GatorOnco significantly outperformed existing open-source large language models (P<0.01), surpassing human experts in readability and completeness while matching oncologists in correctness, timeliness, and safety.

0 citationsRead paper