Institution profile

University of Colorado

Academic institutionnorthamerica · us
Official website
Research library135linked papers
Opportunities0open roles
Selected work

Representative Papers

Generative Modeling by Minimizing the Wasserstein-2 Loss

Jun 19, 2024arXiv.org

This work addresses the challenge of minimizing the Wasserstein-2 (W₂) distance in unsupervised generative modeling. We propose the first explicit, distribution-dependent ordinary differential equation (ODE) characterizing the W₂ gradient flow, and theoretically prove that its time-marginal distributions converge rigorously to the target data distribution. Methodologically, we design a persistent Euler discretization algorithm that couples with the gradient flow structure, circumventing the instability inherent in conventional adversarial training. Our key contributions are: (1) the first explicit ODE formulation of the W₂ gradient flow; and (2) a novel persistence training mechanism that enhances discretization fidelity and convergence robustness. Experiments demonstrate that our approach significantly outperforms WGAN on both high- and low-dimensional benchmarks; moreover, increasing persistence strength further improves generation quality and training stability.

4 citationsRead paper

Decoupling the "What" and "Where" With Polar Coordinate Positional Embeddings

Sep 05, 2025arXiv.org

In Transformer architectures, content (“what”) and position (“where”) representations are deeply entangled in mainstream positional encodings such as RoPE, inducing modeling bias—particularly degrading zero-shot length extrapolation. This work first identifies and formalizes the “what–where” coupling mechanism inherent in RoPE. To address it, we propose Polar Coordinate Positional Encoding (PoPE): it explicitly decouples content and position at the geometric level by encoding relative position as angular coordinates and content-dependent modulation as radial coordinates. PoPE is parameter-free, plug-and-play, and fully compatible with standard Transformers. Experiments across music, genomic, and language modeling tasks demonstrate consistent perplexity reduction across model scales (124M–774M parameters). Crucially, PoPE significantly improves zero-shot length extrapolation—enabling coherent generation far beyond training sequence lengths—without interpolation or fine-tuning.

1 citations1 influentialRead paper

Exploring Interdisciplinary Team Collaboration in Clinical NLP Projects Through the Lens of Activity Theory

Sep 30, 2024arXiv.org

Interdisciplinary collaboration between clinicians and AI researchers—particularly speech-language pathologists (SLPs) and natural language processing (NLP) scientists—is frequently undermined by blurred disciplinary boundaries, fragmented terminologies, and divergent interpretations of clinical data; existing literature lacks systematic analysis of underlying mechanisms and actionable mitigation strategies. Method: This study pioneers the application of activity theory to examine SLP–NLP collaboration, employing semi-structured interviews and thematic analysis across multiple clinical NLP projects. Contribution/Results: We identify three core barriers: (1) professional discourse conflict, (2) data interpretation tension, and (3) absence of effective knowledge mediation. We empirically validate clinical data’s dual role as a “boundary object”—facilitating yet also complicating cross-domain coordination. Building on this, we propose the novel paradigm of “AI as knowledge mediator” and design an AI-driven knowledge-brokering framework, offering a transferable, theory-grounded collaboration model and practical guidelines for clinical NLP initiatives.

1 citationsRead paper

Scientific productivity as a random walk

Sep 08, 2023arXiv.org

While the scientific community widely assumes a canonical “rise-then-decline” productivity trajectory across scholars, empirical analysis reveals that only ~20% of individuals conform to this pattern—highlighting an apparent tension between aggregate trends and individual heterogeneity. Method: We propose a parametric stochastic walk model with time-varying variance, grounded in longitudinal data of 29,119 papers published by 2,085 computer science professors from 205 universities (1980–2016). We combine variance decomposition, trajectory clustering, and model validation to examine interannual output dynamics. Contribution/Results: We demonstrate that the canonical curve is a statistical emergent phenomenon arising from attenuation of early-career output variance—not a shared developmental pattern. Individual annual publication counts follow a parsimonious statistical law; our model not only accurately reproduces the population-level average trajectory but also captures the true trajectory morphology for ~80% of individuals, challenging prevailing linear and stage-based career development paradigms.

1 citationsRead paper
Recent publications

Latest Papers

Robots influencing humans to reveal their goals during collaboration and competition

Aug 31, 2026Autonomous Robots

We propose a unified strategy for fast goal inference in human–robot interaction. The core idea is to drive the human toward Critical Decision Points (CDPs)–states where competing human strategies prescribe different next actions and thus maximally reveal the goal. We formalise CDPs using a goal-conditioned policy divergence measure and incorporate them into a Receding-Horizon Planner that explores future action sequences while optimizing a cost function balancing task progress and information gain. We evaluate this approach in both a collaborative, fully observable cooking task and a competitive, partially observable hide-and-seek game, each in simulation and on real robots. In both scenarios, our method infers human goals more accurately and earlier than baseline strategies.

0 citationsRead paper