Institution profile

New York University

Academic institutionnorthamerica · us
Official website
Research library2,301linked papers
Opportunities0open roles
Selected work

Representative Papers

Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora

Apr 10, 2025Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning

Large language models (LLMs) suffer from low data efficiency, typically requiring trillion-word corpora for effective pretraining. Method: Inspired by child language acquisition, this work proposes a cognitively grounded, highly efficient pretraining paradigm using only a developmentally appropriate corpus of under 100 million tokens. We systematically demonstrate—contrary to prevailing assumptions—that such small-scale data can surpass trillion-parameter models’ performance when combined with short-sequence training, knowledge distillation, and multi-task evaluation (covering syntactic competence, downstream task transfer, and out-of-distribution generalization); notably, curriculum learning proves ineffective in this low-data regime. Contribution/Results: Leveraging the LTG-BERT architecture, our best-performing model achieves state-of-the-art results across diverse benchmarks, significantly outperforming standard large baselines. The project yields over 30 empirically validated guidelines—identifying both viable strategies and dead ends—for efficient pretraining, thereby establishing a novel paradigm for cognitive modeling and environmentally sustainable (“green”) AI.

105 citations18 influentialRead paper

Revisiting the Last-Iterate Convergence of Stochastic Gradient Methods

Dec 13, 2023arXiv.org

This work addresses three fundamental gaps in the theoretical understanding of stochastic gradient descent (SGD): (1) absence of optimal-rate convergence guarantees under non-compact domains and bounded-noise assumptions; (2) scarcity of final-iterate analyses for smooth optimization; and (3) lack of a unified framework for composite objectives, non-Euclidean geometries, and heavy-tailed noise. We propose the first unified analysis framework accommodating non-compact domains, composite regularization, Bregman geometry, heavy-tailed stochastic noise, and general convexity/smoothness conditions. Our approach integrates generalized Bregman divergences, adaptive step sizes, and heavy-tailed robust estimation techniques. We establish optimal $Oig(sqrt{log(1/delta)/T}ig)$ convergence rates both in expectation and with high probability $1-delta$, thereby substantially broadening the theoretical applicability of SGD.

24 citations10 influentialRead paper

Causal Panel Analysis under Parallel Trends: Lessons from A Large Reanalysis Study

Sep 27, 2023

The two-way fixed effects (TWFE) estimator is widely used for causal inference in political science but suffers from sensitivity to heterogeneous treatment effects (HTE) and the parallel trends (PT) assumption, undermining result reliability. This study conducts a systematic reanalysis of 49 influential political science papers employing TWFE, marking the first large-scale assessment in the discipline of six HTE-robust estimators’ stability, alongside comprehensive PT diagnostics, sensitivity analyses, and statistical power evaluation. Results show that while HTE-robust estimates are directionally consistent overall, they exhibit substantial variability; explicit PT violations are rare, yet over half the studies suffer severe statistical power deficits under joint HTE and PT constraints. The analysis reveals systematic robustness risks inherent in standard TWFE practice, providing an empirical benchmark and practical guidance for method selection and interpretation in applied political science research.

16 citations1 influentialRead paper

3D U-KAN Implementation for Multi-modal MRI Brain Tumor Segmentation

Aug 01, 2024arXiv.org

To address the challenges of high intra-tumoral heterogeneity in gliomas and the difficulty of modeling 3D multimodal MRI data, this paper proposes UKAN-SE—the first U-Net variant adapted for 3D medical image segmentation by integrating the Kolmogorov–Arnold Network (KAN). It innovatively incorporates a Squeeze-and-Excitation (SE) module to model channel-wise global attention. The architecture employs 3D convolutions and multimodal feature fusion. Evaluated on BraTS 2024, UKAN-SE significantly outperforms baseline models including U-Net, Attention U-Net, and Swin UNETR in segmentation accuracy. With only 10.6 million parameters, it achieves superior computational efficiency: training time is reduced to one-quarter that of U-Net and one-sixth that of Swin UNETR. Thus, UKAN-SE delivers both state-of-the-art performance and exceptional parameter- and time-efficiency for 3D glioma segmentation.

11 citationsRead paper

The Impact of Large Language Models on Open-source Innovation: Evidence from GitHub Copilot

Sep 12, 2024International Conference on Interaction Sciences

This study investigates the differential impact of large language models (LLMs) on capability innovation (i.e., exploring novel functionalities) versus iterative innovation (i.e., optimizing and maintaining existing code) in open-source collaborative development. Leveraging GitHub Copilot’s phased, language-specific rollout as a quasi-natural experiment, we employ a multi-period difference-in-differences design, using variation in programming language support as an exogenous shock to identify causal LLM effects. Our key contribution is the first empirical evidence that—under unguided, spontaneous collaboration—LLMs significantly boost iterative innovation while exerting limited influence on capability innovation. This effect strengthens with model upgrades (e.g., the 2022 release) and higher project activity, and is especially pronounced in Python and Rust projects. The findings indicate that current LLMs are better aligned with maintenance-oriented development than with exploratory feature creation, offering critical empirical insights into the pathways and boundaries of AI-augmented open-source innovation.

9 citationsRead paper
Recent publications

Latest Papers