Institution profile

University of Michigan

Academic institutionnorthamerica · us
Official website
Research library2,266linked papers
Opportunities0open roles
Selected work

Representative Papers

Acquiring Grounded Representations of Words with Situated Interactive Instruction

Feb 28, 2025

This study addresses the challenge of jointly acquiring perceptual, semantic, and procedural knowledge in embodied lexical learning. We propose a human-robot hybrid active interaction framework enabling a robot to dynamically identify unknown concepts in real-world desktop environments, autonomously initiate contextualized learning requests to human teachers, and unify multimodal knowledge acquisition within a “perception–comprehension–planning–execution” semantic closed loop. The system is built upon the Soar cognitive architecture and integrates visual perception, natural language understanding, task planning, and robotic arm control. Experiments demonstrate significant improvements in rapid generalization to novel lexical items and instruction execution accuracy, along with zero-shot concept transfer capability. Our core contributions are (1) the first hybrid active interaction mechanism for embodied lexical learning, and (2) a unified representational paradigm for heterogeneous knowledge types—perceptual, semantic, and procedural—within a single cognitive architecture.

58 citations1 influentialRead paper

Preconditioned Gradient Descent for Over-Parameterized Nonconvex Matrix Factorization

Apr 13, 2025Neural Information Processing Systems

In overparameterized nonconvex matrix factorization—where the specified rank $r$ exceeds the true rank $r^*$—gradient descent suffers sublinear convergence, severely limiting efficiency. This paper proposes PrecGD, a lightweight preconditioned gradient descent method that restores linear convergence without requiring prior knowledge of $r^*$. Key contributions include: (i) the first theoretical demonstration that $ell_2$ regularization, within a specific damping range, effectively mitigates ill-conditioning of the factor matrices; and (ii) a novel adaptive damping strategy, computed cheaply from current iterates, which robustly handles the conditioning of the ground-truth solution. PrecGD maintains linear convergence even under noise and achieves the information-theoretically optimal estimation error bound. Experiments across diverse overparameterized matrix sensing and factorization tasks confirm substantial improvements in both convergence speed and reconstruction accuracy.

31 citations5 influentialRead paper

Selection and Parallel Trends

Mar 17, 2022Social Science Research Network

This paper addresses how treatment-group selection threatens the parallel trends assumption in Difference-in-Differences (DiD) estimation—a critical yet under-characterized identification challenge. Method: We formally characterize the empirical content of this threat and derive necessary and sufficient conditions for parallel trends to hold under general selection mechanisms. We propose a “selection-driven bias decomposition framework” that systematically partitions DiD estimation bias into selection effects and time-varying heterogeneity effects, and develop operational benchmarking strategies—both with and without covariates—grounded in causal inference theory, selection modeling, and sensitivity analysis. Contribution/Results: Applied to the National Supported Work (NSW) experiment reanalysis, our approach quantifies and corrects selection bias, substantially improving the credibility of DiD estimates and the robustness of causal conclusions.

27 citations4 influentialRead paper

View Selection for 3D Captioning via Diffusion Ranking

Apr 11, 2024European Conference on Computer Vision

To address vision-language hallucinations in 3D object captioning caused by rendering viewpoint mismatch, this paper proposes an unsupervised view-ranking method based on diffusion models. Specifically, we leverage pre-trained text-to-3D models (e.g., Stable Diffusion 3D) to quantify 3D–2D view alignment scores and select the most discriminative 2D views for input to multimodal large language models (e.g., GPT-4V) to generate accurate captions. This approach extends the view-ranking paradigm to 3D visual question answering (3D-VQA), achieving significant improvements over CLIP-based baselines on Objaverse/XL. Furthermore, we correct 200K erroneous captions in Cap3D and construct the first million-scale, high-quality 3D-caption dataset—Cap3D-v2—establishing a robust benchmark for 3D understanding and generation.

21 citationsRead paper

Understanding Deep Representation Learning via Layerwise Feature Compression and Discrimination

Nov 06, 2023arXiv.org

This work investigates the fundamental mechanisms underlying hierarchical representation learning in deep neural networks. Addressing the central question—“how do features evolve across layers”—we propose a joint quantification framework for inter-layer feature compression ratio and discriminability. We theoretically uncover, for the first time, a geometric–linear dual-rate pattern of feature evolution in deep linear networks: intra-class features contract geometrically, while inter-class discriminability increases linearly. This pattern is rigorously established under minimal norm, weight balancing, and near-low-rank assumptions, and extended to nonlinear networks via intermediate-feature modeling for multi-class classification. Numerical experiments validate its robustness across architectures and datasets. Our results provide an interpretable theoretical foundation for representation learning and yield quantitative guidance for layer selection in transfer learning and knowledge distillation.

18 citations2 influentialRead paper
Recent publications

Latest Papers