Institution profile

Massachusetts Institute of Technology

Academic institutionnorthamerica · us
Official website
Research library4,012linked papers
Opportunities0open roles
Selected work

Representative Papers

Time-based Fairness Improves Performance in Multi-Rate WLANs

Jun 27, 2004USENIX ATC, General Track

This study addresses the significant degradation in system throughput caused by conventional throughput-based fairness mechanisms in multi-rate WLANs, where heterogeneous channel conditions lead to inefficient resource allocation. To resolve this issue, the paper introduces the concept of time fairness and proposes TBR (Time-based Regulator), a novel algorithm implemented at the access point that enforces fair scheduling by regulating each node’s channel occupancy time. Operating within the constraints of the existing DCF MAC protocol, TBR achieves a balanced trade-off between fairness and aggregate throughput. Experimental results demonstrate that the proposed mechanism effectively enhances overall network throughput while preserving fundamental per-node access performance in single-rate scenarios.

273 citations37 influentialRead paper

On-line Policy Improvement using Monte-Carlo Search

Dec 03, 1996Neural Information Processing Systems

This paper addresses the low decision quality and high error rates of controllers in real-time adaptive control. We propose an online Monte Carlo Policy Improvement (MCPI) algorithm that requires neither an environmental model nor gradient information, relying solely on a simulatable environment and an initial policy. MCPI estimates long-term action returns via parallel multi-step random rollouts and dynamically updates the policy. Its key innovation lies in directly applying a lightweight, scalable Monte Carlo Tree Search (MCTS) for online policy optimization, enabling plug-and-play reinforcement learning enhancement. Evaluated on backgammon, MCPI reduces decision error rates by over fivefold compared to baselines—including random policies and TD-Gammon—demonstrating strong generalization capability and real-time efficacy in practical adaptive control scenarios.

268 citations20 influentialRead paper

Explaining the Success of Nearest Neighbor Methods in Prediction

May 31, 2018Found. Trends Mach. Learn.

Despite widespread empirical success, the theoretical foundations and practical deployment guidelines for k-nearest neighbors (k-NN) in predictive tasks remain inadequately understood. Method: We establish the first non-asymptotic error bound framework tailored for real-world deployment, replacing conventional smoothness or margin assumptions with verifiable cluster structure as the key success criterion. We integrate approximate nearest neighbor techniques—including LSH and graph-based indexing—and unify k-NN theory with emerging paradigms such as random forests, graphon modeling, and crowdsourcing. We further introduce a novel distance-learning perspective, characterizing how ensemble methods implicitly learn neighborhood structure. Contribution/Results: Evaluated on time-series forecasting, recommender systems, and medical image segmentation, our framework demonstrates that high accuracy is guaranteed solely under cluster-structured data—enhancing both theoretical interpretability and engineering tractability. It provides actionable, error-tolerance-driven guidance for data volume and hyperparameter selection.

144 citations10 influentialRead paper

Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora

Apr 10, 2025Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning

Large language models (LLMs) suffer from low data efficiency, typically requiring trillion-word corpora for effective pretraining. Method: Inspired by child language acquisition, this work proposes a cognitively grounded, highly efficient pretraining paradigm using only a developmentally appropriate corpus of under 100 million tokens. We systematically demonstrate—contrary to prevailing assumptions—that such small-scale data can surpass trillion-parameter models’ performance when combined with short-sequence training, knowledge distillation, and multi-task evaluation (covering syntactic competence, downstream task transfer, and out-of-distribution generalization); notably, curriculum learning proves ineffective in this low-data regime. Contribution/Results: Leveraging the LTG-BERT architecture, our best-performing model achieves state-of-the-art results across diverse benchmarks, significantly outperforming standard large baselines. The project yields over 30 empirically validated guidelines—identifying both viable strategies and dead ends—for efficient pretraining, thereby establishing a novel paradigm for cognitive modeling and environmentally sustainable (“green”) AI.

105 citations18 influentialRead paper

Medical Hallucinations in Foundation Models and Their Impact on Healthcare

Feb 26, 2025arXiv.org

Medical foundation models may generate “hallucinations”—factual, logical, or evidence-inconsistent errors—that jeopardize clinical decision-making and patient safety. To address this, we first propose a multidimensional taxonomy of medical hallucinations and establish a real-world, clinician-annotated benchmark dataset derived from authentic clinical cases; we further validate its clinical impact via an international physician survey. Methodologically, we integrate expert annotation, empirical behavioral surveys, and large language model (LLM) evaluation to systematically assess the efficacy of chain-of-thought (CoT) reasoning and retrieval-augmented generation (RAG) in mitigating hallucinations. Results show both techniques significantly reduce hallucination rates, yet residual hallucinations remain clinically hazardous. Building on these findings, we introduce a patient-safety-centered AI governance and ethics framework, offering theoretical foundations and actionable pathways for responsible deployment of medical AI. (149 words)

41 citationsRead paper
Recent publications

Latest Papers