Institution profile

King Abdullah University of Science and Technology

Academic institutionasia · sa
Official website
Research library876linked papers
Opportunities0open roles
Selected work

Representative Papers

A Modern Introduction to Online Learning

Dec 31, 2019arXiv.org

This work addresses worst-case online optimization, unifying the study of online convex and non-convex optimization over Euclidean and non-Euclidean domains—including simplices and matrix manifolds—under a regret minimization framework. We propose a parameter-free, adaptive algorithmic framework that supports unbounded decision sets and unknown gradient magnitudes. Unifying online mirror descent (OMD) and follow-the-regularized-leader (FTRL), we reformulate first- and second-order methods and, for the first time, integrate convex surrogate losses, randomization schemes, and multi-armed bandit feedback—both adversarial and stochastic—into this coherent paradigm. Our theoretical analysis is self-contained, elementary, and accessible without prerequisites; all algorithms achieve tight, optimal regret bounds. The resulting framework establishes a universal, concise, and pedagogically transparent foundation for modern online learning, substantially lowering both theoretical barriers and practical implementation complexity.

372 citations115 influentialRead paper

Annotated History of Modern AI and Deep Learning

Dec 21, 2022arXiv.org

This paper addresses the fragmentation of AI historiography, the obscuration of neural networks’ intellectual origins, and the marginalization of cybernetics in mainstream narratives. It proposes a unified historical reconstruction framework centered on the concept of “credit assignment.” Through rigorous historical document analysis, interdisciplinary knowledge graph construction, and scholarly provenance tracing, the study systematically traces the mathematical and technical lineage—from the 17th-century chain rule and 19th-century linear regression to the first implementation of deep learning in 1965—thereby correcting widespread textbook misconceptions and reaffirming cybernetics’ foundational role in modern AI. The resulting contribution is the most comprehensive chronology of deep learning to date (as of 2022), documenting over one hundred pivotal events, rigorously attributing original contributions, and embedding hundreds of hyperlinked authoritative sources. This chronology has been published as a core chapter in an academic monograph on AI.

43 citationsRead paper

Deeper or Wider: A Perspective from Optimal Generalization Error with Sobolev Loss

Jan 31, 2024International Conference on Machine Learning

This work investigates the optimal generalization error trade-off between deep neural networks (DeNNs) and wide neural networks (WeNNs) under Sobolev-norm losses. Addressing the “depth vs. width” architectural selection problem, we establish, for the first time, a theoretically grounded criterion based on Sobolev regularity: wide networks dominate under high parameter budgets, whereas deep architectures excel with large sample sizes and higher-order Sobolev loss regularization. Methodologically, we integrate Sobolev-space generalization error bounds with the Deep Ritz method and the physics-informed neural network (PINN) framework to derive interpretable, theory-driven design principles. These principles are empirically validated on PDE-solving tasks using both Deep Ritz and PINN approaches. Our results provide the first generalization-error-theoretic foundation for depth–width selection in PDE numerical solvers, bridging theoretical learning guarantees with practical neural PDE discretization.

12 citations1 influentialRead paper

Mamba-FSCIL: Dynamic Adaptation with Selective State Space Model for Few-Shot Class-Incremental Learning

Jul 08, 2024arXiv.org

Few-shot class-incremental learning (FSCIL) confronts a dual challenge: static architectures suffer from overfitting to base classes, while dynamic architectures incur excessive parameter growth. To address this, we propose Dual-SSM—a dual selective state space model framework. It introduces a class-sensitive selective scanning mechanism to decouple feature evolution between base and novel classes, and integrates sequence modeling–driven incremental feature alignment with dynamic weight projection to enable parameter-adaptive expansion and efficient knowledge consolidation. Evaluated on miniImageNet, CUB-200, and CIFAR-100, Dual-SSM consistently surpasses existing state-of-the-art methods. It significantly mitigates catastrophic forgetting and enhances few-shot generalization for novel classes. By jointly achieving parameter efficiency and architectural scalability, Dual-SSM establishes a new paradigm for FSCIL that balances lightweight design with extensibility.

6 citationsRead paper

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation

May 29, 2025

To address the challenges of cross-domain transfer in multi-agent systems—namely, the need for redesign and retraining—we propose Workforce, a hierarchical multi-agent framework that decouples domain-agnostic planning (Planner-Coordinator) from domain-specific execution (modular, plug-and-play Workers), enabling zero-shot cross-domain adaptation. We introduce OWL (Online Weighted Learning), a novel reinforcement learning method that enables the planner to achieve domain-invariant generalization through optimization guided by real-world feedback. Workforce integrates tool invocation, modular Worker architecture, and hierarchical coordination mechanisms. On the GAIA benchmark, Workforce achieves state-of-the-art open-source performance (69.70%). A 32B model trained with OWL attains 52.73% accuracy—16.37 percentage points higher than the baseline—and approaches the performance of GPT-4o.

4 citationsRead paper
Recent publications

Latest Papers