Institution profile

MIM Solutions

Industry researchnorthamerica · us
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Decoupled Relative Learning Rate Schedules

Jul 04, 2025

Conventional Transformer training employs a uniform learning rate across all model components, ignoring their dynamic heterogeneity in parameter sensitivity and update magnitude—leading to suboptimal optimization efficiency. Method: We propose a dynamic decoupled learning rate scheduling framework, introducing— for the first time—the concept of *relative learning rates*, which adaptively scale per-component learning rates based on layer-specific gradient statistics and architectural roles. Our approach is architecture-agnostic within the Transformer family and integrates seamlessly with Mixture of Experts (MoE) configurations. Contribution/Results: The method enables direct hyperparameter transfer across model scales—from small baselines to models 27× larger—without manual retuning. Empirical evaluation demonstrates up to 23% faster convergence for complex models and substantial reductions in computational resource consumption. This work establishes a scalable, efficient, and broadly generalizable optimization paradigm for large-scale neural networks.

0 citationsRead paper

VARSHAP: Addressing Global Dependency Problems in Explainable AI with Variance-Based Local Feature Attribution

Jun 08, 2025

Existing feature attribution methods (e.g., KernelSHAP, LIME) rely on global data distributions, leading to inaccurate characterization of local model behavior and distorted explanations. To address this, we propose VARSHAP—a model-agnostic local feature attribution method that introduces prediction variance reduction as the core Shapley value metric, the first such formulation. VARSHAP rigorously satisfies the efficiency, symmetry, and additivity axioms of Shapley values. It estimates conditional variances via Monte Carlo sampling, eliminating the need for surrogate models or distributional assumptions, and inherently exhibits robustness to data distribution shifts. Experiments on synthetic and real-world datasets demonstrate that VARSHAP improves attribution accuracy by 12–23% over KernelSHAP and LIME. Qualitative evaluations confirm its superior alignment with local decision logic, significantly mitigating the local explanation bias induced by global distribution dependence.

0 citationsRead paper

Since Faithfulness Fails: The Performance Limits of Neural Causal Discovery

Feb 22, 2025

Neural causal discovery methods face fundamental limitations: they struggle to reliably distinguish true from spurious causal edges under finite samples, and the faithfulness assumption—critical for identifiability—frequently fails even at reasonable sample sizes, causing catastrophic collapse in graph recovery accuracy. Method: Through rigorous theoretical analysis and comprehensive simulations, we formally prove that faithfulness violation constitutes an insurmountable performance bottleneck and derive the first theoretical upper bound on structural recovery accuracy for neural causal discovery. Contribution/Results: Experiments demonstrate that state-of-the-art methods substantially deviate from the ground-truth DAG—even on small graphs and large samples. Our work quantifies an inherent ceiling on current paradigm’s performance and calls for a paradigm shift: from assumption-dependent modeling toward robust causal representation learning.

0 citationsRead paper

Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient

Feb 07, 2025

Scaling Mixture-of-Experts (MoE) models under fixed memory and compute budgets remains challenging due to unclear trade-offs among expert count, active parameter count, and dataset size. Method: We propose the first unified scaling law jointly governing dense and MoE architectures, theoretically modeling performance dependence on these key factors. Our analysis is grounded in rigorous theoretical derivation and empirically validated across 280+ large-scale experiments—spanning up to 2.7B active parameters and 5B total parameters. Contribution/Results: We demonstrate that, under identical memory constraints, MoE models consistently outperform dense counterparts—refuting the conventional view that MoE gains stem solely from parameter growth. Furthermore, we introduce a practical, principled framework for MoE configuration selection, enabling efficient large-model training. This work provides both theoretical foundations and actionable guidelines for resource-aware MoE scaling.

0 citationsRead paper
Recent publications

Latest Papers

Decoupled Relative Learning Rate Schedules

Jul 04, 2025

Conventional Transformer training employs a uniform learning rate across all model components, ignoring their dynamic heterogeneity in parameter sensitivity and update magnitude—leading to suboptimal optimization efficiency. Method: We propose a dynamic decoupled learning rate scheduling framework, introducing— for the first time—the concept of *relative learning rates*, which adaptively scale per-component learning rates based on layer-specific gradient statistics and architectural roles. Our approach is architecture-agnostic within the Transformer family and integrates seamlessly with Mixture of Experts (MoE) configurations. Contribution/Results: The method enables direct hyperparameter transfer across model scales—from small baselines to models 27× larger—without manual retuning. Empirical evaluation demonstrates up to 23% faster convergence for complex models and substantial reductions in computational resource consumption. This work establishes a scalable, efficient, and broadly generalizable optimization paradigm for large-scale neural networks.

0 citationsRead paper

VARSHAP: Addressing Global Dependency Problems in Explainable AI with Variance-Based Local Feature Attribution

Jun 08, 2025

Existing feature attribution methods (e.g., KernelSHAP, LIME) rely on global data distributions, leading to inaccurate characterization of local model behavior and distorted explanations. To address this, we propose VARSHAP—a model-agnostic local feature attribution method that introduces prediction variance reduction as the core Shapley value metric, the first such formulation. VARSHAP rigorously satisfies the efficiency, symmetry, and additivity axioms of Shapley values. It estimates conditional variances via Monte Carlo sampling, eliminating the need for surrogate models or distributional assumptions, and inherently exhibits robustness to data distribution shifts. Experiments on synthetic and real-world datasets demonstrate that VARSHAP improves attribution accuracy by 12–23% over KernelSHAP and LIME. Qualitative evaluations confirm its superior alignment with local decision logic, significantly mitigating the local explanation bias induced by global distribution dependence.

0 citationsRead paper

Since Faithfulness Fails: The Performance Limits of Neural Causal Discovery

Feb 22, 2025

Neural causal discovery methods face fundamental limitations: they struggle to reliably distinguish true from spurious causal edges under finite samples, and the faithfulness assumption—critical for identifiability—frequently fails even at reasonable sample sizes, causing catastrophic collapse in graph recovery accuracy. Method: Through rigorous theoretical analysis and comprehensive simulations, we formally prove that faithfulness violation constitutes an insurmountable performance bottleneck and derive the first theoretical upper bound on structural recovery accuracy for neural causal discovery. Contribution/Results: Experiments demonstrate that state-of-the-art methods substantially deviate from the ground-truth DAG—even on small graphs and large samples. Our work quantifies an inherent ceiling on current paradigm’s performance and calls for a paradigm shift: from assumption-dependent modeling toward robust causal representation learning.

0 citationsRead paper

Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient

Feb 07, 2025

Scaling Mixture-of-Experts (MoE) models under fixed memory and compute budgets remains challenging due to unclear trade-offs among expert count, active parameter count, and dataset size. Method: We propose the first unified scaling law jointly governing dense and MoE architectures, theoretically modeling performance dependence on these key factors. Our analysis is grounded in rigorous theoretical derivation and empirically validated across 280+ large-scale experiments—spanning up to 2.7B active parameters and 5B total parameters. Contribution/Results: We demonstrate that, under identical memory constraints, MoE models consistently outperform dense counterparts—refuting the conventional view that MoE gains stem solely from parameter growth. Furthermore, we introduce a practical, principled framework for MoE configuration selection, enabling efficient large-model training. This work provides both theoretical foundations and actionable guidelines for resource-aware MoE scaling.

0 citationsRead paper