multivariate regression

Builds and applies multivariate regression models to analyze relationships among multiple predictors and outcomes, producing fitted models, diagnostics, and inference for multilevel or multivariate data.

multivariateregression

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.22
Aug 01, 2026Aug 01, 2026
Career
Value
No comparison yet
$203K/year
Aug 01, 2026Aug 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Matrix-Variate Regression Model for Multivariate Spatio-Temporal Data

Nov 06, 2025
CA
Carlos A. R. Diniz
🏛️ Federal University of Sao Carlos | University of Connecticut

This paper addresses the challenge of jointly modeling covariate effects, spatial dependence, and temporal dependence in multivariate spatiotemporal data. We propose a matrix-response varying-coefficient regression model that embeds covariates into the mean structure of the response matrix and employs a Kronecker-product-based separable covariance structure to explicitly decouple spatial and temporal correlations. Parameter estimation is conducted via maximum likelihood, ensuring both computational efficiency and statistical accuracy, while enabling robust inference under heterogeneous spatial resolutions. Simulation studies demonstrate excellent parameter recovery performance. Applied to municipal-level agricultural and livestock panel data from Brazil, the method successfully uncovers interpretable spatiotemporal dynamic patterns and reveals heterogeneous impacts of key covariates—including climate variables and policy interventions—across space and time. The framework provides a scalable, principled paradigm for high-dimensional spatiotemporal causal analysis.

Estimating parameters via maximum likelihood for dependenciesModeling multivariate spatio-temporal data with matrix regressionValidating model performance through simulations and agricultural data

Multivariate Temporal Regression at Scale: A Three-Pillar Framework Combining ML, XAI, and NLP

Apr 02, 2025
JK
Jiztom Kavalakkatt Francis
🏛️ Iowa State University

To address weak interpretability, strong redundant interference, and low expert trust in high-dimensional time-series regression, this paper proposes a tripartite ML-XAI-NLP framework. First, it introduces a novel global key feature identification method based on perturbation sensitivity. Second, it establishes a retraining-free, cross-sensor redundancy-aware dimensionality reduction mechanism. Third, it integrates Granger causality graphs, SHAP-LIME hybrid attribution, and temporal embedding semantic alignment to generate verifiable natural language explanations. Evaluated on both real-world and synthetic datasets, the framework achieves 42%–67% dimensionality compression and reduces regression error by 19.3%. Moreover, it significantly enhances domain experts’ comprehension of model decisions and improves debugging efficiency. The approach bridges machine learning, explainable AI, and natural language processing to deliver both statistical performance and human-centered interpretability in time-series modeling.

Analyzing high-dimensional data with complex relationshipsIdentifying global key features for simpler, interpretable modelsReducing computational demands and human effort in data validation

This paper addresses the failure of variable selection in Bayesian multivariate linear regression under strong collinearity and sparse information (weak signals, small sample sizes, high inter-variable correlation) in the design matrix. It demonstrates that jointly estimating regression coefficients and the off-diagonal elements of the error covariance matrix exacerbates estimation bias and degrades predictive performance. To mitigate this, we propose a two-step Bayesian variable selection strategy: first, estimate the mean structure (i.e., regression coefficients) under a diagonal error covariance assumption; second, independently model residual dependence. Simulation studies and empirical analysis on NIR spectroscopy data confirm that the method substantially improves variable selection accuracy, coefficient estimation precision, and out-of-sample prediction in low-information regimes. The key contribution is identifying the “overfitting risk” inherent in full error covariance modeling and establishing that decoupling mean and covariance estimation achieves a favorable trade-off between robustness and statistical efficiency.

Bayesian variable selection in multivariate regression with collinearityPoor estimation in low-information settings with non-diagonal covarianceTwo-step procedure improves accuracy by separating mean and covariance estimation

A Latent Variable Approach to Learning High-dimensional Multivariate longitudinal Data

May 23, 2024
SM
Sze Ming Lee
🏛️ London School of Economics and Political Science | The Chinese University of Hong Kong

This paper addresses the challenges of covariate effect inference and future outcome prediction in high-dimensional multivariate longitudinal data—characterized by complex dependencies among variables and over time, mixed-type outcomes (continuous and discrete), and irregular observation patterns (missingness or right-censoring). We propose a novel latent-variable modeling framework that (i) unifies the treatment of mixed-type responses and irregular temporal observations for the first time; (ii) introduces an information criterion tailored to high-dimensional longitudinal settings for automatic selection of the latent factor dimension; and (iii) establishes a rigorous central limit theorem for regression coefficient estimators, ensuring valid statistical inference. Evaluated on a customer shopping behavior prediction task, our method significantly improves long-term trend modeling accuracy and robustness of personalized forecasting, demonstrating both practical utility and theoretical soundness in real-world high-dimensional longitudinal applications.

Analyzing covariate effects using latent variable approachModeling high-dimensional multivariate longitudinal data dependenciesPredicting future outcomes with mixed-type and missing data

This study addresses the challenge of variable selection in high-dimensional linear regression, where noise and proxy variables often hinder the simultaneous achievement of accuracy and sparsity. To tackle this issue, the authors propose Boosting with Multiple Testing (BMT), a novel method that integrates a multiple hypothesis testing framework into a forward stepwise boosting procedure. At each iteration, only the most statistically significant variable is selected, and candidate variables are screened using family-wise error rate control. This approach effectively curbs greedy selection of noise or proxy variables and enjoys near-oracle properties, enabling consistent recovery of the true underlying model. The theoretical analysis leverages the multiple testing framework of Chudik et al. (2018) and strong mixing process inequalities from Dendramis et al. (2022). Simulation studies demonstrate that BMT outperforms OCMT and Lasso-type methods in both model selection accuracy and coefficient estimation (measured by RMSE), while empirical applications in macro-finance yield sparse, interpretable, and highly predictive models.

high-dimensional regressionmodel selectionmulticollinearity

Latest Papers

What's happening recently
View more

This study addresses regression problems involving functional predictors and multivariate responses by proposing a novel coefficient function decomposition method that explicitly leverages the interdependencies among response variables. By integrating the functional predictor structure with response correlations, the approach introduces a joint smoothing-and-sparsity penalty strategy that enhances both curve selection and estimation accuracy across settings ranging from small to large-scale scenarios—with up to thousands of functional predictors. Theoretical analysis and extensive numerical experiments demonstrate that the proposed method substantially outperforms existing alternatives. An efficient implementation is provided in the R package FRegSigCom, enabling scalable and high-dimensional functional regression modeling.

functional predictorsfunctional regressionhigh-dimensional

This study addresses the limitations of traditional multivariate response regression methods, which often neglect inter-variable dependencies and struggle to simultaneously achieve smooth fitting and dimensionality reduction. The authors propose a novel unified framework that, for the first time, integrates P-spline smoothing into reduced-rank regression by employing B-spline basis expansions with penalized coefficients. This approach jointly models the correlations among multiple responses and their nonlinear trends. An efficient block-relaxation algorithm is developed for parameter estimation, while biplots and partial dependence plots are incorporated to enhance model interpretability. Extensive simulations and analyses of three real-world datasets demonstrate that the proposed method substantially improves both fitting smoothness and the ability to elucidate the underlying multivariate response structure.

Multiple OutcomesMultivariate RegressionP-splines

This study addresses the lack of a unified formulation for scalar, multivariate, and functional regression models, which obscures their intrinsic connections. By leveraging an integral operator defined with respect to general measures, the authors propose a unified framework that subsumes all three regression types as special cases of the same operator under different input and output measures. This framework reveals classical regression forms as measure-dependent manifestations of a single operator, clarifies discretized modeling as operator estimation under specific measures, and explains the efficacy of vectorized multivariate regression in linear settings. Theoretically, the authors prove that discrete representations correspond exactly to operator evaluations under discrete measures and converge to the continuous case as the discretization grid refines; moreover, this estimator is equivalent to standard multivariate regression and inherits its classical statistical properties.

functional regressionintegral operatorsmeasure theory

This study addresses the challenge of regression with high-dimensional responses and covariates when the responses are influenced by both observed covariates and unobserved latent variables, a setting where conventional multivariate regression methods fail to provide effective modeling. The authors propose a generalized latent variable model that accommodates mixed-type high-dimensional responses and allows for flexible dependence structures between covariates and latent factors. By decomposing the non-convex estimation problem into a sequence of convex subproblems through alternating optimization, and integrating debiased estimation with asymptotic normality analysis, the work achieves, for the first time, valid statistical inference on covariate effects within a high-dimensional generalized latent variable framework. The proposed estimator enjoys statistical consistency and guaranteed error bounds, while the debiased version exhibits asymptotic normality, as demonstrated empirically in an application to PISA data for assessing educational equity.

covariate effectsgeneralized latent variable modelshigh-dimensional responses

This study addresses the limitation of conventional approaches that ignore the interdependencies among multiple categorical outcomes—such as PTSD, depression, and pain—leading to information loss and reduced predictive performance. To overcome this, the authors propose a multivariate multinomial logit model based on ANOVA decomposition, which explicitly captures the conditional dependence structure among outcomes to reduce the complexity of the high-dimensional parameter space. Efficient estimation and variable selection are achieved through a composite likelihood framework combined with bridge penalization. Computationally, a Minorization-Maximization (MM) algorithm is employed to ensure numerical stability. Simulation studies demonstrate the method’s superior accuracy in both parameter estimation and variable selection, and its practical utility is further validated through application to real-world data from the AURORA cohort.

correlated categorical outcomeshigh-dimensional parameter spaceinterdependent outcomes

Hot Scholars

HL

Han Lin Shang

Department of Actuarial Studies and Business Analytics, Macquarie University
Functional data analysisnonparametric smoothingnonparametric statisticsmachine learning
NL

Nils Lid Hjort

Professor of Mathematical Statistics, University of Oslo
Theoretical and applied statistics and probability theory
LF

Long Feng

Professor of Nankai University
High Dimensional DataHigh Frequency Data
MK

Masahiro Kato

Mizuho-DL Financial Technology Co., Ltd. / The University of Tokyo
Economics
SF

Stefan Feuerriegel

Professor, LMU Munich
AI in ManagementBusiness AnalyticsComputational Social ScienceAI for Good