🤖 AI Summary
本文解决了多变量线性组合中最大化严格标准化平均差异的问题,通过广义瑞利商求解,并提出多种估计方法和置信区间构建方式。
📝 Abstract
The strictly standardized mean difference (SSMD) is an interpretable, unit-free effect size originally developed for quality control and hit selection in high-throughput screening and subsequently extended to effect-size assessment and error-rate control. Existing SSMD theory considers one variable at a time. Here, we extend SSMD to a linear combination of multiple variables and determine the coefficients that maximize the resulting SSMD. We show that this optimization can be formulated as a generalized Rayleigh quotient with a closed-form solution and identify conditions under which the optimal direction coincides with Fisher's linear discriminant. We derive method-of-moments, bias-corrected, and maximum-likelihood estimators of the maximized SSMD and construct confidence intervals based on Hotelling's T-squared statistic and the noncentral F distribution under different correlation and variance structures. We also characterize classification cutoffs for the resulting linear score and establish its relationship with the area under the receiver operating characteristic curve (AUROC). Simulations show that SSMD maximization coincides with Fisher's linear discriminant and logistic regression under equal covariance but diverges under unequal covariance and class imbalance. In paired designs, SSMD maximization uniquely incorporates cross-covariance information. Linear support vector machines are also evaluated as a modern comparator. This framework provides an interpretable approach to constructing multivariable biomarkers and signatures for high-throughput screening, cytokine profiling, metabolomics, salivary diagnostics, and continuous-monitoring applications.