🤖 AI Summary
In event-related potential (ERP)-based brain–computer interfaces (BCIs), conventional accuracy metrics often fail to adequately reflect users’ actual spelling performance, particularly due to limitations in handling class imbalance and evaluating spelling speed. This study systematically evaluates the correlation of 13 performance metrics—including Brier score, Matthews correlation coefficient (MCC), ROC AUC, PR AUC, average precision (AP), and partial AUC (pAUC)—with spelling rate across varying numbers of trial repetitions, using the publicly available LARESI and OpenBMI datasets. The work reveals, for the first time, that the Brier score, MCC, and several metrics designed to address class imbalance exhibit strong correlations with spelling rate. These findings advocate for prioritizing such metrics in ERP-BCI studies to enhance comparability and practical relevance.
📝 Abstract
For predictive models, the often-reported performance metrics are the loss and accuracy. In synchronous Brain- Computer Interface (BCI) systems, these metrics are informative for most BCI paradigms; however, for Event-Related Potential (ERP) applications the spelling rate, which measures the number of characters correctly selected is more important as it influences the estimation of information transfer rate (ITR) and any related metric measuring spelling performance. Moreover, ERP-based BCIs hold imbalanced data class distributions, which require reporting metrics that can handle the imbalance, such as the area under the receiver operating characteristic curve (ROC AUC). In this work, we study the correlation of the spelling rate with 13 metrics to identify which among them best reflect user spelling performance and how they are affected by trial repetition. The Results of two datasets (a private LARESI ERP dataset and the public OpenBMI ERP dataset) favor the Brier score, Matthews Correlation Coefficient (MCC), and the metrics that account for class imbalance in binary classification: ROC AUC, area under the Precision-Recall curve (PR AUC), Average Precision (AP), and partial AUC (pAUC). These findings encourage researchers and practitioners to report those metrics in ERP-based BCI experiments.