When we talk about brain-computer interfaces (BCI) based on event-related potentials (ERP), the success metric is not always what one might imagine. In systems like visual spellers, the ultimate goal is not simply to correctly classify a signal, but to enable the user to write characters efficiently. This raises a key question: what metric truly reflects writing accuracy? The answer goes beyond traditional accuracy.
Datasets from ERP experiments often present a strong class imbalance: the target signal (the character the user wishes to select) appears very few times compared to non-relevant stimuli. In this context, metrics like accuracy can be misleading, as a classifier that always predicts the majority class will achieve a good percentage but will be useless for real communication. Therefore, the scientific community has begun to recommend more robust indicators such as the area under the ROC curve (ROC AUC), Matthews correlation coefficient (MCC), Brier score, or the area under the precision-recall curve (PR AUC).
A relevant finding from recent studies is that the spelling rate —that is, the number of correctly selected characters per unit of time— correlates strongly with metrics that penalize both false positives and false negatives. This has direct implications for the design of BCI paradigms and for classifier optimization. It is not just about maximizing an abstract number, but about ensuring the user has a fluid and accurate writing experience.
In the field of software development for these types of systems, it is essential to have tools that allow implementing advanced classification models and evaluating them with the appropriate metrics. At Q2BSTUDIO, we work on developing custom applications for research and business environments, integrating artificial intelligence components that efficiently process biomedical signals. Additionally, we offer AI for companies that need real-time classification solutions, adapted to imbalanced data.
The choice of the correct metric not only affects model evaluation but also the configuration of experiments. For example, when repeating trials, some metrics like accuracy tend to stabilize quickly, while ROC AUC or MCC remain sensitive to marginal improvements. This is especially relevant when seeking to optimize the number of stimuli needed to achieve reliable writing, reducing user fatigue.
From a technical perspective, implementing a robust ERP BCI system involves combining signal processing, machine learning algorithms, and an intuitive user interface. At Q2BSTUDIO, we also offer AWS and Azure cloud services to deploy inference models in the cloud, as well as business intelligence services with Power BI to monitor system performance in real time. Cybersecurity is another key pillar, especially when handling sensitive user data.
Ultimately, the answer to the initial question is not unique: writing accuracy in ERP BCIs is best reflected by a set of metrics that combine sensitivity to imbalance, clinical interpretability, and direct correlation with user experience. Adopting these metrics not only improves research but also paves the way for more reliable commercial applications, such as those we develop with AI agents and process automation.

.jpg)



