BEST-RQ-2: Two-step approach for self-supervised audio representations

BEST-RQ-2: two-step approach for audio representations. Outperforms models in transfer without increasing compute. X-ARES results.

miércoles, 1 de julio de 2026 • 2 min read • Q2BSTUDIO Team

Self-supervised learning with contextualization and prediction

Self-supervised learning has transformed the way artificial intelligence models understand audio signals, enabling representations that transfer effectively across different domains and tasks. The recent evolution from BEST-RQ to BEST-RQ-2 introduces a two-phase pretraining scheme —contextualization and prediction— that significantly improves generalization capability without increasing computational load at inference. This advancement replaces the traditional Conformer encoder with a Vision Transformer (ViT) that processes only unmasked regions of the spectrogram, while a lightweight predictor, discarded after pretraining, infers the labels of hidden areas. Results on the X-ARES and X-ARES-LLM benchmarks show that BEST-RQ-2 outperforms its predecessors in global transfer, with slightly lower performance on speech but notably superior performance on music and ambient sounds.

This two-step architecture not only optimizes learning efficiency but also opens new possibilities in the development of AI for businesses that need to analyze large volumes of acoustic data. For example, in applications such as virtual assistants, industrial environment monitoring, or music recommendation systems, having robust audio representations is critical. At Q2BSTUDIO, we understand that implementing these models requires solid, customized infrastructure. That is why we offer custom applications that integrate everything from signal preprocessing to deployment in cloud environments. Our custom software services allow adapting these advanced techniques to specific use cases, while our AWS and Azure cloud services solutions ensure scalability and security.

Additionally, BEST-RQ-2's ability to abstract auditory features facilitates the creation of AI agents that interact with the sound world, from anomaly detection in factories to conversation analysis for business intelligence services. With tools like Power BI, processed audio data can be visualized and correlated with business metrics. Cybersecurity also benefits: self-supervised audio models can identify fraud patterns or intrusions in communications. In short, BEST-RQ-2 represents a step forward in artificial intelligence applied to audio, and at Q2BSTUDIO we are prepared to help organizations adopt these innovations with turnkey solutions, from consulting to development and integration.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.