Phone segmentation and recognition are two fundamental tasks in speech processing that have traditionally been approached separately. However, recent research shows that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms). Using techniques such as Phonological Activation Mapping (SPAM), it is possible to directly extract features like voicing or nasality from each audio frame, and on top of that build lightweight prediction heads that perform both tasks simultaneously. This approach requires less than a minute of labeled phonetic transcriptions and generalizes to unseen phonemes during training, representing a significant advance in efficiency and scalability.
From a technical and business perspective, this innovation opens the door to much more agile and cost-effective applications. For example, in customer service environments, a system capable of segmenting and recognizing phonemes in real time can improve the accuracy of virtual assistants, reduce response times, and enable more detailed sentiment analysis. Moreover, by requiring so little labeled data, companies can quickly adapt models to new languages or dialects without massive investments in manual transcription.
At Q2BSTUDIO, as a software and technology development company, we see in this methodology an opportunity to integrate advanced artificial intelligence capabilities into custom solutions. Our team can design speech recognition systems that leverage these self-supervised models, deploy them on cloud infrastructures such as AWS or Azure, and ensure the cybersecurity of processed data. Additionally, information extracted from phonetic transcriptions can feed Business Intelligence dashboards using Power BI, enabling organizations to make decisions based on conversation patterns.
The integration of AI agents acting on these phonological analysis layers is another area where our capabilities make a difference. For instance, an intelligent agent could automatically detect key moments in a call (such as complaints or requests) and trigger automated workflows, improving operational efficiency. All of this is supported by a foundation of custom applications that perfectly adapt to each client’s needs, whether in the healthcare, financial, or logistics sector.
The combination of phone segmentation and recognition through phonological activation is not only an academic research topic but a technology ready to be transferred to production environments. At Q2BSTUDIO we are prepared to help companies implement these solutions, ensuring that AI, cloud, and cybersecurity work together to deliver tangible and measurable results.




