In the vast universe of artificial intelligence, most multimodal models have focused on what we see and hear, leaving aside one of the most complex and evocative senses: smell. However, an emerging field is demonstrating that it is possible to teach machines to interpret aromas from images, using language as a semantic bridge. This approach, exemplified by architectures like SCENT, opens the door to applications ranging from augmented reality to environmental monitoring, as well as the creation of immersive sensory experiences. The key lies in the fact that images alone are not enough: they need linguistic context to unravel what smells might be associated with a scene, whether it be freshly brewed coffee, ozone after a storm, or the aroma of a damp forest.
The research highlights a fundamental point: smells are not always visible in pixels, but rather depend on contextual and environmental factors. To address this challenge, language and vision models (VLMs) are used to generate semantically rich scene descriptors. These descriptors not only name objects but also infer plausible smells and environmental conditions. From there, an olfactory encoder learns to align signals from electronic noses with that shared representation space, managing to decompose complex mixtures into specific and contextual components. This advance not only allows retrieving images from a smell but also generating textual descriptions of an aroma without the need for physical sensors.
For companies looking to explore these technological frontiers, having a partner that masters both AI for businesses and the development of customized solutions is essential. At Q2BSTUDIO we offer custom applications that integrate artificial intelligence, cloud services aws and azure, and business intelligence services like Power BI, all backed by a robust cybersecurity architecture. Our AI agents and artificial intelligence systems can adapt to domains as diverse as agribusiness, perfumery, or healthcare, where the combination of vision, language, and smell opens new avenues for analysis and automation.
From a technical perspective, the language-guided latent decomposition proposed by this research allows interpreting smell mixtures similarly to how a sommelier describes a wine: separating fruity, mineral, or floral notes coming from different sources. This capability is especially valuable in sectors where air quality, leak detection, or product authentication depend on olfactory signatures. The alignment between electronic sensors, text, and image not only improves information retrieval —as demonstrated in smell-to-image and smell-to-text tasks— but also provides a foundation for building multispectral reasoning systems.
The path toward truly multimodal artificial intelligence involves integrating all sensory channels, and smell has so far been the great forgotten one. With approaches like SCENT, and relying on the custom software we develop at Q2BSTUDIO, it is possible to build platforms that not only see and understand but also smell the world. This has direct implications for creating immersive virtual environments, monitoring industrial processes, and personalizing user experiences based on olfactory context. The technology is already ready; it just takes the right approach to capture what images do not tell.

.jpg)

