Attributes from Images, Not Class Names: Data-Driven Selection for CLIP

Improve zero-shot classification by selecting attributes from actual images instead of LLM descriptors. Outperforms prompt-tuning in minutes.

jueves, 23 de julio de 2026 • 3 min read • Q2BSTUDIO Team

Selecciona atributos de imágenes reales para mejorar CLIP

Artificial intelligence has advanced to the point of enabling image classification without prior examples, a field known as zero-shot classification. However, interpretability remains a critical challenge, especially when large language models (LLMs) are used to generate class descriptions. Recent research reveals a fundamental weakness: these descriptors are often conditioned on the label itself, not on the actual images. For instance, an LLM insists that strawberries are red, but in a dataset like ImageNet-Sketch, where strawberries appear as colorless line drawings, that descriptor becomes misleading. This causes a drastic drop in accuracy: removing the class name from the prompt drops ImageNet accuracy from 59.5% to 15.5%. The proposed solution is a radically different approach: selecting attributes directly from the target images, not from language. In this article, we explore how this conditioned selection technique improves robustness, and how companies like Q2BSTUDIO integrate these concepts into their custom software, artificial intelligence, cybersecurity, and cloud solutions.

The core problem is that LLM-generated descriptors describe the concept in the abstract, not its concrete visual manifestation in the data. When the domain shifts — for example, from real photographs to sketches or satellite images — those attributes become irrelevant. Instead, by scoring a large attribute pool against images in CLIP's shared embedding space and selecting the top-scoring ones per class, we obtain prompts that do not depend on class names. The results speak for themselves: with this selection, ImageNet accuracy rises to 23.8% (versus 15.5% from LLM descriptors), and the gain holds across four shifted ImageNet variants. Even when reusing the same attribute set from the LLM, the selection mechanism is the cause of the improvement. Moreover, with just one image per class, this method outperforms techniques like CoOp (prompt-tuning) by 3 points, and does so in under a minute instead of 14 hours, without introducing opaque soft prompts that hinder interpretability.

From a business and technical perspective, this innovation has profound implications. At Q2BSTUDIO, we understand that artificial intelligence must be not only accurate but also understandable and adaptable. Our custom software development services incorporate data-driven attribute selection techniques to ensure that classification models perform under real-world conditions, where training data rarely matches production data perfectly. For example, in an e-commerce product classification system, relevant visual attributes — such as shape, texture, or dominant color — should be extracted from actual catalog images, not from generic descriptions. The same applies in AI-assisted medical diagnostics, where features from X-ray or MRI images cannot rely on preconceived textual labels.

Conditioned selection not only improves accuracy but also offers an additional advantage: the selected attribute set acts as a readable summary of the dataset. This allows describing in words how a data distribution changes across domains. For companies working with big data and Business Intelligence (BI) — such as Power BI reports — having a textual representation of data changes is invaluable. Imagine monitoring a computer vision system in a factory: when lighting or camera angles vary, automatically selected attributes indicate which features have shifted, enabling quick adjustments without retraining the entire model.

Furthermore, this approach integrates naturally with cloud infrastructure. At Q2BSTUDIO, we offer cloud AWS/Azure services to deploy attribute selection pipelines at scale. Storing CLIP embeddings in the cloud and running attribute scoring in parallel dramatically reduces computation times, allowing real-time updates to classifiers. Cybersecurity also benefits: by not relying on human class names, the risk of an attacker manipulating labels to deceive the model is eliminated. AI agents, increasingly used in process automation, can leverage this technique to adapt their visual perception to the specific context without human intervention.

In summary, selecting attributes from images rather than from class names represents a paradigm shift toward more robust, interpretable, and efficient AI systems. For companies seeking competitive advantages, incorporating these techniques through technology partners like Q2BSTUDIO accelerates the adoption of classification solutions free from textual biases, improving decision-making based on real data. From custom applications to intelligent agents, the key is to let the data itself guide the descriptors, not preconceived language.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.