Tversky Monosemanticity Score: Measuring SAE Interpretability

The Tversky Monosemanticity Score (TMS) measures latent coherence in SAEs. This label-free metric outperforms embedding-based methods for evaluating

sábado, 25 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Cómo TMS Evalúa la Coherencia Latente en Autoencoders Dispersos

Explainable artificial intelligence is one of the most critical areas for the responsible adoption of deep learning models. Within this field, mechanistic interpretability seeks to decompose neural representations into features that are more understandable to humans. A popular technique is Sparse Autoencoders (SAEs), which learn latents that ideally should be monosemantic, i.e., activate exclusively for a specific concept. However, quantitatively measuring that monosemanticity has been a persistent challenge. Existing metrics rely on external labels or pre-trained embedding models, whose geometry can distort the evaluation. To address this limitation, the Tversky Monosemanticity Score (TMS) emerges, a label-free and encoder-free metric that proposes a new approach to evaluate the internal coherence of binarized latents.

TMS is based on the notion of activation set coherence. Instead of comparing features against human-annotated concepts, it measures the internal consistency of each latent's activation patterns when binarized. This avoids the bias introduced by embedding space anisotropy, a common issue in models like CLIP or DINOv3. Experiments with SAEs trained on different architectures (DINOv3, CLIP, BLIP2) and training regimes (TopK, BatchTopK) show that TMS aligns with traditional monosemanticity indicators but is less sensitive to geometric distortions. Moreover, it reveals distinct training dynamics depending on the base model, providing valuable insights for hyperparameter tuning.

From a business perspective, the ability to measure the quality of features extracted by SAEs has direct implications for developing robust and transparent AI systems. For instance, a company developing AI solutions can use TMS to evaluate whether its model latents represent clean concepts or are contaminated with irrelevant information. This is especially relevant in applications where explainability is a regulatory requirement, such as in finance or healthcare. Q2BSTUDIO, as a software and technology development company, integrates these advanced metrics into its workflows to ensure the models it deploys are interpretable and reliable.

Furthermore, the label-free nature of TMS facilitates its adoption in environments where annotated datasets are not available. This is common in custom software projects, where data comes directly from the client and is not always labeled. Combined with other tools like cybersecurity, cloud (AWS/Azure), or Business Intelligence with Power BI, Q2BSTUDIO offers a complete ecosystem for building intelligent solutions. For example, an AI agent system processing business data can benefit from improved interpretability to debug unwanted behaviors and enhance user trust.

One of the problems TMS solves is the anisotropy of embedding encoders. When using models like CLIP to measure similarity between latents and concepts, the non-uniform geometry of the embedding space can favor certain directions, leading to artificially inflated or reduced metrics. TMS, by operating directly on binarized latent activations, avoids this distortion. For a company like Q2BSTUDIO, which develops AI-based cybersecurity solutions, having reliable metrics is crucial: a false positive in interpreting a latent could lead to erroneous conclusions about model behavior, affecting threat detection.

Additionally, TMS's computational simplicity allows its integration into machine learning pipelines with little overhead. In projects using cloud services like AWS or Azure, where each training cycle incurs a cost, having an efficient metric is a competitive advantage. Q2BSTUDIO offers cloud AWS/Azure services that can host these training and evaluation processes, ensuring scalability and performance. Combining TMS with monitoring tools allows data teams to detect degradations in latent quality over time, especially useful in AI agent systems that learn continuously.

Another area where TMS can make a difference is in developing Business Intelligence systems with Power BI. Although Power BI is not directly related to SAEs, the interpretability of models feeding dashboards is increasingly in demand. If an AI model generates predictions visualized in Power BI, business stakeholders need to trust that the model's latent variables make sense. TMS provides a quantitative guarantee that the extracted features are coherent, reinforcing the credibility of analytics. Q2BSTUDIO integrates these capabilities into its BI / Power BI solutions, offering added value to its clients.

Finally, it is worth noting that research into monosemanticity is still evolving. TMS is not a perfect metric, but it represents a step forward compared to existing alternatives. As more companies adopt mechanistic interpretability techniques, tools like TMS will become standard. Q2BSTUDIO, as a technology partner, stays at the forefront of these innovations to offer its clients the best automation and cybersecurity solutions, always with a focus on quality and transparency.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.