Context by unique information: Auditable working memory

Learn how a Dirichlet process-based working memory scales with unique insights, reducing tokens and improving auditing over long flows

martes, 14 de julio de 2026 • 6 min read • Q2BSTUDIO Team

Novelty, recency and summary: three types of context memory

In the fast-paced world of artificial intelligence, one of the most persistent bottlenecks has been memory management. Today's models, from large languages to recommender systems, process terabytes of data in continuous flow, but their storage architecture is often tied to a rigid concept: the token as a unit of memory. Every word, every number, every character becomes an input that must be retained or discarded, regardless of whether that information is truly novel or redundant. This limitation has led to deep reflection within the technical community: what if we could scale memory based on the unique information that really matters, rather than by the volume of tokens?

The core idea of a novelty-based approach—where a cache is only opened when an incoming key is truly different—transforms the way we understand context. Instead of compressing the past into an opaque recurring state, or keeping an entry for each token in a sliding window, a working memory is proposed that is addressable by content and, moreover, auditable. This means that each stored item is not a simple number, but a meaningful record: a template, a code, a medicine, a place. Memory becomes an inspectable table, not a black box. For companies looking for transparency in their AI systems, this feature is revolutionary, as it allows you to track why a model made a decision, making it easier to audit and comply with regulations.

From a practical perspective, the application of this concept goes far beyond theoretical research. In business environments where the flow of data is massive and continuous – such as in the management of medical claims, financial transactions or cybersecurity logs – the ability to distinguish the novel from the repetitive allows for a drastic reduction in computational costs. For example, a model trained to predict medical billing codes can keep in its memory only those diagnoses or procedures that actually provide new information, ignoring the repetitions that overwhelm token-based systems. In controlled tests, this technique achieves a performance equivalent to that of a complete attention, but catering to half of the tokens. The advantage is amplified as the sequence lengthens, while on short sections dominated by the locality, a sliding window remains the optimal option.

What implications does this have for custom software development? Companies that need to integrate artificial intelligence into their internal processes are often faced with the trade-off between accuracy and efficiency. A model that consumes less memory and resources can run on edge devices, in cost-controlled cloud environments, or even in real time. The applications we build in Q2BSTUDIO directly benefit from these advances: we can design systems that retain only the information relevant to each task, avoiding the extra cost of maintaining huge windows of tokens that are never used. For example, in an AI-based agent-based customer service assistant, the model can remember the history of key interactions without storing each redundant message, improving speed and personalization without inflating the cloud bill.

The concept of auditable working memory aligns perfectly with modern data governance needs. In regulated sectors such as health or finance, it is not enough for a model to be accurate; it must be explainable. Being able to inspect a table with the novel keys that the system has chosen to keep allows auditors to verify that no sensitive information has been leaked or that decisions are based on fair criteria. In addition, this transparency makes it easier to detect biases: if a model only retains certain demographic patterns, it can intervene before it causes discrimination. The corporate AI we offer at Q2BSTUDIO is not only concerned with performance, but also with ethics and control. Our developments integrate audit mechanisms from the design, allowing each decision of the model to be traceable.

Another fascinating aspect is the differentiation between the types of memory according to the task. The baseline study suggests that tasks can be classified into three large families: those that depend on recall (such as searching for specific information in a long text), those that depend on a summary or global state (such as trend forecasting), and those that depend on locality (such as predicting the next word in a short sentence). For each, the optimal memory mechanism is different: a novelty cache for what needs to be remembered, a recurring state for the summary, and a current window for what is nearby. This segmentation allows you to design hybrid architectures that dynamically adjust your resource consumption. For example, in an AWS and Azure cloud services system, a module could be implemented that decides which type of memory to activate according to the nature of the data flow, optimizing the use of instances and reducing execution costs. At Q2BSTUDIO, we've worked with customers migrating their workloads to the cloud, and in-memory efficiency translates directly into less CPU/GPU consumption and lower monthly bills.

Of course, not everything is immediately applicable. The experiments are still small-scale and use public data, but the principle is well established: context can scale with different information, not tokens. This opens the door to much lighter and more sustainable artificial intelligence systems. For companies that have already invested in big data infrastructure or business intelligence services like Power BI, adaptation is natural. Imagine a Power BI dashboard that doesn't have to reload thousands of redundant rows every time it's updated, but only processes what's new, maintaining an intelligent summary of what you've learned. This is possible by combining novel caching techniques with visualization tools. Our team in Q2BSTUDIO integrates business intelligence services that leverage these concepts to deliver real-time dashboards with low memory consumption, improving the end-user experience.

Cybersecurity is another field where this selective memory shines. Intrusion detection systems generate millions of alerts a day, most of them false positives. An AI model that retains only the truly novel attack patterns can identify emerging threats without being overwhelmed by noise. In addition, the ability to audit memory—knowing exactly which keys were considered suspicious—is vital for incident response teams. At Q2BSTUDIO we offer custom applications for cybersecurity that implement this type of memory, allowing companies to protect their assets without the need for hyperscaled hardware.

Finally, the evolution towards a content-addressable and auditable working memory marks a before and after. The AI agents of the future will not only be more efficient, but more reliable. Organizations that adopt these architectures will be able to scale their language models, recommendation systems, and virtual assistants without fear of skyrocketing computational costs. And they will do so with the peace of mind of knowing that every decision can be reviewed, every memory can be justified. At Q2BSTUDIO, as a software and technology development company, we are committed to bringing these innovations into business practice, transforming cutting-edge concepts into tangible solutions that truly add value.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.