The key-value (KV) cache has become one of the main bottlenecks when deploying large language models (LLMs) with extended context windows. As sequence length grows, the storage required for KV pairs increases linearly, making both inference and production deployment more expensive. Techniques such as vector quantization (VQ) and, in particular, residual quantization (RQ) have shown promise in compressing this memory below 1 bit per element. However, traditional K-means-based methods using Euclidean distance present a subtle problem in high-dimensional spaces: centroid averaging can induce norm shrinkage, which weakens angular alignment and hinders the directional preservation of vectors. To address this, a new variant called Gain-Shape K-means (GSKM) emerges, which separates gain (norm) from shape (direction). By integrating this technique into a residual quantization pipeline, Gain-Shape Residual Quantization (GSRQ) is obtained. In models such as LLaMA-3-8B, GSRQ substantially improves accuracy on LongBench tasks, achieving a jump from 11.34 to 33.54 in average accuracy at 1 bit, an increase of more than 22 percentage points over baselines like VQLLM. This advancement is relevant for any organization seeking to deploy artificial intelligence for businesses with large language models, as it drastically reduces memory consumption without sacrificing performance. At Q2BSTUDIO, we understand that infrastructure optimization is key to efficiently deploying AI solutions. Therefore, we offer AWS and Azure cloud services that facilitate the management of memory-intensive workloads, as well as custom application development and custom software to integrate these technologies into production environments. Additionally, we combine these capabilities with business intelligence through Power BI and with AI agents that automate complex processes. Our cybersecurity expertise ensures that any LLM deployment in the cloud meets the highest data protection standards. If your company needs to harness the potential of language models with long contexts without skyrocketing infrastructure costs, we invite you to explore how our business intelligence services and AI for businesses solutions can be tailored to your needs.

.jpg)



