State Compression in Two-Agent LLM Relays: Constraint Preservation

Study compares compression methods in two-agent LLM relays. JSON extraction achieves 0.96 feasibility, while narrative summarization drops to 0.48. Learn which

jueves, 23 de julio de 2026 • 3 min read • Q2BSTUDIO Team

JSON vs resumen narrativo: preservación de restricciones

In the current artificial intelligence ecosystem, large language model (LLM)-based agents are becoming fundamental components for automating complex processes. However, when these agents execute long-running tasks, they accumulate intermediate traces that include audits, eliminations, and numerical calculations. This state accumulation, if not properly managed, creates an information bottleneck that can break strict constraints, whether numerical or categorical. State compression, or 'hand-off compression,' emerges as a necessary solution, but not all techniques are equally effective. A recent study on a travel planning system with two LLM agents (a Researcher and a Booker) reveals that the way information is compressed and transmitted between agents directly impacts the feasibility of subsequent decisions. Results show that structured JSON extraction achieves a feasibility accuracy of 96%, while narrative summaries, despite being more compact, degrade reliability to 48%. These findings have profound implications for companies developing custom applications and AI-based systems, such as Q2BSTUDIO.

To understand the relevance of this problem, imagine a workflow where one agent investigates fixed inventories of hotels and flights, and another agent must select an optimal combination using only a compressed summary. If the summary omits critical details — such as a budget constraint or exact availability — the final decision may be infeasible. In the mentioned study, four compression techniques were compared: no compression, narrative summarization, schema-constrained JSON extraction, and embedding-based pruning. JSON extraction not only outperformed in accuracy but also provided an auditable and structured representation. This is key in enterprise environments where traceability and constraint compliance are mandatory. Q2BSTUDIO, as a company specialized in AI and software development, understands that state compression must not sacrifice data integrity for mere token savings.

From a technical perspective, embedding-based pruning, which uses vector representations to select the most relevant information, achieved the same accuracy as the no-compression case (88%) without requiring an additional generative call. This suggests that intelligent pruning techniques can maintain constraint fidelity without increasing computational load. For companies operating in the cloud, such as those using cloud AWS/Azure, this efficiency is crucial for scaling multi-agent systems without skyrocketing costs. Additionally, cybersecurity plays a fundamental role: when agents exchange compressed data, it is necessary to ensure that sensitive information is not leaked or corrupted. Q2BSTUDIO integrates cybersecurity practices into its automation solutions, ensuring that every hand-off meets integrity and confidentiality standards.

Another relevant aspect is business analytics. LLM agents generate massive volumes of intermediate data that, if structured correctly, can feed Business Intelligence dashboards. Using tools like Power BI, it is possible to visualize the evolution of agent decisions and detect bottlenecks. Q2BSTUDIO offers BI / Power BI services that help companies transform that data into actionable insights. For example, if a travel planning system shows that 40% of decisions fail due to poor compression, the representation scheme can be adjusted to improve success rates.

Process automation through AI agents is not only a matter of efficiency but also reliability. Companies adopting these technologies must carefully consider how agents communicate with each other. The main lesson from the study is that representation matters more than compactness. A narrative summary may be brief but loses precision; a well-designed JSON, on the other hand, preserves the structure of constraints and allows for later audits. Q2BSTUDIO, with its experience in process automation, helps clients design multi-agent flows that maintain constraint fidelity, whether in cloud, on-premise, or hybrid environments.

In conclusion, state compression in two-agent LLM relays is a technical challenge that directly affects the feasibility of automated decisions. Structured JSON extraction stands out as the most reliable technique, while narrative summaries, though tempting for brevity, can introduce critical errors. For companies looking to implement robust AI solutions, the key is to choose a compression strategy that preserves constraints and enables auditing. Q2BSTUDIO, as a technology partner, offers custom applications that integrate these principles, combining AI, cloud AWS/Azure, cybersecurity, and BI to create intelligent and trustworthy systems. The future of automation lies in agents that are not only fast but also accurate and auditable.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.