Cost-Effective Agent Harnesses for Abstract Reasoning on ARC-AGI-1

Discover how cost-effective DeepSeek agent harnesses reach 67.25% on ARC-AGI-1 for just $0.62 per task — no ARC fine-tuning.

viernes, 31 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Arquitecturas DeepSeek económicas: +52 puntos en ARC-AGI-1

Abstract reasoning has become one of the most demanding frontiers in artificial intelligence. Benchmarks such as ARC-AGI-1 present visual and logical problems where the system must infer hidden rules, generalize them and apply them to new cases. For a long time, the dominant strategies were two: using frontier models with huge amounts of inference compute, through exhaustive search or extended chains of thought; or training small models on benchmark-specific data. Both options work, but at the expense of efficiency and with limited transfer to real environments.

At Q2BSTUDIO we see a third alternative with strong business potential: AI agents that operate on open models, without benchmark-specific fine-tuning and within tight budgets. The key is not model size but agent design. The architecture separates pattern discovery from solution building, and adds a reflection mechanism that allows retrying when hypotheses fail. This philosophy fits the development of custom software, where available resources and interpretability matter as much as accuracy.

The first stage, which we can call the explorer, focuses on identifying regularities in the training examples. It does not seek a final answer, but a set of hypotheses about the underlying rule. The second stage, the definer, turns those hypotheses into concrete, verifiable transformations. This separation prevents the model from having to solve two very different problems at the same time: understanding the pattern and building a function that represents it.

On top of this base, the reflective orchestrator adds a control layer. When proposed transformations fail to reproduce the examples, the orchestrator decides to explore other families of patterns, modify the data representation or reconsider the scope of the rule. In other words, it does not simply run the same process with more attempts; it introduces new search directions based on failure diagnosis. The inclusion of an explicit thinking tool is important: removing it reduces two-attempt accuracy by 5.75 percentage points.

Results on the public ARC-AGI-1 evaluation demonstrate the value of this strategy. A model that responds directly without an agentic architecture barely reaches 15.50 percent success. The explorer-definer pipeline achieves 57.50 percent with two attempts per task, at an estimated cost of 0.25 dollars per task. The reflective orchestrator goes up to 67.25 percent at a cost of 0.62 dollars per task. These figures are remarkable considering that no benchmark-specific training or large compute deployments were used.

One of the most interesting findings is that the main bottleneck is not candidate selection, but candidate generation. When success rate is measured in an unbiased way, it is observed that a selector based on accuracy over training pairs captures around 95 percent of the maximum possible performance given the generated set. This implies that adding a better classifier or a more complex ranking would produce marginal gains. Instead, increasing the diversity of generated transformations would have a real impact. The reflective orchestrator confirms this hypothesis: its adaptive re-exploration produces an improvement of 9.81 percentage points in one-attempt accuracy, a result that matches the selection-mediated gain in two-attempt mode.

This lesson applies directly to business. In AI projects, it is not always worth investing in bigger models or more compute. Often, the real breakthrough lies in designing processes able to generate diverse hypotheses, validate them with real data and decide when to explore a different path. This approach is especially useful in sectors where data is scarce, decisions have consequences or cost per operation must be controlled. The same logic can be applied to cybersecurity: an agent that observes logs, detects anomalous patterns and generates actionable explanations can behave like an explorer-definer of threats.

In organizations, technology infrastructure also plays a decisive role. Deploying this kind of agent on cloud AWS/Azure makes it possible to scale processing without compromising budget, keep traceability of every attempt and manage data access with security policies. In addition, the results of these agents can be integrated into BI/Power BI platforms to give business teams a clear view of which patterns are discovered, how they evolve and what decisions are derived. Combining AI agents and business intelligence turns technical experimentation into competitive advantage.

At Q2BSTUDIO we understand AI as an engineering discipline, not a black box. We have spent years developing software and technology for companies that need to automate processes, protect their data and make better decisions. Our proposal combines custom software with agent architectures, cloud, cybersecurity and analytics. An AI orchestrator does not have to be a research project reserved for large corporations; when well designed, it can deliver measurable results at reasonable costs in any sector.

The ARC-AGI-1 experiment leaves us with a clear conclusion: abstract reasoning is achievable without relying on frontier models or specific training. Agentic architecture, focused reflection and good budget management can compensate for lack of scale. For a company, this is a reminder that AI solutions should be evaluated by their impact, not by their complexity. In artificial intelligence and software development, the goal is not to imitate what everyone does, but to build systems that solve real problems efficiently.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.