RuBench: repository-level benchmark with tasks in Russian

Discover RuBench, the first benchmark that evaluates coding agents with real tasks in Russian. Results and surprises in model substitution.

miércoles, 8 de julio de 2026 • 2 min read • Q2BSTUDIO Team

First realistic benchmark for coding agents in Russian

The evaluation of artificial intelligence tools applied to software development has taken a significant turn with the emergence of benchmarks that reflect real working conditions. Traditionally, test suites were designed in English with highly structured statements, but engineering teams need to measure how AI agents behave when given instructions in the team's native language, as occurs in multinational environments. The new RuBench benchmark precisely addresses this gap: it proposes 25 tasks extracted from real fix commits in active repositories (aiohttp, aiogram, Laravel, NestJS, Fastify), with statements written in Russian mimicking real client requests, not translations. Each task is validated using the regression tests that the original project maintainer already had, ensuring the solution is functional and not merely syntactic.

The relevance of this approach goes beyond academia. For any company developing custom applications or custom software, having a reliable method to evaluate coding agents in multilingual contexts with loosely structured requirements is essential. RuBench also reveals an important practical finding: when auditing the complete trajectories of an off-competition agent, it was discovered that the deployed product silently replaced the underlying model in 20% of the tasks, redirecting HTTP fixes to a more powerful model. This shows that, in practice, the unit actually being measured is not the language model, but the complete system (agent + security configurations + fallback logic).

For organizations integrating enterprise AI into their development processes, these lessons are key. It is not enough to select an advanced language model; the behavior of the final product must be validated in real scenarios, with requests written in the team's language and with regression tests that reflect actual code maintenance. At Q2BSTUDIO, we apply this same principle when designing artificial intelligence solutions for our clients: we combine powerful models with careful orchestration that avoids unwanted substitutions and ensures traceability. Additionally, we offer AWS and Azure cloud services to scale these solutions, business intelligence services with Power BI to visualize agent performance, and cybersecurity to protect continuous integration pipelines. All of this is integrated into a custom applications approach that respects the particularities of each team and each project.

The evolution of benchmarks like RuBench indicates that the industry is moving towards more ecological and realistic evaluations. Companies that are already adopting AI agents for software maintenance and evolution tasks must demand transparency about the real behavior of the system, not just about the model's capabilities in the lab. In this regard, having a technology partner that understands both software engineering and artificial intelligence is a competitive advantage. At Q2BSTUDIO, we help organizations implement AI for businesses in a robust, secure, and aligned manner with real workflows, whether by integrating coding agents, automating processes, or deploying cloud infrastructure. The future of AI evaluation lies in measuring what actually happens when a system faces a client problem, in their language and under their rules. RuBench is a step in that direction, and from professional practice, we must take note and adapt our methodologies.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.