ITPEval: Benchmarking Formal Translation Across ITPs

ITPEval: first benchmark for formal theorem translation across Lean, Rocq, Isabelle, HOL Light. Evaluates LLMs on statement and proof translation.

viernes, 24 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Evaluación de la traducción formal con ITPEval

The ecosystem of interactive theorem provers (ITPs) has grown remarkably over the last decade, driven by the promise of mathematical and software verification with absolute guarantees. However, fragmentation among systems such as Lean 4, Rocq, Isabelle, and HOL Light has created silos of knowledge and code that hinder both the portability of results and the training of machine learning models for automated theorem proving. In this context, ITPEval emerges as the first benchmark specifically designed to evaluate the automatic translation of theorems and proofs across these four assistants, spanning two distinct logical foundations. This article provides an in-depth analysis of what ITPEval entails, how it works, what challenges it reveals, and how companies like Q2BSTUDIO can apply these methodologies in the development of advanced technological solutions.

ITPEval's proposal is far from trivial. With 1,560 source files and 6,848 theorems organized into two tiers (a controlled tier with axiomatized files that isolate foundational translation difficulty, and an ecosystem tier taken from real libraries that exposes API and proof-style mismatches), the benchmark offers a rigorous testing ground. Initial results are revealing: statement translation reaches 29.1% success at pass@1, while proof translation barely hits 10.5%. This indicates that the main bottleneck is not the underlying logic, but the mismatch between libraries and each system's style. This is where the expertise in custom software development provided by Q2BSTUDIO becomes relevant: each ITP has its own conventions, data types, and reasoning patterns, just as each software project requires careful adaptation to the client's context.

One of the most interesting findings of the study is that native type-checking overestimates semantic fidelity. In translations of statements from various sources to Lean 4, a deterministic BEq (boolean equivalence) check confirmed only 54% of those that had passed type-checking. This underscores the need for more sophisticated validation methods, similar to those used in AI agents to ensure that outputs from generative models maintain the expected meaning. In a business context, Q2BSTUDIO integrates verification and validation techniques into its cloud developments, both on AWS and Azure, to ensure that data transformations and automations do not introduce silent errors.

The benchmark also includes a round-trip autoformalization/autoinformalization study, where Rocq and HOL Light turned out to be easier formalization targets than Lean 4 and Isabelle. Furthermore, using multi-ITP context improved pooled Lean 4 success from 4.8% to 10.6%. These data have direct implications for the design of artificial intelligence systems that learn to translate between formal languages. In cybersecurity, where critical properties need to be verified, having tools capable of moving theorems between systems reduces the risk of incompatibility and allows verification reuse in heterogeneous environments. Q2BSTUDIO, with its cybersecurity division, applies similar principles to validate protocols and configurations in cloud infrastructures, ensuring that security policies remain consistent when migrating between providers.

From a technical perspective, ITPEval relies on a unified verification infrastructure called itpeval, which implements warm backends with state isolation to preserve the native semantics of each artifact. This resembles the microservice architectures that Q2BSTUDIO deploys in cloud AWS/Azure projects, where each service maintains its own context and state but integrates via well-defined APIs. The ability to isolate the translation problem from the underlying logic allows researchers to focus on the specific difficulties of each pair, an approach also followed by the company's custom software development methodology.

Another relevant aspect is the evaluation with state-of-the-art language models, both frontier and open-weight, on 12 directed translation pairs. The pass@1 results for proof translation (10.5%) and statement translation (29.1%) show that complexity is high, but the path is open. In the business context, the ability to automatically translate business rules, queries, or BI (Business Intelligence) logic between different platforms can save months of manual work. Q2BSTUDIO has developed BI/Power BI solutions that precisely face the challenge of unifying heterogeneous data sources; applying formal translation methods could ensure that semantic transformations do not lose meaning.

The future of formal verification lies in interoperability. ITPEval is not just a benchmark but a call to the community to standardize intermediate representations, share libraries, and develop intelligent translators. Companies like Q2BSTUDIO, specialized in process automation, see in this line an opportunity to extend formal verification to industrial environments, where software correctness is not a luxury but a requirement. Moreover, the integration of artificial intelligence in this area —AI agents capable of proposing proofs or correcting translations— will be a continuous field of innovation. Q2BSTUDIO already collaborates with startups and research centers to pilot solutions that combine formal verification with machine learning.

In conclusion, ITPEval marks a milestone in the maturity of the formal translation field among theorem provers. Its metrics reveal where the real difficulties lie and offer a testing ground for the next generation of intelligent assistants. For a development company like Q2BSTUDIO, understanding these challenges and applying similar techniques in its custom software and cloud (AWS/Azure) projects not only improves product quality but opens the door to value-added services such as automatic verification of business logic, formal cybersecurity, and reliable system migration. The publication of the benchmark, verification infrastructure, and evaluation pipelines is a gesture of transparency that will foster collaboration between academia and industry, a path that Q2BSTUDIO is already actively walking.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.