NanoZK: Privacy-Preserving Verifiable Inference for LLMs with ZK Proofs

NanoZK introduces layerwise ZK proofs for LLM inference, hiding weights and activations. Efficient, parallelizable, and constant-size proofs. Learn more.

miércoles, 22 de julio de 2026 • 3 min read • Q2BSTUDIO Team

Privacidad en modelos de lenguaje: pruebas capa por capa

Verifying the inference of large language models (LLMs) without exposing sensitive data is one of the most relevant challenges in enterprise artificial intelligence. Recently, systems like NanoZK propose a solution based on zero-knowledge proofs (ZKP) that allows clients and auditors to verify that a provider executed the advertised model on a committed input without learning weights or activations. This advancement opens new possibilities for secure AI adoption in sectors such as banking, healthcare, and public administration. At Q2BSTUDIO, a company specialized in custom software development and cloud solutions, we analyze how these technologies can be integrated into business ecosystems.

NanoZK introduces a layerwise proof framework that decomposes transformer inference into independently provable layers linked by a SHA-256 commitment chain. This yields constant-size sub-circuit proofs (3.5-3.7 KB) and a total of approximately 83 KB for a 12-layer model. Compared to previous monolithic approaches requiring 101-126 KB, NanoZK offers much greater parallelization, facilitating its integration into large-scale verification processes. Additionally, 16-bit lookup-table approximations are designed for softmax, GELU, and normalization, with perplexity degradation below 1e-4 across six model/dataset combinations, ensuring near-identical accuracy to original inference.

In terms of performance, the MLP sub-circuit proof on CPU takes about 6.3 seconds in prove-only mode and about 43 seconds including setup, with verification requiring only about 22 ms regardless of layer width. Attention prove-only time ranges from 0.9 seconds for dimensionality 16 to 184 seconds for d=256. For full blocks, a GPU time of approximately 68 seconds per block at d=768 is projected, leveraging an MSM speedup of 15-30x based on Icicle estimates. This performance makes real-time verification feasible even for large models.

An innovative aspect of NanoZK is the use of a Fisher-information-guided audit budget, which allows prioritizing layers that most impact model output, although full soundness requires verifying every layer. This offers a practical efficiency tool for rapid audits in resource-constrained environments without sacrificing final guarantees.

The privacy scope of NanoZK is notable: it hides weights and activations from both verifiers and external auditors, but does not hide the prompt from the prover, complementing other approaches like homomorphic encryption (HE) or multiparty computation (MPC). This feature is especially useful in environments where trust in the AI provider is partial, such as cloud services or outsourced natural language processing.

For companies seeking to implement trustworthy AI solutions, the ability to audit an LLM's behavior without exposing critical data is a key differentiator. Thanks to NanoZK's architecture, it is possible to build custom applications that integrate real-time verification, for example in virtual assistants or recommendation systems. Combined with AWS or Azure cloud infrastructure, scalable environments are deployed where model integrity is verified through cryptographic proofs. Additionally, in the cybersecurity domain, verifying that a model has not been tampered with protects against poisoning or model replacement attacks. At Q2BSTUDIO we offer consulting and implementation services for these technologies, as well as integration with Business Intelligence platforms like Power BI to add trust layers to AI-generated analytics.

Research in zero-knowledge proofs for LLMs is advancing rapidly. NanoZK represents a step towards verifiable and private inference, and its layerwise architecture is especially promising for integration into autonomous AI agent systems. These agents, combining reasoning, planning, and execution, benefit from cryptographic verification that ensures decisions based on LLMs are authentic. At Q2BSTUDIO we closely follow these innovations to offer our clients software solutions that combine transparency, performance, and privacy. If you are interested in applying these techniques to your business, contact us to explore custom developments in artificial intelligence, cloud, and cybersecurity.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.