HiFA4: 4-bit FlashAttention without training on Ascend NPUs

Discover HiFA4, an operator design that executes FlashAttention in 4 bits on Ascend NPUs HIF4, reducing latency and maintaining precision in inference of

martes, 7 de julio de 2026 • 2 min read • Q2BSTUDIO Team

LLM optimization with 4-bit quantization on Ascend

The advancement of large language models (LLMs) has brought inference efficiency into the spotlight. 4-bit quantization is a promising technique to reduce memory consumption and accelerate processing, but it often introduces precision losses that hinder its adoption in production environments. In this context, HiFA4 emerges, a post-training design that executes key FlashAttention operations (QK^T and PV) as 4-bit matrix multiplications on specialized hardware such as Ascend NPUs, maintaining numerical stability through innovative techniques.

HiFA4 relies on two complementary mechanisms. On one hand, Smooth-QK rescales the Q and K matrices after RoPE to transfer the quantization difficulty from the K matrix to the Q matrix, avoiding costly online reductions. On the other hand, P-Reordering accumulates the softmax normalizer from the same quantized weights used in the PV multiplication, eliminating the need for high-precision reconstructions and merging the normalizer calculation into the matrix operation itself. This approach significantly reduces output errors and enables more efficient execution on the hardware.

Results on standard benchmarks show that HiFA4 recovers a significant portion of the accuracy loss caused by direct quantization. For example, in the Qwen3-8B model, the decision drift induced by quantization was reduced, decreasing accuracy regressions by 57% and achieving a loss of only 0.70 percentage points compared to the original BF16 precision. In other models such as Gemma2-9B, LLaMA3.1-8B, and Mistral-7B, consistent improvements were also observed, with regression reductions between 27% and 52%. These data confirm that it is possible to quantize to 4 bits without compromising prediction quality.

For companies seeking to implement high-performance generative AI, these optimizations are essential. Running large models with fewer resources translates into lower operational costs and faster response times. In this context, Q2BSTUDIO offers comprehensive solutions ranging from artificial intelligence development for businesses to the integration of custom AI agents and business intelligence services with Power BI. Additionally, we have experience in AWS and Azure cloud services to deploy optimized models on scalable and secure infrastructure. Our team develops custom software and custom applications that incorporate these capabilities, also addressing critical aspects such as cybersecurity.

The evolution of quantization and specialized hardware architecture, such as Ascend NPUs, opens a new horizon for efficient LLM inference. Companies that adopt these technologies can gain a significant competitive advantage. At Q2BSTUDIO, we are prepared to accompany this process, combining deep technical knowledge with a practical focus on business results. The key lies in understanding that model optimization is not just an academic exercise, but a real lever to make artificial intelligence viable in production.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.