The optimization of generative models has gained strategic relevance in the current artificial intelligence landscape. Diffusion Transformers (DiTs) stand out for their ability to generate high-quality images, but their high computational cost limits their deployment in real-world environments. Post-training pruning techniques offer a promising way to reduce the processing load without sacrificing performance, although they present specific challenges due to the unique architecture of these models. Unlike traditional methods developed for large language models (LLMs), DiTs distribute their parameters very differently, with magnitudes that hinder the application of importance metrics based on successive approximations. Furthermore, conventional pruning granularity does not respect the internal structure of these transformers, causing significant degradation in the quality of generated images.
Recent research proposes a novel approach that redefines both saliency criteria and pruning granularity. From an energy perspective, a metric is designed that balances the contribution of weights and activations, more precisely identifying the key elements of the network. At the same time, it is observed that weights in the two-dimensional space form clustering patterns, allowing for a cluster-aware granularity to achieve effective sparse allocation. Experimental results show that, even with sparsity levels of 50%, the loss in CLIP score is minimal, clearly outperforming other recent pruning methods. This opens the door to efficient implementations on resource-constrained devices.
For companies looking to incorporate these capabilities into their products, having a specialized technology partner makes the difference. At Q2BSTUDIO we offer AI for businesses that integrates optimized generative models, as well as AI agents capable of operating in real time. Our team develops custom applications that leverage pruning and quantization techniques to reduce costs without losing precision. We also provide AWS and Azure cloud services that facilitate the scaling of these solutions, and cybersecurity to protect sensitive data during training and inference. If your organization needs to transform data into decisions, our business intelligence services with Power BI can visualize model performance. Additionally, we develop custom software for each use case, from network pruning to workflow automation. To learn more about how we apply these innovations, visit our custom applications section.

.jpg)



