SoftMoR: Deeper and More Efficient Recursive Vision Transformers

SoftMoR improves recursive Vision Transformers: by increasing depth, it achieves higher accuracy on ImageNet with only 1.7M extra parameters.

jueves, 2 de julio de 2026 • 2 min read • Q2BSTUDIO Team

New soft recursion technique optimizes vision transformers

In the rapid advancement of artificial intelligence, visual transformers have shown extraordinary potential for image analysis, but their increase in depth often comes with an excessive growth in parameters. The recent development of SoftMixture-of-Recursions (SoftMoR) and its implementation in SR-ViT proposes an alternative path: using recursion to stack layers efficiently without skyrocketing computational resources. The concept is ingenious: instead of adding independent blocks, parameters are reused across multiple recursive steps, but with a soft mixture mechanism that flexibly leverages all intermediate representations by assigning weights at the token level. This resolves the limitation of previous recursive approaches, which underutilized such representations. Results on ImageNet-1K are telling: going from 1 to 4 recursions, SR-ViT-S accuracy increases by nearly three percentage points with only 1.7 million additional parameters, outperforming much larger models like DeiT-B at a fraction of its size. This parametric efficiency opens the door to deeper and more powerful models without requiring excessive infrastructure.

From a business perspective, this type of innovation has direct implications for deploying AI for businesses. At Q2BSTUDIO, we understand that adopting artificial intelligence requires not only advanced algorithms but also an implementation strategy that optimizes costs and resources. A model like SR-ViT, which achieves high accuracy with a reduced parameter footprint, is ideal for integration into custom applications where the balance between performance and scalability is critical. For example, in industrial vision systems or document analysis, where every millisecond counts and hardware budgets are tight, being able to run a deep visual transformer on devices with limited capacity provides a competitive advantage.

Furthermore, the flexibility of the SoftMoR approach fits perfectly with modern architectures of AI agents that require multimodal processing and real-time adaptability. The ability to mix representations from different recursive depths allows the model to learn when to stop or when to go deeper, reminiscent of dynamic attention mechanisms we already explore in our custom software development. At Q2BSTUDIO, we combine these techniques with AWS and Azure cloud services to offer scalable inference pipelines, and we also apply efficiency principles in our business intelligence services with Power BI, where complex data visualization benefits from lightweight yet accurate models.

Of course, any artificial intelligence deployment must consider cybersecurity. Incorporating recursive models into production environments requires validating that they do not introduce vulnerabilities, and at Q2BSTUDIO we integrate security audits as part of our custom application development. Ultimately, SoftMoR is not just an academic advancement; it represents a practical direction toward deeper and more efficient transformers that businesses can leverage today with the right technological support.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.