Z-1: Efficient Reinforcement Learning for VLA Models

Z-1 applies GRPO to improve VLA models, achieving 80.6% success in RoboCasa, surpassing SFT by 13.2%. Efficient reinforcement learning.

miércoles, 1 de julio de 2026 • 2 min read • Q2BSTUDIO Team

VLA policy optimization via GRPO in RoboCasa

In the field of intelligent robotics, Vision-Language-Action (VLA) models are marking a turning point by integrating natural language instructions, visual observations, and continuous control. However, most current systems are limited to cloning behaviors from fixed demonstrations, which prevents the policy from learning from its own mistakes. This is precisely the gap that Z-1 addresses, a post-training reinforcement learning framework specifically designed for flow-based VLA models. Unlike traditional approaches, Z-1 combines an initial supervised fine-tuning (SFT) phase with public demonstrations and subsequently applies a Group Relative Policy Optimization (GRPO) strategy adapted to each task. The result is remarkable: across 24 standard tasks from the RoboCasa benchmark, Z-1 achieves an average success rate of 80.6%, surpassing its SFT initialization by 13.2 percentage points and even improving upon the latest published models.

Behind this advancement are key technical decisions that deserve attention. Efficiency in online deployment is achieved through rollout constructions with shared prefixes, tree-based trajectory branching, and completion-aware reward calibration. Furthermore, the selective joint training of the language and vision model with the action expert allows the system to continuously refine its behavior. This paradigm demonstrates that systematic reinforcement learning can enhance VLA policies without the need for additional private demonstrations, opening the door to more adaptive and robust robotic applications.

For companies looking to integrate AI for business solutions into their processes, approaches like Z-1 illustrate how combining artificial intelligence with reinforcement learning can revolutionize industrial automation. At Q2BSTUDIO, as a company specialized in software development and technology, we understand that implementing VLA models in real-world environments requires not only cutting-edge algorithms but also a solid infrastructure. That is why we offer custom applications that integrate everything from AI agents to control systems, as well as cybersecurity and AWS and Azure cloud services. Our team also works with business intelligence services and tools like Power BI so that the data generated by these systems can be effectively analyzed and visualized.

Looking to the future, the evolution of VLA models will depend on frameworks like Z-1 that enable continuous and autonomous learning. At Q2BSTUDIO, we are prepared to advise and develop AI agents that not only execute predefined tasks but also improve with experience. If your organization seeks to adopt these technologies, our engineering team can help you design the path from proof of concept to production, ensuring performance, security, and scalability through custom software. Intelligent robotics is no longer futurology: it is a tangible opportunity that, when properly implemented, transforms processes and generates real competitive advantages.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.