Robotics has advanced enormously in recent years, but one of the most complex challenges remains object manipulation in cluttered and unstructured environments. Traditional computer vision-based systems often fail when objects do not visually resemble those in their database, or when stacking and occlusion conditions prevent reliable grasping. In this context, the approach known as Agentic RAG-VLM (Retrieval-Augmented Generation with vision-language models and self-reflective planning) proposes an innovative solution that integrates semantic reasoning with physical grasping capabilities.
The key to this architecture lies in three interconnected components. First, a hierarchical affordance retrieval system (HAA-RAG) that, instead of comparing images, evaluates attributes such as object type, material, fragility, and regions suitable for grasping. This allows selecting manipulation strategies based on the object's actual functionality, not just its appearance. Second, a spatial constraint reasoner that builds a scene graph from the vision-language model's perception, translating proximity, occlusion, and support relationships into concrete grasping parameter adjustments. Third, a self-reflective pipeline that classifies failures into 14 categories and applies adaptive retries at three levels, closing the control loop to achieve continuous improvement.
The results obtained in a benchmark of 12 tasks show an overall success rate of 78.3%, surpassing by more than 53 percentage points methods that only use vision-language models without physical reasoning. This demonstrates that the combination of affordance-based retrieval, spatial reasoning, and agentic retrieval is essential for robust robotic manipulation.
For companies looking to bring artificial intelligence to their industrial or logistics processes, this type of solution represents a qualitative leap. It is not just about recognizing objects, but understanding how to interact with them safely and efficiently. At Q2BSTUDIO we develop AI for businesses that integrate contextual reasoning and autonomous learning, adapting to changing environments. Our team combines custom applications with computer vision and automation capabilities, and we can deploy these systems on aws and azure cloud services to ensure scalability and low latency.
Furthermore, the self-reflective architecture of Agentic RAG-VLM fits perfectly with the concept of AI agents that learn from their mistakes and improve with experience. In our implementations, we combine custom software with language and vision models, and apply cybersecurity techniques to protect sensitive data processed at the edge or in the cloud. For business areas, we offer business intelligence services with power bi to visualize robot performance metrics and optimize their operation.
In short, self-reflective planning in robotic grasping is not just cutting-edge research, but a practical necessity for intelligent automation. At Q2BSTUDIO we are prepared to help companies integrate these capabilities into their workflows, providing technical knowledge and robust solutions.



