CEDI: Contextualized Evaluation of MLLMs through Dynamic Interactions

Discover how CEDI framework evaluates Vision Language Models through dynamic, multi-turn interactions, revealing more realistic visual hallucinations and

domingo, 26 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Cómo las interacciones dinámicas revelan más alucinaciones visuales

In the fast-paced evolution of artificial intelligence, multimodal language models (MLLMs) have achieved impressive milestones on controlled benchmarks. However, the gap between those test labs and the real world is becoming increasingly evident. This is where CEDI (Contextualized Evaluations of MLLMs through Dynamic, multi-round Interactions) emerges—a framework that reimagines evaluation as a three-party dialogue between the evaluated model, an automated examiner, and a grader. This approach not only uncovers visual hallucinations that static tests overlook but also reveals how they accumulate over long contexts and self-reinforcing dialogues. For a company like Q2BSTUDIO, which develops custom software and AI solutions, understanding these limitations is key to designing robust systems that operate in real environments, where continuous interaction and contextualization are the norm.

CEDI's proposal is based on a graph-based representation of the task state. The examiner navigates state transitions, combining clarification requests, adversarial probes, and other strategies to gather performance evidence. Instead of a fixed battery of questions, the examiner dynamically adapts to the model's responses. This is analogous to how the development of AI agents requires conversational design that anticipates drifts and corrects errors in real time. Empirical results show that CEDI finds significantly more hallucinations than conventional static evaluation, and these hallucinations more closely resemble those arising in practical use cases, such as virtual assistants or image-based decision support systems.

One of the most relevant findings is that hallucinations tend to accumulate in long contexts. When a model interacts over multiple rounds, the dialogue history can reinforce its own errors, creating a downward spiral of accuracy. This has direct implications for developments like customer service chatbots or real-time visual analysis tools. Q2BSTUDIO, as a company specializing in custom software, integrates contextualized evaluation practices to ensure that the AI solutions it deploys on AWS/Azure cloud not only meet benchmarks but maintain reliable performance under prolonged interactive conditions.

The vulnerability of models to questions that require premise rejection or refusal is another critical point. For example, if asked to evaluate a false statement about an image, the model tends to confirm the premise rather than reject it. This behavior is a risk in applications where factual accuracy is vital, such as medical image-assisted diagnosis or legal document review. Cybersecurity is also affected: a model that cannot say 'no' can be manipulated through deceptive prompts, generating incorrect outputs that compromise system integrity. Q2BSTUDIO addresses these challenges through AI audits and penetration testing, integrating cybersecurity as a pillar in the software development lifecycle.

CEDI is not just an evaluation tool—it is a paradigm shift. By modeling interaction as a state graph, it allows designing test strategies as diverse as real situations. This recalls the logic of Business Intelligence (BI) dashboards: static data is not enough; one needs dynamic exploration, follow-up questions, and anomaly detection. Integrating CEDI with Power BI platforms, for instance, could automate the validation of dashboards that rely on multimodal descriptions. For companies developing custom BI solutions, like those offered by Q2BSTUDIO, this type of contextualized evaluation is a key differentiator to ensure the reliability of AI-generated reports.

From a technical perspective, CEDI implements an examiner that traverses a task's state space, generating questions that can be clarifying, confirmatory, contradictory, or adversarial. The grader then evaluates the coherence and accuracy of responses based on the accumulated context. This process can be computationally intensive, but the availability of cloud infrastructure such as AWS or Azure facilitates its scalability. Q2BSTUDIO leverages these services to deploy continuous evaluation pipelines integrated with CI/CD cycles, ensuring that every update of a multimodal model is tested under realistic conditions. Process automation, through robotic process automation (RPA) tools and AI agents, directly benefits from these evaluation systems, as they detect subtle failures that a static test would never capture.

CEDI's impact goes beyond academic research. In the business sector, where adoption of multimodal assistants is rapidly growing, having an evaluation method that reflects real human interaction is a competitive advantage. For example, in a product recommendation system combining images and text, hallucinations can lead to inappropriate suggestions. With contextualized evaluation, these weaknesses can be identified before putting the system into production. Q2BSTUDIO, with its expertise in custom applications and cloud, helps companies implement such evaluations, tailoring the examiner according to the business domain and expected interaction patterns.

In conclusion, CEDI represents a significant step toward ecologically valid evaluations of MLLMs. By focusing on dynamic and contextualized interaction, it reveals vulnerabilities that static benchmarks hide. For Q2BSTUDIO, integrating these principles into software development and AI solutions is a priority, ensuring that systems are not only accurate under ideal conditions but robust in the real world. Combining CEDI with cloud services, cybersecurity, and BI enables building more trustworthy AI ecosystems, where transparency and adaptability are the norm.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.