How Faithful Are LLM-Generated Clinical Trial Summaries?

Discover how LLMs like GPT-4o, Claude, and Gemini perform in summarizing clinical trials for diverse audiences, and how knowledge graph augmentation improves

martes, 28 de julio de 2026 • 3 min read • Q2BSTUDIO Team

Benchmark de fidelidad en resúmenes de ensayos clínicos

Artificial intelligence has made a strong entry into the healthcare sector, and large language models (LLMs) are increasingly used to generate summaries of clinical trials targeting physicians, patients, and payers. However, the tendency of these models to 'hallucinate' (plausible but false information) poses a critical risk in a context where accuracy can have direct consequences on people's health. A recent academic study has proposed a framework for evaluating the faithfulness of these summaries, analyzing the performance of models such as GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash. The results reveal that the dominant failure mode is unsupported claims, highlighting the need for solutions that guarantee the veracity of AI-generated information.

From a technical and business perspective, this problem is not just a research challenge but an opportunity for software development companies like Q2BSTUDIO to provide robust solutions. Building systems that combine LLMs with verified knowledge sources, such as clinical trial databases, can drastically reduce hallucinations. A promising approach is the integration of knowledge graphs that act as a verification layer, similar to what the study authors propose by developing a graph-augmented retrieval system. This type of architecture not only improves faithfulness but also allows auditing the origin of each claim, a mandatory requirement in regulated sectors such as pharmaceuticals.

For a technology company, implementing these solutions requires expertise in several areas. First, the development of custom software that integrates LLMs with clinical data pipelines. Second, the need to deploy these systems in scalable cloud environments, either AWS or Azure, to handle large volumes of requests and ensure availability. Additionally, cybersecurity is essential: clinical trial data is highly sensitive and must comply with regulations like HIPAA or GDPR. A company offering cybersecurity services, such as those provided by Q2BSTUDIO, can help protect these information flows. Finally, business intelligence (BI) with tools like Power BI allows visualizing faithfulness metrics and model performance, facilitating informed decision-making.

The reference study measures faithfulness using a six-dimension scheme, finding that the average unsupported claim score is 1.55 out of 3 across all models. This indicates that even the most advanced LLMs generate content that does not match the original data. The improvements observed when incorporating the augmented retrieval system (cross-encoder NLI) are statistically significant but modest in absolute terms (entailment increase of 0.0125). This suggests there is still a long way to go and that solutions must be multimodal: combining audience-specific prompting techniques, external verification, and continuous model refinement.

From the perspective of a company like Q2BSTUDIO, the challenge is not only technical but also one of user experience design. A summary aimed at a specialist physician should include statistical details and levels of evidence, while for a patient it must be understandable and empowering. AI agents can dynamically personalize these summaries, but always under a verification framework. Implementing cloud services (AWS/Azure) allows orchestrating these agents efficiently, scaling on demand and ensuring low latency. Furthermore, process automation through custom software can integrate these summaries directly into electronic health records or clinical decision platforms.

Another key aspect is transparency. Current models are essentially black boxes, which hinders trust. Enterprise solutions must include traceability mechanisms, where each generated claim is linked to the original clinical trial source. This is where the combination of structured knowledge bases (such as the Aggregate Analysis of ClinicalTrials.gov database used in the study) and LLMs can make a difference. Q2BSTUDIO, with its expertise in artificial intelligence and software development, can help build these platforms by integrating NLI verification modules and knowledge graphs.

The future of AI-generated clinical trial summaries will depend on companies' ability to go beyond general models and develop specific, auditable, and secure solutions. The study results are a call to action: we cannot blindly delegate to LLMs without a robust validation infrastructure. Collaboration between researchers, developers, and technology providers like Q2BSTUDIO will be essential to make AI a reliable ally in healthcare, rather than a source of misinformation with serious consequences.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.