Introduction
Artificial intelligence (AI) has transformed healthcare information delivery, providing patients new ways to access medical knowledge. AI tools like ChatGPT may increasingly replace traditional sources, such as patient information leaflets. However, while leaflets—such as those created by the International Urogynecological Association (IUGA)—undergo quality assessments, the accuracy of AI-generated content in urogynecology remains largely unexamined. To address this gap, we aimed to enhance AI-generated content in urogynecology by integrating high-quality, vetted patient education materials into an AI framework and compare the performance of a retrieval-augmented ChatGPT model with the standard ChatGPT model.
Methods
We developed a retrieval-augmented ChatGPT model using IUGA patient leaflets. This model prioritized IUGA-sourced content, permitting supplemental information only if non-contradictory, and provided hyperlinks for further reading. Ten common urogynecology-related questions were submitted to both models. Responses were collected and source information redacted for blinded assessment. Six board-certified urogynecology subspecialists rated each response for accuracy, clarity, relevance, completeness, and usefulness using the validated Quality Analysis of Medical Artificial Intelligence (QAMAI) tool [1]. Preferences for model’s response were also recorded. After scoring, references were revealed, and quality assessed. Wilcoxon signed-rank tests compared scores, with p-value<0.05 considered significant.
Results
The retrieval-augmented model achieved a higher median QAMAI score (22 [IQR 19-25]) than the standard model (16 [IQR 13-18], p<0.01), outperforming the standard model in all domains (Table 1). In the sensitivity analysis, after removing the Reference domain scores, the retrieval-augmented model continued to demonstrate a significantly higher total QAMAI score compared to the standard model (18 [IQR 16–20] vs. 14.5 [IQR 11–17], p < 0.01). Under blinded conditions, raters preferred the retrieval-augmented model responses in 81% of cases.
Conclusions
Integrating vetted educational materials into AI significantly improves response quality. Retrieval-augmented models can enhance the of AI-generated medical information.
References
1. Vaira LA, Lechien JR, Abbate V, Allevi F, Audino G, Beltramini GA, Bergonzani M, Boscolo-Rizzo P, Califano G, Cammaroto G, Chiesa-Estomba CM, Committeri U, Crimi S, Curran NR, di Bello F, di Stadio A, Frosolini A, Gabriele G, Gengler IM, Lonardi F, Maglitto F, Mayo-Yáñez M, Petrocelli M, Pucci R, Saibene AM, Saponaro G, Tel A, Trabalzini F, Trecca EMC, Vellone V, Salzano G, De Riu G (2024) Validation of the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool: a new tool to assess the quality of health information provided by AI platforms. Eur Arch Otorhinolaryngol 281 (11):6123-6131. doi:10.1007/s00405-024-08710-0